Abnormity early warning method and device of batch system, electronic equipment and storage medium
By using anomaly prediction models and dynamic threshold comparison technology in batch systems, combined with the classification capabilities of anomaly recognition models, the problem of low detection accuracy in existing technologies is solved, and efficient anomaly detection and timely warning report generation are achieved.
Patent Information
- Application Number
- CN202510771225.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-12
AI Technical Summary
In the prior art, the batch system anomaly detection method based on statistical principles and threshold comparison has low detection accuracy, resulting in low accuracy of detection results and prone to false alarms.
An anomaly prediction model is used to predict the operating data of the batch system and generate a time series of operating data in the future time period; the predicted time series data is compared with the dynamically updated operating data threshold to identify abnormal moments; the anomaly recognition model is used to classify the abnormal time series data and identify the anomaly type; emergency strategies are matched based on the anomaly type to generate an anomaly warning report.
Through accurate future predictions and dynamic threshold comparisons, the accuracy of anomaly detection is improved, anomalies can be identified and classified in a timely manner, effective early warning reports can be generated, and the stability and reliability of the system are improved.
Smart Images

Figure CN120631707A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence or other related technical fields, and in particular to an abnormality early warning method for a batch system, a device thereof, an electronic device, and a storage medium. Background Art
[0002] In today's era of accelerating digital transformation, particularly with the increasing electronicization and intelligentization of financial services, batch systems have become a critical infrastructure supporting daily operations across various industries. Batch systems are automated systems that perform specific tasks within a predetermined timeframe. These tasks often involve large amounts of data processing and transactions. For example, in the manufacturing industry, batch systems can be used to perform operations such as inventory counting, sales data analysis, and supply chain management. In the financial sector, batch systems can be used for batch payments, account settlements, report generation, and data backups. Batch systems are characterized by their high degree of automation, scheduled tasks, and batch processing, aiming to improve efficiency, reduce costs, and ensure data consistency and accuracy.
[0003] Especially in the financial sector, where financial institutions handle large volumes of data and numerous, often complex, financial operations, batch systems are crucial for ensuring business continuity and operational efficiency. They not only process daily business data but also handle critical financial settlement and compliance audits. Therefore, the stability and reliability of these systems are directly linked to the smooth functioning of banks and the security of customer funds. Furthermore, with the increasing diversification and complexity of financial services, the scale and frequency of batch processing are also increasing, placing higher demands on the system's stable operation and rapid processing capabilities. Therefore, in the daily operation and maintenance of batch systems, monitoring and anomaly detection are essential to ensure their efficient and stable operation. Any unforeseen performance degradation or system failure could result in data loss, transaction delays, or even system crashes, resulting in significant financial losses.
[0004] In related technologies, anomaly detection and early warning are performed on batch systems through statistical principles and threshold comparison. However, when faced with ever-changing system environments and business needs, this method lacks flexibility and has low detection accuracy, resulting in low accuracy of detection results and prone to false alarms.
[0005] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0006] The embodiments of the present invention provide a batch system abnormality warning method and its device, electronic device and storage medium, so as to at least solve the technical problem of low detection accuracy of batch system abnormality detection methods based on statistical principles and threshold comparison in related technologies.
[0007] According to one aspect of an embodiment of the present invention, a method for abnormality warning of a batch system is provided, comprising: collecting operating data of the batch system at various operating moments, and preprocessing the operating data to obtain operating series data; inputting the operating series data into an abnormality prediction model, outputting operating data of the batch system at various target moments in a future time period, and obtaining predicted time series data, wherein the abnormality prediction model is a pre-built model for predicting operating series data of the batch system in a future time period; comparing the operating data at each target moment in the predicted time series data with a preset operating data threshold to obtain a comparison result, and determining target moments with abnormalities and their corresponding operating data based on the comparison result, wherein the operating data threshold is dynamically updated; constructing abnormal time series data based on the target moments with abnormalities and their corresponding operating data, and inputting the abnormal time series data into an abnormality recognition model to output an abnormality type of the batch system, wherein the abnormality recognition model is a pre-built model for classifying operating abnormalities of the batch system; matching an abnormality emergency strategy based on the abnormality type, and generating an abnormality warning report based on the abnormality type and the abnormality emergency strategy.
[0008] According to another aspect of an embodiment of the present invention, an abnormality warning device for a batch system is also provided, comprising: a collection unit for collecting operating data of the batch system at each operating moment, and preprocessing the operating data to obtain operating time series data; a first output unit for inputting the operating time series data into an abnormality prediction model, outputting the operating data of the batch system at each target moment in a future time period, and obtaining predicted time series data, wherein the abnormality prediction model is a pre-built model for predicting the operating time series data of the batch system in a future time period; a comparison unit for comparing the operating data of each target moment in the predicted time series data with a preset operating time series data; The data threshold is compared to obtain a comparison result, and the target moment of the abnormality and the corresponding operating data are determined based on the comparison result, wherein the operating data threshold is dynamically updated; a second output unit is used to construct abnormal time series data based on the target moment of the abnormality and the corresponding operating data, and input the abnormal time series data into the abnormality recognition model to output the abnormality type of the batch system, wherein the abnormality recognition model is a pre-built model for classifying the operating abnormalities of the batch system; a generation unit is used to match the abnormal emergency strategy based on the abnormality type, and generate an abnormal warning report based on the abnormality type and the abnormal emergency strategy.
[0009] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is further provided, wherein the computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute any of the above-mentioned abnormal warning methods for batch systems.
[0010] According to another aspect of an embodiment of the present invention, an electronic device is also provided, comprising one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned abnormality warning method for a batch system.
[0011] In the present application, the following steps are taken: first, the operating data of the batch system at each operating moment is collected, and the operating data is pre-processed to obtain operating time series data, the operating time series data is input into the anomaly prediction model, and the operating data of the batch system at each target moment in the future time period is output to obtain predicted time series data, wherein the anomaly prediction model is a pre-built model for predicting the operating time series data of the batch system in the future time period, and then the operating data of each target moment in the predicted time series data is compared with the preset operating data threshold to obtain a comparison result, and based on the comparison result, the target moment with the abnormality and its corresponding operating data are determined, wherein the operating data threshold is dynamically updated, and abnormal time series data is constructed based on the target moment with the abnormality and its corresponding operating data, and the abnormal time series data is input into the anomaly recognition model, and the abnormal type of the batch system is output, wherein the anomaly recognition model is a pre-built model for classifying the operating anomalies of the batch system, and finally, the abnormal emergency strategy is matched based on the abnormality type, and an abnormal warning report is generated based on the abnormality type and the abnormal emergency strategy.
[0012] In this application, by analyzing real-time operating data through an anomaly detection model, it is possible to effectively process non-stationary time series data, capture trend changes and seasonal fluctuations in operating data, and make accurate future predictions. By comparing the predicted time series data with the dynamically updated threshold, the system can identify operating data that exceeds the normal range in real time, thereby locating the target moment where the anomaly exists. At the same time, for the target moment where the anomaly exists, the anomaly recognition model is used to carefully classify the anomaly, identify the cause of the anomaly, and generate an anomaly emergency strategy, thereby achieving the purpose of improving the accuracy of anomaly detection, improving the accuracy of predicting the future operating status of the batch system, and achieving the technical effect of timely warning of potential anomalies. This solves the technical problem of low detection accuracy in the batch system anomaly detection method based on statistical principles and threshold comparison in related technologies. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0014] Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing an abnormality early warning method for a batch system is shown;
[0015] Figure 2 is a flow chart of an optional abnormality warning method for a batch system according to an embodiment of the present invention;
[0016] Figure 3 is an architectural diagram of an optional anomaly detection system for a batch system according to an embodiment of the present invention;
[0017] Figure 4 is a schematic diagram of an optional anomaly detection process for a batch system according to an embodiment of the present invention;
[0018] Figure 5 is a schematic diagram of an optional abnormality warning device for a batch system according to an embodiment of the present invention;
[0019] Figure 6 This is a hardware structure block diagram of an optional electronic device (or mobile device) for executing an abnormality warning method for a batch system according to an embodiment of the present invention. DETAILED DESCRIPTION
[0020] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0021] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0022] It should be noted that the abnormality warning method and device for batch systems in this application can be used in the field of artificial intelligence when performing abnormality detection and warning on batch systems based on artificial intelligence, and can also be used in any field other than the field of artificial intelligence when performing abnormality detection and warning on batch systems based on artificial intelligence. This application does not limit the application field of the abnormality warning method and device for batch systems.
[0023] It should be noted that the collected information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in this application are information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation portals for users to choose to authorize or refuse. For example, an interface is set up between this system and relevant users or institutions to provide users with corresponding operation portals for users to choose to agree or refuse the automated decision-making results; if the user chooses to refuse, the expert decision-making process will be entered.
[0024] The following embodiments of the present invention can be applied to various batch system anomaly warning / detection systems / applications / devices. The present invention uses an anomaly prediction model to predict anomalies in batch systems. This model combines the characteristics of differencing, autoregression, and moving average to effectively process non-stationary time series data, capture trend changes and seasonal fluctuations in the data, and intelligently adjust parameters through an automatic parameter optimization mechanism to adapt to the dynamic changes in the batch system state and provide accurate future predictions. At the same time, an anomaly recognition model is used to carefully classify anomalies. Based on the prediction results, anomalies at future times are effectively identified, thereby providing timely warnings of potential anomalies, improving the accuracy of anomaly detection, and effectively providing warnings for batch systems.
[0025] The present invention will be described in detail below with reference to various embodiments.
[0026] Example 1
[0027] According to an embodiment of the present invention, an embodiment of an abnormality warning method for a batch system is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0028] The method embodiment provided in the first embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1The following is a hardware block diagram of a computer terminal (or mobile device) for implementing an abnormality warning method for a batch system. Figure 1 As shown, the computer terminal 10 (or mobile device) may include one or more (illustrated as 102a, 102b, ..., 102n in the figure) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0029] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0030] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the abnormal warning method for batch systems in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the abnormal warning method for batch systems mentioned above. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0031] The transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.
[0032] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).
[0033] Under the above operating environment, this application provides Figure 2 The abnormality warning method for the batch system shown is implemented by the abnormality detection system of the batch system.
[0034] Figure 2 FIG. 1 is a flow chart of an optional abnormality warning method for a batch system according to an embodiment of the present invention. Figure 2 As shown, the method includes the following steps:
[0035] Batch systems undertake critical data processing and task scheduling tasks in numerous industries, including finance, telecommunications, and manufacturing. Any system anomaly can lead to data loss, transaction delays or failures, or even system crashes, severely impacting business operations. Therefore, anomaly detection is a primary step in preventing and mitigating these risks. Batch systems typically process large amounts of data, and their performance directly impacts the speed and efficiency of data processing. Anomaly detection can promptly identify system bottlenecks, optimize resource allocation, avoid data backlogs, and ensure efficient system operation. Therefore, anomaly detection in batch systems is not only a preventative measure but also a key strategy for improving system performance, optimizing resource utilization, ensuring data security, and enhancing service quality. Through continuous monitoring and analysis, batch systems can ensure stability, efficiency, and security when processing massive amounts of data, providing strong support for enterprises' digital transformation and business continuity.
[0036] Step S201 : collecting the operation data of the batch system at each operation moment, and pre-processing the operation data to obtain the operation sequence data.
[0037] In the above step S201, the first step of anomaly detection is to collect data. For batch tasks, during specific periods of time every day, such as after get off work at financial institutions, the business system enters a low-load state and data can be collected safely. The business system will analyze, verify and summarize the batch data to be batch analyzed on that day into processing tasks, and then subdivide the processing tasks of each business into the level of task nodes. Each task node is load balanced on different regional servers. When collecting data, it is necessary to collect virtual machines of specified application groups on each regional server and collect various indicators on the virtual machines, such as network traffic, CPU usage, IO transmission rate, memory usage, etc. at various time points. These data can reflect the resource consumption and performance of the batch system at different time points, thereby obtaining the operation data of the batch system at each moment.
[0038] It's important to note that after acquiring operational data, preprocessing is required to improve data quality. Specific preprocessing operations include missing value filling, noise removal, and time series processing. Through data preprocessing, batch system operational data is converted into smooth, continuous, and low-noise operational series data, providing a higher-quality data foundation for subsequent model training and anomaly detection.
[0039] Furthermore, the steps of preprocessing the running data to obtain the running series data include: detecting missing values in the running data, calculating the missing values based on linear interpolation, and performing missing value processing on the running data; smoothing the noise in the running data through a sliding average filter; and sequentially processing the running data after missing value processing and smoothing according to timestamps to obtain the running series data.
[0040] Specifically, missing values may occur during the operational data collection process. This can be caused by network failures, hardware issues, or data recording errors. The presence of missing values can affect the accuracy of data analysis. Therefore, preprocessing of operational data first involves missing value handling. The operational data is checked and, for missing values, an estimated value is calculated using a linear relationship using the valid data points before and after the missing value. The basic principle of linear interpolation is to assume that the data trend is linear within the time period before and after the missing value. Therefore, a straight line can be used to approximate the true value of the missing point. Furthermore, operational data from batch systems is often subject to interference from various external and internal factors and contains a certain amount of noise. Fluctuations in noisy data can obscure meaningful system behavior, so data smoothing is necessary to reduce the impact of noise. In this embodiment of the present invention, a sliding average filter is used to smooth the noise. First, the size of the sliding window is set based on the characteristics of the operational data and the required accuracy. The window is then moved to different positions in the data sequence. The average value of all data points within the window is calculated. For each center point within the window, the calculated average value is used to replace the original value. This smoothing process effectively reduces the impact of noise. After missing value handling and smoothing, the gaps in the data sequence are filled and noise is effectively suppressed. Next, the processed data is sorted by timestamp to ensure the temporal order of the data, thus forming complete runtime series data. Time series data is the basis for time series forecasting and anomaly detection. By sorting by timestamp, the temporal continuity and sequence integrity of the data can be ensured.
[0041] In step S202 , the operation time series data is input into the abnormality prediction model, and the operation data of the batch system at each target moment in the future time period is output to obtain the predicted time series data.
[0042] In step S202, embodiments of the present invention encode the runtime data and input it into an anomaly prediction model. This anomaly prediction model is a pre-built model for predicting the runtime data of a batch system in a future time period. The model includes three parts: a differential model, an autoregressive model, and a moving average model. The anomaly prediction model combines the characteristics of differential, autoregressive, and moving average to analyze the trend, seasonality, and autocorrelation of the runtime data to predict the runtime data of the batch system in the future (e.g., the next half hour or hour). This process uses the algorithm within the model to predict system indicators at future time points, forming predicted time series data.
[0043] Furthermore, the step of constructing an anomaly prediction model includes: collecting historical running time series data of a batch system, and preprocessing the historical running time series data to obtain sample data; dividing the sample data to obtain a training set and a test set; constructing an initial differential model, an initial autoregressive model and an initial moving average model, and merging the initial differential model, the initial autoregressive model and the initial moving average model to obtain an initial anomaly prediction model; iteratively training the initial anomaly prediction model based on the training set to obtain a trained anomaly prediction model, wherein, during the iterative training process, the model order is determined based on a model order function; testing the trained anomaly prediction model based on the test set to obtain test index data of the trained anomaly prediction model, and when the test index data all meet the preset requirements, obtaining a final anomaly prediction model, wherein the test index data includes at least one of the following: mean absolute error, root mean square error and hit rate.
[0044] Specifically, the anomaly prediction model is pre-built. When building the anomaly prediction model, first, key performance indicator data within the historical time period is collected from the operation records of the batch system, including but not limited to CPU usage, memory usage, disk I / O and network traffic, etc., to obtain historical runtime data, and pre-process it to construct sample data for model training. The sample data is divided according to a preset ratio to obtain a training set and a test set. The training set is used for model learning and parameter adjustment, and the test set is used to evaluate the generalization ability of the model to ensure that the model not only performs well on known data, but also can make accurate predictions on unknown data.
[0045] Next, the initial model is constructed, including an initial difference model, an initial autoregressive model, and an initial moving average model. These initial difference, autoregressive, and moving average models are then merged to form the initial anomaly prediction model. The initial anomaly prediction model is iteratively trained using the training set, adjusting model parameters to minimize prediction error. During training, the optimal order of the anomaly prediction model (d for the difference model, p for the autoregressive model, and q for the moving average model) is dynamically determined using the model order function to balance model complexity and predictive performance. The trained anomaly prediction model is then tested on the test set to verify its predictive ability. Test metrics include mean absolute error (MAE), root mean square error (RMSE), and hit rate, which evaluate the model's predictive accuracy from different perspectives. If the MAE, RMSE, and hit rate obtained from the test meet the preset performance requirements (e.g., MAE is less than a preset MAE threshold, RMSE is within a reasonable range, and hit rate is greater than a preset hit rate threshold), the model is considered trained and the final anomaly prediction model is obtained.
[0046] The difference model in the anomaly prediction model can eliminate the influence of trends and periodicity in non-stationary time series, turning them into stationary time series. The autoregressive model uses data from p historical time points to predict data at future time points, capturing the long-term dependencies of the time series. The moving average model uses the errors from the past q time points to smooth the series and reduce the impact of noise.
[0047] Furthermore, the anomaly prediction model includes three parts: a differential model, an autoregressive model, and a moving average model. The steps of inputting the running series data into the anomaly prediction model and outputting the running data of the batch system at each target moment in the future time period include: inputting the running series data into the anomaly prediction model, and performing smooth processing on the running series data based on the differential model of the anomaly prediction model to obtain smooth time series data; inputting the smooth time series data into the autoregressive model of the anomaly prediction model, and predicting the running data of the batch system at each target moment in the future time period based on the autoregressive relationship mechanism of the autoregressive model to obtain the initial time series data for each target moment in the future time period; inputting the initial time series data into the moving average model of the anomaly prediction model, calculating the random error term at each moment in the future time period through the moving average model, and correcting the running data at each target moment in the future time period based on the random error term to obtain the corrected time series data for each target moment in the future time period; performing differential processing on the corrected time series data based on the differential model of the anomaly prediction model to obtain the running data of the batch system at each target moment in the future time period.
[0048] Specifically, after the runtime series data is input into the anomaly detection model, it first performs periodicity and stationarity tests on the data. Periodicity tests can help determine the appropriate lag operator in the differencing model. For series with periodicity, directly applying the anomaly prediction model to forecast analysis without performing periodicity checks and appropriate processing can lead to biased forecasts. Periodicity checks and adjustments can significantly improve forecast accuracy, especially when early warning or resource allocation planning is required. If the data exhibits significant periodicity, selecting an appropriate lag operator can eliminate seasonal effects and stabilize the series.
[0049] Through stationarity verification, for non-stationary time series data, a differential operator is introduced to perform differential processing to eliminate trends in the time series data and eliminate the non-stationarity of the data, so that subsequent autoregressive and moving average models can make accurate predictions on stationary data. Non-stationary time series usually include trend or seasonal components, which will affect the predictive ability of the model. Through stationarity verification, these components can be identified and appropriate preprocessing (such as differential processing) can be taken to eliminate them, thereby improving the accuracy of model predictions. If the data is stationary, then complex differential operations can be avoided when building the model, thereby simplifying the model structure, reducing the model's computational complexity and the required training time, and also helping to avoid overfitting problems.
[0050] After periodicity and stationarity checks, running series data with significant periodicity and / or nonstationarity are fed into the differencing model. The differencing model in the anomaly prediction model processes the input time series data to eliminate nonstationarity. This step typically involves differencing the data until the series becomes stationary. Differencing removes trends and seasonal variations from the series, making autocorrelations easier to capture in a stationary series. This allows subsequent autoregressive and moving average models to accurately predict on stationary data, thereby improving the accuracy of forecast results.
[0051] Next, the autoregressive model in the anomaly prediction model is used to predict the stationary time series data after the stationary processing. The autoregressive model analyzes the correlation between historical data and predicts the operating data of each target time in the future time period based on the p-order autoregressive relationship mechanism, thereby obtaining the initial time series data.
[0052] The initial time series data is then fed into the moving average model of the anomaly prediction model. The moving average model calculates random error terms for each future moment using a q-order mechanism. This error term reflects the uncertainty between the predicted time series data and the actual system operating state. Using the calculated random error terms, the moving average model corrects the initial time series data to more accurately reflect the actual future operating state of the batch system. The resulting corrected time series data improves the robustness and accuracy of the prediction.
[0053] After the forecast is complete, the differencing model within the anomaly prediction model re-engages, performing an inverse differencing operation on its forecast output. This operation reincorporates the trend and seasonality eliminated during the differencing process to restore the data's original characteristics. This inverse processing step ensures that the final output, the forecasted time series data—the operational data for the batch system at each target moment in the future time period—truly reflects the system's operational status.
[0054] In the above steps, the combined application of differential, autoregressive, and moving average models enables accurate prediction of the future operating status of the batch system. The differential model eliminates data non-stationarity, laying the foundation for subsequent predictions; the autoregressive model captures long-range dependencies between data, providing a preliminary framework for prediction; and the moving average model improves the robustness and accuracy of the model by correcting prediction errors. Ultimately, through differential inverse processing, the future operating data output by the model not only takes into account trend and seasonal changes, but also effectively reduces prediction bias, providing strong support for anomaly detection and early warning of batch systems, significantly enhancing system stability.
[0055] Step S203 , comparing the operating data of each target moment in the predicted time series data with a preset operating data threshold to obtain a comparison result, and determining the target moment with abnormalities and its corresponding operating data based on the comparison result.
[0056] In step S203, the operating data at each target moment in the predicted time series data is compared with a pre-set operating data threshold. Abnormal data is determined based on the comparison results. If the operating data at a particular moment exceeds the pre-set operating data threshold, it indicates a potential abnormality in the batch system at that moment. The aforementioned operating data threshold is dynamically updated, meaning it can self-adjust based on the system's recent operating status to accommodate dynamic changes in system status, including but not limited to changes in system load during specific periods such as holiday peaks and promotional events. This dynamic threshold mechanism can improve the sensitivity and accuracy of anomaly detection and avoid false positives or negatives caused by static thresholds.
[0057] By dynamically updating thresholds, it can flexibly adapt to changes in the operating environment, reducing false positives and missed negatives, and ensuring timely response when true anomalies occur. This mechanism is particularly suitable for complex and dynamic environments such as batch systems in financial institutions. It can consider the impact of specific periods such as holidays and promotions on system operating indicators, thereby more accurately identifying anomalies and providing strong guarantees for stable system operation and effective resource management.
[0058] Furthermore, the step of updating the operating data threshold includes: obtaining the predicted value and actual value of the operating data of the batch system at each operating moment; calculating the residual value of the batch system at each operating moment based on the predicted value and the actual value; drawing a control chart based on the residual value at each operating moment, and identifying the residual fluctuation anomaly of the batch system based on the control chart and the residual fluctuation anomaly judgment rule; in the case of residual fluctuation anomaly in the batch system, calculating a new operating data threshold based on the residual value at each moment.
[0059] Specifically, when updating the operating data threshold, the actual operating data values at each runtime are obtained from the batch system's operating records. These values may include key performance indicators such as CPU (Central Processing Unit) utilization, memory utilization, disk I / O (Input / Output) read / write rates, and network traffic. Simultaneously, the predicted operating data values for the same runtime are obtained from the anomaly prediction model. Next, the residual values for the batch system at each runtime are calculated based on the predicted and actual values. The residual value is the difference between the actual operating data and the model's predicted value. It reflects the accuracy of the model's prediction and serves as the basis for subsequent control chart creation and anomaly detection. The residual values at each runtime are used to create a control chart, which displays how the residual values change over time. The control chart and pre-defined residual fluctuation anomaly determination rules are used to determine whether the batch system's residual fluctuations are abnormal. These rules are designed to identify data points that fall outside the normal fluctuation range, indicating potential system issues. Once the control chart indicates abnormal residual fluctuations in the batch system, it means that the current operating data threshold may no longer be applicable and needs to be updated to more accurately reflect the system's actual operating status. Based on the residual values at each moment, new running data thresholds are calculated through statistical analysis. This usually involves recalculating the mean and standard deviation of the residual values and adjusting the upper and lower thresholds based on this.
[0060] The purpose of updating the operating data threshold is to ensure the dynamic adaptability and reliability of the anomaly detection system. By continuously monitoring and analyzing the residual values, the threshold can be adjusted in real time to better reflect the current operating status of the batch system, reducing false positives and false negatives caused by changes in the system environment or model prediction bias, and improving the accuracy of anomaly detection.
[0061] Furthermore, the residual fluctuation abnormality judgment rule includes at least one of the following: the residual values of M consecutive data points in the control chart exceed the residual threshold, N consecutive data points in the control chart are in an upward or downward state, K consecutive data points in the control chart are in an alternating upward and downward fluctuation state, G consecutive points in the control chart are located on the same side of the center line of the control chart, and Y points out of X consecutive points in the control chart exceed the residual critical value, wherein M, N, K, G, X, and Y are all positive integers, X is greater than Y, and the residual critical value is less than the residual threshold.
[0062] Specifically, when determining abnormal residual fluctuations, the batch system is considered to have abnormal residual fluctuations if the control chart meets at least one of the following conditions: First, M consecutive data points exceed the residual threshold. Specifically, in the control chart, observe whether the residual values of M consecutive data points all exceed the preset residual threshold. Here, M is a positive integer representing a set of consecutive time points. This rule aims to detect whether abnormal fluctuations in batch system operating data have a certain degree of persistence and intensity. The longer the duration of the fluctuation exceeds the threshold, the greater the likelihood of an abnormality.
[0063] Second, monitor whether the residual values of N consecutive data points in the control chart show a continuous upward or downward trend. N is also a positive integer, indicating the number of consecutive data points. This rule focuses on identifying directional anomalies in data fluctuations. Even if a single data point does not exceed the threshold, a continuous upward or downward trend may indicate that system performance is deteriorating or improving beyond normal expectations.
[0064] Third, check for K consecutive data points that fluctuate alternately up and down. This means checking whether the residual values of the control chart alternate between rising and falling for K consecutive data points. K is a positive integer representing the number of data points in a group. This rule aims to detect abnormal periodic or oscillatory patterns that may indicate periodic failures or instability within the system.
[0065] Fourth, observe whether G consecutive residual data points are always on the same side of the control chart centerline, where G is a positive integer. This may indicate that the system operating data has deviated from the average state. Even if it does not significantly exceed the upper or lower limits, the persistent deviation still requires attention to prevent potential small anomalies from accumulating into serious problems.
[0066] Fifth, check whether Y consecutive points exceed the residual threshold in a sequence of X data points. Specifically, in a sequence of X data points, check whether at least Y data points have residual values exceeding the residual threshold. Here, X and Y are both positive integers, X>Y, and the residual threshold is less than the residual threshold. This rule provides another perspective for determining the concentration of anomalies. Even if discrete data points exceed a lower threshold, if they occur frequently within a short period of time, they may still indicate system instability.
[0067] Each residual fluctuation anomaly determination rule is designed to capture abnormal patterns in batch system operation data from different perspectives. Whether it's sustained extreme fluctuations, directional trend changes, periodic oscillations, centerline deviations, or short-term high-frequency anomalies, each can be sensitively identified. The combined application of these rules improves the comprehensiveness and accuracy of anomaly detection, promptly identifies potential batch system performance issues, and provides early warning and decision-making support for batch system operations and maintenance.
[0068] Furthermore, the step of determining the target moment with an abnormality and its corresponding operating data based on the comparison result includes: when the comparison result indicates that the operating data of the target moment exceeds a preset operating data threshold, determining that the target moment has an abnormality; and adding the target moment with an abnormality and the operating data corresponding to the target moment to the abnormal data set.
[0069] Specifically, the system determines whether there are any anomalies in the batch system. If the comparison indicates that the operating data at the target time exceeds the preset operating data threshold, this indicates a potential anomaly in the batch system at that target time. The target time and corresponding operating data are recorded and added to the anomaly dataset. The anomaly dataset serves as a critical information repository for subsequent anomaly analysis, root cause tracing, and emergency response.
[0070] Step S204 : construct abnormal time series data based on the target time at which the abnormality occurs and its corresponding operating data, input the abnormal time series data into the abnormality recognition model, and output the abnormality type of the batch system.
[0071] The anomaly recognition model is a pre-built model used to classify abnormal operation conditions of a batch system.
[0072] In step S204, the present embodiment introduces an anomaly recognition model to conduct in-depth analysis of the target time at which an anomaly occurs and identify the anomaly type. This allows for proactive emergency response and provides an emergency solution. After confirming the target time at which the anomaly occurs and its corresponding operational data, this data is combined into a new anomaly time series data sequence. The construction of anomaly time series data aims to comprehensively reflect the system's operational status at the time of the anomaly. The constructed anomaly time series data is input into the anomaly recognition model. The anomaly recognition model is a machine learning or statistical analysis model that has been pre-trained to identify patterns or associations between anomaly data and thus classify operational anomalies in the batch system. The core of the anomaly recognition model is training based on a large amount of historical data and labeled anomaly cases to learn the characteristic manifestations of different anomaly types. When receiving new anomaly time series data, the model can match it against known anomaly patterns and characteristics to identify the anomaly type. After analyzing the anomaly time series data, the anomaly recognition model outputs the anomaly type for the batch system. Anomaly types may include, but are not limited to, hardware failures, software bugs, network congestion, and resource bottlenecks.
[0073] The output of the exception type provides specific guidance for subsequent fault diagnosis and emergency response. It can help operation and maintenance personnel quickly locate the problem based on the exception type and take appropriate measures to address it, thereby shortening fault recovery time and reducing the risk of system interruption.
[0074] Step S205 : matching an abnormality emergency response strategy based on the abnormality type, and generating an abnormality warning report based on the abnormality type and the abnormality emergency response strategy.
[0075] In the above step S205, after determining the abnormal type of the batch system abnormal situation, the preset abnormal emergency strategy is matched based on this type. The matching process is achieved by querying the association rules between the abnormal type and the emergency strategy. For example, if the identified abnormal type is "network congestion", the system will automatically retrieve the emergency strategy related to the network congestion, which may include flow control, line switching, or increasing network bandwidth and other countermeasures. And based on the determined abnormal type and the corresponding emergency strategy, an abnormal warning report is generated. This report is not only a description of the current abnormal situation, but also contains a recommended action guide for the abnormality, clarifying how to respond to the abnormality quickly and effectively. By automatically matching emergency strategies and generating warning reports, the abnormal response time is greatly accelerated, the decision delay is reduced, and timely support is provided for the stable operation of the system.
[0076] Through the above steps, the operation data of the batch system at each operation moment is first collected, and the operation data is preprocessed to obtain the operation series data, which is input into the anomaly prediction model, and the operation data of the batch system at each target moment in the future time period is output to obtain the predicted time series data, wherein the anomaly prediction model is a pre-built model for predicting the operation series data of the batch system in the future time period, and then the operation data of each target moment in the predicted time series data is compared with the preset operation data threshold to obtain the comparison result, and based on the comparison result, the target moment with the abnormality and its corresponding operation data are determined, wherein the operation data threshold is dynamically updated, and the abnormal time series data is constructed based on the target moment with the abnormality and its corresponding operation data, and the abnormal time series data is input into the anomaly recognition model to output the anomaly type of the batch system, wherein the anomaly recognition model is a pre-built model for classifying the operation anomalies of the batch system, and finally the abnormal emergency strategy is matched based on the anomaly type, and an abnormal warning report is generated based on the anomaly type and the abnormal emergency strategy.
[0077] In this embodiment, by analyzing real-time operating data through an anomaly detection model, it is possible to effectively process non-stationary time series data, capture trend changes and seasonal fluctuations in operating data, and make accurate future predictions. By comparing the predicted time series data with dynamically updated thresholds, the system can identify operating data that exceeds the normal range in real time, thereby locating the target moment where the anomaly occurs. At the same time, for the target moment where the anomaly occurs, the anomaly recognition model is used to carefully classify the anomaly, identify the cause of the anomaly, and generate an anomaly emergency strategy, thereby achieving the purpose of improving the accuracy of anomaly detection, improving the accuracy of predicting the future operating status of the batch system, and achieving the technical effect of timely warning of potential anomalies. This solves the technical problem of low detection accuracy in the batch system anomaly detection method based on statistical principles and threshold comparison in related technologies.
[0078] The following describes in detail another optional specific implementation.
[0079] Figure 3 is an architectural diagram of an optional anomaly detection system for a batch system according to an embodiment of the present invention, such as Figure 3 As shown, the anomaly detection system is used to detect and warn of abnormal situations in batch systems, and specifically includes three parts: data collection system, anomaly detection system, and early warning system.
[0080] The data collection system collects the operation data of the batch system based on the data monitoring module. The operation data includes the operation data of each virtual machine in the batch system, including network traffic, CPU utilization, IO transmission rate, memory utilization, etc. at each time point.
[0081] The anomaly detection system includes a data preprocessing module, an operating status prediction module and an anomaly identification module. Based on the data collection system, an anomaly detection system independent of the batch system is built to predict the operating status of the batch system, identify abnormal signals, and classify abnormal signals.
[0082] The early warning system includes an abnormal warning report generation module, which is used to match abnormal situations with abnormal emergency strategies, such as emergency back-cutting system, turning on the flow switch, starting batch standby jobs, etc., thereby generating abnormal warning reports and reporting them to the operation and maintenance terminal.
[0083] Figure 4 FIG. 1 is a schematic diagram of an optional anomaly detection process of a batch system according to an embodiment of the present invention. Figure 4 As shown in Figure 1, the anomaly detection process for batch systems specifically includes:
[0084] Step 1, start;
[0085] Step 2: data preprocessing;
[0086] After a specific time each day, the business system will analyze, verify, and summarize the batch business of the day to construct a batch analysis task. Each batch analysis task is sent to the task node level, and each task node is load balanced on different servers. In order to realize anomaly detection of the batch system, when collecting data, it is necessary to collect the virtual machines of the specified application group on each server and collect various indicators on the virtual machines, such as network traffic, CPU utilization, IO transmission rate, memory utilization, etc. at each time point to obtain the original time series, that is, the running data. Then, the data is preprocessed, including noise removal and missing value filling, to obtain the running series data.
[0087] When filling missing values, linear interpolation is used to fill data. The linear interpolation formula is expressed as:
[0088]
[0089] in, Indicates the missing value to be filled, and k is the interval from the missing value to the next valid value.
[0090] To process the noise of the data set, we can use sliding average filtering to smooth the noise. First, set the window size ω:
[0091]
[0092] in, Represents the denoised data.
[0093] Step 3: Model building;
[0094] Model building involves establishing an anomaly prediction model and determining its parameters. This model allows for more accurate predictions for subsequent time series system analysis. It also includes building an anomaly recognition model.
[0095] The anomaly prediction model integrates three parts: autoregressive model, moving average model and difference model, among which,
[0096] The autoregressive model predicts future data based on the weighting of long historical data, but it is not suitable for processing data with large noise, periodicity, and mutation. The mathematical definition of the autoregressive model is as follows: (assuming x t The sequence is a zero-mean sequence, and the autoregressive model describes {x t The relationship between the sequence at time t and the previous p moments)
[0097]
[0098] Among them, x t represents {x t}sequence value, t uses historical label values during training and t uses predicted label values during prediction. represents the sum of the weighted effects of the time series values of the past p time points on the current time t, a i is the weight coefficient of the lag term of the i-th original data sequence, ε t represents the noise term at time t.
[0099] The moving average model predicts future data based on the weighted past noise, and the noise signal at each time point follows the same distribution, with the same mean and variance. This model can handle data with large noise, periodicity, and sudden changes well, but for data with a long historical trend, the moving average model cannot analyze trend changes. The mathematical definition of the moving average model is:
[0100]
[0101] in, represents the sum of the weighted effects of the noise sequence values of the model at the past q time points on the current time t, where b j is the weight coefficient of the jth lag noise term, ε t The mean is 0 and the variance is σ 2 noise sequence.
[0102] Combining the autoregressive model and the moving average model can improve the accuracy of prediction and save computational effort. It can predict future data based on the weighting of longer historical data. It can also handle data with high noise, periodicity, and mutations. The model formula is defined as follows:
[0103]
[0104] Simplified:
[0105]
[0106] Let α p (B) = 1-a1B-a2B 2 -…-a p B p β q (B) = 1 + b1B + b2B 2 +…+b q B q , where B is the lag operator, and B p x t =x t-p 、B q ε t =ε t-q .
[0107] The combined model can be written as:
[0108]
[0109] α p (B)x t =β q (B)ε t ;
[0110] Among them, α p (B) and β q (B) non-linear correlation; α p (B)≠0,β q (B)≠0; {ε t} has a mean of 0 and a variance of σ 2 Noise sequence E[x t e s ]=0,t<s, that is, the noise at time s is the same as the noise at the previous moment x t Linearly independent.
[0111] However, it is difficult to have a stationary time series in the anomaly detection of various indicators in a real batch system. Therefore, it is necessary to introduce a differential model to eliminate the influence of trends and periodicity in non-stationary sequence data through d-order differential processing and D-order lag processing, so as to make the non-stationary sequence into a stationary sequence.
[0112] So we introduce the difference operator and the hysteresis operator Δ D ,in,
[0113] Represents x t The n-order difference sequence of the sequence, the sequence length is td;
[0114] Δ D x t =(1-B D )x t , indicating x t Sequence and x t-D The difference sequence of the sequence, the sequence length is tD;
[0115] The above operation steps are required to merge the model α p (B)x t =β q (B)ε t {x t The sequence is processed with a d-order difference and a D-order lag. This gives the anomaly prediction model:
[0116] α p (B)(1-B) d (1-B D )x t =β q (B)ε t ;
[0117]
[0118] Among them, {x t}sequence is the prediction target sequence of the anomaly prediction model, p is the order of the autoregressive model, q is the order of the moving average model, d is the difference order, and D is the lag order. In order to transform the non-stationary series into a stationary series, add Δ D This is to remove the trending effect caused by the periodicity of the sequence.
[0119] When building the model, it is also necessary to perform periodicity and stationarity tests on the collected operating data to determine whether the data needs to be differentially processed and to determine the differential operator and lag operator for the differential processing.
[0120] Through periodic verification, data with insignificant periodicity is lagged. First, the following indicator data of the batch system (such as CPU utilization, memory utilization, disk I / O, and network traffic) are expressed in the form of time series. Because the values of the disk I / O and network traffic series are generally large, the values of the series are logarithmically processed.
[0121] Then use the power spectrum density to test its periodicity, the sequence {x t The power spectral density of} can be expressed as the discretized sequence {x k}'s Fourier transform is squared modulo the sequence length:
[0122]
[0123] The above process is to convert {x t The time domain t of} is converted to the frequency domain f. In the power spectrum density diagram, the frequency axis (x-axis) represents different frequency components, and the power axis (y-axis) represents the energy of each frequency component. Periodic components will show obvious peaks at specific frequencies. These peaks correspond to periodic fluctuations in the time series. The frequency of the peak corresponds to the period in the time series. For example, if a peak appears at a frequency Then T is the period of the periodic component, and the period is expressed as If there are multiple peaks, it means that there may be multiple periodic components in the time series. The strength of the periodicity can be judged by the height of the peak. The higher the peak, the more significant the corresponding periodic component in the time series. Therefore, {x t The inverse of the frequency corresponding to the highest peak of the power spectrum density of the sequence Determine the hysteresis operator Δ D The lag order D, that is, the lag operator Δ D The lag order D is set to That's it.
[0124] The non-stationary running series data is differentially processed through the stationary test. The test method used is the unit root test method, that is, whether the test time contains a unit root. If a unit root exists, it is judged to be a non-stationary series. For the autoregressive model:
[0125]
[0126] Remove the white noise term ε t Then we get the homogeneous equation:
[0127] x t -a1x t-1 -a2x t-2 -…-a p x t-p =0;
[0128] Assume the solution is in exponential form x t =r t , substitute into the homogeneous equation:
[0129] r t -a1rt-1 -a2r t-2 -…-a p r t-p =0;
[0130] Divide both sides by r t-p have to:
[0131] r p -a1r p-1 -a2r p-2 -…-a p =0;
[0132] The above polynomial is the characteristic equation of the autoregressive model, and its root r i Determines the stationarity of the data in the model:
[0133] If all roots | r i |<1, it means that the sequence data {x t} is stationary; if |r i |=1 (unit root), it means that the series is not stationary.
[0134] If the model is not stationary, then by introducing the difference operator Set d to a suitable value to eliminate non-stationarity in the sequence data.
[0135] Then we enter the model order determination stage. The Akaike Information Criterion function can be used to determine the model order. This function is a function of the model fit and the model order. When the Akaike Information Criterion function reaches a minimum value, it can be considered that the abnormal prediction model has reached the optimal value, and the model order is determined. This function is expressed as:
[0136] AIC=2k-2lnL(a,b,σ 2 );
[0137] Where k is the order of the model, k = p + q + 1, and the maximum likelihood function of the model is lnL(a, b, σ 2 ), which is used to represent the fitting accuracy of the model.
[0138] Based on the abnormal prediction model of this application, the maximum likelihood function of the model is solved by recursive calculation, and the Akaike information criterion function is rewritten as:
[0139] AIC=2k+nln(σ 2 );
[0140] Where n represents {x t}Number of samples in the sequence.
[0141] Assume a parameter range p∈[0,2], q∈[0,2], and calculate the Akaike Information Criterion value under each set of model orders, and select the model order with the smallest Akaike Information Criterion value as the final model order.
[0142] The model parameters are then estimated.
[0143] The log-joint likelihood function of the anomaly prediction model is:
[0144]
[0145] Where m = max(p,q) is the initial lag order, and the effective sample size is n = Tm.
[0146] The calculation formula of the abnormal prediction model residual is:
[0147]
[0148] The initial residual is ε0,ε -1 ,...,ε -1-n = 0. Calculation starts from t = m + 1 to avoid the influence of the initial value.
[0149] By maximizing lnL(a,b,σ 2 ) About a, b, σ 2 To obtain the coefficients of the abnormal prediction model. Gradient Need to be calculated by chain rule, involving ε t Derivative of the parameter. 2 ) About a, b, σ 2 Take the partial derivative and set it to zero to determine all model coefficients:
[0150] to a j Partial derivative:
[0151]
[0152] to b j Partial derivative:
[0153]
[0154] σ 2 Partial derivative:
[0155]
[0156] Step 4: Model verification;
[0157] The constructed model is tested for validity, including white noise test and parameter test. If the model passes the test, it enters the anomaly detection stage. Otherwise, if the model fails the test, it returns to the model order determination stage to re-determine the order and set the model parameters.
[0158] Step 5: Anomaly detection;
[0159] According to the aforementioned verified anomaly prediction model, the real-time collected batch system operation time series data is input into the anomaly prediction model to predict the predicted operation data at each moment in the future. The abnormal state of the batch system is determined by threshold comparison, and the abnormal time series data is constructed.
[0160] In this embodiment, the threshold data for threshold comparison (i.e., the operating data threshold) is dynamically adjusted. The statistical process control method is adopted to determine whether the residual of the batch system is in a stable state by analyzing the statistical characteristics of the process data. The statistical process is achieved by monitoring the predicted residual (the difference between the actual value and the predicted value) and identifying the residual fluctuation of the batch system based on the control chart, thereby updating the threshold.
[0161] Assume that the predicted value of the anomaly prediction is The actual value corresponding to the predicted value is x t (actual observed values of network traffic, CPU usage, and memory usage), the residual is:
[0162]
[0163] The mean residual is:
[0164]
[0165] The residual standard deviation is:
[0166]
[0167] A control chart is constructed based on the residual value to show the fluctuation of the residual. The moving range is expressed as:
[0168] MR t =|e t -e t-1 |, t is the time series coordinate of the prediction part.
[0169] The center line of the control chart is represented as:
[0170]
[0171] The operating data threshold is expressed as:
[0172]
[0173] If the control chart satisfies any of the following conditions, it is determined that the batch system has abnormal residual fluctuations: First, M consecutive data points exceed the residual threshold, that is, e t > Up or e t <Down; Second, N consecutive data points are in an upward or downward state; Third, K consecutive data points are in an alternating up and down fluctuation state; Fourth, G consecutive points are on the same side of the center line of the control chart; Fifth, Y points out of X consecutive points exceed the residual critical value, that is, or
[0174] If a control chart reveals abnormal residual fluctuations in a batch system, the current operating data thresholds may no longer be applicable and need to be updated to more accurately reflect the system's true operating status. Based on the residual values at each moment, new operating data thresholds are calculated through statistical analysis. This typically involves recalculating the mean and standard deviation of the residual values, and adjusting the upper and lower thresholds based on these values.
[0175] Finally, the anomaly recognition model is used to identify abnormal features in the abnormal time series data and determine the anomaly type.
[0176] Step six, emergency warning;
[0177] Based on the above anomaly types, the abnormal emergency strategy is matched and it is determined whether it is an emergency warning. If so, an email notification is sent. If not, the time series is moved forward to perform the next stage of anomaly prediction and anomaly identification, and at the same time, an abnormal warning report is generated.
[0178] Step seven, end.
[0179] In an embodiment of the present invention, an anomaly prediction model is used to predict abnormal situations of a batch system. The model combines the characteristics of differencing, autoregression and moving average, can effectively process non-stationary time series data, capture trend changes and seasonal fluctuations in the data, and intelligently adjust parameters through an automatic parameter optimization mechanism to adapt to the dynamic changes in the state of the batch system, providing accurate future predictions. At the same time, an anomaly recognition model is used to classify abnormal situations in detail, and abnormal situations at future moments are effectively identified based on the prediction results, thereby timely warning of potential anomalies, improving the accuracy of anomaly detection, and effectively warning the batch system.
[0180] The following describes it in detail with reference to another embodiment.
[0181] Example 2
[0182] An abnormality warning device for a batch system provided in this embodiment includes multiple implementation units, each implementation unit corresponds to each implementation step in the above-mentioned embodiment 1. Its specific implementation method and beneficial effects can refer to the above-mentioned method embodiment and will not be repeated here.
[0183] Figure 5 is a schematic diagram of an optional abnormal warning device for a batch system according to an embodiment of the present invention, such as Figure 5 As shown, the abnormality warning device of the batch system may include: a collection unit 51, a first output unit 52, a comparison unit 53, a second output unit 54, and a generation unit 55, wherein:
[0184] The collection unit 51 is used to collect the operation data of the batch system at each operation time and pre-process the operation data to obtain the operation sequence data;
[0185] The first output unit 52 is configured to input the runtime series data into the anomaly prediction model and output the runtime data of the batch system at each target time in the future time period to obtain predicted time series data. The anomaly prediction model is a pre-built model for predicting the runtime series data of the batch system in the future time period.
[0186] A comparison unit 53 is configured to compare the operating data at each target moment in the predicted time series data with a preset operating data threshold to obtain a comparison result, and determine the target moment with an abnormality and its corresponding operating data based on the comparison result, wherein the operating data threshold is dynamically updated;
[0187] The second output unit 54 is used to construct abnormal time series data based on the target time when the abnormality occurs and the corresponding operating data, and input the abnormal time series data into the abnormality recognition model to output the abnormality type of the batch system, wherein the abnormality recognition model is a pre-built model for classifying the operating abnormalities of the batch system;
[0188] The generating unit 55 is configured to match an abnormality emergency response strategy based on the abnormality type, and generate an abnormality warning report based on the abnormality type and the abnormality emergency response strategy.
[0189] The above-mentioned abnormality warning device for the batch system collects the operating data of the batch system at each operating moment through the collection unit 51, and pre-processes the operating data to obtain operating time series data; inputs the operating time series data into the abnormality prediction model through the first output unit 52, outputs the operating data of the batch system at each target moment in the future time period, and obtains predicted time series data, wherein the abnormality prediction model is a pre-built model for predicting the operating time series data of the batch system in the future time period; compares the operating data of each target moment in the predicted time series data with a preset operating data threshold through the comparison unit 53 to obtain a comparison result, and determines the target moment with an abnormality and its corresponding operating data based on the comparison result, wherein the operating data threshold is dynamically updated; constructs abnormal time series data based on the target moment with an abnormality and its corresponding operating data through the second output unit 54, and inputs the abnormal time series data into the abnormality recognition model to output the abnormality type of the batch system, wherein the abnormality recognition model is a pre-built model for classifying the operating abnormalities of the batch system; matches the abnormal emergency strategy based on the abnormality type through the generation unit 55, and generates an abnormality warning report based on the abnormality type and the abnormal emergency strategy.
[0190] In this embodiment, by analyzing real-time operating data through an anomaly detection model, it is possible to effectively process non-stationary time series data, capture trend changes and seasonal fluctuations in operating data, and make accurate future predictions. By comparing the predicted time series data with dynamically updated thresholds, the system can identify operating data that exceeds the normal range in real time, thereby locating the target moment where the anomaly occurs. At the same time, for the target moment where the anomaly occurs, the anomaly recognition model is used to carefully classify the anomaly, identify the cause of the anomaly, and generate an anomaly emergency strategy, thereby achieving the purpose of improving the accuracy of anomaly detection, improving the accuracy of predicting the future operating status of the batch system, and achieving the technical effect of timely warning of potential anomalies. This solves the technical problem of low detection accuracy in the batch system anomaly detection method based on statistical principles and threshold comparison in related technologies.
[0191] Furthermore, the anomaly prediction model includes three parts: a differential model, an autoregressive model, and a moving average model. The first output unit includes: a first processing module, which is used to input the running series data into the anomaly prediction model, and perform smooth processing on the running series data based on the differential model of the anomaly prediction model to obtain smooth time series data; a first prediction module, which is used to input the smooth time series data into the autoregressive model of the anomaly prediction model, and predict the running data of the batch system at each target moment in the future time period based on the autoregressive relationship mechanism of the autoregressive model, and obtain the initial time series data of each target moment in the future time period; a first calculation module, which is used to input the initial time series data into the moving average model of the anomaly prediction model, calculate the random error term at each moment in the future time period through the moving average model, and correct the running data at each target moment in the future time period based on the random error term, and obtain the corrected time series data of each target moment in the future time period; a second processing module, which is used to perform differential processing on the corrected time series data based on the differential model of the anomaly prediction model to obtain the running data of the batch system at each target moment in the future time period.
[0192] Furthermore, the comparison unit includes: a first acquisition module, used to obtain the predicted value and actual value of the operation data of the batch system at each operation moment; a second calculation module, used to calculate the residual value of the batch system at each operation moment based on the predicted value and the actual value; a first identification module, used to draw a control chart based on the residual value at each operation moment, and identify the residual fluctuation anomaly of the batch system based on the control chart and the residual fluctuation anomaly judgment rule; a third calculation module, used to calculate a new operation data threshold based on the residual value at each moment when there is a residual fluctuation anomaly in the batch system.
[0193] Furthermore, the residual fluctuation abnormality judgment rule includes at least one of the following: the residual values of M consecutive data points in the control chart exceed the residual threshold, N consecutive data points in the control chart are in an upward or downward state, K consecutive data points in the control chart are in an alternating upward and downward fluctuation state, G consecutive points in the control chart are located on the same side of the center line of the control chart, and Y points out of X consecutive points in the control chart exceed the residual critical value, wherein M, N, K, G, X, and Y are all positive integers, X is greater than Y, and the residual critical value is less than the residual threshold.
[0194] Furthermore, the comparison unit also includes: a first determination module, which is used to determine that there is an abnormality at the target moment when the comparison result indicates that the operating data at the target moment exceeds a preset operating data threshold; a first adding module, which is used to add the target moment with the abnormality and the operating data corresponding to the target moment to the abnormal data set.
[0195] Furthermore, the abnormality warning device for the batch system also includes: a first acquisition module, which is used to acquire historical running time series data of the batch system, and pre-process the historical running time series data to obtain sample data; a first partitioning module, which is used to partition the sample data to obtain a training set and a test set; a first construction module, which is used to construct an initial differential model, an initial autoregressive model and an initial moving average model, and merge the initial differential model, the initial autoregressive model and the initial moving average model to obtain an initial abnormality prediction model; a first training module, which is used to iteratively train the initial abnormality prediction model based on the training set to obtain a trained abnormality prediction model, wherein, during the iterative training process, the model order is determined based on the model order function; a first testing module, which is used to test the trained abnormality prediction model based on the test set to obtain test index data of the trained abnormality prediction model, and when the test index data all meet the preset requirements, a final abnormality prediction model is obtained, wherein the test index data includes at least one of the following: mean absolute error, root mean square error and hit rate.
[0196] Furthermore, the acquisition unit includes: a third processing module, used to detect missing values in the operating data, and calculate the missing values based on linear interpolation to perform missing value processing on the operating data; a fourth processing module, used to smooth the noise in the operating data through sliding average filtering; a fifth processing module, used to sequentially process the operating data after missing value processing and smoothing according to timestamps to obtain operating series data.
[0197] It should be noted that the acquisition unit 51, the first output unit 52, the comparison unit 53, the second output unit 54, and the generation unit 55 correspond to steps S201 to S205 in the first embodiment. The examples and application scenarios implemented by the above units and the corresponding steps are the same, but are not limited to the contents disclosed in the first embodiment. It should be noted that the above modules or units can be hardware components or software components stored in a memory (e.g., the memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The above modules or units can also be part of the device and can be run in the computer terminal 10 provided in the first embodiment.
[0198] The present invention is described below in conjunction with another optional embodiment.
[0199] Example 3
[0200] An embodiment of the present invention may further provide an electronic device, Figure 6 is a hardware structure block diagram of an electronic device (or mobile device) for executing an abnormality warning method for a batch system according to an embodiment of the present invention, such as Figure 6As shown, the electronic device may include: one or more ( Figure 6 Only one is shown) processor 602, memory 604, storage controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.
[0201] Among them, the memory can be used to store software programs and modules, such as program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implementing the above-mentioned method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.
[0202] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: collect the operating data of the batch system at each operating moment, and pre-process the operating data to obtain operating time series data; input the operating time series data into the abnormality prediction model, output the operating data of the batch system at each target moment in the future time period, and obtain predicted time series data, wherein the abnormality prediction model is a pre-built model for predicting the operating time series data of the batch system in the future time period; compare the operating data of each target moment in the predicted time series data with the preset operating data threshold to obtain a comparison result, and determine the target moment with abnormality and its corresponding operating data based on the comparison result, wherein the operating data threshold is dynamically updated; construct abnormal time series data based on the target moment with abnormality and its corresponding operating data, and input the abnormal time series data into the abnormality recognition model, and output the abnormality type of the batch system, wherein the abnormality recognition model is a pre-built model for classifying the operating abnormalities of the batch system; match the abnormal emergency strategy based on the abnormality type, and generate an abnormal warning report based on the abnormality type and the abnormal emergency strategy.
[0203] By adopting the embodiment of the present invention, a batch system anomaly detection and early warning scheme is provided. By analyzing the real-time operation data through the anomaly detection model, it is possible to effectively process non-stationary time series data, capture trend changes and seasonal fluctuations in the operation data, and make accurate future predictions. By comparing the predicted time series data with the dynamically updated threshold, the system can identify the operation data that exceeds the normal range in real time, thereby locating the target moment where the anomaly exists. At the same time, for the target moment where the anomaly exists, the anomaly recognition model is used to carefully classify the anomaly, identify the cause of the anomaly and generate an anomaly emergency strategy, thereby achieving the purpose of improving the accuracy of anomaly detection, improving the accuracy of the prediction of the future operation status of the batch system, and achieving the technical effect of timely early warning of potential anomalies. This solves the technical problem of low detection accuracy of the batch system anomaly detection method based on statistical principles and threshold comparison in the related technology.
[0204] It can be understood by those skilled in the art that Figure 6 The structure shown is for illustration only, and the electronic device may also be a smart phone, a tablet computer, a PDA, a mobile Internet device (MID), a PAD or other terminal device. Figure 6 It does not limit the structure of the above electronic device. For example, the electronic device may also include Figure 6 More or fewer components (such as network interfaces, display devices, etc.) shown in, or with Figure 6 Different configurations shown.
[0205] A person skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0206] The present invention is described below in conjunction with another optional embodiment.
[0207] Example 4
[0208] The embodiment of the present invention further provides a computer-readable storage medium. Optionally, in the embodiment of the present invention, the computer-readable storage medium can be used to store the program code executed by the abnormality warning method for batch systems provided in the first embodiment.
[0209] Optionally, in an embodiment of the present invention, the above-mentioned storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.
[0210] An embodiment of the present invention also provides a computer program product, which, when executed on a data processing device, is suitable for executing the steps of the abnormality warning method for a batch system: collecting operating data of the batch system at each operating moment and preprocessing the operating data to obtain operating series data; inputting the operating series data into an abnormality prediction model, outputting the operating data of the batch system at each target moment in a future time period, and obtaining predicted time series data, wherein the abnormality prediction model is a pre-built model for predicting the operating series data of the batch system in a future time period; comparing the operating data at each target moment in the predicted time series data with a preset operating data threshold to obtain a comparison result, and determining the target moment with an abnormality and its corresponding operating data based on the comparison result, wherein the operating data threshold is dynamically updated; constructing abnormal time series data based on the target moment with an abnormality and its corresponding operating data, and inputting the abnormal time series data into an abnormality recognition model to output the abnormality type of the batch system, wherein the abnormality recognition model is a pre-built model for classifying the operating abnormalities of the batch system; matching the abnormal emergency strategy based on the abnormality type, and generating an abnormality warning report based on the abnormality type and the abnormal emergency strategy.
[0211] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.
[0212] In the above embodiments of the present invention, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0213] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0214] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0215] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0216] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, etc. Various media that can store program codes.
[0217] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A batch system abnormality warning method, characterized in that: include: Collecting operation data of the batch system at each operation moment and preprocessing the operation data to obtain operation sequence data; Inputting the running time series data into an anomaly prediction model, outputting the running data of the batch system at each target moment in a future time period, and obtaining predicted time series data, wherein the anomaly prediction model is a pre-built model for predicting the running time series data of the batch system in a future time period; Comparing the operating data at each target moment in the predicted time series data with a preset operating data threshold to obtain a comparison result, and determining the target moment with abnormalities and its corresponding operating data based on the comparison result, wherein the operating data threshold is dynamically updated; Constructing abnormal time series data based on the target time at which the abnormality exists and its corresponding operating data, and inputting the abnormal time series data into an abnormality recognition model to output the abnormality type of the batch system, wherein the abnormality recognition model is a pre-constructed model for classifying the operational abnormalities of the batch system; An abnormality emergency response strategy is matched based on the abnormality type, and an abnormality warning report is generated based on the abnormality type and the abnormality emergency response strategy.
2. The method according to claim 1, characterized in that The abnormality prediction model includes three parts: a differential model, an autoregressive model, and a moving average model. The steps of inputting the running series data into the abnormality prediction model and outputting the running data of the batch system at each target time in the future time period include: Inputting the running time series data into the abnormality prediction model, and performing smooth processing on the running time series data based on the differential model of the abnormality prediction model to obtain smooth time series data; Inputting the stationary time series data into the autoregressive model of the abnormality prediction model, predicting the operating data of the batch system at each target moment in a future time period based on the autoregressive relationship mechanism of the autoregressive model, and obtaining the initial time series data at each target moment in the future time period; Inputting the initial time series data into the moving average model of the abnormality prediction model, calculating the random error term at each moment in the future time period through the moving average model, and correcting the operating data at each target moment in the future time period based on the random error term, and the corrected time series data at each target moment in the future time period; The modified time series data is differentially processed based on the differential model of the abnormality prediction model to obtain the operation data of the batch system at each target moment in the future time period.
3. The method according to claim 1, characterized in that The step of updating the operating data threshold comprises: Obtaining predicted values and actual values of operation data of the batch system at each operation moment; Calculating residual values of the batch system at each operating moment based on the predicted value and the actual value; Drawing a control chart based on the residual values at each operating moment, and identifying the residual fluctuation anomaly of the batch system based on the control chart and the residual fluctuation anomaly determination rule; In the case where abnormal residual fluctuation exists in the batch system, a new operating data threshold is calculated based on the residual values at each moment.
4. The method according to claim 3, characterized in that The residual fluctuation abnormality judgment rule includes at least one of the following: the residual values of M consecutive data points in the control chart exceed the residual threshold value, N consecutive data points in the control chart are in an ascending or descending state, K consecutive data points in the control chart are in an alternating up and down fluctuation state, G consecutive points in the control chart are located on the same side of the center line of the control chart, and Y points out of X consecutive points in the control chart exceed the residual critical value, wherein M, N, K, G, X, and Y are all positive integers, X is greater than Y, and the residual critical value is less than the residual threshold value.
5. The method according to claim 1, wherein The step of determining the target time at which an abnormality exists and its corresponding operating data based on the comparison result includes: If the comparison result indicates that the operating data at the target moment exceeds the preset operating data threshold, determining that an abnormality exists at the target moment; The target time at which an anomaly occurs and the operating data corresponding to the target time are added to the anomaly data set.
6. The method according to claim 1, characterized in that The steps of constructing the abnormality prediction model include: Collecting historical runtime data of the batch system and preprocessing the historical runtime data to obtain sample data; Dividing the sample data into a training set and a test set; Constructing an initial difference model, an initial autoregressive model, and an initial moving average model, and merging the initial difference model, the initial autoregressive model, and the initial moving average model to obtain an initial anomaly prediction model; Iteratively training the initial anomaly prediction model based on the training set to obtain the trained anomaly prediction model, wherein, during the iterative training process, the model order is determined based on a model order function; The trained anomaly prediction model is tested based on the test set to obtain test index data of the trained anomaly prediction model. When the test index data all meet preset requirements, the final anomaly prediction model is obtained, wherein the test index data includes at least one of the following: mean absolute error, root mean square error and hit rate.
7. The method according to claim 1, characterized in that The step of preprocessing the operating data to obtain operating sequence data includes: detecting missing values in the operating data, calculating the missing values based on linear interpolation, and performing missing value processing on the operating data; Smoothing out noise in the operating data by using a sliding average filter; The running data after the missing value processing and the smoothing processing are processed sequentially according to the timestamps to obtain the running series data.
8. An abnormality warning device for a batch system, characterized in that: include: A collection unit, configured to collect operation data of the batch system at each operation moment, and pre-process the operation data to obtain operation sequence data; a first output unit, configured to input the running series data into an anomaly prediction model, and output the running data of the batch system at each target moment in a future time period to obtain predicted time series data, wherein the anomaly prediction model is a pre-built model for predicting the running series data of the batch system in the future time period; a comparing unit, configured to compare the operating data of each target moment in the predicted time series data with a preset operating data threshold to obtain a comparison result, and determine the target moment with an abnormality and its corresponding operating data based on the comparison result, wherein the operating data threshold is dynamically updated; a second output unit, configured to construct abnormal time series data based on the target time at which the abnormality occurs and the corresponding operating data, input the abnormal time series data into an abnormality recognition model, and output the abnormality type of the batch system, wherein the abnormality recognition model is a pre-constructed model for classifying operational abnormalities of the batch system; A generating unit is configured to match an abnormality emergency response strategy based on the abnormality type, and generate an abnormality warning report based on the abnormality type and the abnormality emergency response strategy.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the abnormality warning method for a batch system according to any one of claims 1 to 7.
10. An electronic device, characterized in that: The system comprises one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the abnormality warning method for a batch system according to any one of claims 1 to 7.
Citation Information
Cited By
Play switching method and device under server exception, terminal and medium
CN121531188A