Fault detection method and device and electronic equipment
By applying long-term memory network models to analyze multi-dimensional performance index data in big data processing systems, the problems of low fault detection accuracy and poor positioning efficiency in the existing technology are solved, and fast and accurate fault detection and positioning are achieved, system stability is improved and operation and maintenance costs are reduced.
Patent Information
- Application Number
- CN202510280556.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, big data processing systems have problems with low abnormal detection accuracy and poor fault positioning efficiency in fault detection, resulting in extended fault detection and recovery cycles, reduced system stability and increased operation and maintenance costs.
By obtaining the multi-dimensional performance index data generated during data processing, using long-term and short-term memory network models for analysis, predicting performance index values, and determining the abnormal detection results based on the predicted value and preset threshold values, thereby determining the cause of the failure.
It realizes rapid detection and precise positioning of potential faults in the data processing system, improves system stability, and reduces fault processing time and operation and maintenance costs.
Smart Images

Figure CN120216237A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and more particularly, to a fault detection method, apparatus, and electronic device. Background Art
[0002] In the current digital age, data processing has become a crucial part of the core business processes in all industries. As the infrastructure for data processing, big data clusters are facing unprecedented challenges. Especially with the explosive growth of data volume and the increasing complexity of processing tasks, the stability and reliability of the system have become the primary goals in the design of all data processing systems.
[0003] However, there are many limitations in the fault management of big data processing systems in related technologies. For example, the real-time response defect of the monitoring mechanism, the insufficient accuracy of anomaly detection and the inability to recognize complex patterns, the low efficiency of the fault location process, and the static and non-intelligent characteristics of the recovery strategy. These problems together lead to an extended fault detection and recovery cycle, reduced system stability, and increased operation and maintenance costs, thus severely restricting the efficiency of data processing tasks and business continuity.
[0004] In response to the above problems, no effective solution has been proposed yet. Summary of the Invention
[0005] Embodiments of this application provide a fault detection method, apparatus, and electronic device to at least solve the technical problems that the fault detection methods in related technologies mostly rely on manual monitoring and manual intervention, resulting in relatively low anomaly detection accuracy and poor fault location efficiency.
[0006] According to one aspect of the embodiments of this application, a fault detection method is provided, including: obtaining data information generated during data processing, and determining multi-dimensional performance metric data corresponding to the data information, where the performance metric data is multi-dimensional time series data including timestamps; analyzing the performance metric data through a long short-term memory network model to obtain a performance metric prediction value, where the performance metric prediction value is the expected performance metric value of the data processing system in a future time period; determining an anomaly detection result based on the performance metric prediction value and a preset threshold, where the preset threshold includes multiple performance metric thresholds corresponding to the performance metric data, and the anomaly detection result is used to reflect the fault probability and fault type of the data processing system; and determining the fault cause of the data processing system based on the anomaly detection result.
[0007] Optionally, after determining the multi-dimensional performance metric data corresponding to the data information, the method further includes: normalizing the performance metric data; dividing the normalized performance metric data according to a preset time window to obtain multiple window data corresponding to the performance metric data, where the window data is used to reflect the operating state of the data processing system within any time period.
[0008] Optionally, determining the anomaly detection result based on the performance metric prediction value and the preset threshold includes: determining the true value of the performance metric of the data processing system at the target time, where the target time is any time within the future time period; determining the deviation value between the true value of the performance metric and the performance metric prediction value, where the deviation value includes the performance deviations of the data processing system in multiple dimensions; determining the anomaly detection result based on the deviation value and the preset threshold.
[0009] Optionally, determining the anomaly detection result based on the deviation value and the preset threshold includes: comparing the target deviation value and the target threshold to obtain a comparison result, where the target deviation value is the performance deviation value corresponding to the target dimension in the deviation value, the target threshold is the performance metric threshold corresponding to the target dimension in the preset threshold, and the target dimension is any one of the multiple dimensions; determining that the data processing system is in an abnormal state when the comparison result indicates that the target deviation value exceeds the target threshold.
[0010] Optionally, determining the fault cause of the data processing system based on the anomaly detection result includes: obtaining the abnormal features in the anomaly detection result, where the abnormal features at least include the abnormal performance metrics and abnormal time points of the data processing system; processing the abnormal features through a decision tree model to determine the fault type of the data processing system; querying the historical knowledge database and determining the fault cause corresponding to the fault type from the historical knowledge database, where the historical knowledge database contains historical fault information and historical fault repair information.
[0011] Optionally, processing the abnormal features through a decision tree model to determine the fault type of the data processing system includes: performing a conditional test on the abnormal features at the root node of the decision tree model to obtain a test result, where the test conditions in the root node include all potential fault types of the data processing system, and the test result is used to represent the matching degree between the abnormal features and the test conditions in the root node; continuing to perform a conditional test on the abnormal features along the decision path of the decision tree model according to the test result until the conditional test stops after reaching the leaf node of the decision tree model, where the decision path is a moving path including multiple decision nodes, and the leaf node identifies the fault type that best matches the abnormal features.
[0012] Optionally, the method further includes: determining the degree of fault of the fault type, where the degree of fault is used to represent the degree of influence of the fault type on the data processing system; generating a fault repair suggestion based on the fault cause, the degree of fault, and the historical knowledge database; and repairing the data processing system according to the fault repair suggestion.
[0013] According to another aspect of the embodiments of the present application, there is also provided a fault detection device, including: an acquisition module, configured to acquire data information generated during data processing and determine multi-dimensional performance index data corresponding to the data information, where the performance index data is multi-dimensional time series data including timestamps; a prediction module, configured to analyze the performance index data through a long short-term memory network model to obtain a performance index prediction value, where the performance index prediction value is the expected performance index value of the data processing system in a future time period; a first determination module, configured to determine an anomaly detection result based on the performance index prediction value and a preset threshold, where the preset threshold includes multiple performance index thresholds corresponding to the performance index data, and the anomaly detection result is used to reflect the fault probability and fault type of the data processing system; and a second determination module, configured to determine the fault cause of the data processing system based on the anomaly detection result.
[0014] According to yet another aspect of the embodiments of the present application, there is also provided an electronic device, including: a memory and a processor, where the memory is configured to store program instructions; and the processor is connected to the memory and configured to execute to implement the above-mentioned fault detection method.
[0015] According to still another aspect of the embodiments of the present application, there is also provided a non-volatile storage medium, where the non-volatile storage medium includes a stored computer program, and the device where the non-volatile storage medium is located executes the above-mentioned fault detection method by running the computer program.
[0016] According to still another aspect of the embodiments of the present application, there is also provided a computer program product, including computer instructions, and when the computer instructions are executed by a processor, the above-mentioned fault detection method is implemented.
[0017] In the embodiments of the present application, by obtaining the data information generated during the data processing process and determining the multi-dimensional performance index data corresponding to the data information, where the performance index data is multi-dimensional time series data including timestamps; analyzing the performance index data through a long short-term memory network model to obtain a performance index prediction value, where the performance index prediction value is the expected performance index value of the data processing system in a future time period; determining an anomaly detection result based on the performance index prediction value and a preset threshold, where the preset threshold includes multiple performance index thresholds corresponding to the performance index data, and the anomaly detection result is used to reflect the failure probability and failure type of the data processing system; determining the cause of the failure of the data processing system based on the anomaly detection result, the purpose of quickly detecting and accurately locating potential failures in the data processing system is achieved, thereby realizing the technical effects of improving system stability, reducing the failure handling time, and reducing the operation and maintenance costs, and further solving the technical problems that the failure detection methods in the related technologies mostly rely on manual monitoring and manual intervention, resulting in low anomaly detection accuracy and poor failure location efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:
[0019] Figure 1 is a hardware structure diagram of a computer terminal for implementing a failure detection method according to an embodiment of the present application;
[0020] Figure 2 is a flowchart of a failure detection method according to an embodiment of the present application;
[0021] Figure 3 is a schematic architecture diagram of a failure detection system according to an embodiment of the present application;
[0022] Figure 4 is a structure diagram of a failure detection device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0024] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0025] First, some nouns or terms that appear in the process of explaining the embodiments of this application are applicable to the following explanations:
[0026] LSTM (Long Short-Term Memory): A special recurrent neural network structure specifically used to process and predict time series data, capable of learning long-term dependencies in the data. In fault detection, the LSTM model can analyze historical performance metric data, predict future states, and detect abnormal patterns in a timely manner.
[0027] Decision tree: A tree-structured classification model that classifies data into different categories (leaf nodes) through a series of conditional judgments (nodes). In this application, the decision tree is used to determine the fault type based on the abnormal features in the performance metric data, and then execute the corresponding recovery strategy.
[0028] Agent technology: An intelligent entity running in a system that can automatically execute tasks, such as responding to trigger events, performing recovery operations, etc. In this application, the Agent is used to execute the recovery strategy output by the decision tree to achieve automatic recovery of faults.
[0029] Sliding window technique: Refers to moving a fixed-size window over time series data to analyze the data trend over a period of time. In this application, the sliding window technique is used to collect and process performance metric data in real time to ensure the coherence and timeliness of the analysis.
[0030] In order to solve the problem of poor fault detection efficiency in the related art, the embodiments of this application provide a fault detection method, which can run on Figure 1 the computer terminal shown below, and the computer terminal will be described below.
[0031] The fault detection method embodiments provided by the embodiments of this application can be executed on a mobile terminal, a computer terminal or a similar computing device. Figure 1The figure shows a hardware block diagram of a computer terminal for implementing a fault detection method. As Figure 1 shown, the computer terminal 10 may include one or more processors (the processors may include, but are not limited to, processing devices such as a microprocessor MCU or a programmable logic device FPGA, shown as 102a, 102b, ……, 102n in the figure), a memory 104 for storing data, and a transmission module 106 for communication functions connected by wired and / or wireless networks. In addition, it may further include: a display, a keyboard, a cursor control device, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, and a BUS bus. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components than Figure 1 shown, or have a different configuration from Figure 1 shown.
[0032] It should be noted that the above one or more processors and / or other data processing circuits are generally referred to as "data processing circuits" in this article. The data processing circuit may be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the computer terminal 10. As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistor terminal path connected to an interface).
[0033] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the fault detection method in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned fault detection method. The memory 104 may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely set relative to the processor, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0034] The transmission module 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission module 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission module 106 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0035] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables a user to interact with the user interface of the computer terminal 10.
[0036] It should be noted here that, in some alternative embodiments, the above-mentioned Figure 1 illustrated computer terminal may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware elements and software elements. It should be pointed out that Figure 1 is only an example of a specific specific instance and is intended to illustrate the types of components that may exist in the above-mentioned computer terminal.
[0037] Under the above operating environment, an embodiment of a fault detection method is provided in an embodiment of the present application. It should be noted that the steps illustrated in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is illustrated in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0038] Figure 2 is a flowchart of a fault detection method according to an embodiment of the present application. As Figure 2 shown, the method includes the following steps:
[0039] Step S202, obtain data information generated during data processing, and determine multi-dimensional performance index data corresponding to the data information, where the performance index data is multi-dimensional time series data including timestamps.
[0040] In the above step S202, first, data information generated during the data processing process can be collected in real time through the data monitoring module in the data processing system, such as server monitoring information, offline task operation logs, real-time data processing logs, etc., to evaluate the system health status. Secondly, the corresponding key KPI data (i.e., the above multi-dimensional performance index data) is calculated, such as CPU utilization rate, memory usage, disk I / O activity, disk space growth, etc. Finally, it is output as a multi-dimensional time series data set: (t, v1, v2, v3, v4), providing a data basis for subsequent fault analysis. Among them, t is the timestamp, v1 is the CPU utilization rate, v2 is the memory usage, v3 is the disk I / O, and v4 is the disk space growth.
[0041] Step S204, analyze the performance index data through the long short-term memory network model to obtain the performance index prediction value, where the performance index prediction value is the expected performance index value of the data processing system in the future time period.
[0042] In the above step S204, the long short-term memory (LSTM) network model is applied to the collected multi-dimensional performance index data. Among them, the LSTM model can effectively capture the long-term dependencies in the time series and predict the performance index values in the future time period through learning from historical performance index data.
[0043] Step S206, determine the anomaly detection result based on the performance index prediction value and the preset threshold. Among them, the preset threshold includes multiple performance index thresholds corresponding to the performance index data, and the anomaly detection result is used to reflect the failure probability and failure type of the data processing system.
[0044] In the above step S206, the preset threshold is set based on the statistical analysis and business requirements during the normal operation of the system. Each performance index has its specific threshold range. When the deviation value between the predicted value and the actual value exceeds these thresholds, the system will consider that a failure may occur in the data processing process. This anomaly detection result not only reflects the possibility of failure but also can initially indicate the failure type of the data processing system.
[0045] Step S208, determine the cause of the failure of the data processing system based on the anomaly detection result.
[0046] In the above step S208, through information summary extraction technology and decision tree model, the anomaly detection result can be transformed into fault characteristics and matched with the historical (anomaly) knowledge database to determine the most likely cause of the failure. Among them, the historical knowledge database contains past failure cases, fault characteristics, and solution strategies, which is an important basis for fault location and recovery.
[0047] Through the above steps S202 to S208, the purpose of quickly detecting and accurately locating potential faults in the data processing system is achieved, thereby realizing the technical effects of improving system stability, reducing fault handling time, and lowering operation and maintenance costs. Furthermore, it solves the technical problems that in the related art, fault detection methods mostly rely on manual monitoring and manual intervention, resulting in relatively low anomaly detection accuracy and poor fault location efficiency. The following is a detailed description.
[0048] In the above step S202, after determining the multi-dimensional performance index data corresponding to the data information, it may further include: performing normalization processing on the performance index data; dividing the normalized performance index data according to a preset time window to obtain multiple window data corresponding to the performance index data, where the window data is used to reflect the operating state of the data processing system within any time period.
[0049] In the embodiment of the present application, it is also crucial to preprocess the obtained multi-dimensional performance index data. The specific process can be as follows:
[0050] First, through normalization processing, the original performance index data is converted into a unified numerical range, such as the interval [0, 1], to eliminate the influence brought by the dimensional difference between different indicators and provide standardized data input for subsequent model training and prediction.
[0051] Secondly, the sliding window technique is used to divide the normalized performance index data according to a preset time window (such as 30 minutes) to generate a series of window data. Each window data actually reflects the operating state of the data processing system within a specific time period, including the changes in multiple key indicators such as CPU utilization rate, memory usage rate, and disk I / O activities within that time period. The generation of window data not only provides sequential continuous data input for the LSTM model, facilitating the model to capture the dynamic change trend and potential abnormal patterns of performance indicators, but also provides a structured time series analysis basis for real-time detection and location of faults.
[0052] Through normalization and sliding window division, the system can more finely monitor and understand the performance fluctuations in the data processing process, laying a solid data foundation for timely discovery and solution of faults.
[0053] In the above step S204, the LSTM model can be trained in the following way:
[0054] Specifically, the model structure is designed as an LSTM network, and this network structure is particularly suitable for processing data with time series characteristics and can effectively capture long-term dependencies.
[0055] During the training process, the LSTM model will use multi-dimensional historical performance metric data as learning materials. By iterating and optimizing the model parameters multiple times, it can more accurately predict future performance metrics. Specifically, the training input is historical performance metric data divided by a sliding window. Each window data contains performance metrics at multiple past time points, such as historical CPU utilization, historical memory usage, etc. These data form a continuous sequence on the time axis. The output is the predicted multi-dimensional performance metric values at the next time point predicted by the model, that is, the expected performance metrics of the future state. Among them, the loss function can adopt the mean squared error (MSE). MSE measures the gap between the model's predicted values and the actual values. The goal of model training is to minimize this error and make the prediction results as close as possible to the actual data.
[0056] Through the above process, the LSTM model can learn from historical multi-dimensional performance metric data, continuously adjust and optimize its parameters to improve the prediction ability of future data change trends. This training process provides strong model support for subsequent real-time data analysis and anomaly detection, and is the key to realizing fast fault warning and location.
[0057] In step S206 above, determining the anomaly detection result based on the performance metric prediction value and the preset threshold includes: determining the true value of the performance metric of the data processing system at the target moment, where the target moment is any moment within the future time period; determining the deviation value between the true value of the performance metric and the performance metric prediction value, where the deviation value includes the performance deviations of the data processing system in multiple dimensions; and determining the anomaly detection result based on the deviation value and the preset threshold.
[0058] Optionally, determining the anomaly detection result based on the deviation value and the preset threshold includes: comparing the target deviation value and the target threshold to obtain a comparison result, where the target deviation value is the performance deviation value corresponding to the target dimension in the deviation value, the target threshold is the performance metric threshold corresponding to the target dimension in the preset threshold, and the target dimension is any one of the multiple dimensions; and determining that the data processing system is in an abnormal state when the comparison result indicates that the target deviation value exceeds the target threshold.
[0059] In the embodiments of the present application, not only the prediction of the future performance state of the data processing system is realized, but also by comparing the deviation between the predicted value and the true value and combining the preset threshold, the operating state of the system in multiple dimensions is accurately judged, providing timely and accurate fault warning information for subsequent fault location and recovery strategy formulation. The specific process can be as follows:
[0060] First, determine the true value of the performance metrics of the data processing system at the target moment, which can be any time point within a future time period. Among them, the true value of the performance metrics can be obtained by the real-time data monitoring module, including multi-dimensional real-time data such as CPU utilization, memory usage, and disk I / O activities, reflecting the actual operating state of the system at the target moment.
[0061] Secondly, calculate the deviation value between the true value of the performance metrics and the predicted value of the LSTM model. Among them, the deviation value reflects the difference between the predicted value and the actual value, and it exists specifically in each dimension, such as CPU utilization deviation, memory usage deviation, etc.
[0062] Finally, determine the anomaly detection result based on the deviation value and the preset threshold. The following strategy can be adopted: for each dimension (i.e., the above-mentioned target dimension), compare the target deviation value under the target dimension with the target threshold; if the target deviation value exceeds the target threshold, it is determined that the data processing system is in an abnormal state in the target dimension. Among them, the preset threshold is set based on the statistical analysis and expert knowledge during the normal operation of the system, and they define the normal operation range for each dimension. Any change beyond this range will be regarded as a possible fault signal.
[0063] For example, if the preset CPU utilization threshold is 10%, and the deviation between the CPU utilization predicted by the LSTM model and the actual value reaches 15%, this indicates that the sudden change in CPU utilization is larger than the preset normal fluctuation range. The system will, based on this comparison result, determine that the data processing system may be abnormal in terms of CPU utilization. This process independently compares and judges the performance metrics of each dimension, ensuring the comprehensiveness and accuracy of fault detection.
[0064] In the above step S208, determine the fault cause of the data processing system based on the anomaly detection result, including: obtaining the abnormal features in the anomaly detection result, where the abnormal features at least include the abnormal performance metrics and abnormal time points of the data processing system; processing the abnormal features through a decision tree model to determine the fault type of the data processing system; querying the historical knowledge database and determining the fault cause corresponding to the fault type from the historical knowledge database, where the historical knowledge database contains historical fault information and historical fault repair information.
[0065] Optionally, the abnormal features are processed through a decision tree model to determine the fault type of the data processing system, including: performing a conditional test on the abnormal features at the root node of the decision tree model to obtain a test result, where the test condition in the root node includes all potential fault types of the data processing system, and the test result is used to represent the matching degree between the abnormal features and the test condition in the root node; continuing to perform a conditional test on the abnormal features along the decision path of the decision tree model according to the test result until the conditional test stops after reaching the leaf node of the decision tree model, where the decision path is a moving path including multiple decision nodes, and the leaf node identifies the fault type that best matches the abnormal features.
[0066] In the embodiment of the present application, through the logical judgment of the decision tree and the intelligent query of the historical knowledge database, the fault type and the fault cause can be quickly and accurately identified, laying a solid foundation for the subsequent formulation and execution of the fault recovery strategy. The specific process can be as follows:
[0067] First, through the abstract extraction technology, abnormal features are extracted from the anomaly detection results. These abnormal features include the abnormal performance metrics of the data processing system at a specific time point (such as abnormal increase in CPU utilization, abnormal memory usage, etc.) and the exact time point when the anomaly occurs.
[0068] Secondly, the abnormal features are processed through a decision tree model. A decision tree is a commonly used machine learning model that can classify data based on a series of rules. In the embodiment of the present application, the root node of the decision tree model includes all potential fault types of the data processing system. By performing a conditional test on the abnormal features, a test result is obtained, which reflects the matching degree between the abnormal features and the test condition in the decision tree. For example, the root node may include multiple test conditions such as "high CPU utilization", "abnormal memory usage", "disk I / O anomaly", etc. The system will perform matching according to the abnormal features and continue to perform conditional tests along the decision path of the decision tree model. Among them, the decision path includes multiple decision nodes, and each decision node represents a fault detection rule or condition. The system will perform tests on the abnormal features at each node until a leaf node is reached, which identifies the fault type that best matches the abnormal features. This process is similar to a step-by-step troubleshooting of the fault. Through a series of logical judgments, the system can quickly locate the most likely fault type from the complex abnormal information, such as "resource exhaustion", "data skew", "software defect", etc.
[0069] Finally, by querying the historical knowledge database, the fault causes corresponding to the determined fault type are obtained. Among them, the historical knowledge database is accumulated over a long period and contains the fault information and repair experience that occurred in the past. By querying the database, the system can compare the current fault with historical cases and find the closest fault cause. For example, if the fault type determined by the decision tree model is "resource exhaustion", the system will search the historical knowledge database for problem descriptions and solutions related to resource exhaustion, so as to provide specific fault cause analysis and possible recovery strategies for the current fault.
[0070] Optionally, the above method further includes: determining the fault degree of the fault type, where the fault degree is used to represent the impact degree of the fault type on the data processing system; generating a fault repair suggestion based on the fault cause, the fault degree and the historical knowledge database; and repairing the data processing system according to the fault repair suggestion.
[0071] In the embodiment of the present application, by intelligently evaluating the fault degree, automatically generating repair suggestions and automatically executing recovery strategies, combined with the transparent display and marking function of the user interface, an efficient, intelligent and transparent data processing system fault recovery mechanism is constructed, which significantly improves the operation and maintenance efficiency of the system and the user experience.
[0072] Specifically, based on the fault cause, the fault degree and the historical knowledge database, the data processing system can automatically generate a fault repair suggestion, and the formulation of this fault repair suggestion integrates rule-based strategies and intelligent decisions of the decision tree model. For example, when a service memory leak is detected, considering the fault degree and historical experience, the system may suggest restarting the relevant service and at the same time adjusting the resource allocation strategy to prevent similar problems from occurring again. Among them, the determination of the fault degree is based on in-depth analysis of the anomaly detection results, which quantifies the impact of the fault on the data processing system, including the degree of decline in multiple dimensions such as resource consumption, service response time, and data processing efficiency.
[0073] Furthermore, the execution of the recovery strategy can be achieved by integrating Agent technology, which automatically calls the corresponding API for operations such as task rerun, service restart, and resource adjustment. This automated execution ensures the efficiency of the recovery process, reduces the delay and risk of misoperation caused by manual operation, and improves the speed and success rate of fault recovery. At the same time, the application of Agent technology enhances the adaptability and flexibility of the system, can quickly respond to different types of faults, and effectively improves the system performance.
[0074] The data processing system also provides a record display of the fault history and automatic recovery operations, so that operation and maintenance personnel can view the time, type, degree of the fault occurrence, and the recovery measures taken by the system. At the same time, it supports marking the results of automatic processing, such as marking the fault type, the effectiveness of the recovery strategy, etc. This function not only improves the transparency of the fault handling process, but also provides rich historical data for subsequent model optimization and algorithm iteration. Through the user interface, the system can accumulate experience in fault handling, continuously optimize the fault detection and recovery strategies, form a closed-loop of continuous learning and improvement, and thus further enhance the system's fault response ability and long-term operation stability.
[0075] In the embodiments of the present application, a series of quantitative indicators are also provided to ensure that each link of the fault detection and recovery process has clear performance goals, providing a direction for the continuous improvement of the system. Among them, the quantitative indicators are, for example:
[0076] Fault detection time: The time from the occurrence of an anomaly to its detection, with a target of <5 minutes;
[0077] Fault location time: The time from the detection of an anomaly to the location of the fault cause, with a target of <2 minutes;
[0078] Recovery operation time: The time from the execution of the recovery strategy to the resolution of the fault, with a target of <10 minutes;
[0079] System stability improvement: The frequency of fault occurrence is reduced, with a target reduction of >50%;
[0080] Reduction of manual intervention: The number of times of manual fault handling is reduced, with a target reduction of >70%.
[0081] By monitoring and optimizing the above indicators, continuous improvement of the fault detection and recovery efficiency can be achieved. At the same time, the significant enhancement of system stability and operation and maintenance efficiency demonstrates the improvement of the response speed and automatic processing ability of the offline data processing system in the face of faults in the present application, thereby enhancing the overall operation stability and reliability. The achievement of quantitative indicators marks that the system has reached a high level of automation and intelligence in fault management, which is of great significance for reducing the impact of faults on the data processing process and improving system availability and operation and maintenance efficiency.
[0082] Figure 3 It is a schematic diagram of the architecture of a fault detection system according to an embodiment of the present application. As Figure 3As shown, the system collects abnormal information such as CPU, memory, IO, processing delay, timeout exception, and logs from the CNC monitoring of the big data cluster, the operation logs of offline tasks, and the real-time data processing logs through the data monitoring module. Subsequently, this abnormal information is sent to the large model analysis module, which conducts in-depth analysis using time series analysis, neural networks, LSTM algorithms, and combined with historical knowledge prompting engineering and large model calls to obtain the abnormal detection results. Further, the abnormal detection results are used for fault location and fault recovery. Among them, fault location includes information summary extraction and fault classification to obtain the fault type; while fault recovery can be based on historical experience and strategies, decision tree models and redo mechanisms, as well as fault recovery routing. Finally, the fault history, recovery history, policy configuration, abnormal classification, and tagging of the historical (abnormal) knowledge database are displayed on the user interface. It should be noted that during this process, the historical (abnormal) knowledge database needs to be maintained in real time, recording abnormal marks, fluctuation ranges, equipment abnormalities, data abnormalities, emergency strategies, and recovery records, etc., to support the full process of fault detection and fault recovery.
[0083] In the embodiment of this application, taking the example of multi-dimensional resource abnormalities occurring in any task within a short period of time, the specific implementation steps and detection data can be as follows:
[0084] S1. Data collection.
[0085] Real-time collect multi-dimensional performance index data, for example:
[0086] (12:00:00, CPU: 50%, memory: 60%, disk I / O: 30MB / s, disk space: 40GB);
[0087] (12:01:00, CPU: 55%, memory: 65%, disk I / O: 35MB / s, disk space: 41GB);
[0088] (12:02:00, CPU: 60%, memory: 70%, disk I / O: 40MB / s, disk space: 42GB).
[0089] S2. Data preprocessing.
[0090] Normalize (standardize) and slide the window for the collected performance index data to obtain window data:
[0091] Window: [CPU: 50%, 55%, 60%]; [memory: 60%, 65%, 70%]; [disk I / O: 30, 35, 40]; [disk space: 40, 41, 42].
[0092] S3. LSTM model prediction.
[0093] Predict the input window data to obtain the predicted performance metric values:
[0094] CPU: 70%, Memory: 80%, Disk I / O: 50 MB / s, Disk Space: 44 GB.
[0095] S4. Anomaly detection.
[0096] Obtain the actual performance metric values in real time:
[0097] CPU: 85%, Memory: 90%, Disk I / O: 60 MB / s, Disk Space: 50 GB.
[0098] Calculate the deviation values between the predicted performance metric values and the actual performance metric values:
[0099] CPU deviation: 15% (>10%), Memory deviation: 10% (<15%), Disk I / O deviation: 10% (<20%), Disk Space deviation: 6 GB (>25%).
[0100] Conclusion: The CPU utilization rate and disk space growth exceed the preset thresholds, and it is determined as an anomaly.
[0101] S5. Generate fault repair suggestions.
[0102] Call the large model to generate a prompt: "It is detected that the CPU utilization rate has increased abnormally (85%), and the possible reason is the increase in task computing-intensive operations; the disk space has grown abnormally (50 GB), and the possible reason is that the log files have not been cleared in time. It is recommended to optimize the task computing logic and clear the log files."
[0103] In the embodiment of the present application, a comprehensive, intelligent and efficient fault management framework is constructed by integrating real-time data monitoring, deep learning anomaly detection, knowledge graph intelligent decision-making, fault rapid location, and automated recovery strategies. In particular, applying deep learning technologies such as LSTM for accurate identification of multi-dimensional resource anomalies, combining historical knowledge databases with decision tree models for intelligent classification and cause inference of faults, and using Agent technology to automatically execute fault recovery strategies significantly improve the real-time response ability, fault handling accuracy and automation level of the system, thereby effectively reducing the time for fault detection and recovery, reducing manual intervention, and enhancing the overall stability and reliability of the data processing system. In addition, by introducing quantitative metrics to measure system performance, the continuous optimization and iteration of the fault management process are ensured, reflecting its innovative value and application potential in the field of big data operation and maintenance.
[0104] According to an embodiment of the present application, a fault detection device is provided. It should be noted that the fault detection device in the embodiment of the present application can be used to execute the fault detection method provided in the embodiment of the present application. The following introduces the fault detection device provided in the embodiment of the present application.
[0105] Figure 4 It is a structural diagram of a fault detection device provided according to an embodiment of the present application. As Figure 4 shown, the device includes:
[0106] An acquisition module 40, configured to acquire data information generated during data processing, and determine multi-dimensional performance index data corresponding to the data information, where the performance index data is multi-dimensional time series data including timestamps;
[0107] A prediction module 42, configured to analyze the performance index data through a long short-term memory network model to obtain a performance index prediction value, where the performance index prediction value is the expected performance index value of the data processing system in a future time period;
[0108] A first determination module 44, configured to determine an anomaly detection result based on the performance index prediction value and a preset threshold, where the preset threshold includes multiple performance index thresholds corresponding to the performance index data, and the anomaly detection result is used to reflect the fault probability and fault type of the data processing system;
[0109] A second determination module 46, configured to determine the fault cause of the data processing system based on the anomaly detection result.
[0110] Through the acquisition module, prediction module, first determination module, and second determination module in the above-mentioned fault detection device, the purpose of quickly detecting and accurately locating potential faults in the data processing system is achieved, thereby realizing the technical effects of improving system stability, reducing fault handling time, and reducing operation and maintenance costs, and further solving the technical problems that the fault detection methods in the related art mostly rely on manual monitoring and manual intervention, and there are low anomaly detection accuracy and poor fault location efficiency.
[0111] In the fault detection device provided in the embodiment of the present application, the acquisition module is further configured to perform normalization processing on the performance index data; divide the normalized performance index data according to a preset time window to obtain multiple window data corresponding to the performance index data, where the window data is used to reflect the operating state of the data processing system in any time period.
[0112] In the fault detection device provided by the embodiment of the present application, the first determination module is further configured to determine the true value of the performance index of the data processing system at the target time, where the target time is any time within a future time period; determine the deviation value between the true value of the performance index and the predicted value of the performance index, where the deviation value includes the performance deviations of the data processing system in multiple dimensions; and determine the anomaly detection result according to the deviation value and the preset threshold.
[0113] In the fault detection device provided by the embodiment of the present application, the first determination module is further configured to compare the target deviation value with the target threshold to obtain a comparison result, where the target deviation value is the performance deviation value corresponding to the target dimension in the deviation value, the target threshold is the performance index threshold corresponding to the target dimension in the preset threshold, and the target dimension is any one of the multiple dimensions; and determine that the data processing system is in an abnormal state when the comparison result indicates that the target deviation value exceeds the target threshold.
[0114] In the fault detection device provided by the embodiment of the present application, the second determination module is further configured to obtain the abnormal features in the anomaly detection result, where the abnormal features at least include the abnormal performance indexes and abnormal time points of the data processing system; process the abnormal features through a decision tree model to determine the fault type of the data processing system; query the historical knowledge database, and determine the fault cause corresponding to the fault type from the historical knowledge database, where the historical knowledge database contains historical fault information and historical fault repair information.
[0115] In the fault detection device provided by the embodiment of the present application, the second determination module is further configured to perform a conditional test on the abnormal features at the root node of the decision tree model to obtain a test result, where the test condition in the root node includes all potential fault types of the data processing system, and the test result is used to represent the matching degree between the abnormal features and the test condition in the root node; and continue to perform a conditional test on the abnormal features along the decision path of the decision tree model according to the test result until the conditional test stops after reaching the leaf node of the decision tree model, where the decision path is a moving path including multiple decision nodes, and the leaf node identifies the fault type that best matches the abnormal features.
[0116] In the fault detection device provided by the embodiment of the present application, a repair module 48 is further included, and the repair module is configured to determine the fault degree of the fault type, where the fault degree is used to represent the influence degree of the fault type on the data processing system; generate a fault repair suggestion according to the fault cause, the fault degree and the historical knowledge database; and repair the data processing system according to the fault repair suggestion.
[0117] The embodiment of the present application further provides an electronic device, including: a memory and a processor, where the memory is used to store program instructions; the processor is connected to the memory and is used to execute to implement the above-mentioned fault detection method.
[0118] It should be noted that the above electronic device is used to execute Figure 2 the fault detection method shown, so the relevant explanations in the above fault detection method also apply to this electronic device, and will not be elaborated here.
[0119] The embodiment of the present application also provides a non-volatile storage medium, which includes a stored computer program. Among them, the device where the non-volatile storage medium is located executes the above fault detection method by running the computer program.
[0120] It should be noted that the above non-volatile storage medium is used to execute Figure 2 the fault detection method shown, so the relevant explanations in the above fault detection method also apply to this non-volatile storage medium, and will not be elaborated here.
[0121] The embodiment of the present application also provides a computer program product, including computer instructions, which implement the above fault detection method when executed by a processor.
[0122] It should be noted that the above computer program product is used to execute Figure 2 the fault detection method shown, so the relevant explanations in the above fault detection method also apply to this computer program product, and will not be elaborated here.
[0123] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.
[0124] In the above embodiments of the present application, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0125] In the several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0126] The unit described as a separating component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0127] In addition, each functional unit in various embodiments of the present application may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0128] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs and other various media that can store program codes.
[0129] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A fault detection method, characterized in that: include: Acquire data information generated during data processing, and determine multi-dimensional performance indicator data corresponding to the data information, wherein the performance indicator data is multi-dimensional time series data including a timestamp; Analyzing the performance indicator data through a long short-term memory network model to obtain a performance indicator prediction value, wherein the performance indicator prediction value is an expected performance indicator value of the data processing system in a future time period; Determining an abnormality detection result according to the performance indicator prediction value and a preset threshold, wherein the preset threshold includes a plurality of performance indicator thresholds corresponding to the performance indicator data, and the abnormality detection result is used to reflect the failure probability and failure type of the data processing system; The cause of the failure of the data processing system is determined according to the abnormality detection result.
2. The method according to claim 1, characterized in that After determining the multi-dimensional performance indicator data corresponding to the data information, the method further includes: Normalizing the performance indicator data; The normalized performance indicator data is divided according to preset time windows to obtain a plurality of window data corresponding to the performance indicator data, wherein the window data is used to reflect the operating status of the data processing system in any time period.
3. The method according to claim 1, characterized in that Determining anomaly detection results based on the performance indicator prediction value and the preset threshold value includes: Determining a true value of a performance indicator of the data processing system at a target time, wherein the target time is any time within the future time period; Determine a deviation value between the actual value of the performance indicator and the predicted value of the performance indicator, wherein the deviation value includes performance deviations of the data processing system in multiple dimensions; The abnormality detection result is determined according to the deviation value and the preset threshold.
4. The method according to claim 3, characterized in that Determining the abnormality detection result according to the deviation value and the preset threshold value includes: Compare the target deviation value and the target threshold value to obtain a comparison result, wherein the target deviation value is a performance deviation value corresponding to the target dimension in the deviation value, the target threshold value is a performance indicator threshold corresponding to the target dimension in the preset threshold value, and the target dimension is any one of the multiple dimensions; When the comparison result indicates that the target deviation value exceeds the target threshold value, it is determined that the data processing system is in an abnormal state.
5. The method according to claim 1, characterized in that Determining the cause of the failure of the data processing system according to the abnormality detection result includes: Acquire abnormal features in the abnormal detection result, wherein the abnormal features at least include abnormal performance indicators and abnormal time points of the data processing system; Processing the abnormal features through a decision tree model to determine the fault type of the data processing system; A historical knowledge database is queried, and a fault cause corresponding to the fault type is determined from the historical knowledge database, wherein the historical knowledge database includes historical fault information and historical fault repair information.
6. The method according to claim 5, characterized in that Processing the abnormal feature by a decision tree model to determine the fault type of the data processing system includes: Performing a conditional test on the abnormal feature at the root node of the decision tree model to obtain a test result, wherein the test condition in the root node includes all potential fault types of the data processing system, and the test result is used to indicate the degree of matching between the abnormal feature and the test condition in the root node; The abnormal feature is conditionally tested along the decision path of the decision tree model according to the test result until the leaf node of the decision tree model is reached and the conditional test is stopped, wherein the decision path is a moving path including a plurality of decision nodes, and the leaf node identifies the fault type that best matches the abnormal feature.
7. The method according to claim 5, characterized in that The method further comprises: Determining a fault degree of the fault type, wherein the fault degree is used to indicate the degree of influence of the fault type on the data processing system; generating a fault repair suggestion based on the fault cause, the fault degree and the historical knowledge database; The data processing system is repaired according to the fault repair suggestion.
8. A fault detection device, characterized in that: include: An acquisition module, used to acquire data information generated during data processing, and determine multi-dimensional performance indicator data corresponding to the data information, wherein the performance indicator data is multi-dimensional time series data including a timestamp; A prediction module, used to analyze the performance indicator data through a long short-term memory network model to obtain a performance indicator prediction value, wherein the performance indicator prediction value is an expected performance indicator value of the data processing system in a future time period; A first determination module, configured to determine an abnormality detection result according to the performance indicator prediction value and a preset threshold, wherein the preset threshold includes a plurality of performance indicator thresholds corresponding to the performance indicator data, and the abnormality detection result is used to reflect the failure probability and failure type of the data processing system; The second determination module is used to determine the cause of the failure of the data processing system according to the abnormality detection result.
9. An electronic device, characterized in that: include: A memory and a processor, wherein the memory is used to store program instructions; The processor is connected to the memory and is used to execute the fault detection method described in any one of claims 1 to 7.
10. A non-volatile storage medium, characterized in that: The non-volatile storage medium includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the fault detection method according to any one of claims 1 to 7 by running the computer program.
11. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by a processor, the fault detection method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Bus anomaly prediction and processing method, equipment, medium and product
CN120670254A
Communication base station fault analysis processing method and system
CN120916185A
Regional pipe network leakage analysis and early warning method based on water big data
CN121257854A
A communication anomaly detection method based on multi-dimensional feature fusion and progressive judgment
CN122621459A