Fault root cause analysis method based on continuous system performance historical data

By continuously collecting and storing performance indicator data in computer systems, combining timestamps and visual analysis, and building a chain of evidence, the problem of locating occasional system failures is solved, and the accuracy of troubleshooting and system stability are improved.

CN120704929APending Publication Date: 2025-09-26ZHEJIANG ZHIWEI CLOUD TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510847220.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing technologies make it difficult to accurately locate non-crash, occasional system performance failures that have occurred in computer systems. Especially in high-demand scenarios such as online games and real-time data processing, system freezes and performance degradation problems are difficult to reproduce and locate.

Method used

By continuously acquiring performance indicator data on the target terminal device and marking it with timestamps, storing it in the data storage system, retrieving and visualizing the data during the fault period, analyzing indicator change trends and abnormal behaviors, and building an evidence chain to locate the root cause of the fault.

Benefits of technology

It achieves precise positioning of occasional system performance failures, improves the accuracy and efficiency of troubleshooting, and improves system operation stability and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704929A_ABST
    Figure CN120704929A_ABST
Patent Text Reader

Abstract

The invention relates to a fault root cause analysis method based on continuous system performance historical data, and the method comprises the steps: continuously collecting multi-dimensional system performance data at high frequency, precisely marking a timestamp, and combining historical data retrieval, multi-dimensional correlation analysis, evidence chain construction and system state snapshot backtracking. According to the invention, the method achieves the precise positioning of the occurred, non-crash and accidental system performance faults, clarifies the process behaviors or resource bottlenecks causing the faults, provides an objective fault diagnosis basis, improves the accuracy and efficiency of troubleshooting, and facilitates the improvement of the system operation stability and user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of fault analysis, and in particular to a method for fault root cause analysis based on continuous system performance historical data. Background Art

[0002] With the rapid development of information technology, computer systems are increasingly used in various fields. Their stability and performance are crucial to user experience and business continuity, especially in application scenarios with high requirements for real-time and smoothness, such as online games, online competitions (common in Internet cafes), real-time data processing, etc. Users are extremely sensitive to small fluctuations in system performance. Occasional system freezes, slow application responses, game frame drops, etc., even if they do not lead to a complete crash of the system or application, will seriously affect the user experience and even cause user complaints and economic losses.

[0003] In traditional computer system fault diagnosis and performance analysis, operations personnel use performance monitoring tools such as Windows Task Manager, Performance Monitor, or third-party tools to observe system resource usage in real time when a problem occurs or after receiving a report, attempting to capture anomalies. However, many performance issues are sudden and transient. By the time operations personnel intervene, the problem symptoms may have disappeared, making it difficult to reproduce and locate the root cause. Furthermore, operating systems and applications typically record error logs or crash dumps, which are very useful for analyzing serious errors that cause system or application crashes. However, for issues that don't reach the point of a crash and only manifest as performance degradation or brief freezes—for example, a background process that instantly consumes a large amount of network bandwidth, causing a game to freeze—such logs often fail to provide sufficient information, or the log information is too scattered to be easily correlated and analyzed. Summary of the Invention

[0004] In order to solve the above problems, the present application provides a fault root cause analysis method based on continuous system performance historical data that can accurately locate existing, non-crash, and occasional system performance failures.

[0005] To achieve the above objectives, the present application designs a fault root cause analysis method based on continuous system performance historical data, which includes the following steps: S1. On the target terminal device, through a preset data collection agent, continuously obtains the performance indicator data of the target terminal device at a preset time interval, and marks the acquisition timestamp of each performance indicator data collected; S2. The performance indicator data with a timestamp is transferred to a data storage system, which uses a device identifier, a collection timestamp, and an indicator type to perform structured storage and indexing on the collected performance indicator data; S3. When a fault report is received, the target terminal device identifier and the time window in which the fault occurs are determined, and the performance indicator data within the corresponding time period is retrieved from the data storage system according to the device identifier and time window; S4. Visualize the retrieved performance indicator data on the same timeline to generate a snapshot of the overall system status of the target terminal device at the time of the failure; S5. Analyze the changing trends of different indicators within the fault time window, identify the temporal synchronization correlation between indicators, and detect abnormal behavior by comparing with data from normal time periods; S6. Based on the temporal synchronization relationship between indicators and the results of abnormal behavior detection, build an evidence chain from abnormal behavior to fault phenomenon; S7. Generate an analysis report including the evidence chain and a snapshot of the overall status of the system.

[0006] Preferably, the acquisition timestamp is accurate to millisecond level, and the preset time interval is 3 seconds.

[0007] Preferably, the data acquisition agent program obtains performance indicator data in the following manner: Query and obtain performance data through the PDH computer performance database interface, including system-level and process-level CPU utilization, memory utilization, network card transceiver rate, network card transceiver operation number, disk I / O read / write rate, disk I / O read / write operation number, memory page fault number, CPU hard interrupt frequency, and GPU functional unit utilization rate; Query and obtain hardware information through the WMI query interface, including temperature sensor information of the CPU, motherboard, and GPU, number of memory sticks, their brand, frequency, and channel information, CPU voltage information, and real-time power information of the power supply; By calling the system API, you can query and obtain process-level information, including the list of active processes, process handle information, working set memory information, the number of user objects, the number of handles, and the number of threads.

[0008] Preferably, the specific steps of obtaining performance data include: Call the PdhOpenQuery function to open a system performance data query; construct a query statement for a specific performance indicator and add the query statement to the query through the PdhAddCounter API; After waiting for the preset delay time, call the PdhCollectQueryData function to collect the current performance data; Call the PdhGetFormattedCounterValue function to obtain the formatted performance data value.

[0009] Preferably, the specific steps of obtaining hardware information include: Call the CoInitializeEx function to initialize the component object model library; Call the CoCreateInstance function to create an instance of the WMI service interface; Constructing a WMI query language statement for specific hardware information and calling ExecQuery to execute the query statement; Parse and extract the required hardware information from the result set returned after executing the query.

[0010] Preferably, the overall system status snapshot generated in step S4 includes the following items within the time window of the fault occurrence: Specific values ​​of various performance indicators; A list of process resource consumption and their respective resource consumption values; Process network connection list; Logged system errors and application events. Application events.

[0011] Preferably, in step S5, identifying the temporal synchronization association relationship between different indicators includes: Identify the temporal synchronization between the resource usage peak of background processes and the performance degradation of target applications; identify the correlation between the increase in disk I / O queue length and the slow response of applications that rely on disk reading and writing; and identify the process distribution when the total CPU utilization is normal but the utilization of a single core is saturated.

[0012] Preferably, in step S5, the abnormal behavior detection includes: The performance indicator data within the fault time window is compared with at least one of the performance indicator data of the target terminal device in the adjacent normal operating time period or the preset historical performance indicator baseline data; and based on the comparison result, the performance indicators of abnormal peaks, abnormal troughs or persistent abnormal states that deviate from the normal operating time period data or the baseline data within the fault time window are identified.

[0013] Preferably, the evidence chain includes the process behavior type that causes the fault and the corresponding fault phenomenon, and the process behavior type includes at least one of excessive CPU usage, network contention, disk bottleneck, memory leak and system error triggering.

[0014] The root cause analysis method based on continuous system performance historical data designed in this application, through continuous and high-frequency collection of multi-dimensional system performance data and precise timestamp marking, combined with historical data retrieval, multi-dimensional correlation analysis, evidence chain construction and system status snapshot backtracking, can achieve accurate positioning of existing, non-crash, sporadic system performance failures, clarify the process behavior or resource bottleneck that caused the failure, provide an objective basis for fault diagnosis, improve the accuracy and efficiency of troubleshooting, and help improve system operation stability and user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a flowchart of a fault root cause analysis method based on continuous system performance historical data provided by an embodiment of the present application. DETAILED DESCRIPTION

[0016] The preferred embodiments of the present application are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application and are not used to limit the present application.

[0017] like Figure 1 As shown, the fault root cause analysis method based on continuous system performance historical data described in this embodiment includes the following steps: S1. On the target terminal device, a preset data acquisition agent is used to continuously acquire performance indicator data of the target terminal device at preset time intervals, and a collection timestamp is assigned to each collected performance indicator data. In this embodiment, the target terminal device may be an Internet cafe client computer installed with a Windows 7 or Windows 10 operating system. A lightweight data acquisition agent is pre-deployed on the client computer, which runs in the background as a system service. Specifically, the collection timestamp is accurate to the millisecond level, for example, recorded in the format of "YYYY-MM-DD HH:MM:SS.mmm", such as "2023-10-27 10:30:15.123". The preset time interval is set to 3 seconds, meaning that the data acquisition agent executes a data acquisition cycle every 3 seconds.

[0018] In specific implementation, the data collection agent program obtains performance indicator data in the following ways: i. Query and obtain performance data through the PDH (Performance Data Helper) computer performance database interface. This performance data includes system-level and process-level CPU utilization, memory utilization, network interface card (NIC) transmit / receive rate, number of NIC transmit / receive operations, disk I / O read / write rate, number of disk I / O read / write operations, number of memory page faults, CPU hard interrupt frequency, and GPU functional unit utilization. This provides data support for in-depth analysis of complex performance issues.

[0019] Specifically, the steps for obtaining performance data include: A). Call the PdhOpenQuery function to open the system performance data query.

[0020] B) Construct a query statement for a specific performance metric and add it to the query using the PdhAddCounter API. For each performance metric you want to monitor, construct a corresponding query statement, such as "\\Processor(_Total)\\Interrupts / sec" to query the total CPU interrupt rate. Then, call PdhAddCounter(hQuery, szCounterPath, 0, &hCounter) for each query statement to add it to the query.

[0021] C) After waiting for the preset delay time, for example, 1000 milliseconds after the first addition, call the PdhCollectQueryData function to collect the current performance data. This function collects the current raw data of all added counters.

[0022] D) Call the PdhGetFormattedCounterValue function to obtain the formatted performance data value for temporary storage or subsequent processing.

[0023] ⅱ. Query and obtain hardware information through the Windows Management Instrumentation (WMI) query interface, including temperature sensor information of the CPU, motherboard, and GPU, number of memory sticks and their brand, frequency, and channel information, CPU voltage information, and real-time power information of the power supply. This provides analytical support for certain performance issues that are caused by hardware status or configuration rather than software. For example, when performing root cause analysis, if the system performance history data shows that the CPU or GPU temperature continues to rise abnormally within a specific time period, or even exceeds the warning threshold, this can be further correlated and analyzed to see whether there are correlations such as automatic CPU frequency reduction within the time period, and whether the performance problem is caused by hardware overheating. The troubleshooting direction will then be focused on physical hardware aspects such as whether the cooling fan has stopped due to a fault, whether the cooling fins have accumulated too much dust, or whether the silicone grease has aged and failed.

[0024] Specifically, the steps for obtaining hardware information include: a). Call the CoInitializeEx function to initialize the Component Object Model (COM) library.

[0025] b). Call the CoCreateInstance function to create an instance of the WMI service interface.

[0026] c) Construct a WMI query language statement for specific hardware information and call ExecQuery to execute it. For example, to query memory speed information, the WQL statement might be "SELECT Speed ​​FROM Win32_PhysicalMemory". Then, call IWbemServices::ExecQuery to execute the query and obtain an enumerator, pEnumerator.

[0027] d) Parse and extract the required hardware information from the result set returned after executing the query. For example, by iterating pEnumerator, parse and extract the required hardware information field values ​​from the returned result set.

[0028] ⅲ. Query and obtain process-level information by calling the Windows system API, including the list of active processes, process handle information, working set memory information, number of user objects, number of handles, and number of threads.

[0029] S2. The performance indicator data with the timestamp is transmitted to a data storage system, and the data storage system uses the device identifier, the collection timestamp and the indicator type to perform structured storage and indexing on the collected performance indicator data.

[0030] In practice, the performance metric data for each dimension collected in step S1, with millisecond-level timestamps, is transmitted periodically (e.g., every minute, when the data volume reaches a certain threshold, or in real time) from the data collection agent on the target terminal device via an encrypted network connection to a data storage system. This data storage system can be deployed in a cloud server cluster or a local data center. The data storage system can use a time series database, such as InfluxDB or TimescaleDB, for data storage, allowing efficient retrieval of historical performance data based on device ID, time range, and specific metric type.

[0031] S3. When a fault report is received, the target terminal device identifier and the time window in which the fault occurs are determined, and performance indicator data within a corresponding time period is retrieved from the data storage system according to the device identifier and the time window.

[0032] For example, suppose an internet cafe user, whose client device ID is CLIENT-001, reports to the network administrator: "Between 2:30 PM and 2:35 PM on October 27, 2023, I experienced significant lag and low FPS while playing game XX." The operation and maintenance personnel can enter the following information through the operation and maintenance management platform interface: Target terminal device identifier: CLIENT-001 Fault occurrence time window: start time 2023-10-27 14:30:00.000, end time 2023-10-27 14:35:00.000.

[0033] Based on this information, the system then initiates a query request to the data storage system in step S2 to retrieve all stored multi-dimensional performance indicator data of the CLIENT-001 device in the time period from 2023-10-27 14:30:00.000 to 2023-10-27 14:35:00.000.

[0034] S4. Visualize the retrieved performance indicator data on the same timeline to generate a snapshot of the overall system status of the target terminal device when the fault occurs.

[0035] Specifically, the system visualizes the performance indicator data for CLIENT-001 retrieved in step S3 within the fault time window on an interactive chart interface, with time as the horizontal axis and the selected performance indicator values ​​as the vertical axis. For example, the following curves can be displayed simultaneously: the FPS value of the XX game process, the CPU usage of the XX game process, the total CPU usage of the entire machine, the total memory usage of the entire machine, the percentage of active time of Disk C, and the total bandwidth usage of the network interface.

[0036] Based on the retrieved data, the operation and maintenance personnel can select one or more key moments when the fault phenomenon is most obvious, for example, the moment when the FPS curve reaches its lowest point at 2023-10-27 14:32:25.000, to generate a snapshot of the overall system status of the target terminal device at that moment.

[0037] In this embodiment, the system overall status snapshot includes the following items within the time window when the fault occurs: 1) Specific values ​​of various performance indicators. For example, at 14:32:25.000, the total CPU usage was 75%, the total memory usage was 80%, the disk C: active time was 90%, and the total network bandwidth was 5 Mbps.

[0038] 2) A list of process resource consumption and their respective resource consumption values. For example, the CPU usage of XX game.

[0039] 3) List of process network connections. For example, the TCP connection between the XX game and the game server, including the instantaneous sending and receiving rates at that time 4) Recorded system errors and application events. For example, if a disk controller error or a crash report event of the game itself is recorded in the system event log between 14:32:00 and 14:33:00.

[0040] S5. Analyze the changing trends of different indicators within the fault time window, identify the temporal synchronization correlation between indicators, and detect abnormal behavior by comparing with the data in the normal time period.

[0041] In specific implementation, identifying the temporal synchronization relationship between different indicators includes: Identify the temporal synchronization between the resource usage peak of background processes and the performance degradation of target applications; identify the correlation between the increase in disk I / O queue length and the slow response of applications that rely on disk reading and writing; and identify the process distribution when the total CPU utilization is normal but the utilization of a single core is saturated.

[0042] Specifically, during analysis, operations personnel discovered that the FPS curve for the XX game began to drop sharply at 14:32:15,000. Simultaneously, they observed that the CPU usage curve for the WindowsYY.exe process rapidly climbed to a peak at 14:32:10,000, and its disk write rate also peaked simultaneously. This indicates that the performance degradation of the XX game and the resource usage of WindowsYY.exe were highly synchronized. For another example, they observed a sudden increase in in-game latency (ping value), coinciding with a sudden peak in the network upload / download rate of a certain P2P downloader.

[0043] The abnormal behavior detection includes comparing the performance indicator data within the fault time window with at least one selected from the performance indicator data of the target terminal device during an adjacent normal operating time period or preset historical performance indicator baseline data. Based on the comparison result, the abnormal peak, abnormal trough, or persistent abnormal state of the performance indicator within the fault time window is identified, which deviates from the normal operating time period data or the baseline data.

[0044] Specifically, during analysis, the O&M personnel compared the performance indicator data of CLIENT-001 within the fault time window (14:30:00-14:35:00) with the following data: similar performance indicator data of CLIENT-001 in an adjacent operating time period with normal user feedback (for example, 13:00:00-13:30:00 on the same day); or with the performance indicator baseline data for that period obtained by the system based on the long-term historical data of CLIENT-001.

[0045] Based on this comparison, abnormal behavior can be identified. For example, the average CPU usage of WindowsYY.exe during the fault period (e.g., 25%) is much higher than its average CPU usage during normal periods (e.g., 2%), which is an abnormal spike. Another example is the average FPS of game XX (e.g., 25 frames during the fault period) is significantly lower than its average FPS during normal periods (e.g., 60 frames), which is a persistent abnormal trough.

[0046] S6. Based on the temporal synchronization correlation between indicators and the results of abnormal behavior detection, a chain of evidence from abnormal behavior to fault phenomenon is constructed.

[0047] Specifically, the evidence chain includes the process behavior type that causes the fault and the corresponding fault phenomenon, and the process behavior type includes at least one of excessive CPU usage, network contention, disk bottleneck, memory leak and system error triggering.

[0048] In specific implementation, based on the analysis results of step S5, for example, if the FPS drop of the XX game is synchronized with the peak of WindowsYY.exe resource usage, and the resource usage of WindowsYY.exe far exceeds the normal baseline, the following chain of evidence is constructed: "During the time window from 2023-10-27 14:32:10.000 to 14:32:40.000, the XX game running on the target terminal CLIENT-001 experienced severe lag (FPS dropped from a stable 60 frames to an average of 25 frames). This phenomenon is highly correlated with the abnormal behavior of the process WindowsYY.exe.

[0049] Specifically, during this period, the CPU usage of WindowsYY.exe exceeded the normal baseline by 2% to 25%, and the disk write rate exceeded the normal baseline by 5MB / s to 80MB / s, indicating excessive CPU usage and continuous disk I / O bottleneck behavior.

[0050] The peak in WindowsYY.exe's resource usage coincided with the sudden drop in the FPS of the XX game. This chain of evidence clearly identifies the fault phenomenon, the suspicious process, the specific type of abnormal behavior, the abnormal performance data supporting the determination, and their temporal correlation with the fault phenomenon. Therefore, it is determined that the background update activity of WindowsYY.exe is the root cause of the XX game lag.

[0051] S7. Generate an analysis report containing the evidence chain and a snapshot of the overall system status. In specific implementations, the report can be exported to PDF or other formats for network management to archive, explain reasons to users, or use for subsequent system optimization and policy adjustments.

[0052] The root cause analysis method based on continuous system performance historical data provided in the embodiment of the present application, through continuous and high-frequency collection of multi-dimensional system performance data and precise timestamp marking, combined with historical data retrieval, multi-dimensional correlation analysis, evidence chain construction and system status snapshot backtracking, can achieve accurate positioning of existing, non-crash, sporadic system performance failures, clarify the process behavior or resource bottleneck that causes the failure, provide an objective basis for fault diagnosis, improve the accuracy and efficiency of troubleshooting, and help improve system operation stability and user experience.

[0053] In the description of this application, it should be noted that the terms "vertical", "up", "down", "horizontal", etc. indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, they cannot be understood as limitations on this application.

[0054] In the description of this application, it should also be noted that, unless otherwise clearly specified and limited, the terms "set," "install," "connect," and "connect" should be understood in a broad sense. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections or electrical connections; they can refer to direct connections or indirect connections through an intermediate medium, or they can refer to internal connections between two components. For ordinary operation and maintenance personnel in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances.

[0055] Finally, it should be noted that the above description is merely a preferred embodiment of the present application and is not intended to limit the present application. Although the present application has been described in detail with reference to the aforementioned embodiments, operators skilled in the art may still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein with equivalents. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of protection of the present application.

Claims

1. A method for root cause analysis of failures based on continuous system performance historical data, characterized in that: The following steps are involved: S1. On the target terminal device, through a preset data collection agent, continuously obtains the performance indicator data of the target terminal device at a preset time interval, and marks the acquisition timestamp of each performance indicator data collected; S2. The performance indicator data with a timestamp is transferred to a data storage system, which uses a device identifier, a collection timestamp, and an indicator type to perform structured storage and indexing on the collected performance indicator data; S3. When a fault report is received, the target terminal device identifier and the time window in which the fault occurs are determined, and the performance indicator data within the corresponding time period is retrieved from the data storage system according to the device identifier and time window; S4. Visualize the retrieved performance indicator data on the same timeline to generate a snapshot of the overall system status of the target terminal device at the time of the failure; S5. Analyze the changing trends of different indicators within the fault time window, identify the temporal synchronization correlation between indicators, and detect abnormal behavior by comparing with data from normal time periods; S6. Based on the temporal synchronization relationship between indicators and the results of abnormal behavior detection, build an evidence chain from abnormal behavior to fault phenomenon; S7. Generate an analysis report including the evidence chain and a snapshot of the overall status of the system.

2. The method for root cause analysis of failures based on continuous system performance historical data according to claim 1, characterized in that: The acquisition timestamp is accurate to millisecond level, and the preset time interval is 3 seconds.

3. The method for root cause analysis of failures based on continuous system performance historical data according to claim 1, characterized in that: The data collection agent obtains performance indicator data in the following ways: Query and obtain performance data through the PDH computer performance database interface, including system-level and process-level CPU utilization, memory utilization, network card transceiver rate, network card transceiver operation number, disk I / O read / write rate, disk I / O read / write operation number, memory page fault number, CPU hard interrupt frequency, and GPU functional unit utilization rate; Query and obtain hardware information through the WMI query interface, including temperature sensor information of the CPU, motherboard, and GPU, number of memory sticks, their brand, frequency, and channel information, CPU voltage information, and real-time power information of the power supply; By calling the system API, you can query and obtain process-level information, including the list of active processes, process handle information, working set memory information, the number of user objects, the number of handles, and the number of threads.

4. The method for root cause analysis of failures based on continuous system performance historical data according to claim 3, characterized in that: The specific steps to obtain performance data include: Call the PdhOpenQuery function to open a system performance data query; construct a query statement for a specific performance indicator and add the query statement to the query using the PdhAddCounter API; After waiting for the preset delay time, call the PdhCollectQueryData function to collect the current performance data; Call the PdhGetFormattedCounterValue function to obtain the formatted performance data value.

5. The method for root cause analysis of faults based on continuous system performance historical data according to claim 3, characterized in that: The specific steps to obtain hardware information include: Call the CoInitializeEx function to initialize the component object model library; Call the CoCreateInstance function to create an instance of the WMI service interface; Constructing a WMI query language statement for specific hardware information and calling ExecQuery to execute the query statement; Parse and extract the required hardware information from the result set returned after executing the query.

6. The method for root cause analysis of failures based on continuous system performance historical data according to claim 3, characterized in that: The system overall status snapshot generated in step S4 includes the following items within the time window of the fault: Specific values ​​of various performance indicators; A list of process resource consumption and their respective resource consumption values; Process network connection list; Logged system errors and application events. Application events.

7. The method for root cause analysis of failures based on continuous system performance historical data according to claim 3, characterized in that: In step S5, identifying the temporal synchronization association relationship between different indicators includes: Identify the temporal synchronization between the resource usage peak of background processes and the performance degradation of target applications; Identify the correlation between the increase in disk I / O queue length and the slow response of applications that rely on disk reading and writing; Identify the process distribution when the total CPU utilization is normal but the utilization of a single core is saturated.

8. The method for root cause analysis of failures based on continuous system performance historical data according to claim 3, characterized in that: In step S5, the abnormal behavior detection includes: The performance indicator data within the fault time window is compared with at least one of the performance indicator data of the target terminal device in the adjacent normal operating time period or the preset historical performance indicator baseline data; and based on the comparison result, the performance indicator of abnormal peaks, abnormal troughs or persistent abnormal states that deviate from the normal operating time period data or the baseline data within the fault time window is identified.

9. The method for root cause analysis of failures based on continuous system performance history data according to any one of claims 6 to 8, characterized in that: The evidence chain includes the process behavior type that causes the fault and the corresponding fault phenomenon, and the process behavior type includes at least one of excessive CPU usage, network contention, disk bottleneck, memory leak and system error triggering.