Computer fault alarm system and method

By deeply analyzing user operation logs and system resources, identifying load fluctuations and event associations, providing accurate fault warning indicators and making resource adjustments, it solves the problems of inaccurate fault diagnosis and inaccurate event analysis in the existing technology, and improves the fault handling efficiency and recovery capabilities of computer systems.

CN120315971AActive Publication Date: 2025-07-15ZHENGZHOU UNIVERSITY OF LIGHT INDUSTRY

Patent Information

Application Number
CN202510380470.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-15
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

The existing computer fault alarm system lacks the ability to handle large-scale real-time data, resulting in inaccurate fault diagnosis and response, lack of accurate user behavior and performance fluctuations analysis, affecting system maintenance complexity and operational costs, and inaccurate event correlation and propagation path analysis, extending system downtime.

Method used

Through the behavioral pattern analysis module, performance fluctuation detection module, event association analysis module and fault warning module, combined with the resource adjustment response module, in-depth analysis of user operation logs and system resources is realized, load fluctuations and event associations are identified, accurate fault warning indicators are provided, and resource adjustments are carried out to improve system recovery capabilities.

Benefits of technology

It improves fault detection capabilities, enhances the response speed and accuracy of early warnings, optimizes the system's fault handling efficiency, reduces system maintenance complexity and operational costs, and improves the fault tolerance of computer systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120315971A_ABST
    Figure CN120315971A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, in particular to a computer fault alarm system and method.The computer fault alarm system comprises a behavior pattern analysis module, a performance fluctuation detection module, an event association analysis module, a fault early warning module and a resource adjustment response module. According to the method, the accurate behavior deviation index is provided by comparing the high-frequency operation type with the historical fault data, the index is used for optimizing the monitoring process of the performance fluctuation, the fault detection capability of the system is improved especially when the system performance abnormity is isolated and recognized, and the fault detection capability of the system is improved through the system analysis of the load fluctuation. Compared with a conventional load change trend, the response speed and early warning accuracy of early warning are enhanced, event correlation analysis is more accurate by comprehensively utilizing the data, an event propagation path and correlation degree can be effectively mapped, the timeliness and accuracy of fault early warning are improved, and the fault early warning efficiency is improved through targeted resource adjustment. And the coping recovery capability and the fault-tolerant rate of the computer system are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and particularly to a computer fault alarm system and method. Background Art

[0002] The field of computer technology is a broad technological field that encompasses various aspects such as computer hardware, software, networks, and data processing. This field explores how to design, construct, and optimize computer systems and their interacting components to improve processing speed, data storage capacity, and information security. Computer technology plays a central role in daily life, supporting not only business operations but also driving the development of emerging technologies such as artificial intelligence, big data analysis, and cloud computing.

[0003] Among them, a computer fault alarm system is a system used to detect potential errors in a computer system. The main purpose of the system is to monitor the running status of computer hardware and software in real time. Once an abnormality or fault is detected, it can immediately notify the system administrator or other relevant personnel, which can help reduce the duration of system errors and optimize maintenance responses, thus ensuring the efficient and stable operation of the information technology infrastructure.

[0004] The prior art fails to provide sufficient data parsing depth, resulting in slow or inaccurate fault diagnosis and response. The defects mainly stem from the system's insufficient ability to process large-scale real-time data, especially in terms of the effectiveness of performance monitoring and anomaly detection. In addition, traditional technologies cannot accurately respond in event correlation and propagation path analysis, affecting the speed and efficiency of fault handling. The lack of accurate user behavior and performance fluctuation analysis makes it difficult for system administrators to accurately identify the fault source, not only increasing the complexity of system maintenance but also resulting in longer system downtime or higher operating costs. Summary of the Invention

[0005] In order to solve the technical problems existing in the prior art, namely, the failure to provide sufficient data parsing depth, resulting in slow or inaccurate fault diagnosis and response, the defects mainly stemming from the system's insufficient ability to process large-scale real-time data, especially in terms of the effectiveness of performance monitoring and anomaly detection. In addition, traditional technologies cannot accurately respond in event correlation and propagation path analysis, affecting the speed and efficiency of fault handling. The lack of accurate user behavior and performance fluctuation analysis makes it difficult for system administrators to accurately identify the fault source, not only increasing the complexity of system maintenance but also resulting in longer system downtime or higher operating costs, embodiments of the present invention provide a computer fault alarm system and method. The technical solution is as follows:

[0006] On the one hand, a computer fault alarm system is provided, and the system includes:

[0007] The behavior pattern analysis module obtains user operation log data, parses the operation time, operation type, and operation frequency, filters out high-frequency operation types, and then analyzes the operation consistency to obtain the behavior deviation index;

[0008] The performance fluctuation detection module, based on the behavior deviation index, identifies the load fluctuation amplitude, isolates the time period of abnormal fluctuations, calculates the CPU utilization rate fluctuation within the time period, and compares it with the average load change trend to obtain the load fluctuation index;

[0009] The event correlation analysis module, based on the load fluctuation index, analyzes error events, timeout events, and resource access conflict events in the computer log, counts the occurrence frequency and distribution pattern of the events within the abnormal time period, identifies the correlation degree between the events, and obtains the event correlation intensity;

[0010] The fault warning module, based on the event correlation intensity, analyzes the occurrence frequency and influence range of the current abnormal events, calculates the cumulative influence degree, evaluates the consistency with the known fault trigger mode, determines the fault risk level, and obtains the fault warning index;

[0011] The resource adjustment response module, based on the fault warning index, identifies computer components with a high current risk level, allocates additional computing resources or restarts service processes to obtain computing resource adjustment information.

[0012] On the other hand, the behavior deviation index includes operation time deviation, operation frequency deviation, and operation type deviation; the load fluctuation index includes CPU load fluctuation, memory usage fluctuation, and network response fluctuation; the event correlation intensity includes error event correlation degree, timeout event correlation degree, and resource conflict event correlation degree; the fault warning index includes event occurrence frequency, event influence range, and event risk level; and the computing resource adjustment information includes resource allocation effect, process restart effect, and performance adjustment result.

[0013] On the other hand, the behavior pattern analysis module includes:

[0014] The log parsing sub-module obtains user operation log data, parses the operation time, operation type, and operation frequency, sorts the logs according to time tags, analyzes the operation activities within each time period, and extracts the correlation distribution data between time points and operation types to obtain operation time series data;

[0015] The frequency calculation sub-module, based on the operation time series data, calculates the occurrence frequency of the operation type within different time periods, statistically analyzes the frequency change trend, filters out high-frequency operation types associated with known faults, and calculates their relative proportion in all operations to obtain the high-frequency operation proportion index;

[0016] The abnormal behavior recognition sub-module, based on the high-frequency operation ratio index, combines the user's authentication data and historical operation patterns to compare the current operation with historical data for consistency, analyzes the deviation degree of the operation behavior, and obtains a behavior deviation index.

[0017] On the other hand, the performance fluctuation detection module includes:

[0018] The system resource collection sub-module, based on the behavior deviation index, collects the computer CPU load, memory usage rate, and network response time, monitors the computer resource utilization in the differential time window, analyzes the time series change situation, and obtains an overview of the resource usage situation;

[0019] The load fluctuation recognition sub-module, based on the overview of the resource usage situation, calculates the fluctuation amplitude of the CPU load and memory usage rate within the time window, compares the resource occupancy trend within the time window, filters out the time periods with the load fluctuation amplitude exceeding the standard, and obtains an abnormal time interval;

[0020] The load trend calculation sub-module calls the abnormal time interval, analyzes the CPU utilization fluctuation within the time period, compares the average load fluctuation of the time period with the overall data, determines the long-term trend and short-term variation of the fluctuation, and obtains a load fluctuation index.

[0021] On the other hand, the event correlation analysis module includes:

[0022] The abnormal event statistics sub-module, based on the load fluctuation index, analyzes the error events, timeout events, and resource access conflict events in the computer log, counts the occurrence frequency of each type of event in the abnormal time period, evaluates the event concentration degree, and identifies the event distribution pattern, and obtains an overview of the abnormal event distribution;

[0023] The event correlation sub-module, based on the overview of the abnormal event distribution, calculates the co-occurrence probability of different types of events in the abnormal time period, analyzes the time interval and frequency between events, determines the dependency relationship between events, and obtains an event correlation degree index;

[0024] The event propagation analysis sub-module calls the event correlation degree index, analyzes the event occurrence order, calculates the propagation path length and influence between events, counts the cascade influence range of events, determines the key trigger event, and obtains the event correlation strength.

[0025] On the other hand, for analyzing the time interval and frequency between events, the formula:

[0026]

[0027] Analyze the time interval and frequency between events, determine the dependency relationship between events, and obtain the event correlation degree index, where PZ uoRepresents the co-occurrence probability of event u and event o, τ uk Represents the start time of event u in the k-th abnormal time period, τ ok Represents the start time of event o in the k-th abnormal time period, T k Represents the total duration of the k-th abnormal time period, N PZ Represents the total number of abnormal time periods.

[0028] On the other hand, the fault warning module includes:

[0029] The abnormal event screening sub-module analyzes the occurrence frequency of abnormal events based on the event correlation strength, calculates the cumulative impact value of each type of event, counts the number of abnormal events exceeding the threshold, and identifies the abnormal events that are key to the impact scope, obtaining a list of key abnormal events;

[0030] The fault mode matching sub-module calls the list of key abnormal events, calculates the matching degree between the events and the known fault trigger modes, analyzes the consistency between the event occurrence order and the known fault modes, and identifies the key impact components for the matching sequence, obtaining the fault mode matching degree;

[0031] The risk level assessment sub-module counts the number of computer components involved in the matching events according to the fault mode matching degree, evaluates the fault risk level, and determines the computer fault risk level in combination with the fault impact scope, obtaining the fault warning index.

[0032] On the other hand, the formula for counting the number of computer components involved in the statistical matching events is:

[0033]

[0034] Evaluates the fault risk level and determines the computer fault risk level, where N Z Represents the total number of computer components involved, C i Represents the absolute value of the set of computer components involved in the i-th fault event, |C i | represents taking the number of elements in the set of components involved in the i-th event, M represents the total number of matching events, f j Represents the fault frequency of the j-th component, and n represents the total number of components within the evaluation period.

[0035] On the other hand, the resource adjustment response module includes:

[0036] The high-risk component identification sub-module analyzes the risk levels of computer components based on the fault warning index, counts the resource usage of high-risk components, evaluates the current load status, and determines the resource allocation objects that need to be adjusted, obtaining the key risk components;

[0037] The resource adaptation and adjustment sub-module calls the key risk components, analyzes the status of the computer's available resource pool, evaluates the allocation ability of computing resources, calculates the current resource utilization rate, filters the computing resources available for adjustment, and replenishes resources or restarts processes for high-risk components to obtain the resource allocation details;

[0038] The performance change recording sub-module monitors the running status of the adjusted computing resources according to the resource allocation details, records the change trends of CPU, memory, and network usage rates, and evaluates the impact of resource adjustment on computer stability to obtain the computing resource adjustment information.

[0039] On the other hand, a computer fault alarm method is provided. The method is applied to a computer fault alarm system and includes the following steps:

[0040] S1: Obtain user operation log data, parse the operation time, operation type, and operation frequency, arrange the operation time series, calculate the operation frequency change rate, filter out the high-frequency operation types associated with known faults, and then analyze the operation consistency based on the user identity and historical operation behavior to obtain the behavior deviation index;

[0041] S2: Based on the behavior deviation index, collect the computer's CPU load, memory usage rate, and network response time. For each time window data, identify the load fluctuation range, calculate the CPU utilization rate fluctuation, and compare it with the average load change trend to obtain the load fluctuation index;

[0042] S3: Based on the load fluctuation index, count the occurrence frequency and distribution pattern of events in the abnormal time period, identify the correlation degree between events, and analyze the occurrence order and propagation path of events to obtain the event correlation strength;

[0043] S4: Based on the event correlation strength, calculate the cumulative impact degree of events, filter out abnormal events exceeding the threshold, evaluate the consistency between events and known fault trigger patterns, and perform correlation analysis on events that conform to the fault pattern and computer components to obtain the fault warning index;

[0044] S5: Based on the fault warning index, identify the computer components with high current risk levels, analyze the status of the available resource pool, perform resource feasibility adjustments, allocate additional computing resources or restart service processes for high-risk components to obtain the computing resource adjustment information.

[0045] The beneficial effects brought by the technical solution provided in the embodiments of the present invention at least include:

[0046] By comparing high-frequency operation types with historical fault data, precise behavior deviation indicators are provided. These indicators are used to optimize the monitoring process of performance fluctuations. Especially when isolating and identifying system performance anomalies, the system's fault detection ability is improved. Through the systematic analysis of load fluctuations and comparison with the conventional load change trend, the reaction speed and accuracy of early warning are enhanced. By comprehensively utilizing these data, event correlation analysis becomes more accurate, capable of effectively mapping the event propagation path and correlation degree, improving the timeliness and accuracy of fault early warning. Through targeted resource adjustment, the computer system's response and recovery ability and fault tolerance rate are enhanced. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0048] Figure 1 System schematic diagram of the present invention;

[0049] Figure 2 System framework schematic diagram of the present invention;

[0050] Figure 3 Flowchart of the behavior pattern analysis module of the present invention;

[0051] Figure 4 Flowchart of the performance fluctuation detection module of the present invention;

[0052] Figure 5 Flowchart of the event correlation analysis module of the present invention;

[0053] Figure 6 Flowchart of the fault early warning module of the present invention;

[0054] Figure 7 Flowchart of the resource adjustment response module of the present invention;

[0055] Figure 8 Flowchart of the method steps of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] The following will describe the technical solutions in the present invention in conjunction with the drawings.

[0057] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to give examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of the word "example" is intended to present concepts in a specific manner. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.

[0058] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same. "Of", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same.

[0059] In the embodiments of the present invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meanings they express are the same.

[0060] To make the technical problems to be solved, technical solutions and advantages of the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.

[0061] The embodiments of the present invention provide a computer fault alarm system, as Figure 1 shown, the system includes:

[0062] The behavior pattern analysis module obtains user operation log data, parses the operation time, operation type and operation frequency, arranges the operation time series, calculates the operation frequency change rate, screens out the high-frequency operation types associated with known faults, calculates their frequency ratios, and then analyzes the operation consistency based on the user identity and historical operation behavior to obtain the behavior deviation index;

[0063] The performance fluctuation detection module, based on the behavior deviation index, collects the computer CPU load, memory usage rate and network response time. For each time window data, it identifies the load fluctuation amplitude, isolates the time period of abnormal fluctuation, and calculates the CPU utilization rate fluctuation within the time period, and compares it with the average load change trend to obtain the load fluctuation index;

[0064] The event correlation analysis module, based on the load fluctuation index, analyzes the error events, timeout events and resource access conflict events in the computer log, counts the occurrence frequency and distribution pattern of the events during the abnormal time period, identifies the correlation degree between the events, and analyzes the occurrence order and propagation path of the events to obtain the event correlation intensity;

[0065] The fault warning module analyzes the occurrence frequency and influence scope of the current abnormal event based on the event correlation intensity, calculates the cumulative influence degree of the event, filters out the abnormal events exceeding the threshold, evaluates the consistency between the event and the known fault trigger mode, and conducts a correlation analysis on the events conforming to the fault mode and the computer components to determine the fault risk level of the computer and obtain the fault warning indicators.

[0066] Based on the fault warning indicators, the resource adjustment response module identifies the computer components with a high current risk level, analyzes the status of the available resource pool, makes feasible resource adjustments, allocates additional computing resources or restarts the service process for the high-risk components, and records the performance changes after the adjustment to obtain the computing resource adjustment information.

[0067] The behavior deviation indicators include operation time deviation, operation frequency deviation, and operation type deviation. The load fluctuation index includes CPU load fluctuation, memory usage fluctuation, and network response fluctuation. The event correlation intensity includes error event correlation degree, timeout event correlation degree, and resource conflict event correlation degree. The fault warning indicators include event occurrence frequency, event influence scope, and event risk level. The computing resource adjustment information includes resource allocation effect, process restart effect, and performance adjustment result.

[0068] As Figure 2 and Figure 3 shown, the behavior pattern analysis module includes:

[0069] The log parsing sub-module obtains the user operation log data, parses the operation time, operation type, and operation frequency, sorts the logs according to the time tags, analyzes the operation activities in each time period, and extracts the correlation distribution data between the time points and the operation types to obtain the operation time series data.

[0070] An access log storage system reads the operation log records of users from the system in chronological order, including the operation time, operation type, and operation frequency in the log, and parses the data into a structured format for storage. Each log data needs to parse out the user ID, operation instruction, and timestamp of the operation occurrence, and sort them in ascending order of the timestamp. During the parsing process, read the log entries, extract the timestamp field, convert it to the standard time format (such as YYYY-MM-DD HH:MM:SS), and store it as the time label in the time series data. Then parse the operation type field, extract the specific instruction executed by the user according to the different formats of the operation log, and mark the operation type identifier. For example, the operation type can be divided into "login", "modify configuration", "access file", "execute command", etc. When parsing the operation frequency, it is necessary to count the number of times the same user executes the same operation within a unit time (such as every 5 minutes, every hour), and generate the operation frequency data. Store the parsed data into the database and sort it according to the time label, and the sorting method uses ascending order based on the timestamp, so that the operation logs of the same user can be analyzed in chronological order. On this basis, analyze the operation activities within each time period. For example, divide a day into 24 hours, and each hour is used as a time period to count the operation behaviors of all users within this time period, generate the operation distribution data within the time period, and extract the correlation data between each time point and the operation type. For example, a certain user accesses the database 10 times, accesses the file 5 times, and modifies the configuration 2 times between 12:00 and 13:00. In this way, the operation characteristics of different time periods can be obtained and stored as operation time series data.

[0071] The frequency calculation sub-module calculates the occurrence frequency of the operation type in different time periods based on the operation time series data, statistically analyzes the frequency change trend, filters out the high-frequency operation types associated with known faults, calculates their relative proportion in all operations, and obtains the high-frequency operation proportion index;

[0072] First, read the operation time series data, classify and count it according to the operation type, and calculate the occurrence times of this operation type within each time period. For example, a certain user executes the "access database" operation 20 times between 10:00 and 11:00, and 30 times between 11:00 and 12:00. Then the frequencies of this operation type in these two time periods are 20 and 30 respectively. Next, calculate the relative proportion of each operation type, that is, the ratio of the execution times of a certain operation type within this time period to the total operation times of this user. For example, this user executes 50 operations between 10:00 and 11:00, among which "access database" accounts for 40% (20 / 50), and "access file" accounts for 30% (15 / 50). Then calculate the proportion index of high-frequency operation types. When screening high-frequency operations, set a benchmark frequency threshold Fth. For example, set Fth = 25 times. Then if the frequency of a certain operation type in a certain time period is greater than or equal to 25 times, it is marked as a high-frequency operation type. Normalize the occurrence frequencies of all high-frequency operation types, that is, divide the frequency of each high-frequency operation type by the total operation times of this user in this time period to obtain the normalized proportion index. For example, a certain user executes 80 operations between 11:00 and 12:00, among which "access database" is 50 times, then the proportion index is 50 / 80 = 0.625. This index is used for subsequent abnormal behavior recognition and analysis.

[0073] The abnormal behavior recognition sub-module, based on the high-frequency operation proportion index, combines the user's identity verification data and historical operation patterns, compares the current operation with the historical data for consistency, analyzes the deviation degree of the operation behavior, and obtains the behavior deviation index;

[0074] Read user authentication data, including the user's identity identifier, login device, IP address, etc., and construct the normal behavior baseline of the user in combination with the user's historical operation mode. Specifically, extract the high-frequency operation proportion index of each time period within a certain period of time in the past (such as the most recent 7 days) of the user, and calculate its mean μ and standard deviation σ. For example, if the mean of the proportion index of a certain operation type of a user in the past 7 days is 0.45 and the standard deviation is 0.05, then the normal operation range of this user can be set as [μ - 2σ, μ + 2σ] = [0.35, 0.55]. Then, perform a consistency comparison on the current operation to calculate whether the high-frequency proportion index of the current operation falls within this interval. For example, if the high-frequency proportion index of the current operation is 0.6, which exceeds the normal range, then calculate the behavior deviation index. This index is defined as the deviation degree D = |(current proportion index - mean) / standard deviation|. In the above example, D = |(0.6 - 0.45) / 0.05| = 3. If D exceeds the set threshold (such as D > 2), then mark this operation as abnormal and further analyze the user's authentication data. For example, check whether the login device has changed and whether the IP address is an address that has not appeared in the historical record. Finally, make a comprehensive judgment based on the characteristics to determine whether the user has abnormal behavior.

[0075] As Figure 2 and Figure 4 shown, the performance fluctuation detection module includes:

[0076] The system resource collection sub-module monitors the computer CPU load, memory usage rate, and network response time based on the behavior deviation index, monitors the utilization of computer resources within the differential time window, analyzes the time series change situation, and obtains an overview of the resource usage situation;

[0077] Call the resource monitoring interface of the operating system (such as / proc / stat in Linux and WMI queries in Windows) to obtain the CPU occupancy rate. Accumulate the total running time of the CPU (including user time, system time, and idle time), and calculate the usage percentage of the CPU per unit time. For example, if the unit time is 5 seconds, read the initial total CPU time T1, and read the new total CPU time T2 after 5 seconds. Then the CPU load calculation formula is (change in CPU usage time) / (change in time) = (T2 - T1) / 5. The memory usage rate is obtained by calling the system interface to get the ratio of the current available memory to the total memory. For example, if a computer has a total memory of 16GB and the current available memory is 4GB, the memory usage rate calculation formula is (total memory - available memory) / total memory = (16 - 4) / 16 = 75%. The network response time is obtained by sending test packets to a specified server and measuring the return time difference. For example, if 10 packets are sent and the return times are 50ms, 55ms, 48ms, 53ms, etc., the network response time is calculated as the average value (50 + 55 + 48 + 53 +...) / 10. The system resource collection sub-module needs to monitor the above resource conditions within different time windows, such as 5 minutes, 1 hour, 1 day, etc., and store the corresponding data. Subsequently, analyze the time series changes, extract the resource utilization conditions for each time period. For example, within the time period from 10:00 to 11:00, the average CPU load is 60%, the memory usage rate is 70%, and the network response time is 52ms. Then construct a resource usage time series based on the data to form an overview of the resource usage situation.

[0078] Based on the overview of the resource usage situation, the load fluctuation identification sub-module calculates the fluctuation amplitude of the CPU load and the memory usage rate within the time window, compares the resource occupancy trends within the time window, filters out the time periods with a load fluctuation amplitude exceeding the standard, and obtains the abnormal time intervals.

[0079] Extract the resource usage data within the set time window. For example, if the set time window is 10 minutes and the CPU load is sampled 6 times within this window to obtain the load data sequence (55%, 60%, 58%, 62%, 59%, 61%), calculate the amplitude of load fluctuation, which is the sum of the absolute values of the difference between the CPU usage rate at the current moment and the previous moment. For example, calculate (|60 - 55| + |58 - 60| + |62 - 58| + |59 - 62| + |61 - 59|) / (the number of samplings - 1) = (5 + 2 + 4 + 3 + 2) / 5 = 3.2%. Similarly, calculate the amplitude of the memory usage rate fluctuation. If the memory usage rate within a certain time window is (65%, 70%, 68%, 73%, 69%, 72%), then calculate the amplitude of fluctuation (|70 - 65| + |68 - 70| + |73 - 68| + |69 - 73| + |72 - 69|) / 5 = (5 + 2 + 5 + 4 + 3) / 5 = 3.8%. Then, compare the resource occupancy trends within the time window, analyze the average load change rate of this time window. The calculation method is the difference between the average resource usage within the window and the average resource usage of the previous window. For example, if the average CPU usage rate of the previous window is 55% and the average CPU usage rate of the current window is 60%, then the change rate is (60 - 55) / 55 = 9.1%. Set the fluctuation standard threshold Th. If Th = 7%, then the CPU load fluctuation in the current window exceeds the standard range. Screen out all time periods that exceed this threshold and record their timestamps to obtain the abnormal time intervals.

[0080] The load trend calculation sub-module calls the abnormal time intervals, analyzes the CPU utilization rate fluctuation within the time period, compares the average load fluctuation of the time period with the overall data, determines the long-term trend and short-term changes of the fluctuation, and obtains the load fluctuation index;

[0081] Analyze the CPU utilization fluctuations during this time period, extract the CPU utilization data within the abnormal time interval, calculate its mean and standard deviation. For example, within the abnormal time period from 10:00 to 10:30, the CPU utilization data is (65%, 70%, 72%, 68%, 75%, 73%). Then calculate the mean μ = (65 + 70 + 72 + 68 + 75 + 73) / 6 = 70.5%, and the standard deviation σ = sqrt(((65 - 70.5)^2 + (70 - 70.5)^2 + (72 - 70.5)^2 + (68 - 70.5)^2 + (75 - 70.5)^2 + (73 - 70.5)^2) / 6) ≈ 3.5%. Subsequently, compare the load fluctuations in this time period with the average load fluctuations of the overall data. Assume that the CPU mean of the overall data is 60% and the standard deviation is 2.5%. Calculate the deviation degree D of the abnormal time interval = |(70.5 - 60) / 2.5| = 4.2. If D exceeds the set threshold (e.g., D > 3), then it is determined that this time period belongs to long-term trend fluctuations. If 2 ≤ D ≤ 3, then it is determined as short-term changes. Calculate the load fluctuation index, define the index LBI = D × σ. For example, LBI = 4.2 × 3.5 = 14.7, and output the load fluctuation index of this time period.

[0082] As Figure 2 and Figure 5 shown, the event correlation analysis module includes:

[0083] The abnormal event statistics sub-module analyzes error events, timeout events, and resource access conflict events in the computer log based on the load fluctuation index, counts the occurrence frequency of each type of event within the abnormal time interval, evaluates the event concentration, and identifies the event distribution pattern to obtain an overview of the abnormal event distribution;

[0084] Read relevant data from the log file or database. Each log contains information such as timestamp, event type, error code, user ID, and scope of impact. Filter out the log records within the time period when the load fluctuation index exceeds the benchmark range according to the timestamp. For example, if the load fluctuation index of a certain system is 15.2 from 10:00 to 11:00, which is higher than the set benchmark value of 10, then extract the logs within this time period, and count the occurrence frequencies of different types of events. For example, within this time period, the error event occurs 10 times, the timeout event occurs 8 times, and the resource access conflict event occurs 5 times. Then calculate the event concentration. Define the concentration index C = (highest frequency of a single event) / (total number of all events within this time period). For example, C = 10 / (10 + 8 + 5) = 0.43. If C is higher than the set threshold (such as 0.4), it indicates that a certain type of event accounts for a relatively high proportion within this time period. Subsequently, analyze the event distribution pattern and count the time distribution of each event. For example, the error event occurs at four time points: 10:05, 10:15, 10:30, and 10:50; the timeout event occurs at 10:07, 10:20, and 10:45; the resource access conflict occurs at 10:25 and 10:40, thus obtaining an overview of the abnormal event distribution.

[0085] Based on the overview of the abnormal event distribution, the event correlation sub-module calculates the co-occurrence probability of different types of events within the abnormal time period, analyzes the time interval and frequency between events, determines the dependency relationship between events, and obtains the event correlation index;

[0086] Analyze the time interval and frequency between events, using the formula:

[0087]

[0088] Analyze the time interval and frequency between events, determine the dependency relationship between events, and obtain the event correlation index, where PZ uo represents the co-occurrence probability of event u and event o, τ uk represents the start time of event u in the k-th abnormal time period, τ ok represents the start time of event o in the k-th abnormal time period, T k represents the total duration of the k-th abnormal time period, N PZ represents the total number of abnormal time periods;

[0089] The specific values obtained from data monitoring are as follows:

[0090] The start times of event u and event o in three time periods are τ u1 = 2 hours, τ o1 = 3 hours, and τ u2 = 5 hours, τ o2 = 7 hours, τ u3 = 8 hours, and τo3 = 9 hours;

[0091] The total duration of each time period is T1 = 2 hours, T2 = 3 hours, and T3 = 1.5 hours;

[0092] N PZ = 3, which is the number of time periods determined through the event log.

[0093] Calculate the absolute value of the difference in the start times of event u and event o within each time period:

[0094] |τ u1 - τ o1 | = 1;

[0095] |τ u2 - τ o2 | = 2;

[0096] |τ u3 - τ o3 | = 1;

[0097] Find the sum of the absolute values of each time difference:

[0098] 1 + 2 + 1 = 4;

[0099] Find the sum of the total durations of each time period:

[0100] 2 + 3 + 1.5 = 6.5;

[0101] Substitute the above results into the formula:

[0102]

[0103] This result indicates that within the three monitored time periods, the co-occurrence probability of event u and event o is approximately 41%. The probability reflects the relative frequency of the simultaneous occurrence of two events within a given time period and is an important indicator for determining the dependency relationship between the two events.

[0104] The event propagation analysis sub-module calls the event correlation index, analyzes the event occurrence order, calculates the propagation path length and influence between events, statistically analyzes the cascading influence range of events, determines the key trigger events, and obtains the event correlation strength;

[0105] Extract timestamp data and construct an event sequence. For example, error events occur at 10:05, 10:15, 10:30, 10:50, timeout events occur at 10:07, 10:20, 10:45, and resource access conflicts occur at 10:25, 10:40. Then establish an event propagation path: error event (10:05) → timeout event (10:07) → error event (10:15) → timeout event (10:20) → resource access conflict (10:25) → error event (10:30) → resource access conflict (10:40) → timeout event (10:45) → error event (10:50). Calculate the length L of the event propagation path, that is, the average propagation interval between events. For example, L = (2 + 8 + 5 + 5 + 5 + 10 + 5 + 5) / 8 = 5.625 minutes. Subsequently, calculate the cascading influence range of the events. Define the influence range R = event propagation path length × number of events. For example, R = 5.625 × 9 = 50.6 minutes. Finally, screen key trigger events, and judge whether an event will trigger multiple subsequent events. For example, if the error event (10:05) subsequently triggers the timeout event (10:07) and the error event (10:15), then the error event (10:05) is a key trigger event. Finally, calculate the event correlation strength. Define S = E × R. For example, S = 0.083 × 50.6 = 4.2, and obtain the event correlation strength.

[0106] Such as Figure 2 and Figure 6 As shown, the fault warning module includes:

[0107] The abnormal event screening sub-module analyzes the occurrence frequency of abnormal events based on the event correlation strength, calculates the cumulative impact value of each type of event, counts the number of abnormal events exceeding the threshold, and identifies the abnormal events that are key to the influence range, obtaining a list of key abnormal events;

[0108] Traverse all records and count the occurrence frequency of each type of event within the abnormal time period. For example, within the time period from 10:00 to 11:00, the error event occurs 10 times, the timeout event occurs 8 times, and the resource access conflict event occurs 6 times. Then record the cumulative frequency of each type of event. Next, calculate the cumulative impact value of each type of event. Define the impact value I = event correlation strength S × number of event occurrences N. For example, if the correlation strength S of a certain type of error event is 4.2 and the number of occurrences N is 10, then the cumulative impact value I of this event type is 4.2 × 10 = 42. Subsequently, count the number of abnormal events with a cumulative impact value exceeding the threshold. Set the impact threshold Th = 30, and screen abnormal events with I > Th. For example, if the impact value I = 42 exceeds the threshold, then determine that this event type is a key abnormal event. Finally, identify the abnormal events that are key to the influence range, and judge whether an event is associated with multiple sub-events. For example, if the error event triggers multiple timeout events after occurrence, then this error event belongs to the abnormal events that are key to the influence range, forming a list of key abnormal events.

[0109] The fault mode matching sub-module calls the list of key abnormal events, calculates the matching degree between the events and the known fault triggering modes, analyzes the consistency between the event occurrence order and the known fault modes, and identifies the key impact components for the matching sequence to obtain the fault mode matching degree;

[0110] Calculate the matching degree between each abnormal event and the known fault triggering mode, extract historical fault data, obtain the event sequences when various faults occur. For example, the event sequence of the known database fault mode is: resource access conflict → timeout event → error event. Extract the same type of event sequence from the current list of key abnormal events, and calculate the matching ratio R = (number of matching events) / (total number of known fault events). Suppose resource access conflict and timeout event are found in the current abnormal events, but the error event is missing, then the matching ratio R = 2 / 3 = 0.67. Subsequently, analyze the consistency between the abnormal event occurrence order and the known fault mode, and calculate the sequence deviation D = sum of squares of (actual sequence position - theoretical sequence position) / number of matching events. For example, the theoretical order is A → B → C, and the actual order is A → C → B, then calculate D = ((1 - 1)^2 + (3 - 2)^2 + (2 - 3)^2) / 3 = 0.67. If D is less than the set threshold (such as 0.5), it is determined that the abnormal sequence matches the known fault mode. Finally, identify the key impact components in the matching sequence, extract all the computer components involved in the abnormal events. For example, resource access conflict involves the database, timeout event involves the network, and error event involves the application server, then record all the involved component information, calculate the fault mode matching degree, and define the matching degree M = R × (1 - D). For example, M = 0.67 × (1 - 0.67) = 0.22, to obtain the fault mode matching degree.

[0111] The risk level assessment sub-module, based on the fault mode matching degree, counts the number of computer components involved in the matching events, assesses the fault risk level, and determines the computer fault risk level in combination with the fault impact scope to obtain the fault warning index;

[0112] Count the number of computer components involved in the matching events, using the formula:

[0113]

[0114] Assess the fault risk level and determine the computer fault risk level, where N Z represents the total number of computer components involved, C i represents the absolute value of the set of computer components involved in the i-th fault event, |C i | represents the number of elements in the component set involved in the i-th event, M represents the total number of matching events, f jrepresents the failure frequency of the j-th component, and n represents the total number of components within the evaluation period;

[0115] There are the following data: the total number of matching events M = 3, the number of computer components involved in each event i being |C1| = 10, |C2| = 15, and |C3| = 5, and the failure frequency f j is the frequency obtained from the monitoring data. For example, f1 = 0.02, f2 = 0.05, and f3 = 0.03, the total number of components n within the evaluation period is 3 (equal to the number of failure frequency terms), and the monitoring and collection of data are carried out through a dedicated computer hardware monitoring system to ensure the accuracy and timely update of the data. Substitute into the formula for calculation, and calculate the average of the failure frequencies:

[0116]

[0117] Calculate the square root of the average failure frequency:

[0118]

[0119] Calculate N Z :

[0120]

[0121] This result indicates that the calculated total number of computer components involved is close to 5.5, reflecting the degree of influence of the components on the system stability under the current failure frequency. Through this value, it is used to quantify and evaluate the potential failure risk of the entire computer system.

[0122] As Figure 2 and Figure 7 shown, the resource adjustment response module includes:

[0123] The high-risk component identification sub-module analyzes the risk levels of computer components based on the failure warning indicators, counts the resource usage of high-risk components, evaluates the current load status, determines the objects of resource allocation that need to be adjusted, and obtains the key risk components;

[0124] Extract the risk level data of computer components, traverse all components, calculate the risk score R of each component, and define R = warning indicator W × component impact coefficient C. For example, for a certain database server, W = 27.6 and C = 0.44, then R = 27.6 × 0.44 = 12.14. Then, count the resource usage of high-risk components, extract CPU, memory, and network resource data, and calculate the resource occupancy rate at the current moment. For example, the CPU usage rate of a certain server is 85%, the memory usage rate is 90%, and the network bandwidth usage rate is 75%. Subsequently, evaluate the current load status, and define the load index L = (CPU usage rate + memory usage rate + network bandwidth usage rate) / 3. If L = (85 + 90 + 75) / 3 = 83.3%, set the load threshold Th = 80%. If L > Th, then determine that the component is in a high-load state. Subsequently, determine the resource allocation object that needs to be adjusted, screen all high-load components, and components with CPU usage rate > 80% and memory usage rate > 85% are determined as adjustment targets to obtain key risk components.

[0125] The resource adaptation and adjustment sub-module calls the key risk components, analyzes the status of the computer's available resource pool, evaluates the allocation ability of computing resources, calculates the current resource utilization rate, screens the computing resources available for adjustment, and supplements resources or restarts processes for high-risk components to obtain the resource allocation details;

[0126] Obtain all underutilized computing resources and calculate their available capacities. For example, the available CPU cores in the current computing resource pool are 10, the free memory capacity is 32GB, and the remaining bandwidth is 100Mbps. Then, evaluate the allocation ability of computing resources and calculate the maximum scalable resource M = min(available CPU cores / target component CPU usage rate, free memory capacity / target component memory usage rate, remaining bandwidth / target component network usage rate). If M = min(10 / 0.85, 32 / 0.90, 100 / 0.75) = min(11.76, 35.56, 133.33) = 11.76. Subsequently, calculate the current resource utilization rate and define U = allocated resources / total available resources. For example, U = (allocated CPU cores + allocated memory + allocated bandwidth) / (total CPU cores + total memory + total bandwidth). Set the adjustment threshold Th = 70%. If U < Th, then determine that there are sufficient adjustable resources. Subsequently, screen the computing resources available for adjustment, traverse all resource pools, and select a server with a CPU occupancy rate below 50% for resource supplementation. For example, select a computing node N1 with a CPU occupancy rate of 45% and allocate some of its CPU cores to high-risk components, and perform resource supplementation or process restart. When the CPU occupancy rate of the high-risk component drops below 80%, stop resource supplementation to obtain the resource allocation details.

[0127] The performance change recording sub-module monitors the running status of the adjusted computing resources according to the resource allocation details, records the changing trends of CPU, memory, and network usage rates, and evaluates the impact of resource adjustment on computer stability to obtain computing resource adjustment information;

[0128] Monitor the running status of the adjusted computing resources, extract the CPU, memory, and network usage rates, record the change values before and after allocation. For example, before resource adjustment, the CPU usage rate of the target server was 85%, and after adjustment, it dropped to 70%, the memory usage rate dropped from 90% to 75%, and the network bandwidth usage rate dropped from 75% to 65%. Subsequently, record the data, calculate the change rate before and after computing resource adjustment. Define the change rate ΔX = (value after adjustment - value before adjustment) / value before adjustment × 100%. For example, the CPU change rate ΔC = (70 - 85) / 85 × 100% = -17.65%, the memory change rate ΔM = (75 - 90) / 90 × 100% = -16.67%, the network change rate ΔN = (65 - 75) / 75 × 100% = -13.33%. Subsequently, evaluate the impact of resource adjustment on computer stability, calculate the stability index S = 1 - (|ΔC| + |ΔM| + |ΔN|) / 3. For example, S = 1 - (17.65 + 16.67 + 13.33) / 3 = 1 - 15.88 = 0.84. Store all the calculated data in the log to obtain computing resource adjustment information.

[0129] As Figure 8 shown, a computer fault alarm method includes the following steps:

[0130] S1: Obtain user operation log data, parse the operation time, operation type, and operation frequency, arrange the operation time series, calculate the operation frequency change rate, and screen the high-frequency operation types associated with known faults. Then, analyze the operation consistency based on the user identity and historical operation behavior to obtain the behavior deviation index;

[0131] S2: Based on the behavior deviation index, collect the computer CPU load, memory usage rate, and network response time. For each time window data, identify the load fluctuation amplitude, calculate the CPU utilization fluctuation, and compare it with the average load change trend to obtain the load fluctuation index;

[0132] S3: Based on the load fluctuation index, count the occurrence frequency and distribution pattern of events in the abnormal time period, identify the correlation degree between events, and analyze the occurrence order and propagation path of events to obtain the event correlation intensity;

[0133] S4: Based on the event correlation intensity, calculate the cumulative impact degree of events, screen the abnormal events exceeding the threshold, evaluate the consistency between events and known fault trigger patterns, and conduct correlation analysis on the events conforming to the fault pattern and computer components to obtain the fault warning index;

[0134] S5: Based on the fault warning indicators, identify computer components with a high current risk level, analyze the status of the available resource pool, perform resource feasibility adjustment, allocate additional computing resources to high-risk components or restart service processes, and obtain computing resource adjustment information.

[0135] It should be understood that the term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context before and after.

[0136] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0137] It should be understood that in various embodiments of the present invention, the magnitude of the sequence numbers of the above processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0138] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different systems for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.

[0139] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing system embodiments, and will not be elaborated here.

[0140] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and systems can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0141] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0142] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0143] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the system described in each embodiment of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, and other media that can store program codes.

[0144] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A computer fault alarm system, characterized in that, The system includes: The behavior pattern analysis module obtains user operation log data, parses the operation time, operation type, and operation frequency, filters the high-frequency operation types, and then analyzes the operation consistency to obtain the behavior deviation index; The performance fluctuation detection module, based on the behavior deviation index, identifies the load fluctuation amplitude, isolates the time period of abnormal fluctuation, calculates the CPU utilization rate fluctuation within the time period, and compares it with the average load change trend to obtain the load fluctuation index; The event correlation analysis module, based on the load fluctuation index, analyzes error events, timeout events, and resource access conflict events in the computer log, counts the occurrence frequency and distribution pattern of the events within the abnormal time period, identifies the correlation degree between the events, and obtains the event correlation strength; The fault warning module, based on the event correlation strength, analyzes the occurrence frequency and influence range of the current abnormal events, calculates the cumulative influence degree, evaluates the consistency with the known fault trigger mode, determines the fault risk level, and obtains the fault warning index; The resource adjustment response module, based on the fault warning index, identifies the computer components with a high current risk level, allocates additional computing resources or restarts the service process, and obtains the computing resource adjustment information.

2. The computer fault alarm system according to claim 1, wherein The behavior deviation index includes operation time deviation, operation frequency deviation, and operation type deviation. The load fluctuation index includes CPU load fluctuation, memory usage fluctuation, and network response fluctuation. The event correlation strength includes error event correlation degree, timeout event correlation degree, and resource conflict event correlation degree. The fault warning index includes event occurrence frequency, event influence range, and event risk level. The computing resource adjustment information includes resource allocation effect, process restart effect, and performance adjustment result.

3. The computer failure alarm system according to claim 1, characterized in that, The behavior pattern analysis module includes: The log parsing sub-module obtains user operation log data, parses the operation time, operation type, and operation frequency, sorts the logs according to the time tags, analyzes the operation activities within each time period, and extracts the correlation distribution data between the time points and the operation types to obtain the operation time series data; The frequency calculation sub-module, based on the operation time series data, calculates the occurrence frequency of the operation type within the different time periods, statistically analyzes the frequency change trend, filters the high-frequency operation types associated with known faults, and calculates their relative proportion in all operations to obtain the high-frequency operation proportion index; The abnormal behavior identification sub-module, based on the high-frequency operation proportion index, combines the user's authentication data and historical operation patterns, compares the current operation with the historical data for consistency, and analyzes the deviation degree of the operation behavior to obtain the behavior deviation index.

4. The computer failure alarm system according to claim 1, wherein The performance fluctuation detection module includes: The system resource collection sub-module, based on the behavior deviation index, collects the computer CPU load, memory usage rate, and network response time, monitors the computer resource utilization within the different time windows, analyzes the time series change situation, and obtains an overview of the resource usage situation; Based on the overview of resource usage, the load fluctuation identification sub-module calculates the fluctuation amplitude of CPU load and memory usage rate within a time window, compares the resource occupancy trends within the time window, filters out the time periods with load fluctuation amplitude exceeding the standard, and obtains the abnormal time intervals. The load trend calculation sub-module calls the abnormal time intervals, analyzes the CPU utilization rate fluctuation within the time periods, compares the average load fluctuation of the time periods with the overall data, determines the long-term trend and short-term changes of the fluctuation, and obtains the load fluctuation index.

5. The computer fault alarm system according to claim 1, wherein The event correlation analysis module includes: Based on the load fluctuation index, the abnormal event statistics sub-module analyzes error events, timeout events, and resource access conflict events in the computer logs, counts the occurrence frequency of each type of event within the abnormal time periods, evaluates the event concentration, and identifies the event distribution pattern to obtain an overview of abnormal event distribution. Based on the overview of abnormal event distribution, the event correlation sub-module calculates the co-occurrence probability of different types of events within the abnormal time periods, analyzes the time interval and frequency between events, determines the dependency relationship between events, and obtains the event correlation index. The event propagation analysis sub-module calls the event correlation index, analyzes the order of event occurrence, calculates the propagation path length and impact between events, counts the cascade impact range of events, determines the key trigger events, and obtains the event correlation strength.

6. The computer fault alarm system according to claim 5, characterized in that, The time interval and frequency between events are analyzed using the formula: Analyze the time intervals and frequencies between events, determine the dependency relationships between events, and obtain the event correlation index, where PZ uo represents the co-occurrence probability of event u and event o, τ uk represents the start time of event u in the k-th abnormal time period, τ ok represents the start time of event o in the k-th abnormal time period, T k represents the total duration of the k-th abnormal time period, N PZ represents the total number of abnormal time periods.

7. The computer fault alarm system according to claim 1, wherein, The fault warning module includes: Based on the event correlation strength, the abnormal event screening sub-module analyzes the occurrence frequency of abnormal events, calculates the cumulative impact value of each type of event, counts the number of abnormal events exceeding the threshold, and identifies the abnormal events with key impact scope to obtain a list of key abnormal events. The fault mode matching sub-module calls the list of key abnormal events, calculates the matching degree between the events and the known fault trigger modes, analyzes the consistency between the order of event occurrence and the known fault modes, and identifies the key impact components for the matching sequences to obtain the fault mode matching degree. Based on the fault mode matching degree, the risk level assessment sub-module counts the number of computer components involved in the matching events, evaluates the fault risk level, and determines the computer fault risk level in combination with the fault impact scope to obtain the fault warning index.

8. The computer failure alarm system according to claim 7, wherein The number of computer components involved in the matching events is counted using the formula: Evaluate the failure risk level and determine the computer failure risk level, where N Z represents the total number of computer components involved, C i represents the absolute value of the set of computer components involved in the i-th failure event, |C i | represents taking the number of elements in the set of components involved in the i-th event, M represents the total number of matching events, f j represents the failure frequency of the j-th component, and n represents the total number of components within the evaluation period.

9. The computer fault alarm system according to claim 1, characterized in that, The resource adjustment response module includes: Based on the fault warning index, the high-risk component identification sub-module analyzes the risk level of computer components, counts the resource usage of high-risk components, evaluates the current load status, and determines the resource allocation objects that need to be adjusted to obtain the key risk components. The resource adaptation and adjustment sub-module calls the key risk components, analyzes the status of the computer's available resource pool, evaluates the allocation ability of computing resources, calculates the current resource utilization rate, filters out the computing resources available for adjustment, and replenishes resources or restarts processes for high-risk components to obtain the resource allocation details. The performance change recording sub-module monitors the running status of the adjusted computing resources according to the resource allocation details, records the change trends of CPU, memory, and network usage rates, and evaluates the impact of resource adjustment on computer stability to obtain computing resource adjustment information.

10. A computer fault alarm method, which is used to implement the computer fault alarm system according to any one of claims 1-9, characterized in that, It includes the following steps: S1: Obtain user operation log data, parse the operation time, operation type, and operation frequency, arrange the operation time series, calculate the operation frequency change rate, screen out the high-frequency operation types associated with known faults, and then analyze the operation consistency based on the user identity and historical operation behaviors to obtain the behavior deviation index. S2: Based on the behavior deviation index, collect the computer CPU load, memory usage rate, and network response time. For each time window data, identify the load fluctuation amplitude, calculate the CPU utilization rate fluctuation, and compare it with the average load change trend to obtain the load fluctuation index. S3: Based on the load fluctuation index, count the occurrence frequency and distribution pattern of events in the abnormal time period, identify the correlation degree between events, and analyze the occurrence order and propagation path of events to obtain the event correlation intensity. S4: Based on the event correlation intensity, calculate the cumulative impact degree of events, screen out the abnormal events exceeding the threshold, evaluate the consistency between the events and the known fault trigger patterns, and conduct correlation analysis on the events that conform to the fault patterns and computer components to obtain the fault warning index. S5: Based on the fault warning index, identify the computer components with high current risk levels, analyze the status of the available resource pool, make feasible resource adjustments, allocate additional computing resources or restart service processes to the high-risk components to obtain computing resource adjustment information.

Citation Information

Patent Citations

  • Voltage sag estimation model establishment method, medium and system

    CN117578481A

  • Remote control distributed energy station intelligent operation and maintenance management system and method

    CN118348878A

  • Operation state monitoring method and system based on computer management

    CN118349415A

  • Software security monitoring system and method based on Internet of Things

    CN118349989A

  • Automatic control system and method based on big data analysis

    CN119310918A

Cited By

  • MES (Manufacturing Execution System) for automatically adjusting server manufacturing process based on current network fault

    CN121365309A