A computer fault alarm system and method
By using behavioral pattern analysis, performance fluctuation detection, event correlation, and resource adjustment, the shortcomings of computer fault diagnosis systems in large-scale real-time data processing are addressed, enabling rapid and accurate fault detection and system recovery.
Patent Information
- Application Number
- CN202510380470.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-03-28
AI Technical Summary
Existing computer fault diagnosis systems are insufficient in processing large-scale real-time data, resulting in slow or inaccurate fault diagnosis, inability to accurately respond to event correlation and propagation path analysis, and increased system maintenance complexity and downtime.
The system utilizes a behavior pattern analysis module, a performance fluctuation detection module, an event correlation analysis module, a fault early warning module, and a resource adjustment response module to acquire and analyze user operation logs, computer resource utilization, and event correlation, thereby identifying anomalies and adjusting resources to improve the accuracy and timeliness of fault detection and early warning.
It enables accurate fault detection and rapid response in computer systems, improves the accuracy of fault warnings and system recovery capabilities, and reduces system downtime and operating costs.
Smart Images

Figure CN120315971B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular to a computer fault alarm system and method. BACKGROUND
[0002] The field of computer technology is a broad and diverse field of science and technology, including computer hardware, software, networks, and data processing, among other aspects. This field explores how to design, build, and optimize computer systems and their interacting components to improve processing speed, data storage capacity, and information security. Computer technology plays a central role in daily life, supporting business operations and driving the development of emerging technologies such as artificial intelligence, big data analysis, and cloud computing.
[0003] Among them, the computer fault alarm system is a system for detecting potential errors in computer systems. The main purpose of the system is to monitor the running state of computer hardware and software in real time. Once an anomaly or fault is detected, the system administrator or other relevant personnel can be notified immediately, which can help reduce the duration of system errors and optimize maintenance response, thereby ensuring the efficient and stable operation of information technology infrastructure.
[0004] The existing technology fails to provide sufficient data analysis depth, resulting in slow or inaccurate fault diagnosis and response. The main defect lies in the system's inability to handle large-scale real-time data, particularly in terms of performance monitoring and anomaly detection effectiveness. Additionally, traditional technology cannot accurately respond to event correlation and propagation path analysis, affecting the speed and efficiency of fault handling. The lack of precise user behavior and performance fluctuation analysis makes it difficult for system administrators to accurately identify fault sources, increasing the complexity of system maintenance and leading to longer system downtime or higher operating costs. SUMMARY
[0005] To solve the technical problems of the existing technology that fails to provide sufficient data analysis depth, resulting in slow or inaccurate fault diagnosis and response, the main defect lies in the system's inability to handle large-scale real-time data, particularly in terms of performance monitoring and anomaly detection effectiveness. Additionally, traditional technology cannot accurately respond to event correlation and propagation path analysis, affecting the speed and efficiency of fault handling. The lack of precise user behavior and performance fluctuation analysis makes it difficult for system administrators to accurately identify fault sources, increasing the complexity of system maintenance and leading to longer system downtime or higher operating costs, the present application provides a computer fault alarm system and method. The technical solution is as follows:
[0006] On the one hand, a computer fault alarm system is provided, which comprises:
[0007] The behavior pattern analysis module obtains user operation log data, analyzes operation time, operation type and operation frequency, screens high-frequency operation types, analyzes operation consistency again, and obtains a behavior deviation index;
[0008] The performance fluctuation detection module identifies a load fluctuation amplitude based on the behavior deviation index, isolates a time period of abnormal fluctuation, calculates CPU utilization fluctuation in the time period, compares the CPU utilization fluctuation with an average load change trend, and obtains a load fluctuation index;
[0009] The event correlation analysis module analyzes error events, timeout events and resource access conflict events in computer logs based on the load fluctuation index, counts occurrence frequency and distribution mode of the events in the abnormal time period, identifies correlation degree between the events, and obtains an event correlation strength;
[0010] The fault early warning module analyzes occurrence frequency and influence range of a current abnormal event based on the event correlation strength, calculates cumulative influence degree, evaluates consistency with a known fault triggering mode, determines a fault risk level, and obtains a fault early warning index;
[0011] The resource adjustment response module identifies computer components with a high current risk level based on the fault early warning index, allocates additional computing resources or restarts service processes, and obtains computing resource adjustment information.
[0012] On the other hand, the behavior deviation index includes operation time deviation, operation frequency deviation and operation type deviation, the load fluctuation index includes CPU load fluctuation, memory usage fluctuation and network response fluctuation, the event correlation strength includes error event correlation degree, timeout event correlation degree and resource conflict event correlation degree, the fault early warning index includes event occurrence frequency, event influence range and event risk level, and the computing resource adjustment information includes resource allocation effect, process restart effect and performance adjustment result.
[0013] On the other hand, the behavior pattern analysis module includes:
[0014] The log analysis submodule obtains user operation log data, analyzes operation time, operation type and operation frequency, sorts the logs according to time tags, analyzes operation activities in each time period, and extracts associated distribution data of time points and operation types, and obtains operation time sequence data;
[0015] The frequency calculation submodule calculates occurrence frequency of operation types in different time periods based on the operation time sequence data, statistically analyzes frequency change trends, screens high-frequency operation types associated with known faults, calculates relative proportions of the high-frequency operation types in all operations, and obtains a high-frequency operation proportion index;
[0016] The abnormal behavior recognition submodule is based on the high-frequency operation proportion index, combines the user's identity verification data and historical operation mode, compares the consistency of the current operation and the historical data, analyzes the deviation degree of the operation behavior, and obtains a behavior deviation index.
[0017] In another aspect, the performance fluctuation detection module includes:
[0018] The system resource collection submodule is based on the behavior deviation index, collects computer CPU load, memory usage and network response time, monitors the computer resource utilization in the difference time window, analyzes the time series change, and obtains a resource usage overview;
[0019] The load fluctuation recognition submodule is based on the resource usage overview, calculates the fluctuation amplitude of CPU load and memory usage in the time window, compares the resource occupation trend in the time window, filters the time period with load fluctuation amplitude exceeding the standard, and obtains an abnormal time interval;
[0020] The load trend calculation submodule calls the abnormal time interval, analyzes the CPU utilization fluctuation in the time period, compares the average load fluctuation of the time period and the overall data, determines the long-term trend and short-term change of the fluctuation, and obtains a load fluctuation index.
[0021] In another aspect, the event correlation analysis module includes:
[0022] The abnormal event statistics submodule is based on the load fluctuation index, analyzes error events, timeout events and resource access conflict events in the computer log, counts the occurrence frequency of each type of event in the abnormal time period, evaluates the event concentration, and identifies the event distribution pattern, and obtains an abnormal event distribution overview;
[0023] The event correlation submodule is based on the abnormal event distribution overview, calculates the common occurrence probability of difference type events in the abnormal time period, analyzes the time interval and frequency between events, determines the dependency relationship between events, and obtains an event correlation degree index;
[0024] The event propagation analysis submodule calls the event correlation degree index, analyzes the event occurrence order, calculates the propagation path length and influence between events, counts the cascading influence range of events, determines the key trigger event, and obtains the event correlation strength.
[0025] In another aspect, the time interval and frequency between events are analyzed by using the formula:
[0026]
[0027] The time interval and frequency between events are analyzed, the dependency relationship between events is determined, and an event correlation degree index is obtained, wherein PZ uoThe probability of the co-occurrence of events u and o, τ uk The start time of event u at the kth abnormal time period, τ ok The start time of event o at the kth abnormal time period, T k The total duration of the kth abnormal time period, N PZ Represent the total number of abnormal time periods.
[0028] In another aspect, the fault warning module comprises:
[0029] The abnormal event screening submodule analyzes the frequency of occurrence of abnormal events based on the event correlation strength, calculates the cumulative impact value of each type of event, counts the number of abnormal events exceeding the threshold, and identifies abnormal events that are critical to the impact range to obtain a list of critical abnormal events;
[0030] The fault mode matching submodule calls the list of critical abnormal events, calculates the matching degree of events and known fault trigger patterns, analyzes the consistency of event occurrence order and known fault patterns, and identifies key impact components for matching sequences to obtain a fault mode matching degree;
[0031] The risk level assessment submodule counts the number of computer components involved in matching events according to the fault mode matching degree, assesses the fault risk level, and determines the computer fault risk level in combination with the fault impact range to obtain a fault warning indicator.
[0032] In another aspect, the number of computer components involved in matching events is calculated using the formula:
[0033]
[0034] Assess the fault risk level and determine the computer fault risk level, where N Z represents the total number of computer components involved, C i represents the absolute value of the set of computer components involved in the ith fault event, |C i | represents the number of elements in the set of components involved in the ith event, M represents the total number of matching events, f j represents the failure frequency of the jth component, and n represents the total number of components in the evaluation period.
[0035] In another aspect, the resource adjustment response module comprises:
[0036] The high-risk component identification submodule analyzes the risk level of computer components based on the fault warning indicator, counts the resource usage of high-risk components, assesses the current load state, and determines the resource allocation object that needs to be adjusted to obtain key risk components.
[0037] The resource adaptation adjustment submodule calls the key risk component, analyzes the state of the computer available resource pool, evaluates the allocation ability of the computing resource, calculates the current resource utilization rate, screens the computing resource available for adjustment, and performs resource supplement or process restart on the high-risk component, to obtain resource allocation details;
[0038] The performance change recording submodule monitors the running state of the adjusted computing resource according to the resource allocation details, records the change trend of the CPU, memory and network usage rate, and evaluates the influence of resource adjustment on the stability of the computer, to obtain computing resource adjustment information.
[0039] In another aspect, a computer fault alarm method is provided, which is applied to a computer fault alarm system and includes the following steps:
[0040] S1: Obtain user operation log data, parse operation time, operation type and operation frequency, arrange the operation time sequence, calculate the operation frequency change rate, screen the high-frequency operation type associated with the known fault, and then analyze the operation consistency according to the user identity and historical operation behavior to obtain a behavior deviation index;
[0041] S2: Based on the behavior deviation index, collect the computer CPU load, memory usage rate and network response time, identify the load fluctuation amplitude for each time window data, calculate the CPU utilization rate fluctuation, compare with the average load change trend, and obtain a load fluctuation index;
[0042] S3: Based on the load fluctuation index, count the occurrence frequency and distribution mode of events in the abnormal time period, identify the correlation degree between events, analyze the occurrence sequence and propagation path of events, and obtain an event correlation strength;
[0043] S4: Based on the event correlation strength, calculate the cumulative influence degree of the event, screen the abnormal events exceeding the threshold, evaluate the consistency of the event and the known fault triggering mode, and perform correlation analysis on the events and the computer components that meet the fault mode, to obtain a fault early warning index;
[0044] S5: Based on the fault early warning index, identify the computer components with high current risk level, analyze the available resource pool state, perform resource feasibility adjustment, allocate additional computing resources or restart service processes to the high-risk components, and obtain computing resource adjustment information.
[0045] The technical scheme provided by the embodiment of the application has at least the following beneficial effects:
[0046] By comparing the high-frequency operation type with historical fault data, an accurate behavior deviation index is provided, which is used to optimize the monitoring process of performance fluctuation, especially in isolating and identifying system performance anomalies, thereby improving the fault detection capability of the system, and by analyzing the system load fluctuation and comparing it with the conventional load change trend, the response speed and accuracy of the early warning are enhanced, and by comprehensively utilizing these data, the event correlation analysis is more accurate, the event propagation path and correlation degree can be effectively mapped, the timeliness and accuracy of fault early warning are improved, and by targeted resource adjustment, the recovery capability and fault tolerance of the computer system are enhanced. BRIEF DESCRIPTION OF DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0048] Figure 1 The system schematic diagram of the present application;
[0049] Figure 2 The system framework schematic diagram of the present application;
[0050] Figure 3 The flowchart of the behavior mode analysis module of the present application;
[0051] Figure 4 The flowchart of the performance fluctuation detection module of the present application;
[0052] Figure 5 The flowchart of the event correlation analysis module of the present application;
[0053] Figure 6 The flowchart of the fault early warning module of the present application;
[0054] Figure 7 The flowchart of the resource adjustment response module of the present application;
[0055] Figure 8 The method step flowchart of the present application. DETAILED DESCRIPTION
[0056] The technical solutions in the present application will be described in detail below with reference to the drawings.
[0057] In the embodiments of the present application, the words such as "example", "for example" are used to represent an example, illustration, or description. Any embodiment or design scheme described as "example" in the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the word "example" is intended to present the concept in a specific manner. In addition, in the embodiments of the present application, the meaning expressed by "and / or" can be both, or can be one of the two.
[0058] In the embodiments of the present application, "image" and "picture" can be used interchangeably at times, and it should be pointed out that the meanings expressed are consistent when the distinction is not emphasized. "Of", "corresponding" and "corresponding" can be used interchangeably at times, and it should be pointed out that the meanings expressed are consistent when the distinction is not emphasized.
[0059] In the embodiments of the present application, sometimes the subscript such as W1 can be written in the form of non-subscript such as W1, and the meanings expressed are consistent when the distinction is not emphasized.
[0060] In order to make the technical problems, technical schemes and advantages to be solved by the present application more clear, the following will be described in detail in conjunction with the drawings and specific embodiments.
[0061] The embodiments of the present application provide a computer fault alarm system, as shown in Figure 1 The system comprises:
[0062] The behavior pattern analysis module obtains user operation log data, parses operation time, operation type and operation frequency, arranges operation time sequence, calculates operation frequency change rate, and screens high-frequency operation types associated with known faults, calculates frequency proportion, and then analyzes operation consistency according to user identity and historical operation behavior to obtain behavior deviation index;
[0063] The performance fluctuation detection module collects computer CPU load, memory usage and network response time based on the behavior deviation index, identifies load fluctuation amplitude for each time window data, isolates abnormal fluctuation time period, and calculates CPU utilization fluctuation in the time period, compares with average load change trend to obtain load fluctuation index;
[0064] The event correlation analysis module analyzes error events, timeout events and resource access conflict events in the computer log based on the load fluctuation index, counts the occurrence frequency and distribution mode of events in the abnormal time period, identifies the correlation degree between events, and analyzes the occurrence sequence and propagation path of events to obtain event correlation strength;
[0065] The fault early warning module analyzes the occurrence frequency and influence range of the current abnormal event based on the event correlation strength, calculates the cumulative influence degree of the event, screens abnormal events exceeding the threshold, evaluates the consistency of the event and the known fault trigger mode, and performs correlation analysis on the event and the computer components that meet the fault mode to determine the fault risk level of the computer and obtain a fault early warning index;
[0066] The resource adjustment response module identifies computer components with high current risk levels based on the fault early warning index, analyzes the state of the available resource pool, performs resource feasibility adjustment, allocates additional computing resources or restarts service processes to high-risk components, and records the performance changes after adjustment to obtain computing resource adjustment information.
[0067] The behavior deviation index includes operation time deviation, operation frequency deviation, and operation type deviation, the load fluctuation index includes CPU load fluctuation, memory usage fluctuation, and network response fluctuation, the event correlation strength includes error event correlation degree, timeout event correlation degree, and resource conflict event correlation degree, the fault early warning index includes event occurrence frequency, event influence range, and event risk level, and the computing resource adjustment information includes resource allocation effect, process restart effect, and performance adjustment result.
[0068] As shown in Figure 2 and Figure 3 , the behavior mode analysis module includes:
[0069] The log analysis submodule obtains user operation log data, analyzes operation time, operation type, and operation frequency, sorts the logs according to time tags, analyzes operation activities in each time period, and extracts correlation distribution data of time points and operation types to obtain operation time sequence data.
[0070] The access log storage system is read in chronological order to read the user's operation log records from the system, including the operation time, operation type and operation frequency in the log, and the data is parsed into a structured format for storage. Each log data needs to parse the user ID, operation instruction, and operation timestamp, and is arranged in ascending order of timestamp. In the parsing process, the log entry is read, the timestamp field is extracted and converted into a standard time format (such as YYYY-MM-DDHH: MM: SS) and stored as a time label in the time series data. Then the operation type field is parsed, the specific instructions executed by the user are extracted according to the different formats of the operation log, and the operation type identifier is labeled. For example, the operation type can be divided into "login", "modify configuration", "access file", "execute command", etc. When parsing the operation frequency, the number of times the same operation is performed by the same user within a unit time (such as every 5 minutes, every hour) needs to be counted, and operation frequency data is generated. The parsed data is stored in the database and sorted according to the time label in ascending order based on the timestamp, so that the operation logs of the same user can be analyzed in chronological order. On this basis, the operation activities in each time period are analyzed, for example, a day is divided into 24 hours, each hour is taken as a time period, the operation behavior of all users in the time period is counted, and operation distribution data in the time period is generated. The correlation data between each time point and the operation type is extracted, for example, a user accesses the database 10 times between 12:00-13:00, accesses the file 5 times, and modifies the configuration 2 times. In this way, the operation characteristics of different time periods can be obtained and stored as operation time series data.
[0071] The frequency calculation submodule calculates the occurrence frequency of the operation type in the different time periods based on the operation time series data, analyzes the frequency change trend, selects high-frequency operation types associated with known faults, calculates their relative proportion in all operations, and obtains a high-frequency operation proportion index.
[0072] First, read the operation time series data, classified statistics by operation type, calculate the number of occurrences of the operation type in each time period, for example, a user performs "access database" operation 20 times in 10:00-11:00, 30 times in 11:00-12:00, then the frequency of the operation type in the two time periods is 20 and 30 respectively, then calculate the relative proportion of each operation type, that is, the number of times of a certain operation type in the time period accounts for the proportion of the total operation times of the user, for example, the user performs 50 operations in 10:00-11:00, among which "access database" accounts for 40%(20 / 50), "access file" accounts for 30%(15 / 50), then calculate the proportion index of high-frequency operation type, when screening high-frequency operation, set a benchmark frequency threshold Fth, for example, set Fth=25 times, if the frequency of a certain operation type in a certain time period is greater than or equal to 25 times, it is marked as a high-frequency operation type, and the frequency of all high-frequency operation types is normalized, that is, the frequency of each high-frequency operation type is divided by the total operation times of the user in the time period, to get the normalized proportion index, for example, a user performs 80 operations in 11:00-12:00, among which "access database" is 50 times, the proportion index is 50 / 80=0.625, which is used for subsequent abnormal behavior identification analysis.
[0073] The abnormal behavior recognition submodule compares the current operation with the historical data based on the high-frequency operation proportion index, combines the user's identity verification data and historical operation mode, analyzes the deviation degree of operation behavior, and obtains the behavior deviation index.
[0074] Read user authentication data, including user identity, login device, IP address, etc., and combine user historical operation mode to build user normal behavior baseline, specifically, extract user's high-frequency operation proportion index in each time period in the past certain time (such as the last 7 days), and calculate its mean μ and standard deviation σ, for example, the mean of the proportion index of a certain operation type of a certain user in the past 7 days is 0.45, and the standard deviation is 0.05, then the normal operation range of the user can be set as [μ-2σ, μ+2σ] = [0.35, 0.55], then the consistency of the current operation is compared, whether the high-frequency proportion index of the current operation falls within the interval, for example, the high-frequency proportion index of the current operation is 0.6, which is beyond the normal range, then calculate the behavior deviation index, which is defined as the deviation degree D = |(current proportion index-mean) / standard deviation|, in the above example, D = |(0.6-0.45) / 0.05| = 3, if D exceeds the set threshold (such as D>2), mark the operation as abnormal, and further analyze the user's identity authentication data, such as checking whether the login device has changed, whether the IP address is an address that has not appeared in the historical record, and finally according to the characteristics, determine whether the user has abnormal behavior.
[0075] As shown in Figure 2 and Figure 4 , the performance fluctuation detection module includes:
[0076] The system resource collection submodule collects computer CPU load, memory usage and network response time based on the behavior deviation index, monitors the computer resource utilization in the difference time window, analyzes the time series change, and obtains the resource usage overview;
[0077] The resource monitoring interface of the operating system (such as / proc / stat of Linux, WMI query of Windows) is called to obtain the CPU occupancy rate, the total CPU running time (including user mode time, system mode time and idle time) is accumulated, and the CPU usage ratio in a unit time is calculated, for example, the unit time is 5 seconds, the initial CPU total time T1 is read, the new CPU total time T2 is read after 5 seconds, then the CPU load calculation formula is (CPU usage time change amount) / (time change amount) = (T2-T1) / 5, the memory usage rate is obtained by calling the system interface to obtain the ratio of the current available memory to the total memory, for example, the total memory of a computer is 16GB, the current available memory is 4GB, then the memory usage rate calculation formula is (total memory-available memory) / total memory = (16-4) / 16 = 75%, the network response time is measured by sending test data packets to the specified server, and measuring the return time difference, for example, 10 data packets are sent, and the returned times are 50ms, 55ms, 48ms, 53ms, etc., then the network response time is calculated as (50+55+48+53+...) / 10, the system resource collection submodule needs to monitor the above resource conditions in different time windows, such as 5 minutes, 1 hour, 1 day, etc., and store the corresponding data, then analyze the time series change, extract the resource utilization in each time period, for example, the CPU average load is 60% in the 10:00-11:00 time period, the memory usage rate is 70%, and the network response time is 52ms, and the resource usage time series is constructed according to the data to form a resource usage overview.
[0078] The load fluctuation identification submodule calculates the fluctuation amplitude of the CPU load and the memory usage rate in the time window based on the resource usage overview, compares the resource occupation trend in the time window, filters the time period with a load fluctuation amplitude exceeding the standard, and obtains an abnormal time interval;
[0079] Extract resource usage data within a set time window, for example, set the time window to 10 minutes, sample the CPU load 6 times within the window, get the load data sequence (55%, 60%, 58%, 62%, 59%, 61%), calculate the load fluctuation amplitude, that is, the sum of the absolute values of the difference between the current time CPU usage rate and the previous time, for example, calculate (|60-55|+|58-60|+|62-58|+|59-62|+|61-59|) / (sampling times-1)=(5+2+4+3+2) / 5=3.2%, similarly, calculate the fluctuation amplitude of the memory usage rate, if the memory usage rate in a certain time window is (65%, 70%, 68%, 73%, 69%, 72%), then calculate the fluctuation amplitude(|70-65|+|68-70|+|73-68|+|69-73|+|72-69|) / 5=(5+2+5+4+3) / 5=3.8%, then compare the resource occupation trend within the time window, analyze the average load change rate of the time window, the calculation method is the difference between the resource usage mean value in the window and the resource usage mean value in the previous window, for example, the average CPU usage rate in the previous window is 55%, the average CPU usage rate in the current window is 60%, the change rate is (60-55) / 55=9.1%, set the fluctuation standard threshold Th, if Th=7%, then the CPU load fluctuation of the current window exceeds the standard range, select all time periods that exceed the threshold, and record the time stamp, get the abnormal time interval.
[0080] The load trend calculation submodule calls the abnormal time interval, analyzes the CPU utilization fluctuation in the time period, compares the average load fluctuation of the time period and the overall data, determines the long-term trend and short-term change of the fluctuation, and gets the load fluctuation index;
[0081] The CPU utilization fluctuation in the time period is analyzed, the CPU utilization data in the abnormal time interval is extracted, the mean and standard deviation are calculated, for example, in the abnormal time period 10:00-10:30, the CPU utilization data is (65%, 70%, 72%, 68%, 75%, 73%), the mean μ=(65+70+72+68+75+73) / 6=70.5% is calculated, the standard deviation σ=sqrt(((65-70.5)^2+(70-70.5)^2+(72-70.5)^2+(68-70.5)^2+(75-70.5)^2+(73-70.5)^2) / 6)≈3.5%, then the load fluctuation of the time period is compared with the average load fluctuation of the overall data, assuming that the CPU mean of the overall data is 60% and the standard deviation is 2.5%, the deviation degree D of the abnormal time interval is calculated as |(70.5-60) / 2.5|=4.2, if D exceeds the set threshold (such as D>3), it is determined that the time period belongs to long-term trend fluctuation, if 2≤D≤3, it is determined as short-term change, the load fluctuation index is calculated, the index LBI=D×σ is defined, for example, LBI=4.2×3.5=14.7, and the load fluctuation index of the time period is output.
[0082] As shown in Figure 2 and Figure 5 The event correlation analysis module includes:
[0083] The abnormal event statistics submodule analyzes the error events, timeout events and resource access conflict events in the computer log based on the load fluctuation index, counts the occurrence frequency of each type of event in the abnormal time period, evaluates the event concentration, and identifies the event distribution pattern to obtain an abnormal event distribution overview;
[0084] Read relevant data from log files or databases, each log contains timestamp, event type, error code, user ID and impact range, etc. According to the timestamp, filter out the log records in the time period when the load fluctuation index exceeds the benchmark range, for example, the load fluctuation index of a certain system is 15.2 from 10:00 to 11:00, which is higher than the set benchmark value 10, then extract the logs in this time period, count the frequency of different types of events, for example, in this time period, error events occur 10 times, timeout events occur 8 times, and resource access conflict events occur 5 times, then calculate the event concentration, define the concentration index C=(single event highest frequency) / (total number of all events in this time period), for example, C=10 / (10+8+5)=0.43, if C is higher than the set threshold value (such as 0.4), it means that a certain type of event accounts for a high proportion in this time period, then analyze the event distribution pattern, count the time distribution of each event, for example, error events occur at 10:05, 10:15, 10:30 and 10:50, timeout events occur at 10:07, 10:20 and 10:45, and resource access conflict occurs at 10:25 and 10:40, thus obtaining the abnormal event distribution overview.
[0085] The event correlation submodule calculates the common occurrence probability of difference type events in the abnormal time period based on the abnormal event distribution overview, analyzes the time interval and frequency between events, determines the dependency relationship between events, and obtains the event correlation index;
[0086] The time interval and frequency between events are analyzed, and the formula is:
[0087]
[0088] The time interval and frequency between events are analyzed, and the formula is: uo The common occurrence probability of event u and event o, τ uk The start time of event u in the kth abnormal time period, τ ok The start time of event o in the kth abnormal time period, T k The total duration of the kth abnormal time period, N PZ The total number of abnormal time periods;
[0089] The specific values obtained by data monitoring are as follows:
[0090] The start times of event u and event o in the three time periods are τ u1 =2 hours, τ o1 =3 hours and τ u2 =5 hours, τ o2 =7 hours, τ u3 =8 hours and τo3 = 9 hours;
[0091] The total duration of each time period is T1 = 2 hours, T2 = 3 hours and T3 = 1.5 hours;
[0092] N PZ = 3, the number of time periods determined by the event log.
[0093] Calculate the absolute value of the difference between the start time of event u and event o in each time period:
[0094] |τ u1 -τ o1 | = 1;
[0095] |τ u2 -τ o2 | = 2;
[0096] |τ u3 -τ o3 | = 1;
[0097] Sum the absolute values of each time difference:
[0098] 1 + 2 + 1 = 4;
[0099] Sum the total duration of each time period:
[0100] 2 + 3 + 1.5 = 6.5;
[0101] Substitute the above results into the formula:
[0102]
[0103] The result shows that the probability of the co-occurrence of event u and event o in the three monitored time periods is about 41%, which reflects the relative frequency of the occurrence of the two events simultaneously in a given time period, and is an important indicator for determining the dependence relationship between the two events.
[0104] The event propagation analysis submodule calls the event correlation index, analyzes the event occurrence order, calculates the propagation path length and influence between events, counts the cascade influence range of events, determines the key trigger event, and obtains the event correlation strength;
[0105] The timestamp data is extracted, the event sequence is constructed, for example, the error event occurs at 10:05, 10:15, 10:30, and 10:50, the timeout event occurs at 10:07, 10:20, and 10:45, and the resource access conflict occurs at 10:25 and 10:40, and then the event propagation path is established: error event (10:05) -> timeout event (10:07) -> error event (10:15) -> timeout event (10:20) -> resource access conflict (10:25) -> error event (10:30) -> resource access conflict (10:40) -> timeout event (10:45) -> error event (10:50), the event propagation path length L, i.e. the average propagation interval between events, is calculated, such as L = (2+8+5+5+5+10+5+5) / 8 = 5.625 minutes, then the cascading influence range of the event is calculated, the influence range R = event propagation path length x event quantity is defined, such as R = 5.625 x 9 = 50.6 minutes, finally the key triggering event is screened, whether an event will trigger multiple subsequent events is judged, such as the error event (10:05) triggers the timeout event (10:07) and the error event (10:15), and the error event (10:05) is the key triggering event, and finally the event correlation strength is calculated, S = E x R is defined, such as S = 0.083 x 50.6 = 4.2, and the event correlation strength is obtained.
[0106] As shown in Figure 2 and Figure 6 , the fault early warning module comprises:
[0107] The abnormal event screening submodule analyzes the occurrence frequency of abnormal events based on the event correlation strength, calculates the cumulative influence value of each type of event, counts the number of abnormal events exceeding the threshold value, and identifies the abnormal events critical to the influence range to obtain a list of key abnormal events;
[0108] All records are traversed, the occurrence frequency of each type of event in the abnormal time period is counted, for example, in the 10:00-11:00 time period, the error event occurs 10 times, the timeout event occurs 8 times, and the resource access conflict event occurs 6 times, the cumulative frequency of each type of event is recorded, then the cumulative influence value of each type of event is calculated, the influence value I = event correlation strength S x event occurrence number N is defined, for example, the correlation strength S = 4.2 of a certain type of error event, the occurrence number N = 10, and then the cumulative influence value I = 4.2 x 10 = 42 of the event type, then the number of abnormal events with a cumulative influence value exceeding a threshold value is counted, the influence threshold Th = 30 is set, the abnormal events with I > Th are screened, for example, the influence value I = 42 exceeds the threshold value, and it is determined that the event type is a key abnormal event, and finally the abnormal events critical to the influence range are identified, whether an event is associated with multiple sub-events is judged, for example, the error event triggers multiple timeout events after occurring, and the error event belongs to the abnormal events critical to the influence range, and a list of key abnormal events is formed.
[0109] The fault mode matching sub-module calls the key exception event list, calculates the matching degree of the event and the known fault trigger mode, analyzes the consistency of the event occurrence sequence and the known fault mode, and identifies the key affected components for the matching sequence to obtain the fault mode matching degree;
[0110] The matching degree of each exception event and the known fault trigger mode is calculated, historical fault data is extracted, the event sequence when each type of fault occurs is obtained, for example, the event sequence of the known database fault mode is: resource access conflict→timeout event→error event, the same type of event sequence is extracted from the current key exception event list, and the matching ratio R=(number of matching events) / (total number of known fault events) is calculated. Assuming that resource access conflict and timeout event are found in the current exception event, but the error event is missing, then the matching ratio R=2 / 3=0.67. Then the consistency of the exception event occurrence sequence and the known fault mode is analyzed, and the square sum of the sequence deviation D=(actual sequence position-theoretical sequence position) / matching event number is calculated. For example, the theoretical sequence is A→B→C, and the actual sequence is A→C→B, then D=((1-1)^2+(3-2)^2+(2-3)^2) / 3=0.67. If D is less than a set threshold (such as 0.5), it is determined that the abnormal sequence matches the known fault mode. Finally, the key affected components in the matching sequence are identified, and all computer components involved in the abnormal events are extracted, such as resource access conflict involving database, timeout event involving network, and error event involving application server. Then all the involved component information is recorded, the fault mode matching degree is calculated, and the matching degree M=R×(1-D) is defined, such as M=0.67×(1-0.67)=0.22, and the fault mode matching degree is obtained.
[0111] The risk level evaluation sub-module evaluates the fault risk level according to the fault mode matching degree, counts the number of computer components involved in the matching events, determines the computer fault risk level in combination with the fault influence range, and obtains the fault warning index;
[0112] The number of computer components involved in the matching events is counted, and the formula is:
[0113]
[0114] The fault risk level is evaluated, and the computer fault risk level is determined, wherein, N Z represents the total number of computer components involved, C i represents the absolute value of the computer component set involved in the i-th fault event, |C i | represents the number of elements of the component set involved in the i-th event, M represents the total number of matching events, and f jThe failure frequency of the jth component is represented by f, and n represents the total number of components within the evaluation period;
[0115] The following data is provided: the total number of matching events M = 3, the number of computer components involved in each event |C1| = 10, |C2| = 15, and |C3| = 5, and the failure frequencies f j The failure frequencies f are obtained from monitoring data, for example, f1 = 0.02, f2 = 0.05, and f3 = 0.03, and the total number of components n = 3 (equal to the number of failure frequency terms) within the evaluation period. The monitoring and collection of data are performed through a dedicated computer hardware monitoring system to ensure the accuracy and timely updating of the data. The average failure frequency is calculated by substituting the formula:
[0116]
[0117] The square root of the average failure frequency is calculated:
[0118]
[0119] N Z is calculated:
[0120]
[0121] The result shows that the total number of computer components involved in the calculation is close to 5.5, which reflects the degree of impact of the components on the stability of the system under the current failure frequency. This value is used to quantify and evaluate the potential failure risk of the entire computer system.
[0122] As shown in Figure 2 and Figure 7 , the resource adjustment response module includes:
[0123] The high-risk component identification submodule analyzes the risk level of the computer components based on the failure warning indicators, counts the resource usage of the high-risk components, evaluates the current load state, determines the resource allocation objects that need to be adjusted, and obtains the key risk components.
[0124] Extract the risk level data of the computer components, traverse all components, calculate the risk score R of each component, define R = warning index W x component influence coefficient C, for example, W = 27.6 and C = 0.44 for a certain database server, then R = 27.6 x 0.44 = 12.14, then statistics of the resource usage of high-risk components, extract CPU, memory and network resource data, calculate the resource occupancy rate at the current time, for example, the CPU usage rate of a certain server is 85%, the memory usage rate is 90%, and the network bandwidth usage rate is 75%, then evaluate the current load state, define load index L = (CPU usage rate + memory usage rate + network bandwidth usage rate) / 3, such as L = (85 + 90 + 75) / 3 = 83.3%, set the load threshold Th = 80%, if L > Th, then determine that the component is in a high load state, then determine the resource allocation object that needs to be adjusted, filter all high load components, such as components with CPU usage rate > 80% and memory usage rate > 85% are determined as adjustment targets, and key risk components are obtained.
[0125] The resource adaptation adjustment submodule calls the key risk components, analyzes the available resource pool state of the computer, evaluates the allocation ability of the computing resources, calculates the current resource utilization rate, filters the adjustable computing resources, and performs resource replenishment or process restart on the high-risk components to obtain resource allocation details.
[0126] Get all underutilized computing resources, calculate their available capacity, for example, the available CPU core number in the current computing resource pool is 10, the idle memory capacity is 32GB, and the remaining bandwidth is 100Mbps, then evaluate the allocation ability of the computing resources, calculate the maximum scalable resource M = min(available CPU core number / target component CPU usage rate, idle memory capacity / target component memory usage rate, remaining bandwidth / target component network usage rate), such as M = min(10 / 0.85, 32 / 0.90, 100 / 0.75) = min(11.76, 35.56, 133.33) = 11.76, then calculate the current resource utilization rate, define U = allocated resources / total available resources, for example, U = (allocated CPU core number + allocated memory + allocated bandwidth) / (total CPU core number + total memory + total bandwidth), set the adjustment threshold Th = 70%, if U < Th, then determine that the adjustable resources are sufficient, then filter the adjustable computing resources, traverse all resource pools, select servers with CPU occupancy rate below 50% for resource replenishment, such as selecting a computing node N1 with CPU occupancy rate of 45% to allocate part of the CPU core to the high-risk components, and performing resource replenishment or process restart, such as stopping resource replenishment when the CPU occupancy rate of the high-risk components is reduced to below 80%, and obtaining resource allocation details.
[0127] The performance change recording sub-module monitors the running state of the adjusted computing resources according to the resource allocation details, records the change trend of the CPU, memory and network usage, and evaluates the impact of resource adjustment on the stability of the computer to obtain computing resource adjustment information.
[0128] The running state of the adjusted computing resources is monitored, the CPU, memory and network usage are extracted, the change values before and after allocation are recorded, for example, before resource adjustment, the CPU usage of the target server is 85%, after adjustment, it is reduced to 70%, the memory usage is reduced from 90% to 75%, and the network bandwidth usage is reduced from 75% to 65%, then the change rate before and after the adjustment of the computing resources is calculated, the change rate ΔX=(adjusted value-adjusted value before adjustment) / adjusted value before adjustment×100%, such as CPU change rate ΔC=(70-85) / 85×100%=-17.65%, memory change rate ΔM=(75-90) / 90×100%=-16.67%, network change rate ΔN=(65-75) / 75×100%=-13.33%, then the impact of resource adjustment on the stability of the computer is evaluated, the stability index S=1-(|ΔC|+|ΔM|+|ΔN|) / 3, such as S=1-(17.65+16.67+13.33) / 3=1-15.88=0.84, all the calculated data is stored in the log, and the computing resource adjustment information is obtained.
[0129] As shown in Figure 8 , a computer fault alarm method comprises the following steps:
[0130] S1: Obtain user operation log data, parse operation time, operation type and operation frequency, arrange the operation time sequence, calculate the operation frequency change rate, and filter high-frequency operation types associated with known faults, then analyze the operation consistency according to the user identity and historical operation behavior to obtain a behavior deviation index;
[0131] S2: Based on the behavior deviation index, collect computer CPU load, memory usage and network response time, identify the load fluctuation amplitude for each time window data, calculate the CPU utilization fluctuation, compare it with the average load change trend, and obtain a load fluctuation index;
[0132] S3: Based on the load fluctuation index, count the occurrence frequency and distribution pattern of events in the abnormal time period, identify the correlation degree between events, analyze the occurrence sequence and propagation path of events, and obtain an event correlation strength;
[0133] S4: Based on the event correlation strength, calculate the cumulative influence degree of the event, filter abnormal events exceeding the threshold, evaluate the consistency of the event with the known fault triggering mode, and perform correlation analysis on the events and computer components that meet the fault mode to obtain a fault warning index;
[0134] S5: Based on the fault early warning index, identify the computer component with high current risk level, analyze the available resource pool state, make resource feasibility adjustment, allocate additional computing resources or restart service process to the high-risk component, and obtain computing resource adjustment information.
[0135] It should be understood that the term "and / or" herein only describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after it, but it can also represent an "and / or" relationship, which can be understood in the context before and after it.
[0136] In the present application, "at least one" means one or more, and "multiple" means two or more. "At least one of the following" or the like means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b, or c can represent a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple.
[0137] It should be understood that in various embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution, and the execution order of the processes should be determined by its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0138] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different systems to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0139] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working process of the above-mentioned devices, apparatuses and units can refer to the corresponding process in the foregoing system embodiments, which will not be repeated here.
[0140] In several embodiments provided by the present application, it should be understood that the disclosed devices, apparatuses and systems can be implemented in other manners. For example, the division of the apparatus embodiments is merely an example, and the units can be combined or integrated into another apparatus, or some features can be ignored or not implemented. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, and can be in electrical, mechanical or other forms.
[0141] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one place or distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0142] In addition, the functional units in each embodiment of the present application can be integrated into a processing unit, or each unit can be physically present separately, or two or more units can be integrated into one unit.
[0143] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the parts of the present application that essentially contribute to the prior art or the parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the system described in the embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program code storage media.
[0144] The above description is merely a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A computer failure alarm system, characterized by The system comprises: The behavior pattern analysis module obtains user operation log data, analyzes operation time, operation type and operation frequency, filters high-frequency operation types, analyzes operation consistency again, and obtains behavior deviation indicators; The performance fluctuation detection module identifies load fluctuation amplitude based on the behavior deviation indicators, isolates time periods of abnormal fluctuations, calculates CPU utilization fluctuation in the time periods, compares with average load change trend, and obtains load fluctuation index; The event correlation analysis module analyzes error events, timeout events and resource access conflict events in computer logs based on the load fluctuation index, counts occurrence frequency and distribution pattern of events in abnormal time periods, identifies correlation degree between events, and obtains event correlation strength; The fault early warning module analyzes occurrence frequency and impact range of current abnormal events based on the event correlation strength, calculates cumulative impact degree, evaluates consistency with known fault triggering patterns, determines fault risk level, and obtains fault early warning indicators; The resource adjustment response module identifies computer components with high current risk level based on the fault early warning indicators, allocates additional computing resources or restarts service processes, and obtains computing resource adjustment information; The event correlation analysis module comprises: The abnormal event statistics submodule analyzes error events, timeout events and resource access conflict events in computer logs based on the load fluctuation index, counts occurrence frequency of each type of event in abnormal time periods, evaluates event concentration, identifies event distribution pattern, and obtains abnormal event distribution overview; The event correlation submodule calculates common occurrence probability of difference type events in abnormal time periods based on the abnormal event distribution overview, analyzes time interval and frequency between events, determines dependence relationship between events, and obtains event correlation degree index; The event propagation analysis submodule calls the event correlation degree index, analyzes event occurrence order, calculates propagation path length and impact between events, counts cascading impact range of events, determines key trigger events, and obtains event correlation strength; The formula for analyzing time interval and frequency between events is: analyzing the time interval and frequency between events, determining the dependency between events, and obtaining an event correlation index, wherein, PZ uo represents the joint occurrence probability of event u and event o, τ uk represents the start time of event u in the kth abnormal time period, τ ok represents the start time of event o in the kth abnormal time period, T k represents the total duration of the kth abnormal time period, N PZ represents the total number of abnormal time periods.
2. The computer failure alarm system of claim 1, wherein, The behavior deviation indicators include operation time deviation, operation frequency deviation and operation type deviation, the load fluctuation index includes CPU load fluctuation, memory usage fluctuation and network response fluctuation, the event correlation strength includes error event correlation degree, timeout event correlation degree and resource conflict event correlation degree, the fault early warning indicators include event occurrence frequency, event impact range and event risk level, and the computing resource adjustment information includes resource allocation effect, process restart effect and performance adjustment result.
3. The computer failure alarm system of claim 1, wherein, The behavior pattern analysis module comprises: The log analysis submodule obtains user operation log data, analyzes operation time, operation type and operation frequency, sorts logs according to time tags, analyzes operation activities in each time period, and extracts associated distribution data of time points and operation types, and obtains operation time sequence data; The frequency calculation submodule calculates the occurrence frequency of the operation type in the difference time period based on the operation time series data, statistically analyzes the frequency change trend, screens the high-frequency operation type associated with the known fault, calculates the relative proportion of the high-frequency operation in all operations, and obtains a high-frequency operation proportion index; The abnormal behavior recognition submodule analyzes the deviation degree of the operation behavior by comparing the current operation with the historical data based on the high-frequency operation proportion index, in combination with the identity verification data and the historical operation mode of the user, and obtains a behavior deviation index.
4. The computer failure alarm system of claim 1, wherein, The performance fluctuation detection module includes: The system resource collection submodule collects computer CPU load, memory usage, and network response time based on the behavior deviation index, monitors the computer resource utilization in the difference time window, analyzes the time series change, and obtains a resource usage overview; The load fluctuation recognition submodule calculates the fluctuation amplitude of the CPU load and the memory usage in the time window based on the resource usage overview, compares the resource occupation trend in the time window, screens the time period with a load fluctuation amplitude exceeding the standard, and obtains an abnormal time interval; The load trend calculation submodule calls the abnormal time interval, analyzes the CPU utilization fluctuation in the time period, compares the average load fluctuation of the time period and the overall data, determines the long-term trend and short-term change of the fluctuation, and obtains a load fluctuation index.
5. The computer failure alarm system of claim 1, wherein, The fault early warning module includes: The abnormal event screening submodule analyzes the occurrence frequency of the abnormal event based on the event association strength, calculates the cumulative influence value of each type of event, counts the number of abnormal events exceeding the threshold, and identifies the key abnormal events, and obtains a key abnormal event list; The fault mode matching submodule calls the key abnormal event list, calculates the matching degree of the event and the known fault triggering mode, analyzes the consistency of the event occurrence sequence and the known fault mode, and identifies the key affected components for the matching sequence, and obtains a fault mode matching degree; The risk level evaluation submodule counts the number of computer components involved in the matching event according to the fault mode matching degree, evaluates the fault risk level, and determines the computer fault risk level in combination with the fault influence range, and obtains a fault early warning index.
6. The computer failure alarm system of claim 5, wherein, The number of computer components involved in the statistical matching event uses the formula: evaluating the failure risk level, determining a computer failure risk level, wherein NZ represents the total number of computer components involved, C i represents the absolute value of the set of computer components involved in the i-th failure event, |Ci| represents the number of elements of the set of components involved in the i-th event, M represents the total number of matching events, fj represents the failure frequency of the j-th component, and n represents the total number of components within the evaluation period.
7. The computer failure alerting system of claim 1, wherein, The resource adjustment response module includes: The high-risk component identification submodule analyzes the risk level of the computer component based on the fault early warning index, counts the resource usage of the high-risk component, evaluates the current load state, determines the resource allocation object that needs to be adjusted, and obtains a key risk component; The resource adaptation adjustment submodule calls the key risk component, analyzes the computer available resource pool state, evaluates the allocation ability of the computing resources, calculates the current resource utilization rate, screens the available computing resources for adjustment, and performs resource supplement or process restart on the high-risk component, and obtains resource allocation details; The performance change recording submodule monitors the running state of the adjusted computing resources according to the resource allocation details, records the change trend of the CPU, memory, and network usage, and evaluates the impact of resource adjustment on the stability of the computer, and obtains computing resource adjustment information.
8. A computer failure alarm method for implementing the computer failure alarm system according to any one of claims 1 to 7, characterized by, The method comprises the following steps: S1: obtaining user operation log data, parsing operation time, operation type and operation frequency, arranging operation time sequence, calculating operation frequency change rate, screening high-frequency operation types associated with known faults, and analyzing operation consistency according to user identity and historical operation behavior to obtain behavior deviation index; S2: based on the behavior deviation index, collecting computer CPU load, memory usage and network response time, identifying load fluctuation amplitude for each time window data, calculating CPU utilization fluctuation, comparing with average load change trend to obtain load fluctuation index; S3: based on the load fluctuation index, counting the occurrence frequency and distribution mode of events in the abnormal time period, identifying the correlation degree between events, analyzing the occurrence sequence and propagation path of events, and obtaining the event correlation strength; S4: based on the event correlation strength, calculating the cumulative influence degree of the event, screening abnormal events exceeding the threshold, evaluating the consistency of the event and the known fault triggering mode, and performing correlation analysis on the events and computer components that meet the fault mode to obtain the fault warning index; S5: based on the fault warning index, identifying computer components with high current risk level, analyzing the state of available resource pool, adjusting resource feasibility, allocating additional computing resources or restarting service process to high-risk components, and obtaining computing resource adjustment information.
Citation Information
Patent Citations
Operation state monitoring method and system based on computer management
CN118349415A
Virtualized network fault rapid positioning method fusing multi-source information
CN119520229A