Fault positioning method and device, electronic equipment and storage medium
By acquiring and analyzing various historical operation data of the server and performing causal analysis, the problem of low fault positioning accuracy in the existing technology is solved, and higher fault positioning accuracy is achieved.
Patent Information
- Application Number
- CN202510669886.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-23
AI Technical Summary
The accuracy of fault location in the prior art is low, and a single type of fault data cannot effectively reveal the root cause of the fault.
By obtaining the historical operation data of the server, including at least two of hardware operation data, software operation data and network operation data, and performing causal analysis on different types of data to determine the fault location results.
Through multi-dimensional historical operation data causality analysis, the root cause of server failure can be more accurately identified, thereby improving the accuracy of fault location.
Smart Images

Figure CN120196514A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a fault location method, apparatus, electronic device, and storage medium. Background Art
[0002] When a server fails, the services provided by the server will be interrupted or the performance will decline. In order to quickly repair the server failure, it is necessary to locate the server failure.
[0003] In the related art of fault location, usually a single type of data (for example, hardware metrics in the server) is used for fault location. Since there is an interaction between different types of data, a single type of fault data is not necessarily the root cause of the fault, resulting in a low accuracy rate of fault location. Summary of the Invention
[0004] This application provides a fault location method, apparatus, electronic device, and storage medium to at least solve the problem of low accuracy rate of fault location in the related art.
[0005] This application provides a fault location method, including: When a server fails, obtaining the historical operation data of the server; the historical operation data includes at least two of hardware operation data, software operation data, and network operation data; Performing a causal relationship analysis between different historical operation data to obtain the causal relationship between different historical operation data; Performing a fault location analysis on the causal relationship to obtain the fault location result of the server.
[0006] This application also provides a fault location apparatus, including: An obtaining unit, configured to obtain the historical operation data of the server when the server fails; the historical operation data includes at least two of hardware operation data, software operation data, and network operation data; A first analysis unit, configured to perform a causal relationship analysis between different historical operation data to obtain the causal relationship between different historical operation data; A second analysis unit, configured to perform a fault location analysis on the causal relationship to obtain the fault location result of the server.
[0007] This application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any one of the above fault location methods when executing the computer program.
[0008] The present application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the steps of any one of the above-mentioned fault location methods are implemented.
[0009] The present application also provides a computer program product including a computer program, wherein when the computer program is executed by a processor, the steps of any one of the above-mentioned fault location methods are implemented.
[0010] Through the present application, by obtaining historical operation data of a server, where the historical operation data includes at least two of hardware operation data, software operation data, and network operation data, and performing a causal relationship analysis on the historical operation data, the mutual influence between different types of historical operation data can be revealed, helping to perform fault location more comprehensively. Through the causal relationship analysis of multi-dimensional historical operation data, the root cause leading to the server fault can be identified more accurately, thereby improving the accuracy of fault location. Therefore, the technical problem of low accuracy rate of fault location can be solved, and the technical effect of improving the accuracy of fault location can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0012] Figure 1 It is a schematic flowchart of a fault location method provided by an embodiment of the present application; Figure 2 It is a schematic flowchart of a method for obtaining historical operation data provided by an embodiment of the present application; Figure 3 It is a causal diagram of an undetermined causal relationship provided by an embodiment of the present application; Figure 4 It is a causal diagram of a determined causal relationship provided by an embodiment of the present application; Figure 5 It is a schematic structural diagram of a fault location device provided by an embodiment of the present application; Figure 6 It is a schematic structural diagram of another fault location device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0013] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0014] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0015] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0016] Describe the specific application environment architecture or specific hardware architecture on which the execution of the fault location method depends.
[0017] The embodiments of the present application provide a fault location method, and the method will be described in detail in combination with the execution process of the fault location method.
[0018] Figure 1 It is a schematic flow chart of a fault location method provided by an embodiment of the present application.
[0019] As Figure 1 shown, the method includes the following steps: Step 101, in the case of a server failure, obtain the historical operation data of the server; the historical operation data includes at least two of hardware operation data, software operation data, and network operation data.
[0020] Historical operation data refers to various types of recorded information generated by the server during past normal or abnormal fault operations, covering the hardware, software, and network operation conditions of the server. Historical operation data can be used to analyze the server's operation status and diagnose faults. Hardware operation data refers to the working parameters and status information of the server's hardware devices. Hardware operation data includes, but is not limited to, the temperature, utilization rate, and frequency degradation status of the Central Processing Unit (CPU), the usage rate of memory, Error Checking and Correcting (ECC), the Self-Monitoring Analysis and Reporting Technology (SMART) status of the storage hard disk, the input voltage of power supply and heat dissipation, the fan speed, and the chassis inlet temperature. Among them, the hardware operation data is collected through the Baseboard Management Controller (BMC), and the hardware operation data is collected at a second-level frequency through the Intelligent Platform Management Interface (IPMI) or the Application Programming Interface (API). Software operation data is the operation status information of various software applications running on the server, including, but not limited to, system logs and application logs. Among them, the system logs collect kernel logs and hardware-trace logs through the dynamic message buffer (dmesg), and the application logs output structured application operation logs (in the JavaScript Object Notation (JSON) format) through the log viewing tool (journalctl), including error codes, thread identity identifiers (ID), and stack traces.Network operation data refers to the relevant data related to the server in network communication, including but not limited to switch data and host network metrics. Among them, switch data obtains port status (such as Cyclic Redundancy Check (CRC) errors and packet loss rate) through the Simple Network Management Protocol (SNMP), and host network metrics capture the Transmission Control Protocol (TCP) retransmission rate and connection count through Transmission Control Protocol Top (TCP TOP) or Extended Berkeley Packet Filter (EBPF).
[0021] By obtaining at least two types of historical operation data, the running state of the server can be comprehensively understood from multiple perspectives. A single data type may not reveal the root cause of server failures, while comprehensive analysis of hardware, software, and network data can improve the accuracy of fault location.
[0022] Step 102: Conduct a causal relationship analysis between different historical operation data to obtain the causal relationship between different historical operation data.
[0023] Causal relationship analysis refers to studying whether there is a relationship of influence and being influenced between different historical operation data of the server. For example, when the CPU utilization rate of the server increases (cause), it may lead to an extended response time of software applications (result). Through causal relationship analysis, this potential influence relationship can be determined, which helps to deeply understand the interrelationship between server operation data and provides a basis for accurately finding the root cause of faults.
[0024] Use a monitoring system to collect historical operation data in multiple dimensions in real time, including but not limited to: Hardware operation data: such as CPU usage, memory usage, hard disk read and write speed, device temperature, etc.; Software operation data: such as operating system load, application performance metrics, process execution status, system error logs, etc.; Network operation data: such as network bandwidth, network latency, packet loss rate, number of connections, etc. Preprocess the collected historical operation data, including: Data cleaning: Remove missing values, outliers, and noise data; Data normalization: Unify the dimensions of different data items for cross-dimensional comparative analysis; Time synchronization: Align the time of data from different sources to ensure the time sequence of data during the analysis process. Use causal inference techniques to establish a causal relationship model between historical operation data. The causal relationship analysis methods include but not limited to: Causality test, Bayesian network, and structural equation model: The causality test is used to detect causal relationships in time series data. By checking whether the past values of one variable can help predict the future values of another variable, the causal relationship between variables is inferred; The Bayesian network represents the dependence relationship between data by constructing a conditional probability model to identify possible causal paths; The structural equation model models the relationships between multiple variables by constructing structural equations to identify the causal connections between variables. According to the established causal relationship model, analyze the causal relationships between different historical operation data. For example, through the causality test, it can be identified whether the fluctuation of CPU usage causes changes in memory usage, or whether network latency is related to hard disk read performance.
[0025] Causal relationships include but are not limited to the causal relationship between high CPU load and memory consumption, the relationship between high memory usage and disk read and write speed, the association between network bandwidth occupancy and response time delay, etc.
[0026] Through the analysis of causal relationships between different historical operation data, it is possible to go beyond surface phenomena and deeply explore the real root causes of server failures. For example, when it is found that the service response time is extended, simply looking at the software operation data may not accurately determine the reason. However, through causal relationship analysis, it is found that there is a causal relationship with the CPU utilization and memory usage of the hardware. Thus, the hardware resources can be optimized or upgraded specifically to fundamentally solve the fault problem, avoiding the blindness and one-sidedness of fault troubleshooting.
[0027] Step 103, perform fault location analysis on the causal relationships to obtain the fault location results of the server.
[0028] Causality refers to an association existing among different operation data of a server, that is, a change in one operation data (cause) will lead to a change in another operation data (result). For example, an increase in the CPU utilization rate of the server hardware (cause) may result in an extended response time of the software service (result). Fault location analysis refers to the process of determining the specific location or root cause of a server failure by applying specific analysis methods and models based on the causal relationships among various operation data of the server. The purpose is to accurately identify the source that causes the server performance degradation or service interruption, so as to quickly take repair measures. The fault location result refers to the final output of the fault location analysis, indicating the component, module, historical operation data or specific cause where the server failure occurs. For example, it is determined that the failure is caused by a memory leak in the hardware, a code defect in the software, or insufficient network bandwidth, etc.
[0029] Based on the previously determined causal relationships, a Bayesian network model is constructed. For example, assume that through causal relationship analysis, it is found that there is a causal relationship between the CPU utilization of the hardware and the service response time of the software, and at the same time, the network bandwidth utilization is also related to the service response time. These variables (CPU utilization, network bandwidth utilization, service response time) are used as nodes in the Bayesian network, and the directed edges between the nodes are determined according to the causal relationships, that is, from the cause node to the result node. Define the possible states for each node. For example, the states of CPU utilization can be divided into "normal", "higher", "too high", and the states of service response time can be divided into "normal", "slightly long", "extremely long", etc. According to the historical operation data collected, count and determine the conditional probabilities of each node under different state combinations of its parent nodes. Taking the service response time node as an example, its parent nodes are CPU utilization and network bandwidth utilization. By analyzing the historical data, calculate the probability that the service response time is "slightly long" when the CPU utilization is "higher" and the network bandwidth utilization is "normal"; the probability that the service response time is "extremely long" when the CPU utilization is "too high" and the network bandwidth utilization is "too high", etc., so as to construct a complete conditional probability table. The prior probability refers to the probability estimate of the node state based on past experience or prior knowledge before considering specific observed data. For example, according to the past operation experience of the server, determine that the prior probability of the CPU utilization being in the "normal" state is 0.6, the prior probability of the "higher" state is 0.3, and the prior probability of the "too high" state is 0.1; the prior probability of the network bandwidth utilization being in the "normal" state is 0.7, the prior probability of the "higher" state is 0.2, and the prior probability of the "too high" state is 0.1, etc. When the server fails, collect the current observed data, such as the service response time being in the "extremely long" state. According to Bayes' theorem, use the conditional probability table and the prior probability to calculate the posterior probabilities of each possible cause node (such as CPU utilization, network bandwidth utilization) in different states. For example, calculate the posterior probability that the CPU utilization is in the "too high" state and the network bandwidth utilization is in the "too high" state when the service response time is "extremely long". Compare the posterior probabilities corresponding to the states of each cause node, and determine the state of the cause node with the largest posterior probability as the fault location result. If the calculation result shows that the posterior probability that the CPU utilization is in the "too high" state and the network bandwidth utilization is in the "too high" state is the largest, then it is initially judged that the server failure is jointly caused by the excessive use of the hardware CPU and insufficient network bandwidth. However, it should be clear that this statement is not intended to limit the fault location to only be achieved through the above method, and it can also be achieved through other means.
[0030] Fault location analysis using causal relationships can comprehensively consider the complex associations among various server operation data, accurately locate the root cause of the fault. Compared with traditional fault location methods that rely solely on single data or simple threshold judgments, it improves the accuracy of fault location, avoids blind troubleshooting, reduces the fault repair time, and quickly restores the normal operation of the server.
[0031] Through this application, by obtaining the historical operation data of the server, where the historical operation data includes at least two of hardware operation data, software operation data, and network operation data, and performing causal relationship analysis on the historical operation data, the mutual influence between different types of historical operation data can be revealed, helping to conduct more comprehensive fault location. Through the causal relationship analysis of multi-dimensional historical operation data, the root cause leading to server faults can be more accurately identified, thereby improving the accuracy of fault location. Therefore, the technical problem of low accuracy rate of fault location can be solved, and the technical effect of improving the accuracy of fault location can be achieved.
[0032] As a refinement of step 101, when obtaining the historical operation data of the server, it can be implemented in but not limited to the following ways, such as Figure 2 shown Figure 2 is a flowchart of a method for obtaining historical operation data provided by an embodiment of this application, including: Step 201, obtain the historical operation data to be processed.
[0033] The historical operation data to be processed refers to various historical operation data recorded during the past operation of the server, which are in an unprocessed or original state, including historical hardware operation data to be processed (such as CPU utilization rate, memory usage rate), historical software operation data to be processed (such as process information, transaction processing quantity, error log, service response time, etc.), and historical network operation data to be processed (such as network bandwidth utilization rate, network latency, network packet loss rate, etc.). The historical operation data to be processed is directly obtained from the server's operation records and may contain problems such as noise, missing values, and inconsistent timestamps, and needs to be further processed before it can be used for fault location analysis.
[0034] The acquisition methods of the historical operation data to be processed include but are not limited to monitoring tools, sensors, log recording systems, or API interfaces. It realizes an automated acquisition means, can obtain historical operation data in real time or regularly, reduces manual intervention, and improves the efficiency and accuracy of data acquisition.
[0035] Step 202, preprocess the historical operation data to be processed to obtain historical operation data; the preprocessing includes at least one of time alignment, mean calculation, variance calculation, change rate calculation, and conversion of non-numerical data to numerical data.
[0036] Time alignment refers to converting the running data from different sources, with different timestamp formats or sampling frequencies, into the same timestamp format and arranging them in chronological order, so that the data is consistent and comparable in the time dimension. Mean calculation refers to calculating the average value of the running data over a period of time, which is used to smooth out data fluctuations and reflect the overall trend and average level of the data. Variance calculation refers to measuring the degree of dispersion of the historical running data to be processed, that is, the degree of difference between the historical running data to be processed and the mean value. The larger the variance, the greater the data fluctuation, which may imply the instability of the server operation. Rate of change calculation refers to calculating the rate of change of the historical running data to be processed between adjacent time periods or adjacent data points, which is used to reflect the dynamic change trend of the historical running data to be processed and help identify sudden changes or abnormal fluctuations in the historical running data to be processed. Conversion of non-numerical data to numerical data refers to converting the historical running data to be processed in non-numerical form (such as text status information, category labels, etc.) into numerical form through specific coding methods or mapping rules, so as to perform mathematical operations and data analysis.
[0037] To facilitate a better understanding of the preprocessing, an example is provided. High-precision timestamps are assigned to all historical running data to be processed (synchronized by the Network Time Protocol (NTP), with a precision of <1 ms); the data is aligned according to a fixed time window (such as 1 minute), and the data within the window is aggregated into a feature vector. For example, the second-level sampling of BMC (60 samples / minute) is aligned with the millisecond-level events of application logs (such as 100 errors / minute) in the same time window. Hardware features: Calculate the mean / variance of the CPU temperature and the rate of change of the fan speed (Δ revolutions per minute (RPM) / minute); and convert discrete events (such as "abnormal power input voltage") into binary flag bits. Software features: Log templatization: Using unified processing, the original logs are parsed into specific event types; and calculate the number of error logs per minute and the number of service startups. Network features: Statistically count the port CRC error count and TCP retransmission rate; and the sudden increase in bandwidth utilization.
[0038] Data preprocessing can standardize data, ensure data quality, reduce errors in the analysis process, and improve the reliability of the results.
[0039] As a refinement of step 102, when performing causal relationship analysis between different historical operation data to obtain the causal relationship between different historical operation data, the following methods can be adopted but are not limited to, including: defining variables for different historical operation data to obtain variables of different historical operation data; performing conditional independence tests between different variables to obtain independence results between different variables; determining the causal relationship between different variables according to the independence results, and determining the causal relationship between different variables as the causal relationship between the historical operation data corresponding to different variables.
[0040] Variable definition means designating each data item in the historical operation data as a variable. For example, in the historical operation data, CPU utilization rate can be defined as a variable, and memory usage rate can be defined as another variable. Conditional independence test is a statistical method used to determine whether two variables are independent of each other given other variables. For example, testing whether variables A and B are independent given variable C. The independence result refers to the output of the conditional independence test, indicating whether two variables are independent under the given conditions. If two variables are independent under the given conditions, there is no direct causal relationship between them; if they are dependent, there may be a causal relationship. Causal relationship refers to the relationship in which one historical operation data (cause) causes another historical operation data (effect) to change. In the server operation data, for example, high CPU utilization rate may cause an increase in service response time.
[0041] First, perform variable definition, map the preprocessed features (i.e., historical operation data) to the nodes in the causal graph, and each node represents a variable among them. The observed value at each time point in the historical operation data corresponds to the value of these variables.
[0042] For the sake of easy understanding, an example is provided. Hardware nodes: Cpu temperature (CPU_Temp), Fan speed (Fan_RPM), Ambient temperature (Ambient_Temp); at time t1, Fan_RPM = 3000, CPU_Temp = 65, Ambient_Temp = 25; at time t2, Fan_RPM = 2800, CPU_Temp = 70, Ambient_Temp = 28; at time t3, Fan_RPM = 2500, CPU_Temp = 80, Ambient_Temp = 30; these data points constitute the time series of each variable.
[0043] Then initialize the fully connected graph, assuming that there are potential causal relationships between all variables, that is, each variable is connected to other variables by edges. For example, composed of three variables Fan_RPM, CPU_Temp, Ambient_Temp Figure 1At the beginning, there will be three edges: Fan_RPM - CPU_Temp, Fan_RPM - Ambient_Temp, CPU_Temp - Fan_RPM, forming a fully connected graph.
[0044] However, in reality, there may be bidirectional edges between every variable, or all possible edges exist. But usually in a directed graph, there will be edges in both directions between each pair of variables. However, we first initialize it as an undirected fully connected graph and then determine the directions later.
[0045] Perform conditional independence tests. Test independence in ascending order of the number of variables in the conditional set (zero variables → single variable → two variables...). For each pair of variables X and Y, first test whether they are independent under the condition of zero variables (i.e., without giving any other variables). If they are independent, delete the edge between them; if not, continue to test under the condition of giving one variable... until the smallest conditional set that makes X and Y independent is found. Here, the cases of zero variables and single - variable conditional sets are taken as examples. For how to calculate whether two variables are independent, it can be judged by using a statistical test method based on the likelihood ratio. The calculation formula of the statistical test method can be implemented through formula (1): (1) Where, is the statistic, is the observed frequency (the frequency of X = and Y = in the actual data (historical operation data)), is the expected frequency (the theoretical frequency when assuming X and Y are independent), is the natural logarithm, is the summation symbol, used to traverse all possible combinations of X and Y categories, where X and Y are different variables.
[0046] The formula is based on comparing the likelihood ratio of the observed data with the expected data under the independent hypothesis. If the difference between the two is large, it means the independent hypothesis does not hold. The likelihood - ratio statistic (LR) is defined as: LR = 2 (likelihood of saturated model / likelihood of independent model), so multiplying by 2 is to make the statistic approximately follow the chi - square distribution.
[0047] For the sake of easy understanding, an example is provided. When testing whether X (Fan_RPM) and Y (CPU_Temp) are independent under the condition of the empty set, first construct a contingency table. Assume that at a historical moment, variable X has m data and variable Y has n data. The contingency table is shown in Table 1: Table 1
[0048] Among them, are different observed frequencies, X and Y are variables, is the number of categories of X, is the number of categories of Y, is the sum of the observed frequencies in different rows, is the sum of the observed frequencies in different columns, is the total sum of all observed frequencies.
[0049] Then, in the case of expectation, that is, under the assumption of independence between X and Y, the calculation formula for the expected frequency ( ) can be achieved through formula (2): (2) The expected frequency of each cell is the product of its corresponding row total and column total, divided by the total sample size. Then substituting into formula (1) to obtain value. Subsequently, determine the degrees of freedom (Degrees of Freedom, ), and the calculation formula for the degrees of freedom can be achieved through formula (3): (3) For ease of understanding, an example is provided. For instance, for a 2x2 contingency table (such as comparing CPU temperature (normal / high) with fan speed (high / low)), the degrees of freedom .
[0050] When choosing the significance level α = 0.05 for using the degrees of freedom, since α is the allowable type I error probability (i.e., the risk of wrongly rejecting H0 (assuming X and Y are independent)), so α = 0.05 means accepting a 5% misjudgment risk. Using the table lookup method, in the chi-square distribution table, find the corresponding critical value according to and α. When = 1 and α = 0.05, the critical value is 3.841. Finally, compare the value with the critical value. If ≥ critical value, reject the null hypothesis (keep the edge, X and Y are not independent); if < critical value, accept the null hypothesis (delete the edge, X and Y are independent).
[0051] When testing for independence under univariate conditions, for the edges that have not been eliminated, gradually increase the conditioning variable Z and test whether X and Y are independent given Z. Taking the edge of X (Fan_RPM) and Y (CPU_Temp) as an example, test whether they are independent given Z (Ambient_Temp). First, assume that X and Y are independent. Follow the above steps to calculate the value, bin the data by Z (such as Z ≤ 28 degrees Celsius and Z > 28 degrees Celsius), and calculate the conditional independence of X and Y respectively. Suppose the calculation gives = 1.2, according to the look-up table method <Critical value, accept the original hypothesis → accept H0, delete the edge, X and Y are independent.
[0052] By combining various types of historical operation data for causal relationship analysis, the occurrence mechanism of server failures can be understood more comprehensively, avoiding misjudgments that may be caused by a single data type.
[0053] As a refinement of the above embodiments, when performing conditional independence tests between different variables to obtain the independence results between different variables, the following methods can be used but are not limited to: obtaining the conditional sets of any at least two variables among different variables; the conditional set is all variables other than at least two variables among different variables; in the direction of increasing the number of variables in the conditional set, use the variables in the conditional set to perform conditional independence tests on at least two variables until the target variables with the least number of variables in the conditional set that cause at least two variables to be independent are found; determine the target variables as the independence results of at least two variables.
[0054] The conditional set refers to, when performing conditional independence tests, the set of all other variables that may affect the target variable except the target variable being tested. The variables in the conditional set are used to help determine whether there is an independent relationship between the target variables. The target variables are those variables with the least number selected from the conditional set that can make any at least two variables among different variables satisfy the independent condition during the conditional independence test process; among them, the number of target variables can be zero or the total number of variables in the conditional set, and the embodiments of the present application do not limit the number of target variables.
[0055] Extract the required variables from historical operation data, such as CPU load, memory usage rate, disk read / write rate, etc. Select at least two variables to be analyzed (for example, CPU load and memory usage rate) for conditional independence testing. For the selected at least two variables, first construct a preliminary conditional set, which includes all other variables that may affect these variables except these two variables (such as network bandwidth, disk usage rate, etc.). In the order of increasing the number of variables in the conditional set, gradually add more variables to the conditional set for conditional independence testing. Start from the conditional set with the smallest number of variables (for example, no variables), and conduct conditional independence testing on the variables to be tested. If the test result shows that the target variables are still not independent, gradually increase the variables in the conditional set and continue the conditional independence testing. In each conditional independence test, when through a combination of variables in a certain conditional set, at least two variables to be tested can be made independent, the variables in this conditional set can be regarded as the target variables. Determine the smallest target variables, that is, the variables in the smallest conditional set that can make the target variables independent. Take the finally determined target variables as the independence result, as the conclusion of the independence testing between two or more variables, indicating that these variables are independent under specific conditions and no longer affect each other.
[0056] Testing in the direction of increasing the number of variables in the conditional set can potentially find the target variables that meet the conditions at an earlier stage, avoiding unnecessary complex testing of more variables, improving the testing efficiency, and saving computational resources and time costs.
[0057] As a refinement of the above embodiments, when executing to determine the causal relationship between different variables according to the independence result, the following methods can be used but are not limited to: establish a causal relationship between the target variables and at least two variables; in the causal relationship, the target variables are the cause and at least two variables are the effect.
[0058] Cause refers to the variable that can cause changes in other variables in the causal relationship. Effect refers to the variable that changes under the influence of other variables in the causal relationship.
[0059] For ease of understanding, an example is provided to determine the direction of an edge (causal relationship). To do this, a causal chain (V-structure) needs to be utilized. That is, in a causal graph, if there is a structure X → Z ← Y, and X and Y are not adjacent (i.e., there is no direct edge), then Z is called the collision node (Collider (i.e., the target variable)) of X and Y. At this time, X and Y may be related given Z (even if they were originally independent). That is to say, if there exists a node Z such that X-Z-Y is a potential V-structure, and X and Y are independent under zero variables but related given Z, then the direction is determined as X → Z ← Y. Through calculation, it has been confirmed that X (Fan_RPM) and Y (CPU_Temp) are not independent under zero variables. Therefore, we determine that the ambient temperature Ambient_Temp (Z) may affect the fan speed (X) and the CPU temperature (Y) simultaneously. Logically, in terms of causality, a high-temperature environment (Z↑) causes the fan speed to increase (X↑) to enhance heat dissipation; a high-temperature environment (Z↑) directly increases the CPU temperature (Y↑). So, the final causal graph is obtained. For better understanding of the causal graph, as Figure 3 、 Figure 4 shown, Figure 3 is a causal graph of an undetermined causal relationship provided by an embodiment of the present application, Figure 4 is a causal graph of a determined causal relationship provided by an embodiment of the present application. As can be seen from Figure 3 、 Figure 4 , the ambient temperature is the cause, and the fan speed and the central processing unit temperature are the effects.
[0060] Understanding the causal relationships between variables helps optimize the configuration and performance of the server system.
[0061] As a refinement of the above embodiment, when performing a conditional independence test on at least two variables using the variables in the condition set, it can be implemented in, but not limited to, the following ways, including: obtaining the first historical operation data corresponding to the variables in the condition set in the historical operation data, obtaining the second historical operation data corresponding to at least two variables, and obtaining the critical value of the second historical operation data; based on the first historical operation data, dividing the second historical operation data into at least two data sets; each data set contains at least one piece of the second historical operation data; calculating the statistics of at least two data sets to obtain the statistics of at least two data sets; and determining whether at least two variables are independent according to the comparison between the statistics and the critical value.
[0062] The first historical operation data refers to the historical operation data corresponding to the variables in the condition set, and the second historical operation data refers to the historical operation data corresponding to at least two variables. The critical value is a preset numerical threshold in the conditional independence test, used to determine whether the statistic reaches the standard of independence. The degrees of freedom represent the number of independent values that can vary freely in the data when estimating a statistic or parameter.
[0063] The method for obtaining the critical value of the second historical operation data can be achieved through the following method: obtain the degrees of freedom and significance level of the second historical operation data, and based on the degrees of freedom and significance level, look up the critical value corresponding to the degrees of freedom and significance level in the chi-square distribution table.
[0064] For statistical methods such as the chi-square test, the size of the degrees of freedom will affect the judgment criterion of the test result. For example, in a 2×2 contingency table, the degrees of freedom is 1 because it has one independent constraint (i.e., the total count of each cell is fixed), so only the value of one cell can vary freely.
[0065] Extract the first historical operation data corresponding to the condition set variable Z (server load) and the second historical operation data corresponding to variables X (CPU temperature) and Y (CPU usage rate) from the historical operation data of the server. At the same time, determine the degrees of freedom of the second historical operation data. For example, assume that the server load data of 100 sample points, as well as the corresponding CPU temperature and usage rate data, are collected over a period of time. For the chi-square test, if the CPU temperature is divided into three categories: "low temperature", "medium temperature", and "high temperature", and the CPU usage rate is divided into two categories: "low" and "high", then the constructed contingency table is 3×2, and the degrees of freedom at this time is . According to the first historical operation data (server load), divide the second historical operation data (CPU temperature and usage rate) into at least two data sets. For example, according to the median of the server load, divide the data into two data sets: high load and low load. Each data set contains the CPU temperature and usage rate data at the corresponding moment, so that the relationship between the CPU temperature and usage rate under different load conditions can be analyzed separately. For the at least two divided data sets, calculate the statistic respectively. Compare the calculated statistic with the critical value under the corresponding degrees of freedom. For the case where the degrees of freedom is 2, assume that the significance level is 0.05, and the corresponding critical value in the chi-square distribution table is 5.991. If the calculated statistic is greater than or equal to 5.991, then reject the hypothesis that variables X and Y are independent, and consider that there is an association between them; otherwise, consider them independent.
[0066] Calculate the statistics of at least two data sets separately, sum up the statistics of the at least two data sets to obtain a sum result, determine the sum result as the final statistics of the at least two data sets, and determine whether at least two variables are independent according to the comparison between the final statistics and the critical value.
[0067] By obtaining relevant variable data from historical operation data and partitioning the data based on conditional set variables, the influence of various factors in the server operation environment on variable relationships is considered. Combining the comparison of the statistics with the degrees of freedom can more accurately determine whether two variables are independent, avoid drawing wrong conclusions due to ignoring the interference of other variables, and improve the reliability of causal relationship analysis.
[0068] As a refinement of the above embodiment, when determining whether at least two variables are independent according to the comparison between the statistics and the degrees of freedom, the following methods can be adopted but are not limited to: if the statistics are greater than or equal to the critical value, it is determined that at least two variables are not independent; if the statistics are less than the critical value, it is determined that at least two variables are independent.
[0069] For the sake of understanding, an example is provided. Assume that the statistics are 4 and the critical value is 3.5, then it is determined that at least two variables are not independent.
[0070] The method of using the critical value to judge variable independence can judge the dependence between variables through a quantitative standard, avoid misdiagnosis caused by human bias or contingency, and ensure the accuracy and reliability of fault analysis.
[0071] As a refinement of the above embodiment, when calculating the statistics of at least two data sets to obtain the statistics of the at least two data sets, the following methods can be adopted but are not limited to: obtaining the observed frequencies of the at least two data sets and obtaining the expected frequencies of the at least two data sets; the observed frequency is the frequency of the same second historical operation data in the at least two data sets; the expected frequency is the theoretical frequency of independence between the second historical operation data; calculating the statistics for the observed frequency and the expected frequency to obtain the statistics.
[0072] The observed frequency is the frequency at which the same second historical operation data (such as a specific value or interval of the server response time) actually appears in the data set. For example, under a certain condition, the situation where the server response time exceeds 500ms actually appears 20 times in the data set, and this 20 times is the observed frequency. The expected frequency is the theoretical frequency calculated according to probability theory on the premise of assuming independence between the second historical operation data. For example, if the probability that the server response time exceeds 500ms is 0.2 and there are 100 data points in a certain data set, then the expected frequency is 100×0.2 = 20 times. The calculation formula of the statistics is shown in formula (1).
[0073] Calculate the statistics for the observed frequencies and expected frequencies of at least two data sets separately, obtain the statistics corresponding to each of the at least two data sets, and perform a summation calculation on the statistics corresponding to each of the at least two data sets to obtain the final statistic for the at least two data sets.
[0074] By comparing the observed frequencies and expected frequencies, the independence between two data sets can be more accurately evaluated. This provides a reliable statistical basis for system fault diagnosis and performance analysis, avoiding misdiagnosis caused by human judgment or intuitive errors.
[0075] As a refinement of the above embodiment, when performing the calculation of the statistic for the observed frequency and expected frequency to obtain the statistic, it can be implemented in but not limited to the following ways, including: calculating the quotient of the observed frequency and the expected frequency to obtain the quotient result; calculating the logarithm of the quotient result to obtain the logarithm result; calculating the product of the logarithm result and the observed frequency to obtain the statistic.
[0076] The embodiment of the present application describes the calculation method of the statistics corresponding to each of at least two data sets, and the final statistic also needs to perform a summation calculation on the statistics corresponding to each of the at least two data sets.
[0077] Specifically, the implementation process of this embodiment is a literal description of formula (1).
[0078] As a refinement of step 103, when performing the fault location analysis of the causal relationship to obtain the fault location result of the server, it can be implemented in but not limited to the following ways, including: obtaining the conditional probability table of the causal relationship, and obtaining the marginal probability of the causal relationship and the prior probability of the causal relationship; the conditional probability table includes the conditional probability of the effect when the cause in the causal relationship is in different states; determining the conditional probability as the likelihood probability corresponding to the causal relationship; calculating the posterior probability of the causal relationship for the likelihood probability, prior probability, and the corresponding likelihood probability; determining the cause in the causal relationship corresponding to the maximum value in the posterior probability as the fault location result.
[0079] The conditional probability table is a table that describes the probabilities of various factors in a causal relationship leading to the occurrence of an effect in different states. Through the conditional probability table, the probability of an event occurring under each condition can be determined, and it is commonly used in Bayesian network models. Marginal probability refers to the probability of an event occurring without considering other conditions. It is the result of summing or integrating conditional probabilities under specific conditions and reflects the overall probability of a single event occurring in the system. Prior probability is an estimate of the probability of an event occurring based on existing knowledge or experience, usually an estimated value before any new data or observations. Likelihood probability is the probability of a hypothesis or event occurring derived from existing observed data. In fault location analysis, likelihood probability is used to represent the possibility of a certain causal relationship occurring and is calculated based on the observed results of the current state of the system. Posterior probability is the updated probability of an event occurring after considering the prior probability and new observed data. Bayes' theorem is used to calculate the posterior probability, which reflects the updated probability of an event occurring after new information is added.
[0080] The calculation formula for the posterior probability can be implemented through formula (4): (4) Where, is the posterior probability, is the cause, is the effect, is the likelihood probability, is the prior probability, is the marginal probability.
[0081] By combining conditional probability, marginal probability, prior probability, and posterior probability for comprehensive analysis, the specific factors leading to faults can be identified more accurately. The calculation of the posterior probability is based on multiple information sources and can provide more reliable fault location results.
[0082] As a refinement of the above embodiment, when performing the calculation of the posterior probability for the likelihood probability, prior probability, and the corresponding likelihood probability to obtain the posterior probability of the causal relationship, it can be implemented in but not limited to the following ways, including: calculating the product value of the likelihood probability and the prior probability to obtain the product value result; calculating the quotient value of the product value result and the marginal probability to obtain the posterior probability.
[0083] Specifically, the implementation process of this embodiment is a verbal description of formula (4).
[0084] To facilitate a better understanding of the Conditional Probability Table (CPT), an example is provided. The input historical operation data is time-series numerical type: t1, Fan_RPM(2500), Ambient_Temp(30), CPU_Temp(75); t2, Fan_RPM(3500), Ambient_Temp(40), CPU_Temp(95); t3, Fan_RPM(2000), Ambient_Temp(35), CPU_Temp(90); To represent the corresponding relationship intuitively, it is necessary to learn the discrete data, so the variables are discretized (assuming for simplified calculation): Fan_RPM: low (<2500 RPM), medium (2500 - 3500 RPM), high (>3500 RPM). Ambient_Temp: low (<25 °C), medium (25 to 35 °C), high (>35 °C). CPU_Temp: normal (<70 °C), on the high side (70 to 90 °C), too high (>90 °C).
[0085] The CPT represents the probability distribution of the child node (effect) when the parent node (cause) takes a specific combination. For example, the result probability for CPU_Temp is shown in Table 2 as follows: Table 2
[0086] Then, when performing maximum likelihood estimation, we count the number of occurrences of each value of the child node under each combination of parent nodes. The frequency is divided by the total number of observations to obtain the conditional probability. For example, when Fan_RPM = low and Ambient_Temp = medium, it is observed that: CPU_Temp = normal: 80 times; CPU_Temp = on the high side: 15 times; CPU_Temp = too high: 5 times; Total number of observations: 100 times; So the conditional probabilities are: P(normal) = 80 / 100 = 0.8; P(on the high side) = 15 / 100 = 0.15; P(too high) = 5 / 100 = 0.05.
[0087] The reasoning principle of CPT, forward inference: Given the value of the parent node, predict the probability of the child node. Example: If the fan speed is low and the ambient temperature is moderate, the probability of the CPU temperature being normal is predicted to be 80%. Diagnostic inference: Given the value of the child node, infer the posterior probability of the parent node. Example: If the CPU temperature is too high, calculate the possibility of a fan failure or a high ambient temperature. The forward reasoning of CPT directly queries the child node distribution through the CPT and is applicable to preventing faults from occurring: Scenario: Given the value of the parent node, predict the probability distribution of the child node. Input: Fan_RPM = low, Ambient_Temp = medium Output: CPU_Temp = normal: 80%, on the high side: 15%, too high: 5%. The reverse reasoning of CPT is applicable to root cause localization when a fault occurs: Scenario: Given the value of the child node (such as a fault), infer the posterior probability of the parent node. Input: CPU_Temp = too high; Output: The probability of Fan_RPM = low, the probability of Ambient_Temp = high, etc. In reverse reasoning, for each possible combination of parent nodes, the implementation method for calculating the posterior probability is shown in formula (3). Causal reasoning: Combine the observed evidence (such as too high CPU temperature), and by calculating the posterior probability of all nodes and the marginal probability of a single node, locate the most likely root cause (such as insufficient fan speed).
[0088] Example demonstration of conditional independence test: Suppose to test whether the fan speed - Fan_RPM (X: high / low) and the CPU temperature - CPU_Temp (Y: normal / too high) are independent. The collected data is shown in Table 3: Table 3
[0089] Calculate the expected frequency , = 50×60 / 100 = 30; = 50×40 / 100 = 20; = 50×60 / 100 = 30; = 50×40 / 100 = 20.
[0090] Calculate the contribution value of each cell: .
[0091] .
[0092] .
[0093] .
[0094] Calculate the .
[0095] , referring to the chi-square distribution table, with a significance level α = 0.05 and 1 degree of freedom, the chi-square critical value is 3.841.
[0096] The calculated Reject the null hypothesis and conclude that the fan speed and CPU temperature are not independent.
[0097] A failure occurred during the operation of the machine (server): CPU_Temp = 95 degrees Celsius. The CPU temperature was too high, triggering causal reasoning.
[0098] Taking the causal diagram as an example, the parent nodes of CPU_Temp are Fan_RPM and Ambient_Temp, and the CPT table is shown in Table 4: Table 4
[0099] The goal is to calculate the root cause probability leading to high CPU, that is and .
[0100]
[0101] Among them, f can be any one of low, medium, and high of Fan_RPM, and a can be any one of low, medium, and high of Ambient_Temp. is the posterior probability, is the likelihood probability, and are the prior probabilities respectively.
[0102] Below, use CPU (CPU_Temp), Fan (Fan_RPM), and Ambient (Ambient_Temp) for short.
[0103] 1. Likelihood probability, : Obtained from the CPT table. For example, when Fan = low and Ambient = medium, .
[0104] 2. Prior probability, , statistically obtained from the input historical data. The probability of Fan = low in the historical data is 20%, and the probability of Ambient = medium is 50%.
[0105] 3. Joint likelihood, that is, the numerator part: .
[0106] 4. Calculate the marginal probability. When calculating the marginal probability (denominator), sum over all possible combinations of parent nodes: .
[0107] Example calculations (assuming only the following combinations exist) are shown in Table 5:
[0108] The calculated total contribution value = 0.0166. The posterior probabilities of each combination are shown in Table 6: Table 6
[0109] Verification sum: 0.301 + 0.361 + 0.301 + 0.036 ≈ 1.0.
[0110] The marginalized calculation of the posterior probability of a single parent node is:
[0111] where, is the likelihood probability.
[0112] 0.301 (Ambient = medium) + 0.361 (Ambient = high) = 0.662.
[0113] So when the posterior probabilities of all states of Fan are shown in Table 7:
[0114] Similarly, the posterior probability of Ambient_Temp is calculated and shown in Table 8: Table 8
[0115] Final comparison: The contribution of Fan = low is 66.2%, and the contribution of Ambient = high is 36.1%.
[0116] Improve the accuracy and efficiency of fault location; multi-source data fusion: integrate hardware (BMC), software logs, and network metrics, avoid the limitations of a single data source in traditional methods, and comprehensively capture fault correlations; causality reasoning-driven: mine the causal relationships between variables through causal graphs, distinguish symptoms from root causes, and reduce misjudgments (such as attributing high CPU temperature to fan failure rather than environmental temperature); real-time reasoning ability: respond to fault events in seconds, significantly reducing the mean time to repair. Reduce operation and maintenance costs and risks; reduce false alarms: filter false correlations through causality reasoning, and avoid invalid alarms caused by unreasonable threshold settings; predictive maintenance: predict potential faults (such as fan life attenuation) through probability models and trigger preventive maintenance in advance.
[0117] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0118] An embodiment of the present application also provides a fault location device. Figure 5 As shown in the structure diagram of a fault location device provided by an embodiment of the present application, as Figure 5 shown, it includes: An acquisition unit 31, configured to acquire historical operation data of a server when a fault occurs in the server; the historical operation data includes at least two of hardware operation data, software operation data, and network operation data; A first analysis unit 32, configured to perform a causal relationship analysis between different historical operation data to obtain a causal relationship between different historical operation data; A second analysis unit 33, configured to perform a fault location analysis on the causal relationship to obtain a fault location result of the server.
[0119] Through the present application, since the historical operation data of the server is acquired, the historical operation data includes at least two of hardware operation data, software operation data, and network operation data, and a causal relationship analysis is performed on the historical operation data, the mutual influence between different types of historical operation data can be revealed, which helps to perform a more comprehensive fault location. Through the causal relationship analysis of multi-dimensional historical operation data, the root cause leading to the server fault can be more accurately identified, thereby improving the accuracy of fault location. Therefore, the technical problem of low accuracy of fault location can be solved, and the technical effect of improving the accuracy of fault location can be achieved.
[0120] Furthermore, in a possible implementation manner of this embodiment, as Figure 6 shown, the acquisition unit 31 is further configured to: Acquire to-be-processed historical operation data; Preprocess the to-be-processed historical operation data to obtain historical operation data; the preprocessing includes at least one of time alignment, mean calculation, variance calculation, change rate calculation, and conversion of non-numerical data into numerical data.
[0121] Furthermore, in a possible implementation manner of this embodiment, as Figure 6 shown, the first analysis unit 32 includes: A definition module 321, configured to perform variable definition on different historical operation data to obtain variables of different historical operation data; A test module 322, configured to perform a conditional independence test between different variables to obtain an independence result between different variables; A determination module 323, configured to determine the causal relationship between different variables according to the independence result, and determine the causal relationship between different variables as the causal relationship between the historical operation data corresponding to different variables.
[0122] Further, in a possible implementation manner of this embodiment, the test module 322 is further configured to: Obtain a conditional set of any at least two variables among different variables; the conditional set is all variables except the at least two variables among different variables; In the direction of increasing the number of variables in the conditional set, use the variables in the conditional set to perform a conditional independence test on the at least two variables until the target variable with the least number of variables in the conditional set that causes the at least two variables to be independent is found; Determine the target variable as the independence result of the at least two variables.
[0123] Further, in a possible implementation manner of this embodiment, as Figure 6 shown, the determination module 323 is further configured to: Establish a causal relationship between the target variable and the at least two variables; in the causal relationship, the target variable is the cause and the at least two variables are the effects.
[0124] Further, in a possible implementation manner of this embodiment, as Figure 6 shown, the test module 322 is further configured to: Obtain the first historical operation data corresponding to the variables in the conditional set in the historical operation data, obtain the second historical operation data corresponding to the at least two variables, and obtain the critical value of the second historical operation data; Based on the first historical operation data, divide the second historical operation data into at least two data sets; each data set contains at least one second historical operation data; Perform a statistic calculation on the at least two data sets to obtain the statistics of the at least two data sets; Determine whether the at least two variables are independent according to the comparison between the statistic and the critical value.
[0125] Further, in a possible implementation manner of this embodiment, as Figure 6 shown, the test module 322 is further configured to: When the statistic is greater than or equal to the critical value, determine that the at least two variables are not independent; When the statistic is less than the critical value, determine that the at least two variables are independent.
[0126] Further, in a possible implementation manner of this embodiment, as Figure 6 shown, the test module 322 is further configured to: Obtain the observed frequencies of at least two data sets and obtain the expected frequencies of at least two data sets; the observed frequency is the frequency of the same second historical operation data in at least two data sets; the expected frequency is the theoretical frequency when the second historical operation data are independent. Perform a statistic calculation on the observed frequency and the expected frequency to obtain a statistic.
[0127] Furthermore, in a possible implementation manner of this embodiment, as Figure 6 shown, the test module 322 is further configured to: Perform a quotient calculation on the observed frequency and the expected frequency to obtain a quotient result; Perform a logarithm calculation on the quotient result to obtain a logarithm result; Perform a product calculation on the logarithm result and the observed frequency to obtain a statistic.
[0128] Furthermore, in a possible implementation manner of this embodiment, as Figure 6 shown, the second analysis unit 33 is further configured to: Obtain a conditional probability table of the causal relationship, and obtain the marginal probability of the causal relationship and the prior probability of the causal relationship; the conditional probability table includes the conditional probabilities of the effect when the cause in the causal relationship is in different states; Determine the conditional probability as the likelihood probability of the corresponding causal relationship; Perform a posterior probability calculation on the likelihood probability, the prior probability, and the corresponding likelihood probability to obtain the posterior probability of the causal relationship; Determine the cause in the causal relationship corresponding to the maximum value in the posterior probability as the fault location result.
[0129] Furthermore, in a possible implementation manner of this embodiment, as Figure 6 shown, the second analysis unit 33 is further configured to: Perform a product calculation on the likelihood probability and the prior probability to obtain a product result; Perform a quotient calculation on the product result and the marginal probability to obtain the posterior probability.
[0130] For the description of the features in the embodiments corresponding to the fault location device, reference can be made to the relevant descriptions in the embodiments corresponding to the fault location method, which will not be elaborated here one by one.
[0131] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above embodiments of the fault location method.
[0132] Embodiments of the present application also provide a computer-readable storage medium storing a computer program, where the computer program is configured to execute the steps in any of the above-described embodiments of the fault location method when running.
[0133] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), external hard drives, magnetic disks, or optical discs that can store computer programs.
[0134] Embodiments of the present application also provide a computer program product, where the computer program product includes a computer program that implements the steps in any of the above-described embodiments of the fault location method when executed by a processor.
[0135] Embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program that implements the steps in any of the above-described embodiments of the fault location method when executed by a processor.
[0136] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0137] The above has introduced in detail a fault location method and apparatus, an electronic device, and a storage medium provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A fault location method, characterized in that, Including: In the case of a server failure, obtaining the historical operation data of the server; The historical operation data includes at least two of hardware operation data, software operation data, and network operation data; Performing a causal relationship analysis between different pieces of the historical operation data to obtain the causal relationship between different pieces of the historical operation data; Performing a fault location analysis on the causal relationship to obtain the fault location result of the server.
2. The fault location method according to claim 1, characterized in that The obtaining the historical operation data of the server includes: Obtaining the historical operation data to be processed; Performing preprocessing on the historical operation data to be processed to obtain the historical operation data; the preprocessing includes at least one of time alignment, mean calculation, variance calculation, change rate calculation, and conversion of non-numerical data to numerical data.
3. The fault location method according to claim 1, wherein The performing a causal relationship analysis between different pieces of the historical operation data to obtain the causal relationship between different pieces of the historical operation data includes: Defining variables for different pieces of the historical operation data to obtain the variables of different pieces of the historical operation data; Performing a conditional independence test between different ones of the variables to obtain the independence result between different ones of the variables; According to the independence result, determining the causal relationship between different ones of the variables, and determining the causal relationship between different ones of the variables as the causal relationship between the corresponding historical operation data of different ones of the variables.
4. The fault location method according to claim 3, wherein The performing a conditional independence test between different ones of the variables to obtain the independence result between different ones of the variables includes: Obtaining a conditional set of any at least two variables among different ones of the variables; the conditional set is all variables other than the at least two variables among different ones of the variables; In the direction of increasing number of variables in the conditional set, using the variables in the conditional set to perform a conditional independence test on the at least two variables until the target variable with the least number of variables in the conditional set that causes the at least two variables to be independent is found; Determining the target variable as the independence result of the at least two variables.
5. The fault location method according to claim 4, wherein, The according to the independence result, determining the causal relationship between different ones of the variables includes: Establishing a causal relationship between the target variable and the at least two variables; in the causal relationship, the target variable is the cause and the at least two variables are the effects.
6. The fault location method according to claim 4, characterized in that, The using the variables in the conditional set to perform a conditional independence test on the at least two variables includes: Obtaining the first historical operation data corresponding to the variables in the conditional set in the historical operation data, obtaining the second historical operation data corresponding to the at least two variables, and obtaining the critical value of the second historical operation data; Based on the first historical operation data, dividing the second historical operation data into at least two data sets; each data set contains at least one piece of the second historical operation data; Performing a statistic calculation on the at least two data sets to obtain the statistics of the at least two data sets; According to the comparison between the statistic and the critical value, determining whether the at least two variables are independent.
7. The fault location method according to claim 6, characterized in that Determining whether the at least two variables are independent according to the comparison between the statistic and the critical value includes: If the statistic is greater than or equal to the critical value, it is determined that the at least two variables are not independent; If the statistic is less than the critical value, it is determined that the at least two variables are independent.
8. The fault location method according to claim 6, wherein Calculating the statistic for the at least two data sets, obtaining the statistic of the at least two data sets includes: Obtaining the observed frequencies of the at least two data sets, and obtaining the expected frequencies of the at least two data sets; the observed frequencies are the frequencies of the same second historical operation data in the at least two data sets; the expected frequencies are the theoretical frequencies when the second historical operation data are independent; Calculating the statistic by performing a statistic calculation on the observed frequencies and the expected frequencies.
9. The fault location method according to claim 8, wherein, Calculating the statistic by performing a statistic calculation on the observed frequencies and the expected frequencies includes: Calculating the quotient of the observed frequency and the expected frequency to obtain a quotient result; Calculating the logarithm of the quotient result to obtain a logarithm result; Calculating the product of the logarithm result and the observed frequency to obtain the statistic.
10. The fault location method according to claim 1, wherein, Performing a fault location analysis on the causal relationship to obtain the fault location result of the server includes: Obtaining the conditional probability table of the causal relationship, and obtaining the marginal probability of the causal relationship, the prior probability of the causal relationship; the conditional probability table includes the conditional probabilities of the effect when the cause in the causal relationship is in different states; Determining the conditional probability as the likelihood probability of the corresponding causal relationship; Calculating the posterior probability of the causal relationship by performing a posterior probability calculation on the likelihood probability, the prior probability and the corresponding likelihood probability; Determining the cause in the causal relationship corresponding to the maximum value in the posterior probability as the fault location result.
11. The fault location method according to claim 10, wherein Calculating the posterior probability of the causal relationship by performing a posterior probability calculation on the likelihood probability, the prior probability and the corresponding likelihood probability includes: Calculating the product of the likelihood probability and the prior probability to obtain a product result; Calculating the quotient of the product result and the marginal probability to obtain the posterior probability.
12. A fault location device, characterized in that, Includes: An acquisition unit, configured to acquire the historical operation data of the server when the server fails; The historical operation data includes at least two of hardware operation data, software operation data, and network operation data; A first analysis unit, configured to perform a causal relationship analysis on different historical operation data to obtain the causal relationship between different historical operation data; A second analysis unit, configured to perform a fault location analysis on the causal relationship to obtain the fault location result of the server.
13. An electronic device, characterized in that, Includes: A memory, configured to store a computer program; A processor, configured to implement the steps of the fault location method according to any one of claims 1 to 11 when executing the computer program.
14. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program, when executed by a processor, implements the steps of the fault location method according to any one of claims 1 to 11.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the fault location method according to any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
Fault positioning method for Web application system
CN101394314A
Fault root cause positioning method and device and related product
CN119520247A
System and method for proactive handling of multiple faults and failure modes in an electrical network of energy assets
US20200209841A1