Fault Root Cause Determination Method, Apparatus and Electronic Device
By constructing Bayesian network analysis, determining the root cause parameters of the fault, solving the problem that the root cause of the fault is difficult to accurately locate during the silicon wafer wire cutting process, achieving rapid and accurate troubleshooting and improving production efficiency.
Patent Information
- Application Number
- CN202510020777.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-01-03
AI Technical Summary
During the silicon wafer wire cutting process, due to the complex equipment structure, the root of the fault is difficult to accurately locate. The existing troubleshooting methods rely on experience, are time-consuming and easy to misjudgment or misjudgment.
By constructing the first Bayesian network and the second Bayesian network, the degree of abnormality is analyzed based on the key parameters, the edge probability change and causal change of each node are determined, and the root cause fraction is calculated based on the preset weighting coefficient, and the root cause parameters of the fault are determined.
It shortens the troubleshooting time, improves the accuracy of fault location, reduces the frequency of abnormal occurrence, and improves the production efficiency of the production line.
Smart Images

Figure CN119416071B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of fault analysis, and particularly to a method, device and electronic device for determining the root cause of a fault. Background Art
[0002] Various anomalies usually occur during the wire cutting process of silicon wafers. These anomalies may lead to unqualified product quality or even equipment shutdown, thus affecting production efficiency. Therefore, it is necessary to promptly detect and handle the anomalies during the production process. Some of the anomalies during the production process can be solved by adjusting process parameters, while another part is caused by potential faults of the equipment itself. Due to the complex structure of the wire cutting machine, there are a large number of process parameters and components, so it is usually difficult to accurately locate the root cause of equipment faults.
[0003] Current fault detection methods mainly rely on the experience of engineers. Engineers will preliminarily judge possible fault causes based on the working state of the equipment, historical data, and abnormal phenomena. However, due to the subjectivity of experience and the complexity of the equipment, a large number of inspections or disassembly experiments are usually required to finally determine the problem, which not only takes a long time and is likely to affect the overall efficiency of the production line, but also is prone to misjudgment or missed judgment, thus leading to the recurrence of anomalies. Summary of the Invention
[0004] Embodiments of this application provide a method, device and equipment for determining the root cause of a fault to shorten the fault detection time and improve the accuracy of fault location.
[0005] To solve the above technical problems, embodiments of this application disclose the following technical solutions:
[0006] In a first aspect, a method for determining the root cause of a fault is provided, including:
[0007] Constructing a first Bayesian network based on the key parameters of a first machine tool, and constructing a second Bayesian network based on the key parameters of a second machine tool, where the degree of anomaly of the first machine tool meets a preset condition, and the degree of anomaly of the second machine tool exceeds the preset condition;
[0008] Determining the change amount of the marginal probability and the change amount of the causal relationship of each node based on the first marginal probability and the first conditional probability of each node in the first Bayesian network, and the second marginal probability and the second conditional probability of each node in the second Bayesian network;
[0009] Determining the root cause score of each node based on the change amount of the marginal probability and the change amount of the causal relationship of each node, and a preset weighting coefficient;
[0010] Determining the key parameter corresponding to the node with the highest root cause score among all the nodes as the fault root cause parameter.
[0011] In some embodiments, the step of determining the marginal probability change amount and the causal relationship change amount of each of the nodes includes:
[0012] Based on the first marginal probability of each node in the first Bayesian network and the second marginal probability of the corresponding node in the second Bayesian network, determine the marginal probability change amount of each of the nodes;
[0013] Based on the first conditional probability of each node in the first Bayesian network with respect to any parent node and the second conditional probability of the corresponding node in the second Bayesian network with respect to the corresponding parent node, determine the conditional probability change amount of each of the nodes with respect to the parent node;
[0014] Determine the sum of the conditional probability change amounts of each of the nodes with respect to all the parent nodes as the causal relationship change amount of each of the nodes.
[0015] In some embodiments, the step of determining the causal relationship change amount of each of the nodes includes:
[0016] Determine the causal relationship change amount of each of the nodes by the following formula:
[0017]
[0018] where ΔC(X) is the causal relationship change amount of the node, X is the node, Y is the parent node, Pa(X) is the set of parent nodes of node X, is the conditional probability change amount of node X with respect to parent node Y.
[0019] In some embodiments, the preset weighting coefficients include a first weight coefficient and a second weight coefficient; the step of determining the root cause score of each of the nodes includes:
[0020] Determine the root cause score of each of the nodes by the following formula:
[0021]
[0022] where S(X) is the root cause score of node X, W 1 is the first weight coefficient, W 2 is the second weight coefficient, W1 + W2 = 1; ΔP(X) is the marginal probability change amount of node X, and ΔC(X) is the causal relationship change amount of node X.
[0023] In some embodiments, the first weight coefficient is 0.3 and the second weight coefficient is 0.7.
[0024] In some embodiments, constructing the first Bayesian network based on the key parameters of the first machine tool includes:
[0025] Obtain the historical data of each of the key parameters of the first machine tool, and divide the historical data into a training set and a test set;
[0026] Based on the training set, use the expectation-maximization method to determine the structure of the first Bayesian network and the direction of each edge;
[0027] Use the maximum likelihood estimation method to determine the first conditional probability table of each node in the first Bayesian network;
[0028] Verify the first Bayesian network based on the test set, and correct the structure of the first Bayesian network and the direction of each edge according to the verification result to obtain the trained first Bayesian network.
[0029] In some embodiments, the first machine tool and the second machine tool are determined through the following steps:
[0030] Determine the abnormality rate and the variance of the abnormality rate of each operating machine tool in a preset time period;
[0031] Determine the operating machine tool with an abnormality rate less than a first threshold and a variance of the abnormality rate less than a second threshold as the first machine tool;
[0032] Determine the operating machine tool with an abnormality rate greater than or equal to the first threshold and / or a variance of the abnormality rate greater than or equal to the second threshold as the second machine tool.
[0033] In some embodiments, the steps of determining the key parameters include:
[0034] Determine the equipment parameters of each key control point of the operating machine tool, where the key control point is used to characterize the equipment position that can affect the process;
[0035] Select the key parameters from the equipment parameters of each key control point.
[0036] In a second aspect, a device for determining the root cause of a fault is provided, including:
[0037] A network construction module, configured to construct a first Bayesian network based on the key parameters of the first machine tool and a second Bayesian network based on the key parameters of the second machine tool, where the abnormality degree of the first machine tool meets a preset condition, and the abnormality degree of the second machine tool exceeds the preset condition;
[0038] A network comparison module, configured to determine the marginal probability change amount and causal relationship change amount of each node based on the first marginal probability and first conditional probability of each node in the first Bayesian network, and the second marginal probability and second conditional probability of each node in the second Bayesian network;
[0039] A score calculation module, configured to determine the root cause score of each node based on the marginal probability change amount and causal relationship change amount of each node, and a preset weighting coefficient;
[0040] A root cause determination module, configured to determine the key parameter corresponding to the node with the highest root cause score among all the nodes as the fault root cause parameter.
[0041] In a third aspect, an electronic device is provided, the electronic device includes a processor and a memory; a computer program is stored in the memory, and the processor is configured to execute the computer program stored in the memory to implement the fault root cause determination method according to any one of the first aspects.
[0042] One of the above technical solutions has the following advantages or beneficial effects:
[0043] Compared with the prior art, a fault root cause determination method of the present application includes: constructing a first Bayesian network based on the key parameters of a first machine tool, and constructing a second Bayesian network based on the key parameters of a second machine tool, where the abnormal degree of the first machine tool meets a preset condition, and the abnormal degree of the second machine tool exceeds the preset condition; determining the marginal probability change amount and causal relationship change amount of each node based on the first marginal probability and first conditional probability of each node in the first Bayesian network, and the second marginal probability and second conditional probability of each node in the second Bayesian network; determining the root cause score of each node based on the marginal probability change amount and causal relationship change amount of each node, and a preset weighting coefficient; determining the key parameter corresponding to the node with the highest root cause score among all the nodes as the fault root cause parameter. The fault root cause determination method provided by the present application determines the fault root cause parameter by comparing the Bayesian networks of machine tools with different abnormal degrees, which can not only shorten the fault troubleshooting time, improve the production efficiency of the production line, but also greatly improve the accuracy of fault location.
[0044] A fault root cause determination device of the present application determines the fault root cause parameter by comparing the Bayesian networks of machine tools with different abnormal degrees, which can not only shorten the fault troubleshooting time, improve the production efficiency of the production line, but also greatly improve the accuracy of fault location. Description of the Drawings
[0045] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0046] Figure 1 It is a schematic diagram of the overall process of the method for determining the root cause of a fault in the embodiments of the present application;
[0047] Figure 2 It is a schematic diagram of the initial structure of the first Bayesian network in the embodiments of the present application;
[0048] Figure 3 It is a schematic diagram of the initial structure of the second Bayesian network in the embodiments of the present application;
[0049] Figure 4 It is a schematic diagram of the structure of the device for determining the root cause of a fault in the embodiments of the present application;
[0050] Figure 5 It is a schematic diagram of the structure of the electronic device in the embodiments of the present application.
[0051] Reference numerals:
[0052] 401 - Network construction module; 402 - Network comparison module; 403 - Score calculation module; 404 - Root cause determination module; 501 - Memory; 502 - Processor. Detailed implementation manners
[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0054] In the description of the present application, it should be understood that the orientation or positional relationship indicated by the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present application. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more features. In the description of the present application, the meaning of "a plurality" is two or more, and at least one means it can be one, two or more, unless otherwise specifically defined.
[0055] During the wire cutting process of silicon wafers, various abnormalities usually occur, such as wire breakage, uneven cutting, surface defects, uneven cutting thickness, etc. Some abnormalities can be solved by adjusting process parameters, such as adjusting cutting speed, tension, etc., but another part of the abnormalities is caused by potential failures of the equipment itself. Due to the complex structure of the wire cutting machine, there are a large number of process parameters and components, making it difficult to accurately locate the source of the failure.
[0056] The current fault troubleshooting methods mainly rely on the experience of engineers. Engineers will preliminarily judge the possible causes of faults based on the working state of the equipment, historical data, and abnormal phenomena. However, although the experience of engineers can help locate faults to a certain extent, due to the complex and changeable operating state of the equipment, relying on experience for judgment is often inaccurate. Especially when facing newly emerging abnormalities, misjudgment or missed judgment may occur. At the same time, due to the large number of variables involved in the wire cutting machine, when a fault occurs, a large number of troubleshooting or disassembly experiments need to be carried out to check possible causes one by one. This process not only takes a long time but may also miss the golden time of production, further affecting the overall efficiency of the production line. That is to say, the current fault troubleshooting methods often can only find the superficial causes of problems and are difficult to find the real root causes of faults.
[0057] In view of this, the embodiments of the present application provide a method for determining the root cause of faults. By analyzing the complex relationships between various components and parameters of the equipment and finding the real root cause of the fault according to the complex causal relationship chain inside the equipment, the root cause of the fault can be determined efficiently and accurately, avoiding the recurrence of abnormalities, and thus at least part of the above technical problems can be solved.
[0058] Please refer to Figure 1 , Figure 1 which is the overall flow schematic diagram of the method for determining the root cause of faults in the embodiments of the present application. The method for determining the root cause of faults includes the following steps:
[0059] Step 101: Construct a first Bayesian network based on the key parameters of the first machine tool, and construct a second Bayesian network based on the key parameters of the second machine tool. The abnormal degree of the first machine tool meets the preset conditions, and the abnormal degree of the second machine tool exceeds the preset conditions.
[0060] Specifically, the key parameters refer to the parameters of the key points that have a greater impact on the equipment process. It can be understood that for some equipment parameters, adjusting them will not have an obvious impact on the equipment process, so such equipment parameters do not belong to the key parameters.
[0061] In some examples, the steps to determine the key parameters include:
[0062] Step 1: Determine the equipment parameters of each key control point of the operating machine tool. The key control points are used to represent the equipment points that can affect the process.
[0063] Exemplarily, the operating machine tool can be a silicon wafer wire cutting equipment.
[0064] Step 2: Select the key parameters from the equipment parameters of each key control point.
[0065] Specifically, the equipment parameters that can have a greater impact on the process among the equipment parameters of each key control point are determined as the key parameters. There are various methods to select the key parameters from the equipment parameters of each key control point. For example, the machine tool can be disassembled, and analyzed according to the disassembly results and the mechanism of the machine tool to select the key parameters. Or the key parameters can be determined in advance according to actual needs. The embodiments of the present application do not make specific limitations on this.
[0066] Exemplarily, the key parameters may include the cutting fluid flow rate (unit: l / min), the cutting fluid temperature (unit: °C), the tension of the left tension arm (unit: N), the wire mesh speed (unit: m / s), the temperature of the left tension motor (unit: °C), the torque of the left main shaft motor (unit: Nm), the inlet water temperature of the water-cooled motor (unit: °C), and the feed speed (unit: mm / min). Among them, the cutting fluid can play a role in protecting and lubricating during wire mesh cutting. The cutting fluid temperature can affect the protection and lubrication ability of the cutting fluid. The cutting fluid flow rate can affect the lubrication and cooling effect of the cutting fluid. The tension of the left tension arm can control the tension or tightness of the cutting wire mesh. If it is too loose, the silicon wafer is prone to defects. If it is too tight, the wire is prone to breakage. The wire mesh speed can control the speed of the wire mesh going back and forth, and the speed will affect the cutting ability of the wire mesh. The temperature of the left tension motor can control the stability of the wire mesh tension. The torque of the left main shaft motor can reflect the driving force of the main shaft to ensure sufficient torque output during the cutting process. The inlet water temperature of the water-cooled motor can control the temperature of the system to prevent overheating. The feed speed can affect the cutting depth and speed to ensure smooth and efficient cutting.
[0067] Through the above solution, the selected key parameters have an obvious impact on the process and may have a greater impact on the product quality. Therefore, including these key parameters in the monitoring scope can ensure the comprehensiveness and effectiveness of equipment operation monitoring.
[0068] In some embodiments, the degree of abnormality can be characterized by the abnormality rate and the variance of the abnormality rate. The first machine tool can represent an excellent machine tool, and the second machine tool can represent a defective machine tool. The first machine tool and the second machine tool can be determined through the following steps:
[0069] Step 1: Determine the abnormality rate and the variance of the abnormality rate of each operating machine tool within a preset time period.
[0070] Exemplarily, the preset time period can be set according to actual needs. For example, it can be set to one month. The abnormality rate is used to characterize the proportion of abnormalities of the operating machine tool within the preset time period. The abnormality rate can be determined by the following formula:
[0071]
[0072] where λ i is the abnormality rate of the i-th operating machine tool within the preset time period, i is an integer greater than or equal to 1 and less than or equal to I, I is the total number of operating machine tools, n 异常 is the number of abnormal samples of the i-th operating machine tool within the preset time period, n 总 is the total number of samples of the i-th operating machine tool within the preset time period.
[0073] Exemplarily, taking the wire cutting equipment as the operating machine tool, each time the machine tool makes a cut, it will record whether the cut is abnormal. The total number of cuts within the preset time period is the total number of samples of the operating machine tool, and the number of abnormal cuts is the number of abnormal samples of the operating machine tool.
[0074] The variance of the abnormality rate is used to characterize the stability of the operating machine tool within the preset time period. The variance of the abnormality rate can be determined by the following formula:
[0075]
[0076] where is the variance of the abnormality rate of the i-th operating machine tool within the preset time period, t is an integer greater than or equal to 1 and less than or equal to T, T is the total number of sub-time periods into which the preset time period is divided, λ t is the abnormality rate of the i-th operating machine tool within the t-th sub-time period, is the average value of the abnormality rates of the i-th operating machine tool within each sub-time period.
[0077] Exemplarily, the preset time period is one month, and each sub-time period is one week. In this way, the variance of the abnormality rate for one month can be determined based on the abnormality rate of the operating machine tool each week.
[0078] Step 2: Determine the operating machines with an abnormality rate less than the first threshold and an abnormality rate variance less than the second threshold as the first machines.
[0079] Step 3: Determine the operating machines with an abnormality rate greater than or equal to the first threshold and / or an abnormality rate variance greater than or equal to the second threshold as the second machines.
[0080] Exemplarily, the thresholds for different factories and different devices can be different. The target value expected by the factory can be used as the threshold, or the threshold can also be set by means of clustering. The embodiments of the present application do not make specific limitations in this regard.
[0081] Through the above solution, the operating machines with a low abnormality rate and a small variance are classified as excellent machines, and the operating machines with a high abnormality rate and / or a large variance are classified as extremely poor machines, so that the operating machines can be classified according to the probability of abnormality occurrence and the performance stability, thereby laying a foundation for subsequent inference.
[0082] It can be understood that the key parameters of the first machines and the key parameters of the second machines are key parameters of the same type. For example, they can both be cutting fluid flow rate, cutting fluid temperature, left tension arm tension, wire mesh speed, left tension motor temperature, left spindle motor torque, water-cooled motor inlet temperature, and feed speed. It's just that the relationships between these key parameters may change on machines with different degrees of abnormality.
[0083] It can also be understood that the second machines can be used to represent the operating machines in a fault state, and the first machines can be used to represent the operating machines in a normal state.
[0084] Please refer to Figure 2 , Figure 2 which is the schematic diagram of the initial structure of the first Bayesian network in the embodiments of the present application. Among them, a~l represent the nodes in the first Bayesian network, and each node represents a key parameter. The Bayesian network is a directed acyclic graph (DAG), which can model the causal relationships between various variables. Exemplarily, node a can be the feed speed [mm / min], node b can be the temperature of the left movable bearing box [℃], node c can be the cutting fluid flow rate [l / min], node d can be the temperature of the right movable bearing box [℃], node e can be the temperature of the left tension motor [℃], node f can be the left tension arm tension [N], node g can be the rotational speed of the left tension motor [rpm], node h can be the right tension arm tension [N], node i can be the cutting fluid temperature [℃], node j can be the temperature of the right tension motor [℃], node k can be the rotational speed of the right tension motor [rpm], and node l can be the wire mesh speed [m / s].
[0085] In some embodiments, constructing the first Bayesian network can be achieved through the following steps:
[0086] Step 1: Obtain the historical data of each key parameter of the first machine tool, and divide the historical data into a training set and a test set.
[0087] Specifically, the number of the first machine tools can be one or more, and then the historical data of each key parameter of all the first machine tools is collected. The training set and the test set can be divided by using the cross-validation method.
[0088] In some examples, after obtaining the historical data, the data can be preprocessed first to remove missing values and outliers, and then the continuous variables can be discretized to meet the processing requirements of the Bayesian network. For example, the temperature can be divided into "high", "normal", and "low".
[0089] Step 2: Based on the training set, use the expectation-maximization method to determine the structure of the first Bayesian network and the direction of each edge.
[0090] Specifically, the expectation-maximization method refers to the EM (Expectation-Maximization algorithm) algorithm. The construction of the Bayesian network can be divided into two core parts: structure learning (determining the edges and their directions) and parameter estimation (determining the conditional probabilities). In the EM algorithm, the addition, deletion, or retention of edges is achieved through the repeated iteration of the E step and the M step, and finally the optimal network structure is determined.
[0091] Exemplarily, Step 2 can be specifically achieved through the following steps:
[0092] The first step: Initialize the Bayesian network structure.
[0093] Specifically, each key parameter is a node in the Bayesian network. The initial structure can be assumed to be a graph without edges, and all nodes are initially independent, or some of the initial edges can be defined according to the existing domain knowledge.
[0094] The second step: The E step (expectation step), that is, calculate the conditional probability table of the current network.
[0095] Specifically, under the initial structure, the conditional probability table (CPT) of each node can be calculated, and the distribution of the child nodes can be estimated according to the states of the parent nodes. Then calculate the likelihood value of the current network, that is, the overall probability of the data under the current structure.
[0096] Exemplarily, the current network includes three nodes, namely the device temperature X 1 , the vibration amplitude X 2 and the device operating state X 3 , and through collecting data examples, the following Table 1 is obtained.
[0097] Table 1: Data Example
[0098]
[0099] First, calculate the marginal probabilities of each node: P(X 1 = high) = 3 / 5 = 0.6, P(X 1 = low) = 2 / 5 = 0.4, P(X 2 = high) = 2 / 5 = 0.4, P(X 2 = low) = 3 / 5 = 0.6, P(X 3 = abnormal) = 2 / 5 = 0.4, P(X 3 = normal) = 3 / 5 = 0.6. Then calculate the probability of each sample. The probability of each sample is the product of the node marginal probabilities. For example, for the first sample, P(X 1 = high, X 2 = high, X 3 = abnormal) = 0.6 * 0.4 * 0.4 = 0.096. Finally, determine the total likelihood value of the current network as the product of all sample probabilities. The total likelihood value is characterized by the following formula:
[0100]
[0101] where L is the total likelihood value, n is an integer greater than or equal to 1 and less than or equal to N, N is the total number of samples, is the marginal probability of node X 1 under the nth sample, is the marginal probability of node X 2 under the nth sample, is the marginal probability of node X 3 under the nth sample.
[0102] Third step, the M-step (maximization step), that is, perform edge adjustment.
[0103] Specifically, edge adjustment mainly includes edge addition, deletion, and direction determination. The goal is to maximize the likelihood value of the current structure by adding and deleting edges.
[0104] Exemplarily, the steps of edge addition mainly include:
[0105] Select two node pairs without edges, assuming there is a potential causal relationship between these two nodes.
[0106] If there is prior knowledge, such as knowing that temperature affects the flow rate of the cutting fluid, then an edge can be set from temperature to flow rate. Determine the direction through conditional independence testing: If the conditional independence between node X and node Y cannot be explained by other variables, then assume a direct edge from node X to node Y or from node Y to node X. Use a scoring criterion (such as the Bayesian Information Criterion, BIC) to compare the two directions. Select the direction with a higher score and add this edge. After adding the new edge, recalculate the likelihood value of the network and update the conditional probability table (CPT) of each node. If the likelihood value increases and the increase exceeds the set threshold, it means that this edge helps to improve the explanatory power of the network, and retain this edge. If the increase in the likelihood value does not exceed the set threshold or the likelihood value decreases, then do not add this edge.
[0107] Exemplarily, the steps for deleting an edge mainly include:
[0108] Perform redundant edge detection, that is, for the existing edges, try to temporarily delete them and recalculate the likelihood value of the network. If the likelihood value of the network remains unchanged or slightly increases after deleting a certain edge, it means that this edge may be redundant and can be deleted. If deleting the edge causes the likelihood value to decrease significantly, it means that the existence of this edge significantly increases the likelihood value and deleting this edge will reduce the explanatory power of the network, then retain this edge, and finally traverse all edges to complete the likelihood value evaluation.
[0109] Exemplarily, the steps for optimizing the direction of an edge mainly include:
[0110] When there is an edge between two nodes, the final edge direction can be determined by comparing the likelihood values in different directions. If the conditional probability table in a certain direction better fits the data (i.e., the likelihood value is higher), then set the direction of the edge to this direction. Or domain knowledge can be used to determine the direction of the edge at the same time.
[0111] Step 3, use the maximum likelihood estimation method to determine the first conditional probability table of each node in the first Bayesian network.
[0112] Specifically, the conditional probability can be calculated first, that is, according to different combinations of parent nodes, calculate the conditional probability of each node. For each state of the parent node, calculate the probability distribution of each state of the child node. Then perform maximum likelihood estimation (MLE), that is, maximize the conditional probability of each node to maximize the likelihood value of the entire network, and obtain the final first conditional probability table.
[0113] Step 4, verify the first Bayesian network based on the test set, and correct the structure of the first Bayesian network and the direction of each edge according to the verification result to obtain the trained first Bayesian network.
[0114] Specifically, verify the fitting effect of the first Bayesian network of excellent machines on different datasets to ensure the stability of the structure. If the first Bayesian network performs poorly in the test set, the network can be optimized by modifying the network structure (re-adding or deleting edges, or changing the direction of edges). Among them, the connection and direction of edges can be key points for optimization, and finally converge to the optimal network structure through repeated iteration.
[0115] In the above way, a Bayesian network is introduced to handle probability and uncertainty. Each key parameter is used as a node of the Bayesian network, and the first Bayesian network reflecting the operating state of excellent machines is constructed by calculating its probability distribution. The first Bayesian network can not only predict the occurrence of events through conditional probability reasoning, but also perform reverse reasoning to judge the possible causes of a certain observation result. Moreover, with the addition of new data, the network parameters can be dynamically updated through Bayes' theorem, so that the first Bayesian network can remain effective when the environment or conditions change.
[0116] Please refer to Figure 3 , Figure 3 which is a schematic diagram of the initial structure of the second Bayesian network in the embodiments of the present application. Among them, a~l represent each node in the second Bayesian network, and each node represents a key parameter. It can be seen that due to the different degrees of abnormality of the first machine and the second machine, the causal relationship between each node has changed. For example, there is no direct causal relationship between node e and node f in the first Bayesian network, but there is a direct causal relationship in the second Bayesian network. Exemplarily, node a can be the feed rate [mm / min], node b can be the temperature of the left active bearing box [℃], node c can be the cutting fluid flow rate [l / min], node d can be the temperature of the right active bearing box [℃], node e can be the temperature of the left tension motor [℃], node f can be the tension of the left tension arm [N], node g can be the speed of the left tension motor [rpm], node h can be the tension of the right tension arm [N], node i can be the temperature of the cutting fluid [℃], node j can be the temperature of the right tension motor [℃], node k can be the speed of the right tension motor [rpm], and node l can be the wire mesh speed [m / s].
[0117] Correspondingly, the construction of the second Bayesian network can be achieved through the following steps:
[0118] Step 1: Obtain the historical data of each key parameter of the second machine and divide the historical data into a training set and a test set.
[0119] Step 2: Based on the training set, use the expectation-maximization method to determine the structure of the second Bayesian network and the direction of each edge.
[0120] Step 3: Use the maximum likelihood estimation method to determine the second conditional probability table of each node in the second Bayesian network.
[0121] Step 4: Verify the second Bayesian network based on the test set, and correct the structure of the second Bayesian network and the direction of each edge according to the verification results to obtain the trained second Bayesian network.
[0122] It should be noted that the specific details of the above Steps 1 to 4 can refer to the steps of constructing the first Bayesian network, which will not be elaborated here.
[0123] It can be understood that by using a Bayesian network to represent the operating state of a machine tool, it is possible to handle uncertainties that cannot be directly handled by other acyclic graphs, and it has strong robustness to noisy data. It can infer causal relationships through the conditional probability distribution between nodes and quantify the possibility of each root cause. In addition, the Bayesian network can not only predict the occurrence of events, but also has strong reverse reasoning ability, and can judge the possible causes of a certain observation result. Moreover, with the input of new data, the network parameters can be dynamically updated through Bayes' theorem, so that the network can remain effective when the environment or conditions change, without the need for retraining, which is convenient for dynamic adjustment. Therefore, the Bayesian network has an optimal combined effect.
[0124] Step 102: Determine the marginal probability change amount and causal relationship change amount of each node based on the first marginal probability and first conditional probability of each node in the first Bayesian network, and the second marginal probability and second conditional probability of each node in the second Bayesian network.
[0125] Specifically, the marginal probability refers to the probability of a certain node occurring without considering other conditions. In a Bayesian network, the marginal probability of each node is calculated by cumulatively calculating all combinations of the states of its parent nodes. The conditional probability refers to the probability of a child node occurring given the state of its parent node. In a Bayesian network, an edge represents a causal relationship, and the conditional probability is used to describe the way a child node depends on its parent node.
[0126] In some embodiments, the determination of the marginal probability change amount and causal relationship change amount of each node can be achieved through the following steps:
[0127] Step 1: Determine the marginal probability change amount of each node based on the first marginal probability of each node in the first Bayesian network and the second marginal probability of the corresponding node in the second Bayesian network.
[0128] Specifically, the marginal probability change amount of each node is determined by the following formula:
[0129]
[0130] where ΔP(X) is the marginal probability change amount of node X, P 2(X) is the second marginal probability of node X in the second Bayesian network, which is obtained according to the parent nodes of node X in the second Bayesian network and the second conditional probability table; P 1 (X) is the first marginal probability of node X in the first Bayesian network, which is obtained according to the parent nodes of node X in the first Bayesian network and the first conditional probability table.
[0131] Exemplarily, in the first machine, the marginal probability P 1 (X) = 0.8, while in the second machine, the marginal probability P 2 (X) = 0.6. Then the change in the marginal probability of node X is ΔP(X) = 0.8 - 0.6 = 0.2. This result indicates that in the first machine, the probability of node X occurring is higher than that in the second machine, indicating that the state of node X in the first machine is more stable or common.
[0132] Step 2: Based on the first conditional probability of each node in the first Bayesian network relative to any parent node and the second conditional probability of the corresponding node in the second Bayesian network relative to the corresponding parent node, determine the change in the conditional probability of each node relative to the parent node.
[0133] Specifically, the change in the conditional probability of each node relative to any parent node is determined by the following formula:
[0134]
[0135] where is the change in the conditional probability of node X relative to parent node Y, is the conditional probability of node X when the state of parent node Y is given in the first machine, is the conditional probability of node X when the state of parent node Y is given in the second machine.
[0136] Exemplarily, in the first machine, when the state of parent node Y is given as 1, the conditional probability that the state of node X is 1 is = 0.9. In the second machine, when the state of parent node Y is given as 1, the conditional probability that the state of node X is 1 is = 0.7. Then the change in the conditional probability of node X relative to parent node Y is = 0.9 - 0.7 = 0.2. This result indicates that in the first machine, when the state of parent node Y is 1, the probability that the state of node X is 1 is higher than that in the second machine. This shows that the causal relationship behaves differently in the two machines, and the first machine has a stronger dependence on this causal relationship.
[0137] Step 3: Determine the sum of the changes in the conditional probabilities of each node relative to all parent nodes as the change in the causal relationship of each node.
[0138] Specifically, the change in the causal relationship of each node can be determined by the following formula:
[0139]
[0140] where ΔC(X) is the change in the causal relationship of node X, X is the node, Y is the parent node, and Pa(X) is the set of parent nodes of node X. is the change in the conditional probability of node X relative to the parent node Y. ΔC(X) can be used to characterize the change in the causal relationship of node X before and after the fault occurs.
[0141] Through the above solution, by comparing the changes between the first Bayesian network and the second Bayesian network before and after the fault, the propagation path of the fault can be deeply analyzed, and the change in the causal relationship between variables before and after the fault can also be quantified, so as to facilitate inferring the root cause of the fault by analyzing the change in the causal relationship between different variables.
[0142] Step 103: Determine the root cause score of each node based on the change in the marginal probability and the change in the causal relationship of each node, as well as a preset weighting coefficient.
[0143] Specifically, the root cause score can be used to characterize the role of the node in the equipment failure.
[0144] In some examples, the preset weighting coefficient includes a first weight coefficient and a second weight coefficient; the sum of the first weight coefficient and the second weight coefficient is 1. Exemplarily, the first weight coefficient can be 0.3 - 0.4, and the second weight coefficient can be 0.7 - 0.6. For example, the first weight coefficient can be set to 0.3, and the second weight coefficient can be set to 0.7.
[0145] Specifically, the root cause score of each node can be determined by the following formula:
[0146]
[0147] where S(X) is the root cause score of node X, W 1 is the first weight coefficient, W 2 is the second weight coefficient, W1 + W2 = 1; ΔP(X) is the change in the marginal probability of node X, and ΔC(X) is the change in the causal relationship of node X.
[0148] Step 104: Determine the key parameter corresponding to the node with the highest root cause score among all nodes as the root cause parameter of the fault.
[0149] Exemplarily, please continue to refer to Figure 2 and Figure 3, Assume that in the second Bayesian network, the root cause score of node f is the highest, indicating that the key parameter corresponding to node f plays a core role in the equipment failure. By calculating ΔC(f), it can also be found that the causal relationship change of node f before and after the failure is significant, and it is speculated that the key parameter corresponding to node f is the root cause of the abnormality of other key parameters.
[0150] It can be understood that the embodiments of the present application have been put into use in the silicon wafer cutting production line. In actual applications, they have helped the factory successfully identify the root causes of multiple complex failures, and the occurrence frequency of similar failures can be reduced through targeted maintenance. The Bayesian inference method of the embodiments of the present application can deeply analyze the fault propagation path by comparing the changes in the DAG before and after the fault, and explore the root cause of the equipment failure. Specifically, by quantifying the change in the causal relationship between variables before and after the fault, the root cause score can be calculated, and finally the key variable leading to the fault can be located. This root cause localization method based on Bayesian inference can not only shorten the fault troubleshooting time, but also greatly improve the accuracy of fault localization. Compared with the traditional empirical judgment and troubleshooting methods, the method of the embodiments of the present application is more scientific and efficient, providing strong technical support for the intelligent maintenance of the factory.
[0151] Correspondingly, please refer to Figure 4 , Figure 4 which is a schematic structural diagram of the fault root cause determination device of the embodiments of the present application. The fault root cause determination device provided by the embodiments of the present application includes a network construction module 401, a network comparison module 402, a score calculation module 403, and a root cause determination module 404. The network comparison module 402 is connected to the network construction module 401, the score calculation module 403 is connected to the network comparison module 402, and the root cause determination module 404 is connected to the score calculation module 403. Exemplarily, the modules can be communicatively connected to each other.
[0152] The network construction module 401 is configured to construct a first Bayesian network based on the key parameters of the first machine tool and a second Bayesian network based on the key parameters of the second machine tool, where the abnormality degree of the first machine tool meets a preset condition, and the abnormality degree of the second machine tool exceeds the preset condition.
[0153] The network comparison module 402 is configured to determine the marginal probability change amount and the causal relationship change amount of each node based on the first marginal probability and the first conditional probability of each node in the first Bayesian network, and the second marginal probability and the second conditional probability of each node in the second Bayesian network.
[0154] The score calculation module 403 is configured to determine the root cause score of each node based on the marginal probability change amount and the causal relationship change amount of each node, and a preset weighting coefficient.
[0155] The root cause determination module 404 is configured to determine the key parameter corresponding to the node with the highest root cause score among all nodes as the root cause parameter of the fault.
[0156] In some embodiments, the network comparison module 402 is specifically configured to:
[0157] Based on the first marginal probability of each node in the first Bayesian network and the second marginal probability of the corresponding node in the second Bayesian network, determine the marginal probability change amount of each node.
[0158] Based on the first conditional probability of each node in the first Bayesian network relative to any parent node and the second conditional probability of the corresponding node in the second Bayesian network relative to the corresponding parent node, determine the conditional probability change amount of each node relative to the parent node.
[0159] Determine the sum of the conditional probability change amounts of each node relative to all parent nodes as the causal relationship change amount of each node.
[0160] In some embodiments, the network comparison module 402 is specifically configured to:
[0161] Determine the causal relationship change amount of each node through the following formula:
[0162]
[0163] where ΔC(X) is the causal relationship change amount of the node, X is the node, Y is the parent node, Pa(X) is the set of parent nodes of node X, is the conditional probability change amount of node X relative to parent node Y.
[0164] In some embodiments, the preset weighting coefficients include a first weight coefficient and a second weight coefficient. The score calculation module 403 is specifically configured to:
[0165] Determine the root cause score of each node through the following formula:
[0166]
[0167] where S(X) is the root cause score of node X, W 1 is the first weight coefficient, W 2 is the second weight coefficient, W1 + W2 = 1. ΔP(X) is the marginal probability change amount of node X, and ΔC(X) is the causal relationship change amount of node X.
[0168] In some embodiments, the first weight coefficient is 0.3 and the second weight coefficient is 0.7.
[0169] In some embodiments, the network construction module 401 is specifically configured to:
[0170] Obtain the historical data of each key parameter of the first machine tool, and divide the historical data into a training set and a test set.
[0171] Based on the training set, use the expectation-maximization method to determine the structure of the first Bayesian network and the direction of each edge.
[0172] Use the maximum likelihood estimation method to determine the first conditional probability table of each node in the first Bayesian network.
[0173] Verify the first Bayesian network based on the test set, and modify the structure of the first Bayesian network and the direction of each edge according to the verification results to obtain the trained first Bayesian network.
[0174] In some embodiments, the network construction module 401 is further specifically configured to:
[0175] Determine the abnormality rate and the variance of the abnormality rate of each operating machine tool in a preset time period.
[0176] Determine the operating machine tools with an abnormality rate less than the first threshold and a variance of the abnormality rate less than the second threshold as the first machine tools.
[0177] Determine the operating machine tools with an abnormality rate greater than or equal to the first threshold and / or a variance of the abnormality rate greater than or equal to the second threshold as the second machine tools.
[0178] In some embodiments, the network construction module 401 is further specifically configured to:
[0179] Determine the equipment parameters of each key control point of the operating machine tool, where the key control point is used to characterize the equipment points that can affect the process.
[0180] Select key parameters from the equipment parameters of each key control point.
[0181] Next, an exemplary description of the parameter detection method of the embodiments of the present application is given.
[0182] Taking the left tension arm tension, cutting fluid temperature, cutting fluid flow rate, wire mesh speed, and feed speed of the equipment key parameters as examples, the left tension arm tension is detected by setting tension sensors at the upper line end and the lower line segment, and is displayed through the interaction interface of the wire cutting equipment; the cutting fluid temperature is detected by setting a temperature sensor at the outlet of the flow pipeline to detect the temperature of the coolant flowing out, and is displayed through the interaction interface of the wire cutting equipment; the cutting fluid flow rate is detected by a flow meter set on the wire cutting equipment, and is displayed through the interaction interface of the wire cutting equipment; the wire mesh speed and the feed speed are detected by monitoring the rotation speed of the corresponding motor, and are displayed through the interaction interface of the wire cutting equipment.
[0183] Wire cutting equipment usually has a variety of sensors or detection devices integrated inside, and the key parameters of the equipment can be directly displayed on the interaction interface of the wire cutting equipment for easy access.
[0184] It can be understood that the root cause determination device of the embodiments of the present application determines the root cause parameters of the fault by comparing the Bayesian networks of machines with different abnormal degrees, which can not only shorten the fault troubleshooting time and improve the production efficiency of the production line, but also greatly improve the accuracy of fault location.
[0185] Correspondingly, please refer to Figure 5 , Figure 5 , which is a schematic structural diagram of the electronic device of the embodiments of the present application. An electronic device provided by the embodiments of the present application includes a memory 501 and a processor 502. The memory 501 is used to store computer programs. The processor 502 is used to execute the computer programs stored in the memory 501. When the computer programs stored in the memory 501 are executed, the processor 502 executes the root cause determination method of the foregoing embodiments of the present application.
[0186] Correspondingly, the embodiments of the present application further provide a computer-readable storage medium, which stores computer instructions for causing the processor to implement the root cause determination method as described in the foregoing embodiments of the present application when executed.
[0187] The above has introduced in detail a root cause determination method, device and electronic device provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the technical solution and its core idea of the present application; those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for determining a root cause of a fault, characterized in that: include: A first Bayesian network is constructed based on key parameters of a first machine, and a second Bayesian network is constructed based on key parameters of a second machine, wherein the abnormality of the first machine satisfies a preset condition, and the abnormality of the second machine exceeds the preset condition; Determine the change in edge probability and causal relationship of each node based on the first edge probability of each node in the first Bayesian network and the first conditional probability of each node relative to any parent node, and the second edge probability of each node in the second Bayesian network and the second conditional probability of each node relative to the corresponding parent node; Determine the root cause score of each node based on the edge probability change and the causal relationship change of each node and a preset weighting coefficient; The key parameter corresponding to the node with the highest root cause score among all the nodes is determined as the fault root cause parameter.
2. The method for determining the root cause of a fault according to claim 1, characterized in that: The step of determining the edge probability change amount and the causal relationship change amount of each node comprises: Determining an amount of change in the edge probability of each node based on a first edge probability of each node in the first Bayesian network and a second edge probability of a corresponding node in the second Bayesian network; Determine a change in the conditional probability of each node relative to the parent node based on a first conditional probability of each node in the first Bayesian network relative to any parent node and a second conditional probability of the corresponding node in the second Bayesian network relative to the corresponding parent node; The sum of the conditional probability changes of each node relative to all the parent nodes is determined as the causal relationship change of each node.
3. The method for determining the root cause of a fault according to claim 2, characterized in that: The step of determining the causal relationship change amount of each node comprises: The causal relationship change of each node is determined by the following formula: Wherein, ΔC(X) is the change in the causal relationship of the node, X is the node, Y is the parent node, Pa(X) is the parent node set of node X, is the conditional probability change of the node X relative to the parent node Y.
4. The method for determining the root cause of a fault according to claim 1, characterized in that: The preset weighting coefficient includes a first weighting coefficient and a second weighting coefficient; the step of determining the root cause score of each node includes: The root cause score of each of the nodes is determined by the following formula: Wherein, S(X) is the root cause score of node X, W1 is the first weight coefficient, W2 is the second weight coefficient, W1+W2=1; ΔP(X) is the edge probability change of node X, and ΔC(X) is the causal relationship change of node X.
5. The method for determining the root cause of a fault according to claim 4, characterized in that: The first weight coefficient is 0.3, and the second weight coefficient is 0.
7.
6. The method for determining the root cause of a fault according to claim 1, characterized in that: The step of constructing a first Bayesian network based on the key parameters of the first machine includes: Acquire historical data of each of the key parameters of the first machine, and divide the historical data into a training set and a test set; Based on the training set, using the maximum expectation method to determine the structure of the first Bayesian network and the direction of each edge; Determine a first conditional probability table for each node in the first Bayesian network using a maximum likelihood estimation method; The first Bayesian network is verified based on the test set, and the structure of the first Bayesian network and the direction of each edge are corrected according to the verification result to obtain the trained first Bayesian network.
7. The method for determining the root cause of a fault according to claim 1, characterized in that: The first machine and the second machine are determined by the following steps: Determine the abnormality rate and abnormality rate variance of each operating machine in a preset period of time; Determine the operating machine whose abnormality rate is less than the first threshold and whose abnormality rate variance is less than the second threshold as the first machine; The operating machine whose abnormality rate is greater than or equal to the first threshold and / or whose abnormality rate variance is greater than or equal to the second threshold is determined as the second machine.
8. The method for determining the root cause of a fault according to claim 1, characterized in that: The step of determining the key parameters comprises: Determine the equipment parameters of each critical control point of the operating machine, wherein the critical control point is used to characterize the equipment point that can affect the process; The key parameters are selected from the equipment parameters of each of the key control points.
9. A device for determining a root cause of a fault, characterized in that: include: A network construction module, configured to construct a first Bayesian network based on key parameters of a first machine, and to construct a second Bayesian network based on key parameters of a second machine, wherein the abnormality of the first machine satisfies a preset condition, and the abnormality of the second machine exceeds the preset condition; A network comparison module, for determining an edge probability change amount and a causal relationship change amount of each node based on a first edge probability of each node in the first Bayesian network and a first conditional probability of each node relative to any parent node, and a second edge probability of each node in the second Bayesian network and a second conditional probability of each node relative to a corresponding parent node; A score calculation module, used to determine the root cause score of each node based on the edge probability change and the causal relationship change of each node, and a preset weighting coefficient; The root cause determination module is used to determine the key parameter corresponding to the node with the highest root cause score among all the nodes as the fault root cause parameter.
10. An electronic device, characterized in that: The electronic device includes a processor and a memory; the memory stores a computer program, and the processor is used to execute the computer program stored in the memory to implement the fault root cause determination method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Chemical process fault diagnosis method for Bayesian network based on mechanism correlation analysis
CN110766173A
Boiler fault diagnosis method and device
CN111122199A