System fault diagnosis method and device and computer equipment

CN120476360APending Publication Date: 2025-08-12SIEMENS AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380087715.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-01-16
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

In industrial production, equipment failure leads to system paralysis. Existing technologies are difficult to meet the requirements of rapid diagnosis and high performance at the same time. In particular, methods based on natural language processing are difficult to meet the requirements of low time consumption and high performance at the same time. Enterprises want to understand the fault. Frequency of occurrence and common symptoms for better management and control of production.

Method used

Adopt a hierarchical diagnostic tree method to obtain fault descriptions of system and equipment nodes, extract fault features, and match similar fault features to provide solutions; if the match cannot be matched, a second solution will be provided through the manual system; if the number of fault features exceeds Threshold, merge new fault features and solutions into similar fault features and solutions, and achieve system iterative learning.

Benefits of technology

Fast and accurate fault diagnosis and solution provision are achieved, reducing losses. The system is in an iterative learning state and can handle faults more quickly and accurately, improving the efficiency of production management and the accuracy of troubleshooting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120476360A_ABST
    Figure CN120476360A_ABST
Patent Text Reader

Abstract

The invention discloses a system fault diagnosis method and device, computer equipment and a storage medium. Specifically, the invention discloses a system fault diagnosis method, which comprises the following steps: acquiring a system and an equipment node, the system comprising the equipment node; extracting a fault feature from the fault description of the equipment node, and matching the fault feature to a first similar fault feature; and feeding back a first solution of the first similar fault feature. By means of the mode, specific faults of the system can be quickly and accurately checked out, and corresponding solutions can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

System fault diagnosis method, device, and computer equipment Technical Field

[0001] The present application relates to the fields of deep learning and machine learning, and specifically, to a method, apparatus, computer device, and storage medium for system fault diagnosis using machine learning. Background Art

[0002] Throughout industrial production, equipment failures are inevitable. With the rapid advancement of automation and intelligent manufacturing processes, a single equipment failure can paralyze the entire system, severely impacting the company's normal operations and economic profitability. Therefore, in addition to predictive analysis of equipment failures, rapid response and diagnosis after equipment failures are the final line of defense for increasing production and minimizing losses.

[0003] Summary of the Invention

[0004] Based on this, the present application provides a method for system fault diagnosis, including: obtaining a system and a device node, wherein the system includes the device node; extracting fault characteristics from the fault description of the device node and matching it to a first similar fault characteristic; and feeding back a first solution to the first similar fault characteristic.

[0005] Through the above methods, the causes of equipment failures that occur in the company's production process can be quickly and accurately found and corresponding solutions can be provided.

[0006] Furthermore, if the fault feature cannot be matched with a first similar fault feature, a second solution is provided through an artificial system.

[0007] Through the above methods, a second set of solutions can be quickly provided for equipment failure to reduce losses.

[0008] Furthermore, if the number of the fault features that cannot be matched to the first similar fault features exceeds a first threshold, the fault features and the second solution are merged into the first similar fault features and the first solution.

[0009] Through the above methods, the system can be in an iterative and learning process, which makes troubleshooting faster and more accurate.

[0010] Furthermore, the obtaining of the system and the device node includes: obtaining the system, the device node and the fault description of the device node from the system fault description.

[0011] Through the above methods, the source information and traceability information of the fault can be systematically obtained, which makes the fault diagnosis and resolution more comprehensive and clear.

[0012] Furthermore, the method further includes obtaining a historical fault description from historical data; and extracting the first similar fault feature and the first solution from the historical fault description.

[0013] Through the above method, the occurrence and handling of historical faults can be obtained from historical records, which is conducive to system learning and updating, and has the ability to handle similar or similar faults.

[0014] Furthermore, the method further includes performing semantic clustering on the historical fault descriptions based on a topic model to obtain the first similar fault features and the first solution.

[0015] Through the above methods, clustering can be used to better discover common or shared fault characteristics, and thus discover or summarize common or shared solutions.

[0016] Furthermore, the method also includes establishing a correlation between the first similar fault feature and the first solution based on the historical fault description.

[0017] Through the above method, corresponding, accurate and effective solutions can be provided for subsequent failures.

[0018] Furthermore, the method further includes establishing a correlation between the first similar fault feature and the first solution based on a binary classification model.

[0019] Through the above-mentioned method, the relationship between the first similar fault feature and the first solution can be more closely established, so that when the first similar fault is subsequently found, an accurate first solution can be pushed.

[0020] Furthermore, matching the first similar fault feature includes: matching the first similar fault feature through a hierarchical diagnostic tree.

[0021] Through the above method, the first similar fault feature can be matched accurately and quickly.

[0022] The present application also provides a system fault diagnosis device, comprising:

[0023] An acquisition module, configured to acquire a system and a device node, wherein the system includes the device node;

[0024] A matching module, configured to extract a fault feature from the fault description of the device node and match it to a first similar fault feature;

[0025] A feedback module is configured to provide feedback on a first solution to the first similar fault feature.

[0026] The present application also provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and wherein the processor implements the above method when executing the computer program.

[0027] The present application also provides a computer-readable storage medium having a computer program stored thereon, and the computer program implements the above method when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Implementations of the present disclosure are illustrated by way of example and not limitation in the figures of the accompanying drawings in which like references indicate the same or similar parts.

[0029] FIG1 is a flowchart of a method for system fault diagnosis according to an embodiment of the present application.

[0030] FIG2 is a schematic diagram of a system fault diagnosis apparatus according to an embodiment of the present application.

[0031] FIG3 is a schematic diagram of a tree topology structure for system fault diagnosis according to an embodiment of the present application.

[0032] FIG4 is a schematic diagram of a computer device for system fault diagnosis according to an embodiment of the present application.

[0033] The accompanying drawings are numerals as follows:

[0034] Steps S101-S103

[0035] 200: Installation

[0036] 201: Module

[0037] 202: Module

[0038] 203: Module

[0039] 204: Module

[0040] 205: Module

[0041] 206: Module

[0042] 300: Tree topology

[0043] 301: System layer

[0044] 302: Device Node Layer

[0045] 304: First similar fault feature layer

[0046] 305: First solution layer

[0047] 301-1, 301-2: System

[0048] 302-1, 302-2, 302-3: Device nodes

[0049] 304-1, 304-2, 304-3, 304-4: First similar fault characteristics

[0050] 305-1, 305-2, 305-3: First solution

[0051] 400: Computer equipment

[0052] 402: Processor

[0053] 404: Memory DETAILED DESCRIPTION

[0054] In the following description, for the purpose of explanation, a large number of specific details are set forth. However, it is understood that the present invention can be implemented without these specific details. In other examples, well-known circuits, structures, and technologies are not shown in detail so as not to affect the understanding of the description.

[0055] References throughout this specification to "an implementation," "an implementation," "an exemplary implementation," "some implementations," "various implementations," etc., indicate that the implementations of the invention being described may include particular features, structures, or characteristics. However, it does not imply that every implementation must include those particular features, structures, or characteristics. Furthermore, some implementations may have some, all, or none of the features described for other implementations.

[0056] In the following description, the terms "coupled" and "connected" and their derivatives may be used. It should be understood that these terms are not intended to be synonymous with each other. Rather, in certain implementations, "connected" is used to indicate that two or more components are in direct physical or electrical contact with each other, while "coupled" is used to indicate that two or more components cooperate or interact with each other, but they may or may not be in direct physical or electrical contact. FIG1 is a flow chart of a system fault diagnosis method according to an embodiment of the present application.

[0057] It should be understood that although the various steps in the flowchart of FIG1 are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in FIG1 may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0058] Throughout industrial production, equipment failures are inevitable. With the rapid advancement of automation and intelligent manufacturing processes, a single equipment failure can paralyze the entire system, severely impacting the company's normal operations and economic profitability. Therefore, in addition to predictive analysis of equipment failures, rapid response and diagnosis after equipment failures are the final line of defense for increasing production and minimizing losses.

[0059] In addition to relying on sensor data for fault diagnosis, textual data related to the fault (such as fault descriptions, causes, and solutions) also plays an indispensable role in the fault diagnosis process. A reasonable and feasible approach is to build a knowledge base of equipment faults based on accumulated historical fault cases (including descriptions, causes, solutions, and equipment or systems). Then, for a new fault, we can rely on natural language processing (NLP) technology to find similar or identical faults and their solutions from historical cases, thereby guiding the operator in fault diagnosis.

[0060] However, a practical challenge for NLP-based methods is that it's difficult to simultaneously meet the requirements of low time consumption and high performance. Furthermore, in our experience, when faced with equipment failures, companies not only want to quickly find solutions to the problems, but also want to know which equipment (system, production line, unit) is most prone to failure, which types of failures occur most frequently, and which symptoms often occur together, so as to better manage and control production.

[0061] Given these two considerations, most industrial manufacturing companies utilize asset trees or asset hierarchies to distinguish between different equipment, systems, or workshops. Therefore, this paper proposes a method for constructing a hierarchical diagnostic tree. Based on this diagnostic tree, companies can not only perform fault diagnosis but also intuitively understand common statistical information about faults.

[0062] FIG1 provides an embodiment of a method for system fault diagnosis, comprising:

[0063] First, in step S101 , a system and a device node are acquired, wherein the system includes the device node.

[0064] Acquiring the system and device nodes includes acquiring the names or identifiers of the system and device nodes, which may be symbols representing the system or device nodes. There may be multiple systems, each of which includes multiple device nodes. In other words, the multiple device nodes constitute a system. However, a system may also include multiple other device nodes. In this solution, the analysis focuses on device nodes experiencing problems or failures, and irrelevant device nodes may not be acquired.

[0065] Then, in step S102, a fault feature is extracted from the fault description of the device node, and matched to a first similar fault feature.

[0066] The device node fault description can be a textual description that records the device node fault, including, for example, the time, frequency, operator, specific incident or fault problem, etc. This description can be written by a person, generated by the system, or in other ways. In short, it describes the device node fault. The fault description can also be in any other form, such as images, videos, or audio, and can be converted into text format through some technical means or software. It mainly reflects the device node fault description.

[0067] Extracting fault features from the device node's fault description can be done, for example, by segmenting the textual fault description into sentences and distinguishing their meanings to extract key fault feature descriptions. Alternatively, after segmenting the fault description into sentences, the meaning of each sentence can be analyzed, followed by semantic merging and clustering to ultimately obtain a number of keywords. In some embodiments, after obtaining the keywords, one or more sentences describing the fault features are manually summarized.

[0068] Among them, after the fault feature is extracted, the fault feature is matched to the first similar fault feature. For example, the first similar fault feature is an existing, stored, known, organized, learned and summarized fault feature for each device node. After comparing the fault feature just extracted with several first similar fault features, find the most similar, closest in meaning to the fault description, and most consistent with several known first similar fault features. In this way, a process of matching the new fault feature with the existing fault feature is completed, thereby converting, transforming, and re-understanding the fault into several known fault features. For the sake of convenience, this application refers to the known fault feature as the first similar fault feature; and similar means that for the fault feature just extracted, several known or summarized fault features are closest to it, and thus are called first similar fault features.

[0069] The matching of the first similar fault feature is to compare the newly extracted fault feature with the above-mentioned existing fault feature according to the similarity calculation formula. The comparison can be performed by semantics, context, description, etc. to obtain a similarity value. The larger the similarity value, the more similar the two are. After the ergodic comparison, the existing fault feature with the greatest similarity can be selected as the first similar fault feature. There can be multiple first similar fault features obtained by comparison. This application is not limited and can be set according to actual needs.

[0070] Then, in step S103, a first solution to the first similar fault feature is fed back.

[0071] Among them, since the first similar fault feature is known and has been learned, the solution corresponding to the first similar fault feature is also known, so several corresponding first solutions can be fed back based on the first similar fault feature. It is explained here that it is called the first solution because it corresponds to the first similar fault feature. One similar fault feature can have several first solutions, and multiple similar fault features can have multiple overlapping first solutions. Because in reality, it is possible that a solution corresponds to multiple faults, or to multiple fault features. Conversely, multiple fault features can also correspond to one solution, or one fault feature can correspond to multiple solutions. This is not limited here. The first solution may include the cause of the fault, the solution, specific countermeasures and methods, etc. to solve the fault.

[0072] The technical effect of the above scheme is that, by using the above method, the cause of the fault can be quickly and accurately located, or the fault itself can be quickly understood, and then by summarizing the existing faults and solutions, an effective solution can be quickly provided, which is beneficial to production management and troubleshooting.

[0073] Furthermore, if the fault feature cannot be matched to a first similar fault feature, a second solution is provided through an artificial system.

[0074] As described above, after extracting a new fault feature, no similar fault features are found. In other words, after comparing the existing fault features through unrestricted methods such as machine learning or artificial intelligence, it is found that there are no existing fault features that are completely similar or close to the newly extracted fault features. In this case, a second solution is manually provided through other methods, such as through an artificial system. The second solution can be a completely new solution to the newly extracted fault feature, or an existing solution that is applied to the newly extracted fault feature for the first time. The technical effect of this is that it can provide a corresponding solution to the new fault problem, reflecting the robustness, comprehensiveness, and fault adaptability of the method.

[0075] Furthermore, if the number of the fault signatures that cannot be matched to the first similar fault signature exceeds a first threshold, the new fault signature and the second solution corresponding thereto are merged into the first similar fault signature and the first solution.

[0076] Among them, when the newly extracted fault features cannot find the closest or similar known fault features for many times, through a certain amount of accumulation and precipitation, when a certain threshold or threshold is reached, these unmatched new fault features and the second solution are merged, added, and added to the existing first similar fault features and the first solution. In this way, if a similar new fault feature is encountered again in the future, the corresponding similar fault feature can be matched, thereby further obtaining the second solution.

[0077] In some embodiments, the above steps expand the number of first similar fault characteristics and the number of first solutions, as well as the number of corresponding relationships between first similar fault characteristics and first solutions. As described above, the first similar fault characteristics are known and machine-learned fault characteristics from historical cases, and the first solutions are also obtained from historical cases and correspond to solutions that address the first similar fault characteristics. However, in practice, new fault characteristics may appear for the same device, requiring manual processing and resolution to generate new solutions. Therefore, the new fault characteristics can be referred to as second similar fault characteristics, and the solutions corresponding to the new fault characteristics (which may be solutions that have not appeared before, or solutions that have appeared but have not been previously associated with the new fault characteristics) can be referred to as second solutions. These newly summarized second similar fault characteristics and second solutions are then expanded and added to the aforementioned first similar fault characteristics (or the first set of similar fault characteristics) and first solutions (or the first set of solutions). The corresponding relationships between the second similar fault characteristics and the second solutions are also recorded, documented, and retained. The above steps are equivalent to expanding the number of first similar fault signatures (or the sample size within the first similar fault signature set), and corresponding to providing more effective solutions (or the sample size within the first solution set) when solving the problem. This technical effect enables the system to be in an iterative and learning process, constantly updating new fault signatures and solutions, making it faster and more convenient, facilitating the resolution of production problems, and saving time.

[0078] Furthermore, the obtaining of the system and the device node includes: obtaining the system, the device node and the fault description of the device node from the system fault description.

[0079] The aforementioned system and device nodes can be obtained from system fault descriptions. The system fault description can be a general description of a new fault, and the device node fault description is also extracted from the system fault description. This allows for systematic acquisition of fault source information and traceability, making fault diagnosis and resolution more comprehensive and clear.

[0080] Furthermore, historical fault descriptions are obtained from historical data.

[0081] The historical data may be a description of a previous system fault, or a description with other names, or a case. In short, useful and valuable historical fault descriptions can be obtained from the historical data.

[0082] Furthermore, the historical data may be some whole and complete historical cases, and the historical fault description may be a description of some parts thereof in the form of a sentence or a number of clauses.

[0083] Then, the system, device node, first similar fault feature and first solution are extracted from the historical fault description.

[0084] Among them, from the historical fault description, useful information such as the system, device node, first similar fault, and first solution can be extracted by keywords, sentences, and meanings, etc., and the method is not limited. The obtained system, device node, first similar fault, and first solution can be used for machine learning or artificial intelligence learning and training, so as to establish an overall association relationship or tree relationship from the system to the device node, then to the first fault feature, and finally to the solution. The technical effect of this can be a more intuitive, comprehensive, and systematic learning of the overall context from the system to the solution, which is conducive to the analysis of the failure of specific equipment in the production process and the summary of the cause and solution.

[0085] Through the above method, the occurrence and handling of historical faults can be obtained from historical records, which is conducive to system learning and updating, and has the ability to handle similar or similar faults.

[0086] As shown in Figure 3, the tree-like topology 300 includes a system layer 301, a device node layer 302, a first similar fault characteristic layer 304, and a first solution layer 305. Specifically, the system layer 301 includes systems 301-1 and 301-2; the device node layer 302 includes device nodes 302-1, 302-2, and 302-3; the first similar fault characteristic layer 304 includes first similar fault characteristics 304-1, 304-2, 304-3, and 304-4; and the first solution layer 305 includes first solutions 305-1, 305-2, and 305-3. The system layer is connected to the device node layer, then to the first similar fault characteristic layer, and finally to the first solution layer, forming a tree-like, logically connected, and associated four-layer structure, thereby better demonstrating the logical relationships between the various structures and components.

[0087] Furthermore, a topic model can be used to perform unified semantic clustering on historical fault descriptions to obtain the first similar fault characteristics and the first solution. For example, sentences with similar or similar meanings in the historical fault descriptions can be clustered into one, which can better identify common or shared keywords and further form the first similar fault characteristics and the first solution.

[0088] In some embodiments, the systems or devices with different names in the historical fault descriptions can be normalized to unify the names, which can cluster similar systems or devices.

[0089] Furthermore, in some embodiments, the system and device nodes can be obtained through asset tables, asset series, and asset trees. The correlation and connection between systems and devices can usually be obtained through files such as asset tables, asset series, and asset trees, which are not limited in this application.

[0090] The first solution can be obtained, for example, by having one first solution for each historical case, or by having several first solutions for each historical case. The first solution can be an entire description or several clauses. Similarly, the clauses can be semantically aggregated and keywords extracted to obtain a representative sentence description. This application does not limit this.

[0091] In some embodiments, for example, the first similar fault feature and the first solution can be obtained respectively in the following manners, and the extraction process can be as follows:

[0092] 1) Obtain fault descriptions for all historical cases.

[0093] 2) Sentence segmentation can be done by using commas or periods, but is not limited to the above methods.

[0094] 3) For different systems and devices, a topic model is used to perform unified semantic clustering on the fault descriptions of historical cases.

[0095] 4) Obtain a representative sentence in a topic, that is, an atomic symptom.

[0096] For the selection of representative sentences, a semi-automatic method is adopted in one embodiment of the present invention, for example, first obtaining the top N keywords of all sentences in a topic and then performing manual correction.

[0097] The first similar fault feature and the first solution can be obtained respectively through the above methods.

[0098] Furthermore, the “atomic symptom” represents the name of the “representative sentence”, and further, specifically may be the first similar fault feature or the first solution.

[0099] For example, based on the above description, taking a historical case as an example, the fault description is as follows:

[0100] "On October 12, 2021, at 1:30 p.m., an inspection of the primary fan was conducted. During the inspection, it was discovered that the motor sensor of the primary fan had recently indicated a slightly higher-than-normal temperature, posing a risk of bearing burnout. Over the past several months, the primary fan had also been making unusual noises, which had gradually intensified."

[0101] After sentence segmentation, this passage can be divided into three sentences:

[0102] a) At 1:30 p.m. on October 12, 2021, an inspection of the primary fan was carried out.

[0103] b) During the inspection, it was found that the motor sensor of the primary fan recently showed a temperature slightly higher than normal, posing a risk of bearing burnout.

[0104] c) Over a period of several months, the main fans began to make abnormal noises, and the noise gradually increased.

[0105] After completing the sentence segmentation step, all sentences about the wind turbine are collected and then a topic model is created. This way, these sentences are divided into multiple sentence clusters, each of which can be considered to have potential semantic relevance. In some embodiments, the segmentation and collection process can be implemented by the apparatus described herein, or by a specific method, such as a historical data acquisition module within the apparatus, or by a hardware device that executes the method. This application does not limit this.

[0106] Next, for the selection of representative sentences for each sentence group, we first extract keywords for each sentence group (some commonly used keyword extraction techniques can be used), such as "primary fan, tripped...", and then we can manually construct the corresponding first similar fault features based on these keywords, such as "primary fan tripped".

[0107] Furthermore, based on the historical fault description, a correlation between the first similar fault feature and the first solution is established.

[0108] The connection between the first similar fault feature and the first solution can be learned and established through the description of the isolated first similar fault feature and the corresponding first solution. For example, a binary classification model can be used to obtain a correlation score between a fault feature and a solution to indicate the closeness of the relationship between the first similar fault feature and the solution. Through the study and analysis of multiple historical data, if the first similar fault feature and the solution always appear in correspondence, the relationship value between them will become increasingly higher, indicating that their relationship is becoming increasingly close, direct, causal, and related.

[0109] In some embodiments, for example, one method of connecting (correlating) the first similar fault signature with the first solution may be as follows:

[0110] 1) Obtain descriptions and solution data for all historical cases.

[0111] 2) Perform sentence segmentation to form a set of sentences describing and solving problems.

[0112] 3) Construct positive and negative sentence pairs for the sentence correction model.

[0113] For example, in a positive sentence pair: s1 and s2 come from the same case, and s1 comes from the description, while s2 comes from the solution.

[0114] Negative sentence pairs: s1 and s2 come from different cases, s1 comes from the description of one case, and s2 comes from the solution of another case.

[0115] 4) Based on the above positive sentence pairs and negative sentence pairs, perform binary classification modeling.

[0116] 5) Based on the above two-classification model, the predicted positive value can be used as the correlation score between the two sentences, that is, the correlation score of a sentence pair.

[0117] 6) The correlation score between a symptom node (or the first similar fault feature) and the solution (N sentences) should be Max(model(s0,si)).

[0118] Where s0 is the symptom sentence (or the first similar fault feature), and si is one of the N sentences from a solution.

[0119] In some embodiments, the above-mentioned "historical cases" can be historical data, and the above-mentioned "historical case descriptions and solution data" can be historical fault descriptions, and the "sentence segmentation to form a set of sentences of descriptions and solutions" can be aggregated and keywords can be extracted to ultimately form the first similar fault features and the first solution.

[0120] The correlation score or score established between "s1 from the description" and "s2 from the solution" will ultimately be aggregated into the correlation score or score between the first similar fault feature and the first solution. For example, taking the construction of positive and negative sentence pairs as an example, there are two cases.

[0121] Historical case c1 includes:

[0122] Description d1: At 1:30 pm on October 12, 2021, the main fan was inspected.

[0123] Description d2: During the inspection, it was found that the motor sensor of the main fan recently showed a temperature slightly higher than normal, posing a risk of bearing burnout.

[0124] Solution S1: Change the bearing end cover on the non-load side (rear bearing side) of the motor to an insulating end cover to prevent shaft current.

[0125] Historical case c2 includes:

[0126] Description d1: The boiler was put into operation on October 16, 1994. As of March 13, 2007, it had accumulated 81,443 operating hours, during which the medium-temperature reheater had burst four times.

[0127] Solution s1: Replace all the medium-temperature reheaters from the upper part of the tube bank clamping tube to the elbow at the bottom of the tube sheet with φ60*4 and T91 tubes, and do not replace other parts.

[0128] For simplicity, we abbreviate “description d*” to d*, and the same applies to “solution s*” to s*. Therefore, we can construct multiple positive examples such as (c1_d1,c1_s1), (c1_d2,c1_s1) and (c2_d1,c2_s1).

[0129] Then the negative examples can be (c1_d1,c2_s1), (c1_d2,c2_s1), (c2_d1,c1_s1).

[0130] The positive examples mentioned above are the positive sentence pairs mentioned earlier, and the negative examples are the negative sentence pairs mentioned earlier.

[0131] It should be added that, in some embodiments, the above-mentioned “symptoms” may also mean the first similar fault characteristics.

[0132] Then a binary classification model is trained, so the correlation between the two sentences can be represented by a predicted positive value. Furthermore, in general, the binary classification model will predict the correlation score of the positive sentence pair or a relatively high score, or a positive score. Finally, the correlation score between a first similar fault feature s0 and a solution can be the maximum value of the correlation between s0 and all the sentences in the solution. For example, a solution has three sentences s1, s2 and s3, then we can get the correlation between s0 and these three sentences according to the model, for example, the score (s0, s1) is 0.3, the score (s0, s2) is 0.77, and the score (s0, s3) is 0.6. Then the final correlation score between s0 and the solution is the maximum 0.77.

[0133] In some embodiments, it needs to be explained that, among different cases, if the description in one case (which can be called the description before the first similar fault feature is fitted) and the solution in another case are used to establish a relationship of a binary classification model, the result will usually be negative, or in other words, such a description (which can be called the description before the first similar fault feature is fitted) does not have a particularly large correlation with the solution, so it is called a negative example.

[0134] Among them, in the same case, if the description in one case (which can be called the description before the first similar fault feature is fitted) and the solution in this case are used to establish the relationship of the binary classification model, the result will usually be positive, or in other words, such a description (which can be called the description before the first similar fault feature is fitted) has a particularly large correlation with the solution, so it is called a positive example.

[0135] Among them, through the analysis of multiple positive examples, a closer and more prominent connection can be gradually established between the above descriptions and solutions, thereby establishing a relationship between a description and several solutions. However, these relationships can be used to analyze and use new descriptions when encountering them in the future, so as to find solutions for the new descriptions.

[0136] Similarly, the above relationship can also be used for analysis and use when encountering new fault features in the future, so as to find a matching first similar fault feature for the new fault feature, thereby providing a first solution.

[0137] The relationship between the system and the device can usually be obtained through files such as asset tables, asset series, and asset trees, and this application does not limit this. Of course, the system and device nodes can also be obtained through historical data (historical cases), and the relationship between the system and device nodes can also be obtained from historical data (historical cases), and this invention does not limit this.

[0138] Furthermore, the relationship between the device node and the first similar fault feature can be established by obtaining it through files such as the above-mentioned asset table, asset series, asset tree, etc., or it can be gradually established through the above-mentioned machine learning and positive example learning from historical cases, and this application does not limit it.

[0139] Furthermore, matching the first similar fault feature includes: matching the first similar fault feature through a hierarchical diagnostic tree.

[0140] As described above and shown in FIG3 , after obtaining the relationship between the system, device node, first similar fault feature, and first solution, a hierarchical diagnostic tree can be established. In the hierarchical diagnostic tree, the system, device node, first similar fault feature, and first solution are divided into four layers, and each layer is connected in this way. For example, the system layer is connected to the device node layer, the device node layer is connected to the first similar fault feature layer, and the first similar fault feature layer is connected to the first solution layer. Thus, the four layers are connected in sequence to form a tree-like form, which is called a hierarchical diagnostic tree. The technical effect of this is that it is clear which devices often fail or which first similar fault features often occur, so as to make better decisions, perform preventive maintenance in advance, find out the causes, and provide solutions. In addition, according to the first solution of the above-mentioned hierarchical diagnostic tree, from back to front, or from top to bottom, it can be obtained which first similar fault features often appear at the same time, so as to conduct joint troubleshooting or diagnosis of faults.

[0141] The following examples illustrate some system and equipment failures that occur in actual situations and their corresponding solutions.

[0142] In one embodiment, the system may include a combustion system, and the device node may include two device nodes: a primary fan and an induced draft fan.

[0143] For the device node of a primary wind turbine, the first similar fault feature may include two fault features: a primary wind turbine trip and a primary wind turbine surge.

[0144] For the first similar fault characteristics of a primary fan tripping, there are four first solutions, such as: (1) Add ventilation and purification equipment to the inverter compartment, strengthen the inspection and treatment of the internal cleanliness of the inverter equipment components, and ensure a good operating environment for the inverter. (2) The inspection of the lead connectors and supporting porcelain bottles of the motor should be listed as an important inspection item for the planned maintenance of the motor. If necessary, the performance of the supporting porcelain bottles can be judged by conducting a separate pressure test on the supporting porcelain bottles. (3) Seek better quality supporting porcelain bottles and update and transform existing similar equipment. (4) When monitoring the panel, the operating personnel should detect abnormalities as early as possible, carefully check the relevant alarm signals, and promptly contact maintenance personnel to assist in handling the situation on site to avoid the expansion of the accident.

[0145] For the first similar fault characteristics of primary fan surge, four first solutions can be included, for example, (1) reduce the system air pressure setting value to no more than 8.0kPa to reduce the loss of fan throttling adjustment. (2) monitor the changes in air pressure and air volume. If the fan current drops or the air pressure fluctuates, the adjustment damper should be switched to manual in time for manual intervention. (3) When starting and stopping the coal mill, slowly operate the cold / hot primary air adjustment damper to avoid large fluctuations in the primary air pressure. (4) Monitor the coal mill mixed primary air pressure to prevent the mill from being blocked due to too low air pressure.

[0146] For the device node of the induced draft fan, a first similar fault feature of a broken induced draft fan blade may be included.

[0147] For the first similar fault characteristic of the induced draft fan blade fracture, there are four first solutions, for example, (1) Purchase a new impeller and replace the 1A induced draft fan impeller at an appropriate time. Consult the manufacturer to see if the impeller structure and material can be modified. (2) Use the downtime opportunity to check all induced draft fan impeller blades. Take the opportunity to effectively deal with the local blockage of the 1A air preheater. (3) Communicate with the manufacturer to optimize the material and structure design of the core shaft and slider, and make modifications during the B and C repairs. (4) Communicate with the manufacturer and the Institute of Electrical Engineering to optimize the operation adjustment method and reduce the resistance of the blade adjustment.

[0148] In another embodiment, the system may include a steam-water system, and the device nodes may include four device nodes of a feedwater pump, a reheater, a heater, and a condensate pump.

[0149] For the equipment node of the water supply pump, two first similar fault characteristics may be included: a decrease in the output of the water supply pump and an abnormality in the lubrication and cooling system of the water supply pump.

[0150] For the first similar fault characteristic of a drop in feedwater pump output, two primary solutions can be included, such as: (1) Strengthen the management of flow transmitters such as steam pumps and regularly vent them to avoid pump tripping due to abnormal flow. Executed by: Maintenance Department. (2) Strengthen monitoring of steam pump flow and, if an abnormality is detected, promptly notify thermal control for resolution. Executed by: Power Generation Department.

[0151] For the first similar fault characteristic of abnormal lubrication and cooling system of feed water pump, there are three solutions, such as: (1) During the unit outage, after the feed water pump stops operating, the operator should monitor the temperature drop of each pump component, stop the sealing cooling water in time, and strengthen the inspection of cooling water inlet and return water to prevent water from entering the bearing from the oil stop. (2) Strictly implement the rules and regulations. Before stopping the oil circulation of the equipment, the oil system cooling water should be isolated first. (3) The equipment should be carefully inspected before starting, especially when the equipment has been stopped for a long time, the oil level and oil quality of the oil tank should be checked in particular.

[0152] For the equipment node of the reheater, a first similar fault feature of a high-temperature reheater tube burst may be included.

[0153] According to the fault characteristics of high-temperature reheater tube burst, there are five first-line solutions, for example, (1) send the burst elbow to a third-party scientific research institute for failure analysis. (2) further expand the inspection and replace all unqualified tubes. (3) do a good job of inspecting the four tubes of the boiler every time it is shut down, and formulate detailed inspection items and reward and punishment rules for the four tubes to prevent wear and explosion. (4) Harbin Boiler Plant will calculate all high-temperature reheater tube screens, readjust them and implement them during shutdown and maintenance to reduce the high temperature point. (5) make a detailed list of tubes with high wall temperature in the high-temperature reheater, and take advantage of the shutdown and maintenance to conduct radiographic inspections on all the throttling short tubes of these tubes at the inlet header. If any abnormality is found, arrange for corresponding treatment immediately.

[0154] For a device node of a heater, a fault feature indicating abnormal heater liquid level may be included.

[0155] For the first similar fault characteristic of the abnormal heater liquid level, there are three first solutions, such as: (1) Learn from this and thoroughly check whether other filter pressure reducing valves are installed in the wrong direction. If so, promptly handle and restore them while ensuring safety. (2) Refine the division of labor and boundaries of the compressed air system. (3) Upgrade and modify the spring of the high-pressure emergency drain regulating valve actuator.

[0156] For the device node of the condensate pump, two first similar fault characteristics may be included: the condensate pump current drops to low, tripping, and the condensate pump outlet flow rate decreases.

[0157] For the first similar fault characteristic of the condensate pump tripping due to low current, there are five first solutions, such as: (1) To prevent the second-phase condensate pump inverter door limit switch from tripping due to insufficient switch travel or external vibration, which leads to unreliable contact and causes protection error, causing the second-phase inverter condensate pump to trip, thereby threatening the safe operation of the unit, after consultation with the electrical professionals of the technical department and approval by the company leaders, the inverter protection function of the #31 and #42 condensate pump inverter door limit switch will be released. (2) A conspicuous sign "Equipment is running, it is strictly forbidden to open the cabinet door" is posted on the doors of the transformer cabinet and power cabinet of the #31 and #42 condensate pump inverter. (3) During the operation of the inverter, the inverter room is kept locked, and the management of the inverter room key is strengthened. (4) The on-site civilized construction and safe construction in the desulfurization and denitrification construction areas need to be improved. (5) In view of the large amount of coal stored in the coal yard in summer and the coming rainy season, technical measures and safety measures for coal yard management should be implemented, and emergency plans for spontaneous combustion and flood prevention in the coal yard should be formulated. Relevant units and personnel should be organized to study, familiarize themselves with and conscientiously implement them.

[0158] For the first similar fault characteristic of reduced condensate pump outlet flow, there are three first solutions, such as: (1) Introduce a source of pressurized condensate (DN50 pipe) from the condensate pump outlet to the condensate pump inlet to prevent condensate pump vaporization and air leakage. (2) After cleaning the filter, strictly follow the reinstallation process requirements, replace the filter cover gasket in time, and tighten it tightly. (3) After checking that the condensate quality is qualified and there are no large metal particles, remove the condensate pump inlet filter core.

[0159] In some embodiments, for a certain first similar fault feature, the correlation scores or points of several sentences in several solutions are counted, and then the sentences of several solutions with the largest correlation scores or points are determined as the corresponding first solution.

[0160] In some embodiments, the present invention does not limit the expression form of the above-mentioned solution. The above-mentioned solution can be a sentence, or a complete paragraph composed of multiple sentences, or other forms of representative descriptions or displays, which can all serve as representations of the solution.

[0161] By applying the methods provided in this application to the two aforementioned embodiments, it is possible to establish connections between systems, devices, fault characteristics, and solutions. If similar fault characteristics are encountered later, corresponding solutions can be quickly provided, greatly improving the efficiency, time, and accuracy of system fault troubleshooting and diagnosis.

[0162] In some embodiments, no distinction is made between system failures and device failures. A system is composed of multiple devices, and the characteristics of a system failure can be attributed to the failure of several devices. This application is not strictly limited to this, and the subject matter, summary, and main ideas of this application are protected by this application.

[0163] Figure 2 provides an embodiment of an apparatus 200 for system fault diagnosis. The apparatus 200 includes: an acquisition module 201 for acquiring a system and a device node, wherein the system includes the device node; a matching module 202 for extracting a fault signature from the fault description of the device node and matching it to a first similar fault signature; and a feedback module 203 for providing feedback on a first solution to the first similar fault signature.

[0164] Furthermore, the method further includes a manual solution module 204 for providing a second solution through a manual system if the fault feature cannot be matched with the first similar fault feature.

[0165] Furthermore, it includes a merging module 205, which is used to merge and cluster the fault feature and the second solution into the first similar fault feature and the first solution if the number of the fault features that cannot be matched to the first similar fault feature exceeds a first threshold.

[0166] Furthermore, the above-mentioned merging means adding and adding the newly emerged fault characteristics and the second solution, as well as the relationship between them, to the first similar fault characteristics (or the first similar fault characteristics set) and the first solution (first solution set). The relationship between them is also recorded, recorded, and retained in the corresponding merging module, so as to be further used in subsequent fault judgment and resolution.

[0167] The meaning of the merger includes but is not limited to: clustering, adding, merging, merging, and integrating, etc. Its main meaning is to expand and enrich the first similar fault characteristics and the first solution, so as to provide a basis, reference, and judgment for a more comprehensive and comprehensive fault.

[0168] Furthermore, the acquisition module 201 is configured to acquire the system, the device node, and the device node fault description from the system fault description.

[0169] Furthermore, the method further includes a historical data acquisition module 206 for acquiring historical fault descriptions from historical data; and extracting the first similar fault feature and the first solution from the historical fault descriptions.

[0170] Furthermore, the historical data acquisition module 206 is configured to perform semantic clustering on the historical fault descriptions according to a topic model to obtain the first similar fault features and the first solution.

[0171] Furthermore, the historical data acquisition module 206 is configured to establish a correlation between the first similar fault feature and the first solution based on the historical fault description.

[0172] Furthermore, the historical data acquisition module 206 is configured to establish a correlation between the first similar fault feature and the first solution based on a binary classification model.

[0173] Furthermore, the matching module 202 is configured to match a first similar fault feature through a hierarchical diagnostic tree.

[0174] It should be noted that the apparatus may include more or fewer modules to implement the described functionality. For example, at least one module in FIG. 2 may be further divided into a plurality of different submodules, each of which is configured to perform at least a portion of the operations described herein in conjunction with the corresponding module. Furthermore, in some examples, the apparatus 200 may further include additional modules for performing other operations already described in the specification. Furthermore, those skilled in the art will appreciate that the exemplary apparatus 200 may be implemented using software, hardware, firmware, or any combination thereof.

[0175] Figure 4 provides a computer device for system fault diagnosis. According to one embodiment, the computer device 400 may include a processor 402, and the processor 302 executes a computer program stored in the memory 404. When the computer program is executed by the processor, a method for system fault diagnosis is implemented. Furthermore, the computer program can be stored and run in the cloud to perform the method. Furthermore, the components of the program can be deployed on multiple devices and the cloud. For example, the step of obtaining the system and device nodes can be deployed and run on a local or local computer, the step of extracting the fault characteristics from the fault description of the device node and matching the first similar fault characteristics can be deployed and run in the cloud, and the step of feeding back the first solution to the first similar fault characteristics can be deployed and run on the same cloud device, or on different cloud devices, transmitting signals via a communication connection, or can be deployed and run on a local or local computer. This application does not limit the described methods or approaches, and the corresponding technologies can be flexibly deployed to fully utilize the cloud, big data, supercomputing capabilities and other equipment and technologies to execute and complete the method.

[0176] Those skilled in the art will understand that the structure shown in Figure 4 is merely a block diagram of a portion of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0177] Those skilled in the art will appreciate that all or part of the processes in the methods for implementing the above embodiments can be accomplished by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0178] The present application also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the above steps when executed by a processor.

[0179] Some implementations of the present disclosure may include articles of manufacture. Articles of manufacture may include storage media for storing logic. Examples of storage media may include one or more types of computer-readable storage media capable of storing electronic data, including volatile or non-volatile memory, removable or non-removable memory, erasable or non-erasable memory, writable or rewritable memory, and the like. Examples of logic may include various software units, such as software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, processes, software interfaces, application program interfaces (APIs), instruction sets, computing codes, computer codes, code segments, computer code segments, words, values, symbols, or any combination thereof. In some implementations, for example, articles of manufacture may store executable computer program instructions that, when executed by a processor, cause the processor to perform the methods and / or operations described herein. Executable computer program instructions may include any suitable type of code, such as source code, compiled code, interpreted code, executable code, static code, dynamic code, and the like. Executable computer program instructions can be implemented according to a predefined computer language, method or syntax for commanding a computer to perform a specific function. The instructions can be implemented using any appropriate high-level, low-level, object-oriented, visual, compiled and / or interpreted programming language.

[0180] What has been described above includes examples of the disclosed architecture. It is, of course, not possible to describe every conceivable combination of components and / or methodologies, but those skilled in the art will appreciate that many other combinations and permutations are possible. Therefore, the novel architecture is intended to embrace all such alternatives, modifications, and variations that fall within the spirit and scope of the appended claims.

Claims

1. A method for system fault diagnosis, characterized in that: Obtain a system and device nodes, where the system includes the device nodes; Extract fault features from the fault descriptions of the device nodes and match the first similar fault features; Feedback the first solution for the first similar fault features.

2. The method according to claim 1, characterized in that It further includes: If the fault features cannot match the first similar fault features, provide a second solution through an artificial system.

3. The method according to claim 2, wherein It further includes: If the number of fault features that cannot match the first similar fault features exceeds a first threshold, Then merge the fault features and the second solution into the first similar fault features and the first solution.

4. The method according to claim 1, wherein The obtaining of the system and device nodes includes: Obtain the system, device nodes, and fault descriptions of the device nodes from the system fault description.

5. The method according to claim 1, wherein It further includes: Obtain historical fault descriptions from historical data; Extract the first similar fault features and the first solution from the historical fault descriptions.

6. The method according to claim 5, characterized in that, It further includes: According to the topic model, perform semantic clustering on the historical fault descriptions to obtain the first similar fault features and the first solution.

7. The method according to claim 5, characterized in that It further includes: Establish the correlation between the first similar fault features and the first solution according to the historical fault descriptions.

8. The method according to claim 7, wherein It further includes: Establish the correlation between the first similar fault features and the first solution according to the binary classification model.

9. The method according to claim 1, wherein The matching of the first similar fault features includes: Match the first similar fault features through a hierarchical diagnostic tree.

10. A system fault diagnosis device (200), characterized in that: An obtaining module (201) for obtaining a system and device nodes, where the system includes the device nodes; A matching module (202) for extracting fault features from the fault descriptions of the device nodes and matching the first similar fault features; A feedback module (203) for feedbacking the first solution for the first similar fault features.

11. The device (200) according to claim 10, characterized in that, It further includes: An artificial solution module (204) for providing a second solution through an artificial system if the fault features cannot match the first similar fault features.

12. The device (200) according to claim 11, characterized in that, It further includes: A merging module (205) for merging the fault features and the second solution into the first similar fault features and the first solution if the number of fault features that cannot match the first similar fault features exceeds a first threshold.

13. The device (200) according to claim 10, characterized in that, The obtaining module (201) is used for: Obtain the system, device nodes, and fault descriptions of the device nodes from the system fault description.

14. The device (200) according to claim 10, wherein, It further includes that a historical data obtaining module (206) is used for obtaining historical fault descriptions from historical data; And For extracting the first similar fault features and the first solution from the historical fault descriptions.

15. The device (200) according to claim 14, wherein, The historical data obtaining module (206) is used for: According to the topic model, perform semantic clustering on the historical fault descriptions to obtain the first similar fault features and the first solution.

16. The device (200) according to claim 14, characterized in that, The historical data obtaining module (206) is used for: Establish the correlation between the first similar fault features and the first solution according to the historical fault descriptions.

17. The device (200) according to claim 16, characterized in that, The historical data acquisition module (206) is configured to, Establish the correlation between the first similar fault feature and the first solution according to the binary classification model.

18. The device (200) according to claim 10, characterized in that, The matching module (202) is configured to, Match the first similar fault feature through the hierarchical diagnosis tree.

19. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 9 are implemented.

20. A computer-readable storage medium, having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 9 are implemented.

21. A computer program product, the computer program product being tangibly stored on a computer-readable medium and comprising computer-executable instructions, the computer-executable instructions, when executed, causing at least one processor to execute the method according to any one of claims 1-9.