Equipment fault root cause positioning method

Through the visual operation and maintenance platform developed by low-code orchestration technology, the imaged log fingerprint model and fault tree management are used to solve the problem of equipment fault location lag, and the rapid and accurate fault root cause positioning is achieved, which improves operation and maintenance efficiency.

CN120281626APending Publication Date: 2025-07-08HENAN KUNLUN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510399680.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the prior art, there is a lag in the positioning of equipment failures, which leads to a great impact on business, making it difficult to quickly locate faults and reduce the default situation of service level agreements.

Method used

The visual operation and maintenance platform is developed using low-code orchestration technology, providing multiple imaged log fingerprint models, users can build fault modes through interface operations and diagnose whether equipment failures will be caused, and combine fault tree management to quickly locate the root cause.

Benefits of technology

It improves the efficiency and accuracy of fault diagnosis, realizes the rapid positioning of the root causes of equipment failures, and reduces the lag of operation and maintenance work.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281626A_ABST
    Figure CN120281626A_ABST
Patent Text Reader

Abstract

The invention provides an equipment fault root cause positioning method, which is applied to computing equipment and comprises the following steps: receiving a first operation of a user interface provided by a user on the computing equipment; the first operation indicates a target fault mode; the target fault mode comprises a plurality of target log fingerprint models and a combination relationship among the models; the target log fingerprint model is used for matching a target log fingerprint; receiving a second operation of the user on the user interface to obtain an equipment running log; in response to a trigger operation of a user on the user interface, matching a target log fingerprint with a combination relationship from the equipment operation log according to the target fault mode; the matching result is used for root cause positioning of the target fault. Through the mode, the target fault can be diagnosed by effectively utilizing the equipment operation log and the target fault mode, so that the root cause positioning efficiency of the equipment fault is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of fault root cause location, and particularly to a method for locating the root cause of equipment faults. Background Art

[0002] With the continuous progress of big data and cloud computing technologies, the services supported by data centers have become more complex, and the corresponding equipment and its network structure have also become more complex. Currently, most equipment operation and maintenance work is carried out only after a fault occurs. This method has obvious lag, and problems can usually only be discovered after the business is affected, resulting in greater impacts. In this case, how to quickly locate faults and reduce violations of the service level agreement (SLA) has become the goal that customers continuously pursue and is also the biggest challenge faced by operation and maintenance work. Summary of the Invention

[0003] The embodiments of this application provide a method for locating the root cause of equipment faults and a computing device, which can improve the efficiency of locating the root cause of equipment faults.

[0004] In a first aspect, the embodiments of this application provide a method for locating the root cause of equipment faults, which is applied to a computing device. The method includes: receiving a first operation of a user on a user interface provided by the computing device; the first operation indicating a target fault mode; the target fault mode including a plurality of target log fingerprint models and the combination relationship between each model; the target log fingerprint model being used to match a target log fingerprint; receiving a second operation of the user on the user interface to obtain the device operation log; the second operation indicating the acquisition path of the device operation log, and the device operation log being the log collected when the device has a target fault; in response to the trigger operation of the user on the user interface, matching the target log fingerprints with a combination relationship from the device operation log according to the target fault mode; the matching result being used for locating the root cause of the target fault.

[0005] Thus, a visual operation and maintenance platform for locating the root cause of equipment faults is implemented. The platform provides a plurality of graphical log fingerprint models for users to construct different fault modes and diagnose whether these fault modes will cause equipment faults. In this way, the efficiency and accuracy of fault diagnosis can be improved, thereby realizing the rapid location of the root cause of equipment faults.

[0006] In a possible implementation manner, the target fault corresponds to multiple root causes, and the target fault mode causes at least one of the multiple root causes to occur; the method further includes: if the matching result is a successful match, then at least one of the root causes is the root cause of the target fault.

[0007] Thus, it can help users quickly locate and analyze fault information in the log, and improve the efficiency of post-fault diagnosis.

[0008] In a possible implementation, the first operation includes: copying multiple target tags provided by the user interface to the canvas area of the user interface; the target tags correspond one-to-one with the target log fingerprint models; and establishing a combination relationship between the models based on the multiple target tags in the canvas area to obtain a target failure mode.

[0009] In a possible implementation, establishing a combination relationship between the models based on the multiple target tags in the canvas area includes: setting the input-output relationship between the associated fingerprint models in the multiple target log fingerprint models based on the multiple target tags to obtain several independent sub-failure modes; setting the logical operation relationship between the outputs of the sub-failure modes based on the multiple target tags to establish a combination relationship; the logical operation relationship includes at least one of the following, AND, OR, NOT, EQUAL.

[0010] Thus, according to the user's execution of the first operation, a target failure mode can be constructed in real time.

[0011] In a possible implementation, the first operation includes: importing a target file provided by the computing device; the canvas area of the user interface is used to render the target file to obtain a target failure mode.

[0012] In a possible implementation, the target file is obtained by a large model analyzing multiple historical log files, and the target failure mode is a failure mode that appears more than a preset number of times in the multiple historical log files.

[0013] Thus, according to the user's execution of the first operation, a target failure mode can be obtained by importing a pattern file generated by a large model in the past.

[0014] In a possible implementation, the computing device also manages a fault tree, and the fault tree includes a fault node corresponding to the target fault; the method further includes: if the matching result is a successful match; adding at least one root cause caused by the target failure mode as a child node of the fault node respectively.

[0015] Thus, by managing and maintaining the fault tree, the requirements of multiple root cause location scenarios for a single fault phenomenon can be met.

[0016] In a possible implementation, before receiving the first operation of the user on the user interface provided by the computing device, it further includes: establishing target tags according to the log fingerprints in the historical log files; setting matching rules and a unified input-output format for the target tags to obtain the target log fingerprint models corresponding to the target tags.

[0017] Thus, the user can construct the log fingerprint model corresponding to the tag by operating on the tag.

[0018] In one possible implementation, the target log fingerprint model further includes associated fault information, which is used for a user to determine the target fault mode.

[0019] Thus, the user can determine the target fault mode according to the associated fault information.

[0020] In a second aspect, an embodiment of the present application provides a computing device, including:

[0021] A plurality of memories for storing programs;

[0022] A plurality of processors for executing the programs stored in the memories. When the programs stored in the memories are executed, the processors are used to execute the method described in the first aspect or any possible implementation manner of the first aspect.

[0023] In a third aspect, an embodiment of the present application provides a computer storage medium, in which instructions are stored. When the instructions run on a computer, the computer is caused to execute the method described in the first aspect or any possible implementation manner of the first aspect.

[0024] In a fourth aspect, an embodiment of the present application provides a computer program product containing instructions. When the instructions run on a computer, the computer is caused to execute the method described in the first aspect or any possible implementation manner of the first aspect.

[0025] It can be understood that the beneficial effects of the above second aspect to the fourth aspect can refer to the relevant descriptions in the above first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0027] Figure 1 It is an architecture diagram of a computing device for root cause location of device faults provided by an embodiment of the present application;

[0028] Figure 2 It is a flowchart of a method for root cause location of device faults provided by an embodiment of the present application;

[0029] Figure 3 It is a flowchart of constructing a log fingerprint model provided by an embodiment of the present application;

[0030] Figure 4 It is a schematic diagram of constructing a log fingerprint model provided by an embodiment of the present application;

[0031] Figure 5 A flowchart for locating the root cause of device failures provided by an embodiment of the present application;

[0032] Figure 6 A schematic diagram for constructing a failure mode provided by an embodiment of the present application;

[0033] Figure 7 A schematic diagram for locating the root cause of device failures provided by an embodiment of the present application;

[0034] Figure 8 A schematic diagram of a fault tree provided by an embodiment of the present application;

[0035] Figure 9 A schematic diagram of the structure of a computing device provided by an embodiment of the present application. Detailed implementation manners

[0036] In order to make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings.

[0037] In the description of the embodiments of the present application, any embodiment or design solution with "exemplary", "for example", or "for instance" should not be understood as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, the use of words such as "exemplary", "for example", or "for instance" is intended to present relevant concepts in a specific manner.

[0038] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "include", "comprise", "have" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways. "A plurality" may be one or more, where a plurality is two or more.

[0039] In the field of device failure location, a log fingerprint is a unique identifier that can be constructed by analyzing the feature information in log data. These feature information usually include error codes, timestamps, user IDs, etc., and can represent specific system behaviors or events. The main function of log fingerprints is to help automatically extract useful information from a large amount of log data and perform rapid classification and identification, thereby improving the efficiency and accuracy of event response.

[0040] The low-code orchestration technology provides a visual approach to application development, which can accelerate and simplify the application development process, be equally applicable to both small teams and large complex projects, and support reuse on multiple occasions after one-time development. Based on this technology, developers do not need to rely on traditional manual coding methods, but can complete the application development tasks in a more efficient way by dragging and dropping components and setting parameters through a graphical interface.

[0041] The embodiment of this application uses the low-code orchestration technology to implement a visual operation and maintenance platform for device fault root cause location. This platform provides multiple graphical log fingerprint models for users to build different fault modes and diagnose whether these fault modes will cause device failures. In this way, the efficiency and accuracy of fault diagnosis can be improved, thereby achieving rapid location of the root cause of device failures.

[0042] Exemplarily, Figure 1 Figure 1 shows an architecture diagram of a computing device for device fault root cause location provided by the embodiment of this application.

[0043] As Figure 1 shown, the computing device 10 is implemented using the low-code orchestration technology and can be embodied as any computing unit, server, device, device cluster, etc. with processing and computing functions. The computing device 10 includes a log processing module 100, a model management module 110, a mode management module 120, and a fault management module 130.

[0044] Among them, the log processing module 100 includes the following components:

[0045] Log collection unit 101: Used to collect log data from various sources for further processing and analysis.

[0046] Log access unit 102: Used to access external log data into the system for unified management and processing.

[0047] Log persistence unit 103: Used to ensure that log data can be stored for a long time for subsequent query and analysis. For example, decompressing a log compression package, reading a file, and storing the data in a database.

[0048] Log visualization unit 104: Used to display log data in a graphical way for easy user understanding and analysis. For example, establishing a full-text index of log data, providing a query display function, and supporting multi-file switching.

[0049] The model management module 110 includes the following components:

[0050] Tag management unit 111: Used to generate and manage tags for log analysis, and these tags can help identify and classify log entries.

[0051] Matching rule unit 112: Used to define rules for matching log entries to automatically identify and classify logs, including keyword matching, regular expression matching, etc.

[0052] Classification management unit 113: Used to classify newly added tags according to fault phenomena for quick retrieval when constructing fault modes later.

[0053] Model management unit 114: Used to unify the input and output formats of tags, uniformly model the tags, and obtain the corresponding log fingerprint model.

[0054] The mode management module 120 includes the following components:

[0055] Mode management unit 121: Used to manage multiple sub-fault modes that can be executed end-to-end in the canvas.

[0056] Grouping management unit 122: Used to manage the serial execution of multiple feature tags during complex feature matching in the fault mode.

[0057] Logical operation unit 123: Used to manage the logical operation execution of the operation results of different groups and different feature tags, including matching judgment, equality judgment, size comparison, etc.

[0058] Diagnosis execution unit 124: Used to quickly diagnose the newly constructed fault mode in combination with the original log or the device operation simulation tool to verify its effectiveness.

[0059] The fault management module 130 includes the following components:

[0060] Task management unit 131: Used to build and execute diagnostic tasks for specific logs based on the fault tree, or analyze multiple logs for a certain fault branch, including parallel execution scheduling, etc.

[0061] Fault tree management unit 132: Used to manage the fault tree, including adding or deleting branches, maintaining fault phenomena, etc.

[0062] Full network diagnosis unit 133: Conduct a full network-wide investigation for a certain fault branch of the fault tree or multiple branches of a certain fault phenomenon to determine the operating status of network devices.

[0063] Result evaluation unit 134: Conduct accuracy evaluation based on the execution results of diagnostic tasks, and adjust or remove the constructed fault branches based on the evaluation results to ensure the overall effectiveness of the fault tree.

[0064] Thus, each module collaborates together to achieve efficient log management and fault diagnosis.

[0065] Based on the above content, a method for locating the root cause of equipment failures proposed in this application will be introduced in detail.

[0066] Exemplarily, Figure 2 FIG. 1 shows a flowchart of a method for locating the root cause of equipment failures provided by an embodiment of the present application. This method is applied to a computing device, which can be embodied as any computing unit, server, device, device cluster, etc. with computing and processing capabilities. The method for locating the root cause of equipment failures mainly includes the following execution steps:

[0067] Step S201, receive a first operation of the user on the user interface provided by the computing device. The first operation indicates the target failure mode to be diagnosed. The target failure mode includes a plurality of target log fingerprint models and the combination relationship between each model. The target log fingerprint model is used to match the target log fingerprint.

[0068] In one embodiment, the computing device is implemented using low-code orchestration technology and can provide a user interface for the user to perform the first operation. The first operation indicates the target failure mode. The computing device can be Figure 1 the computing device 10 shown in FIG. 1.

[0069] The target failure mode includes a plurality of target log fingerprint models. Each log fingerprint model is used to identify a phenomenon or feature of the device operation, and this phenomenon or feature is manifested as a log fingerprint in the log file. Taking the memory log fingerprint of a computer system as an example, common log fingerprints include memory warnings, error correction code (ECC), etc. These log fingerprints can accurately reflect the operating state of the device.

[0070] The multiple target log fingerprint models included in the target failure mode are also interrelated through a specific combination relationship. Each target log fingerprint model is responsible for identifying and analyzing specific patterns or features in the device operation log, and the combination relationship between the models defines how these models work together to identify and diagnose faults. The combination relationship can include logical operations (such as AND, OR, NOT, EQUAL), weight assignment, or other rules, etc.

[0071] For example, if both model A and model B detect an anomaly, a fault alarm is triggered. This combination relationship ensures that the fault diagnosis of the system is both comprehensive and accurate, capable of covering multiple fault situations and reducing false alarms.

[0072] Before receiving the first operation of the user on the user interface provided by the computing device and establishing the target failure mode, it is necessary to establish a target label based on the log fingerprint in the historical log file. And set matching rules and a unified input-output format for the target label to obtain the target log fingerprint model corresponding to the target label.

[0073] Exemplarily, the computing device can pre - generate tags using the log fingerprints in the historical log file. Alternatively, the computing device directly obtains the tags generated by other log processing tools. These implementations of generating tags can be executed by calling the log processing module 100.

[0074] By setting matching rules and unified input - output formats for each tag, a log fingerprint model corresponding to the tag can be generated. The log fingerprint model is used to identify and match specific log fingerprints according to the matching rules, and output the representation of the log fingerprint according to the unified input - output format. Tags can be regarded as the visual representation of the log fingerprint model, and each tag corresponds to a specific log fingerprint model. Users can operate on the tags displayed on the user interface to achieve operations on the log fingerprint model corresponding to the tag. The log fingerprint model can be generated by calling the model management module 110.

[0075] Generally, a target fault corresponds to multiple root causes, and users can create a target fault mode based on experience. It can be understood that the target fault mode should be able to cause at least one of the multiple root causes to occur.

[0076] Exemplarily, the target log fingerprint model further includes associated fault information, which is used for users to determine the target fault mode they want to create. For example, if a specific log fingerprint model is associated with memory faults, system faults, etc., then for memory faults, users can consider using this specific log fingerprint model to create a target fault mode.

[0077] Each log fingerprint model is independent and responsible for executing its own data - matching logic without involving the calculations of the entire fault diagnosis process, so it is relatively simple to maintain. In addition, these models can also be reused in different diagnostic tasks, thereby improving the efficiency of executing new diagnostic tasks.

[0078] Users also construct fault modes based on these tags. Then, these fault modes can be used to execute corresponding diagnostic tasks. A diagnostic task refers to the process of identifying, analyzing, and resolving faults or problems that occur in a device.

[0079] The computing device displays multiple tags through the user interface, and each tag corresponds to a log fingerprint model, that is, the tag is the representation of the log fingerprint model on the user interface. Based on this display, the computing device receives a first operation from the user on the user interface provided by the computing device. The first operation indicates the target fault mode.

[0080] In one implementation, the first operation includes: first, copying multiple target tags provided by the user interface to the canvas area of ​​the user interface. The target tags correspond to the target log fingerprint models one by one. Secondly, by constructing a combination relationship between multiple target tags in the canvas area, a combination relationship between multiple target log fingerprint models is established to obtain a target fault mode. The target fault mode corresponds to a diagnostic task. When executing this diagnostic task, these sub-fault modes can be run in parallel to improve diagnostic efficiency.

[0081] Specifically, when constructing a combination relationship between multiple target tags in the canvas area through tags, first, the input-output relationship between the associated fingerprint models in the multiple target log fingerprint models is set based on the multiple target tags to obtain several independent sub-fault modes. Next, the logical operation relationship between the sub-fault mode outputs is set based on the multiple target tags to establish a combination relationship. The logical operation relationship includes at least one of the following: and, or, negation, and equality.

[0082] Thus, the target failure mode can be constructed in real time according to the user performing the first operation.

[0083] In another implementation, the first operation includes: importing a target file provided by a computing device. The canvas area of ​​the user interface is used to render the target file to obtain a target failure mode.

[0084] Specifically, when implementing the first operation, the user selects a target file from a plurality of pattern files provided by the computing device, and imports the target file into the user interface. The canvas area of ​​the user interface renders the imported target file to obtain a target fault pattern.

[0085] Exemplarily, multiple pattern files can be obtained by analyzing multiple historical log files by a large model, and each pattern file corresponds to a fault mode. The fault modes corresponding to these pattern files have a high hit rate in actual applications, so the number of occurrences in multiple historical log files exceeds a preset number or the number of occurrences is ranked first. By selecting them for fault diagnosis, the root cause of the equipment fault can be effectively located.

[0086] Alternatively, multiple pattern files can be obtained by analyzing multiple simulation log files by the large model. Users can use the device operation simulation tool to simulate and monitor various device operation scenarios including various failure modes to obtain various device operation results. Based on the operation results of the device operation scenario, it is determined whether various failure modes cause failures. If a specific failure mode has a high failure hit rate in actual applications, the corresponding simulation log file is input into the large model for analysis to obtain multiple pattern files.

[0087] Thus, according to the user's execution of the first operation, the target failure mode can be obtained by importing a pattern file previously generated using a large model.

[0088] Exemplarily, a target fault model is constructed by invoking the pattern management module 120.

[0089] Step S202: Receive a second operation of the user on the user interface to obtain a device operation log. The second operation indicates a way to obtain the device operation log, and the device operation log is a log collected when the device has a target fault.

[0090] In one embodiment, historical logs generated during the past operation of the device are used for root cause location after a fault occurs. This includes: receiving a second operation of the user on the user interface to obtain a device operation log. The second operation indicates a way to obtain the device operation log, and the device operation log is a log collected when the device has a target fault. The way to obtain can include specifying a local file path on a computing device for storing the device operation log, or remotely accessing the location of the log file through a network address. The target fault is a historical fault that has occurred in the past for the device.

[0091] After the second operation indicates the way to obtain the device operation log, the computing device obtains the device operation log through the way to obtain. The device operation log is a log collected when the device has a target fault, including various phenomena or features during the operation of the device, that is, including various log fingerprints, and thus can be used for root cause location of the target fault.

[0092] Step S203: In response to a trigger operation of the user on the user interface, match target log fingerprints with a combined relationship from the device operation log according to the target fault mode. The matching result is used for root cause location of the target fault.

[0093] In one embodiment, the user interface provides a button to trigger diagnosis, and the diagnosis operation is triggered by the user clicking the button.

[0094] In response to a trigger operation of the user on the user interface, match target log fingerprints with a combined relationship from the device operation log according to the target fault mode. If a combination of target log fingerprints that meets the combined relationship is matched, the result of the diagnosis is that the target fault is caused. Determining at least one root cause caused by the target fault mode is the root cause of the target fault.

[0095] Thus, the user can indicate a target fault mode including a target log fingerprint model and the combined relationship between the models and import a log file in the computing device, and automatically extract and analyze the log fingerprints in the log. Based on this, it can help the user quickly locate and analyze the fault information in the log and improve the efficiency of post - fault diagnosis.

[0096] Exemplarily, the computing device also manages a fault tree, which includes fault nodes corresponding to a target fault. If the diagnosis result indicates that the target fault is triggered, the root cause corresponding to the target fault mode is added as a child node of the fault node. The management of the fault tree and the diagnostic task can be implemented by calling the fault management module 130.

[0097] Thus, by managing and maintaining the fault tree, the requirements of multiple root cause location scenarios for one fault phenomenon can be met.

[0098] In summary, the embodiment of the present application uses low-code orchestration technology to implement a visual operation and maintenance platform for device fault root cause location. The platform provides multiple graphical log fingerprint models for users to construct different fault modes and diagnose whether these fault modes will trigger device faults. In this way, the efficiency and accuracy of fault diagnosis can be improved, thereby achieving rapid location of the root cause of device faults.

[0099] Exemplarily, Figure 3 shows a flowchart for constructing a log fingerprint model provided by an embodiment of the present application.

[0100] As Figure 3 shown, this process includes the following execution steps:

[0101] Step S301, preview the log.

[0102] Exemplarily, the user uses the log processing interface provided by the log processing module 100 to preview the historical log file to understand the content and structure of the log.

[0103] Step S302, select the log.

[0104] Exemplarily, after previewing the log, the user selects the specific log file that needs to be further processed or analyzed.

[0105] Step S303, generate tags.

[0106] Exemplarily, the user uses the log processing interface to select some log fragments in the log file as log fingerprints and generates tags based on the selected log fingerprints.

[0107] Step S304, standardize the modeling.

[0108] Exemplarily, based on the generated tags, the user further performs standardized modeling. This may involve using the log processing interface to set matching rules and unify the input and output formats to obtain a log fingerprint model corresponding to the tags.

[0109] Exemplarily, Figure 4 shows a schematic diagram for constructing a log fingerprint model provided by an embodiment of the present application.

[0110] As shown Figure 4 in the figure, a log processing interface of a computing device is shown. The interface includes a top area, a left area, a middle area, and a right area.

[0111] The top area shows relevant information of historical log file packages, such as the name of the file package: log, and the type of the file package: zip.

[0112] The left area shows the directory structure of historical log file packages, presented in a tree diagram, supporting users to perform quick file searches and capable of showing full-text index information.

[0113] The middle area is the log content area, showing the line numbers and text content of log files. Users can view the text content of corresponding files by switching different files in the left directory tree. After selecting a certain file, users can perform tag addition operations by selecting log fingerprints.

[0114] The right area is displayed only when the selected log fingerprint is right-clicked. Some fields in this area will be automatically filled and are not allowed to be modified. For example, the "File" field shows the selected log file: LogDump / fdm_output, and the "Sample" field shows the selected log fingerprint: DIMM000 uncorrect error. Users can set the "Tag Name" in this panel, such as ECC information. Set the "Tag Classification" through a drop-down menu, such as memory fingerprint. The available classifications in the drop-down menu are pre-set by the computing device.

[0115] Furthermore, the matching rules of the tag are supplemented. A variety of matching modes are provided in the drop-down menu of "Operator", such as regular matching and keyword matching, with keyword matching as the default. If keyword matching is selected, keywords also need to be filled in the value, and the keyword matching rule will perform a full-text search according to the value. If regular matching is selected, a regular expression is filled in the value, such as: / [DIMM000 uncorrect error] / gm, which is used to search in the text according to a specific pattern.

[0116] Among them, " / " is the delimiter of the regular expression, used to identify the start and end of the regular expression. [DIMM000uncorrect error]: This is the text pattern to be matched, indicating to search for the text containing the specific string "DIMM000 uncorrect error". g: This is a modifier, indicating global search. m: This is another modifier, indicating multiline search. Therefore, the meaning of this regular expression is: search for all occurrences of the string "DIMM000 uncorrect error" in the entire text, and each matching instance will be returned, rather than only returning the first matching item. At the same time, the search will be performed independently on each line to ensure that cross-line matches can also be found.

[0117] After clicking the "Test" button, test whether the selected sample can be matched according to the regular expression. If it can be matched, the "Test" button will be displayed in green, indicating that the regular expression meets the requirements. Otherwise, the "Test" button will be displayed in red, and the regular expression needs to be changed until the test passes.

[0118] In the "Associated Faults" section, no or multiple associated faults can be added to this label to facilitate subsequent fault diagnosis. These associated faults are the faults included in the fault tree managed by the computing device.

[0119] The "Model Output" section is generated by the computing device according to the predefined structure, and the dataset related to the currently selected log fingerprint is filled in to achieve the standardized processing of the label input and output. Figure 4 An example of the model output of the newly created label "ECC Information" is given, which is represented in JSON format and includes a regular expression matching rule. When necessary, standardized model inputs are also generated for the label.

[0120] So far, by supplementing the matching rule and standardized input and output of this label, the log fingerprint model corresponding to this label is generated.

[0121] Finally, click the "Save" button to save the label, the log fingerprint model data, and their corresponding relationships.

[0122] Thus, the user interface provided by the computing device can display multiple labels for the user to operate, so as to obtain the fault mode to be diagnosed.

[0123] Exemplarily, Table 1 lists the relevant information of this log fingerprint model construction process, and its meaning is as above.

[0124] Table 1

[0125]

[0126] Optionally, the collected historical log file package can also be uploaded to other text index databases, such as the OpenSearch text index database. Further, use the search interface of OpenSearch Dashboard to perform full-text search on the text and visualize the page. The user quickly searches for relevant logs based on the visualized page, checks and "adds tags" after finding the key log fingerprints in combination with the log context, and sets the matching rules and unified input / output formats during the addition process to complete the form as shown in Figure 1 and send it to the computing device for saving.

[0127] Exemplarily, Figure 5 Figure 1 shows a flowchart for locating the root cause of a device failure provided by an embodiment of the present application.

[0128] As shown in Figure 5 Figure 1, the process includes the following execution steps:

[0129] Step S501, construct a failure mode.

[0130] Exemplarily, the user defines the basic information and objectives of the fault diagnosis task, and determines the failure mode to be constructed based on this information and objectives.

[0131] Step S502, copy the tags to the canvas area.

[0132] Exemplarily, the user selects the target tags required from multiple tags provided by the user interface, and copies the target tags to the canvas area by means of selection or dragging.

[0133] Step S503, add the input / output relationships between the associated log fingerprint models.

[0134] Exemplarily, the user sets the input / output relationships between the associated target log fingerprint models based on the target tags, and obtains several independent sub-failure modes. This helps to clarify the interaction between the log fingerprint models and the parameter transfer process.

[0135] Step S504, add the logical operation expressions for the output of the sub-failure mode.

[0136] Exemplarily, the user sets and adds logical operation expressions based on the target tags, and these expressions are used to define the further logical processing and decision-making process based on the output of the sub-failure mode.

[0137] Step S505, diagnose the failure mode and save it.

[0138] Exemplarily, the user needs to diagnose the fault mode to ensure its correctness and effectiveness, and then save the fault mode to a computing device for subsequent use or further management and optimization.

[0139] Step S506, add the root cause corresponding to the fault mode to the fault tree.

[0140] Exemplarily, if the diagnosis result is that a fault is caused, add the root cause corresponding to the fault mode to the fault tree. Graphically displaying the device faults and diagnosis results helps to systematically manage and diagnose faults.

[0141] Exemplarily, Figure 6 shows a schematic diagram of a fault mode construction provided by an embodiment of the present application.

[0142] As Figure 6 shown, it shows the user interface opened by the user for constructing the fault mode. The title bar contains the trouble ticket and problem description associated with the diagnosis task. The tag area contains all the tags established by the method shown Figure 4 below, such as the tag ECC information under the memory fingerprint classification. The user can also query the hidden tags by keyword in the component panel. The operation area contains a toolbar and a canvas area. After opening the user interface, the "Start" button, "Feature Matching" input box, "Logical Operation" input box, "End" button and the connections between these related elements in the canvas area are initialized by default. The tools in the toolbar are used to operate on the various tags copied to the canvas area, such as deleting and copying. The "Output Area" shows the attribute editing of the elements in the canvas area and the diagnosis results. When a specific element in the canvas area is selected, the corresponding attributes are displayed, and attribute editing and saving after editing are supported.

[0143] Specifically, after a problem of "system abnormal restart caused by memory module failure" with an unknown root cause in the trouble ticket "SR0001" occurs, determine the basic information and objectives of the fault diagnosis task. Open the user interface of the computing device, and construct a fault mode based on this information and objective. Elements such as the "Feature Matching" input box and the "Logical Operation" input box will be initialized in the canvas area.

[0144] In response to the user's operation of selecting or dragging the tag button in the tag area, such as "Memory Basic Information", etc., copy the target tag to the "Feature Matching" box in the canvas area. After dragging it in, the computing device will automatically number the target tag, such as A1, A2, A3, etc.

[0145] In response to the user's selection of tools in the toolbar, multiple association relationships corresponding to the log fingerprint model of the label are added. For multiple log fingerprint models with sequence and attribute dependencies, grouping processing can be performed. Connections can be added between the labels in a group, and the connections are used to specify the input parameters of the subsequent log fingerprint model. Thus, multiple independent sub-fault modes are obtained.

[0146] For example, an independent sub-fault mode B1 is established using A1 and A3, and A2 is used as an independent sub-fault mode B2 alone. Moreover, the output of A1 is set as the input of A3.

[0147] And a logical operation expression between the outputs of the sub-fault modes is added. When there are outputs of multiple independent sub-fault modes in the "Feature Matching" area and further logical determination of these outputs is required, operators and operation elements need to be added in the "Logical Operation" area. The operations include but are not limited to determination logics such as "AND", "OR", "NOT", "EQUAL", etc.

[0148] For example, the result of the "AND" logical judgment between sub-fault mode B1 and sub-fault mode B2 is used as the output of the target fault mode.

[0149] As Figure 6 shown, when a specific component in the "Canvas Area" is selected, that is, the connection Line001 between elements A1 and A3, the log fingerprint models at both ends of the connection can be edited in the "Output Area". And the output of the previous element A1, the output of A3, and information such as obtaining data from the output of the previous element A1 as the input parameter of the next element A3 are displayed in the "Output Area". By clicking the "Save" button in the "Output Area", the constructed fault mode is saved.

[0150] Exemplarily, Figure 7 a schematic diagram of device fault root cause location provided by an embodiment of the present application is shown.

[0151] As Figure 7 shown, a user interface for fault diagnosis based on a test log package is displayed. By selecting the test log package: dump01.zip through the drop-down box and clicking the "Test" button in the operation area, the constructed fault mode is diagnosed based on the test log package. The test log package is imported in advance by the computing device.

[0152] If a log fingerprint combination that conforms to the logical operation rules is matched, the diagnosis result is that a fault is triggered. This fault is the fault in the corresponding scenario indicated by the test log package for this match. If the combination cannot be matched, it means there is a problem with the current fault mode, and relevant elements in the canvas area can be continued to be edited and then retried for diagnosis.

[0153] Finally, the "output area" displays the diagnostic results. In the case where the diagnostic result is a matched fault, the associated fault tree root information is also displayed.

[0154] Exemplarily, Figure 8 FIG. 5 shows a fault tree diagram provided by an embodiment of the present application.

[0155] As Figure 8 shown, a fault tree of a server is presented. The fault tree is divided into three layers. The root node is the server fault. The second-layer nodes are the types of server faults, including operating system restart, memory anomaly, graphics card anomaly, etc., corresponding to Figure 4 the "associated fault" field in Figure 7 and corresponding to Figure 7 the "fault tree root" field in. The third-layer nodes show the root cause of each fault. If a fault is diagnosed according to

[0156] It can be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. In addition, in some possible implementation manners, the steps in the above embodiments can be selectively executed according to actual situations, can be partially executed, or can be all executed, which is not limited herein. Additionally, all or part of any feature in the above embodiments can be freely combined arbitrarily on the premise of not being contradictory. The combined technical solution is also within the scope of the present application.

[0157] Exemplarily, an embodiment of the present application further provides a computing device 1000. As Figure 9 shown, the computing device 1000 includes: a bus 1002, a processor 1004, a memory 1006, and a communication interface 1008. The processor 1004, the memory 1006, and the communication interface 1008 communicate with each other through the bus 1002. It should be understood that the present application does not limit the number of processors and memories in the computing device 1000.

[0158] The bus 1002 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 9In the figure, it is represented by a line, but it does not mean that there is only one bus or one type of bus. The bus 1004 may include a path for transmitting information between various components of the computing device 1000 (for example, the memory 1006, the processor 1004, and the communication interface 1008).

[0159] The processor 1004 may include any one or more of a central processing unit, a graphics processing unit (GPU), a microprocessor (MP), a digital signal processor (DSP), a baseboard management controller, and other processors.

[0160] The memory 1006 may include a volatile memory, such as a random access memory (RAM). The processor 1004 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0161] The communication interface 1008 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1000 or a cluster composed of multiple computing devices 1000 and other devices or a communication network.

[0162] The computing device 1000 includes internal network devices, or the computing device 1000 is externally connected to multiple network devices. The internal network devices communicate with the processor 1004, the memory 1006, and the communication interface 1008 through the bus 1002, and the external network devices communicate with the computing device 1000 through interfaces such as Ethernet, Fibre Channel, and Infiniband.

[0163] The memory 1006 stores executable program code / instructions, and the processor 1004 executes the executable program code / instructions to implement Figure 2 the process shown in, thereby implementing all or part of the steps of the method in the above embodiments. In other words, the memory 1006 stores a program / instruction for executing all or part of the steps of the method in the above embodiments.

[0164] An embodiment of the present application provides a computing device, including: a memory and a processor; the memory and the processor are coupled; the memory is used to store programs; the processor is used to execute the programs stored in the memory, and when the programs stored in the memory are executed, the processor is used to execute the methods in the above embodiments.

[0165] Based on the methods in the above embodiments, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor is caused to execute the methods in the above embodiments.

[0166] Based on the methods in the above embodiments, an embodiment of the present application provides a computer program product. When the computer program product runs on a processor, the processor is caused to execute the methods in the above embodiments.

[0167] The method steps in the embodiments of the present application can be implemented in a hardware manner or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in a random access memory (RAM), flash memory, read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), registers, hard disks, removable hard disks, CD-ROMs, or any other form of storage medium well-known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.

[0168] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0169] It can be understood that the various numerical numbers involved in the embodiments of the present application are only for the convenience of description and are not used to limit the scope of the embodiments of the present application.

Claims

1. A method for locating the root cause of equipment failure, characterized in that, Applied to a computing device, the method includes: Receiving a first operation of a user on a user interface provided by the computing device; the first operation indicates a target failure mode; the target failure mode includes a plurality of target log fingerprint models and combination relationships between the models; the target log fingerprint models are used to match target log fingerprints; Receiving a second operation of the user on the user interface to obtain device operation logs; the second operation indicates the acquisition path of the device operation logs, and the device operation logs are logs collected when the device has a target failure; In response to a trigger operation of the user on the user interface, matching target log fingerprints with the combination relationship from the device operation logs according to the target failure mode; the result of the matching is used for root cause localization of the target failure.

2. The method according to claim 1, wherein The target failure corresponds to multiple root causes, and the target failure mode causes at least one of the multiple root causes to occur; The method further includes: If the result of the matching is successful, at least one of the root causes is the root cause of the target failure.

3. The method according to claim 1, wherein The first operation includes: Copying a plurality of target labels provided by the user interface to a canvas area of the user interface; the target labels correspond to the target log fingerprint models one by one: Establishing combination relationships between the models based on the multiple target labels in the canvas area to obtain the target failure mode.

4. The method according to claim 3, characterized in that The establishing combination relationships between the models based on the multiple target labels in the canvas area includes: Setting input-output relationships between associated fingerprint models among the plurality of target log fingerprint models based on the multiple target labels to obtain several independent sub-failure modes; Setting logical operation relationships between the outputs of the sub-failure modes based on the multiple target labels to establish the combination relationships; the logical operation relationships include at least one of the following, AND, OR, NOT, EQUAL.

5. The method according to claim 1, wherein The first operation includes: Importing a target file provided by the computing device; the canvas area of the user interface is used to render the target file to obtain the target failure mode.

6. The method according to claim 5, wherein The target file is obtained by a large model analyzing multiple historical log files, and the target failure mode is a failure mode that appears more than a preset number of times in the multiple historical log files.

7. The method according to claim 1, characterized in that, The computing device also manages a fault tree, and the fault tree includes a fault node corresponding to the target failure; The method further includes: If the result of the matching is successful; Adding at least one root cause caused by the target failure mode as a child node of the fault node respectively.

8. The method according to claim 3 or 5, characterized in that, Before receiving the first operation of the user on the user interface provided by the computing device, it further includes: Establishing the target labels according to the log fingerprints in the historical log files; Setting matching rules and a unified input-output format for the target labels to obtain the target log fingerprint models corresponding to the target labels.

9. The method according to claim 8, characterized in that, The target log fingerprint model further includes associated fault information, and the associated fault information is used for the user to determine the target failure mode.

10. A computing device, characterized in that, Includes: Multiple memories for storing programs; Multiple processors for executing the program stored in the memory, which, when the program stored in the memory is executed, are used to execute the method according to any one of claims 1-9.