Root cause positioning method and device
By building a root cause troubleshooting tree and using a large language model to analyze it step by step, the problem of inaccurate error root cause positioning in the existing technology is solved, efficient and accurate root cause positioning is achieved, and maintenance costs are reduced.
Patent Information
- Application Number
- CN202510399124.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, the error root cause positioning method cannot accurately locate the true error root cause, especially the methods based on string matching and expert experience describes problems such as difficult to compatible with all situations, difficult to correlate complex relationships, difficult to reuse and high maintenance costs.
Build a root cause investigation tree. The node description adopts a mixed method of natural language and formal language. Use a large language model to analyze the root cause investigation tree step by step, select nodes that meet the judgment rules until the error root cause is determined.
It improves the accuracy and efficiency of root cause positioning, reduces maintenance costs, avoids the expansion of input length of large language model and model hallucination problems, and achieves accurate wrong root cause positioning.
Smart Images

Figure CN120276900A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the field of computer technology, and in particular, to a root cause localization method and apparatus. Background Art
[0002] Error root cause localization (hereinafter referred to as root cause localization) is a key step in fault diagnosis, quality management, and technical problem-solving. It refers to the process of identifying and determining the root cause that leads to a certain error or fault. This process not only includes the analysis of problem manifestations, but more importantly, delving into the deep-seated factors hidden behind the direct problems. Usually, error root cause localization involves methods such as data collection, data analysis (such as log analysis, statistical analysis), logical reasoning, and hypothesis verification, aiming to accurately identify the source of the error, thereby formulating effective solutions to prevent the recurrence of similar errors. In multiple fields such as software development, manufacturing, and system operation and maintenance, error root cause localization is a fundamental task to ensure system stability and product reliability.
[0003] In traditional technologies, error root cause localization is performed by string matching of error logs. However, such methods usually cannot accurately locate the true error root cause. Therefore, a reasonable solution is needed to improve the accuracy of error root cause localization. Summary of the Invention
[0004] One or more embodiments of this specification describe a root cause localization method and apparatus. For a target operation with an error, a large language model is used to perform error root cause localization based on a root cause investigation tree corresponding to the target operation, thereby improving the accuracy of error root cause localization.
[0005] In a first aspect, a root cause localization method is provided, including:
[0006] Obtain a root cause investigation tree corresponding to a target operation with an error, where a single node represents a root cause of the error, and its node description adopts a mixed manner of natural language and formal language, including a decision rule described in natural language and determined based on target parameters;
[0007] Select nodes level by level along the root cause investigation tree, including: respectively extract the corresponding decision rules from each candidate node at the current level; respectively determine the current parameter values of each target parameter involved in each decision rule; input the decision rules and the current parameter values into a large language model, and let it select the nodes that satisfy the corresponding decision rules;
[0008] Determine the error root cause of the target operation at least according to the nodes selected at the leaf level.
[0009] In a second aspect, a root cause localization apparatus is provided, including:
[0010] An acquisition unit for acquiring a root cause troubleshooting tree corresponding to a target operation that has an error, where a single node represents a root cause of the error, and the node description uses a mixed manner of natural language and formal language, including a determination rule described in natural language and determined based on target parameters;
[0011] A selection unit for selecting nodes level by level along the root cause troubleshooting tree;
[0012] The selection unit includes:
[0013] An extraction sub-module for respectively extracting the corresponding determination rules from each candidate node at the current level;
[0014] A determination sub-module for respectively determining the current parameter values of each target parameter involved in each determination rule;
[0015] An input sub-module for inputting the determination rules and the current parameter values into a large language model to make it select the nodes that satisfy the corresponding determination rules;
[0016] A determination unit for determining the root cause of the error of the target operation at least according to the nodes selected at the leaf level.
[0017] In a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method of the first aspect.
[0018] In a fourth aspect, a computing device is provided, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method of the first aspect is implemented.
[0019] The root cause location method provided by one or more embodiments of this specification pre-constructs a root cause troubleshooting tree corresponding to a target operation. After an error occurs in the target operation, the determination rules described in natural language corresponding to each candidate node in each level of the root cause troubleshooting tree are parsed level by level using a large language model, and the nodes that satisfy the corresponding determination rules are selected from them until reaching the leaf level. Finally, the root cause of the error of the target operation is determined at least based on the nodes selected at the leaf level. It should be noted that since the determination rules corresponding to the nodes are described in natural language, this solution can use a large language model to parse the rules, thereby improving the node selection efficiency, that is, improving the root cause location efficiency. In addition, in this solution, the large language model only selects nodes for a single level each time. On the one hand, it can control the input length of the large language model, thus solving the problem of model hallucination; on the other hand, it can also ensure the prediction accuracy of the large language model, that is, accurately locate the root cause of the error of the target operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of this specification, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0021] Figure 1 A schematic diagram showing the root cause troubleshooting tree in an example of this specification;
[0022] Figure 2 A schematic diagram of the implementation scenario of an embodiment disclosed in this specification;
[0023] Figure 3 A flowchart showing the root cause location method according to an embodiment of this specification;
[0024] Figure 4 A schematic diagram showing one of the node selection methods in an example of this specification;
[0025] Figure 5 A schematic diagram showing another node selection method in an example of this specification;
[0026] Figure 6 A schematic diagram showing the root cause summary method in an example of this specification;
[0027] Figure 7 A schematic diagram showing the root cause location device according to an embodiment of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] The following describes the solutions provided in this specification in conjunction with the drawings.
[0029] As mentioned above, there is a problem that the true error root cause cannot be located based on string matching for error root cause location. For this reason, related technical solutions propose to build an expert diagnosis system based on expert experience and then use this expert diagnosis system for root cause location. There are currently mainly the following two ways of describing the expert experience language:
[0030] First, use natural language to describe expert experience, that is, convert expert experience into a natural language document, and then perform root cause location based on the natural language document. However, this method has the following problems:
[0031] 1. The system cannot obtain the information required during the root cause location process in real time. The root cause location process usually requires manual participation and cannot achieve automation.
[0032] 2. The system needs to retrieve key information from natural language documents through retrieval. In essence, it is also a string matching scheme without diagnostic ideas and logical support, resulting in the inability to accurately locate the real root cause of errors.
[0033] 3. As the complexity of the system increases, the maintenance cost and synchronization workload of natural language documents will increase significantly.
[0034] Second, describe expert experience using a formal language (also known as a programming language), that is, convert expert experience into formal language code, and then perform root cause localization based on the formal language code. However, this method has the following problems:
[0035] 1. Difficult to enumerate: That is, it is difficult to cover all abnormal situations.
[0036] For example, assume the expert experience is as follows:
[0037] "If an exception of database connection failure occurs, you can first check whether the number of database connections has reached the limit."
[0038] For this expert experience, if it is described using the Java language, the corresponding Java code can be as follows:
[0039]
[0040]
[0041] It can be seen that when the above Java code determines whether there is a database connection exception, it uses a string matching method. This string matching method has the problem of being difficult to be compatible with all situations. For example, for database connection exceptions, it can be expressed as follows:
[0042] Connection refused;
[0043] Unable to connect to the database;
[0044] Can't connect to MySQL server on'server'(10061);
[0045] Unknown database 'dbname';
[0046] It should be understood that only several common representations of database connection exceptions are listed here. In practice, for different database types (MySQL, Oracle, MongoDB), the error messages for database connection exceptions are different. Therefore, it is usually impossible to enumerate all the representations. Even if all the representations can be enumerated, the error messages may still change due to database version upgrades, so the Java code still needs to be modified. It can be seen that there are problems with difficulty in reuse and compatibility with all situations when using formal languages to describe expert experience.
[0047] 2. Difficult to associate: That is, it is impossible to accurately describe the complex relationship between the code and the error.
[0048] For example, assume the expert experience is as follows:
[0049] "If after a user submits a piece of code, a large number of new exceptions appear on the online machine, and if the exception logs and stack traces are strongly correlated with the code submitted by the user, it can basically be determined that it is caused by the code submitted by the user."
[0050] For this expert experience, if it is described using the Java language, the corresponding Java code can be as follows:
[0051]
[0052] Note that the correlation between the user code and the error log mentioned here actually cannot be described using a formal language because there is no certain field value between them that can represent similarity. It may be a function name, a line of log, or a null pointer thrown by a certain variable. Therefore, the above hasCorrelation function actually cannot be implemented using Java code, which means that this expert experience actually cannot be expressed using a formal language.
[0053] 3. Difficult to reuse: That is, the logic of a specific system is difficult to generalize, increasing the system complexity and maintenance difficulty.
[0054] For example, assume the expert experience is as follows:
[0055] "For system Foo, due to historical reasons, it will request system Bar and there will be a large number of error logs of request failures, but this error does not affect the main process because system Bar will be called for backup, so the errors that occur when calling system Bar can be ignored."
[0056] For the above expert experience, when using a formal language to describe it, an additional if logic needs to be added at the place where the request fails to indicate that the call error of the Bar system can be ignored. However, this if logic is not universal in the entire expert diagnosis system. Such code will increase the complexity of the system, be confusing, and difficult to maintain later.
[0057] In addition, due to differences between systems, there may be a lot of code similar to the above if logic. Therefore, it is not feasible to customize a set of expert experience for system Foo and place it in the Bar system.
[0058] In summary, although using a formal language to describe expert experience can solve some problems, it ultimately does not replace manual work. Root cause localization still requires the participation of roles such as operations and SRE. Moreover, due to the above-mentioned difficult-to-solve problems when using a formal language to describe expert experience, building an expert diagnosis system based on the expert experience described in a formal language has become an almost impossible task.
[0059] Given that there are corresponding defects in describing expert experience based on natural language and formal language, related technical solutions also propose root cause localization based on RAG. For example, based on keywords describing incorrect operations, similar knowledge is recalled from the knowledge base, and then the keywords and the recalled knowledge are input into a large language model for processing. However, this solution has the following problems:
[0060] 1. The construction of the knowledge base requires a high cost.
[0061] 2. The essence of recalling knowledge from the knowledge base is still string matching, and the knowledge has slightly improved generalization ability compared to pure string matching.
[0062] 3. It cannot handle complex logical relationships.
[0063] Therefore, this solution proposes to pre-construct a root cause troubleshooting tree corresponding to the target operation, where a single node represents an error root cause, and the node description uses a mixed method of natural language and formal language. Specifically, the determination rules for the error root cause in the node description can be described in natural language. After that, when an error occurs in the target operation, the large language model is used to hierarchically parse the determination rules corresponding to each candidate node in each level of the root cause troubleshooting tree, and select the nodes for which the corresponding determination rules are satisfied until the leaf level is reached. Finally, the error root cause of the target operation is determined based at least on the nodes selected at the leaf level.
[0064] It should be noted that since the determination rules corresponding to the nodes are described in natural language, this solution can utilize large language models for rule parsing, thereby improving the efficiency of node selection, that is, improving the efficiency of root cause location. In addition, in this solution, organizing node descriptions (i.e., expert experience) into a tree structure can avoid the large language model from learning the entire root cause location logic and only requires node selection for a single level. On the one hand, this can save costs, and on the other hand, it can control the input length of the large language model, thereby solving the model hallucination problem and ensuring the accuracy of the large language model's prediction, that is, accurate positioning of the wrong root cause can be achieved. Finally, in this solution, the node description adopts a mixed method of natural language and formal language, which can avoid the respective disadvantages of the two languages. That is, it will not require precise modification to the code logic every time like formal language, with high modification difficulty and large maintenance costs; nor will it lack a unified logical form like pure natural language.
[0065] The above is the inventive concept provided in this specification. Based on this inventive concept, this solution can be implemented. The following provides a detailed description of this solution.
[0066] As mentioned above, in order to avoid the large language model from learning the entire root cause location logic, this solution proposes to organize expert experience into a tree structure, that is, construct a corresponding root cause troubleshooting tree for different target operations. Among them, the different target operations here can include, but are not limited to, machine deployment operations, Pod startup operations, and remote procedure call (RPC) operations, etc.
[0067] Taking any target operation as an example, the corresponding root cause troubleshooting tree can include multiple levels of nodes. Any one of the nodes represents an error root cause, and any one of the nodes has a node description, which adopts a mixed method of natural language and formal language. Specifically, the determination rule of the error root cause in the node description can be described in natural language, and the determination rule is based on target parameters for determination.
[0068] It should be understood that the above target parameters correspond to the target operations. For example, when the target operation is a machine deployment operation, the above target parameters can include log information recorded during the machine deployment process, environmental information on the machine, or configuration information of the machine, etc. When the target operation is a Pod startup operation, the above target parameters can include the most recent N scheduling events of the Pod. When the target operation is a remote procedure call RPC operation, the above target parameters can include error codes, etc.
[0069] In addition to the determination rule, the above node description can also include a collection tool for the target parameters involved in the determination rule.
[0070] It should be noted that to control the size of the node description, the node description may only include the relevant introduction of the collection tool. For example, the tool name, parameter identifiers corresponding to the tool output, etc. In a more specific embodiment, the relevant introduction of the collection tool may be described using a formal language.
[0071] Of course, in practice, the above node description may also include solutions to the root cause of errors, etc., which are not limited in this specification.
[0072] Regarding the above node description, it can be recorded as a file in any of the following data formats: YAML format, JSON format, XML format, etc.
[0073] Taking the target operation as the Pod startup operation as an example, the node description of a certain node in the corresponding root cause troubleshooting tree can be as follows:
[0074] name:image-pull-failed
[0075] description:Image pull failed
[0076] requirements:
[0077] -name:podEventTool
[0078] outputId:podEventInfo
[0079] rule:
[0080] Please observe the events finally generated by the following pod scheduling:
[0081] ${podEventInfo}
[0082] If the event is not empty and an error of image pull failure occurs, such as "Error:ErrImagePull", then it conforms to this node.
[0083] solution:You can try to reapply for a machine. If it cannot be solved, you can consult the relevant students in the build.
[0084] In this example, the node description is a file in YAML format (abbreviated as YAML file), which contains fields such as name and description. The meanings of these fields are shown in Table 1.
[0085] Table 1
[0086] Field Name Meaning name Node Name description Root Cause of Error requirements Tool Introduction requirements.name Tool Name requirements.outputId Parameter Identification Corresponding to Tool Output rule Judgment Rule solution Solution
[0087] It should be understood that Table 1 is only an exemplary illustration. In practice, the node description may include more or fewer fields. For example, it may also include a sub-node list field, etc. This specification does not limit this.
[0088] Combined with the meanings of the fields in Table 1, the meaning of this YAML file is as follows: For the determination rule corresponding to the node, the podEventTool tool can be used to collect the scheduling events of the Pod. If the scheduling event contains information about failed image pulling, it means that the failure of the Pod to start is caused by failed image pulling. Then, it will suggest that the user try to re-apply for a machine or consult the relevant classmates.
[0089] Figure 1 The schematic diagram of the root cause troubleshooting tree in an example of this specification is shown. Figure 1 In it, the root cause troubleshooting tree includes nodes at 3 levels. Among them, node A in the first level represents root cause 1, and node A corresponds to YAML file 1, which describes the determination rule of root cause 1, the collection tool of the target parameters involved in the determination rule, and the solution, etc. Similarly, node B and node C in the second level represent root cause 1.1 and root cause 1.2 respectively, and these two nodes correspond to two YAML files respectively. And nodes D - G in the third level represent root cause 1.1.1, root cause 1.1.2, root cause 1.2.1, and root cause 1.2.2 respectively, and these 4 nodes correspond to 4 YAML files respectively.
[0090] Figure 1 In it, node B has two sub-nodes: node D and node E, indicating that when it is determined at the second level that the error root cause of the target operation is root cause 1.1, there are two situations of root cause 1.1.1 and root cause 1.1.2, and it is necessary to determine whether it is root cause 1.1.1 or root cause 1.1.2 according to the current parameter values collected by the collection tools corresponding to node D and node E.
[0091] It should be noted that for the above root cause troubleshooting tree, the nodes can be flexibly increased or decreased according to the actual situation, and the YAML file corresponding to each node can be modified. It can be seen that the root cause troubleshooting tree constructed by this solution has good maintainability.
[0092] In summary, the proposed solution organizes expert experience into a tree structure, which can provide strong logical support for root cause localization and has a low maintenance cost. In addition, the tree structure can be reused multiple times, and special logic can be added by simply adding corresponding nodes. Moreover, based on this tree structure, the specific process of root cause localization can be intuitively understood, and the decision rules in the node descriptions are described in natural language, which not only simplifies implementation but also improves the generalization of the decision rules and can solve complex logics that are difficult to handle with formal languages. Finally, when organizing expert experience into a tree structure, a large amount of documents and data do not need to be used to build a vector library or knowledge graph, and the large language model can be used extremely quickly and efficiently to achieve accurate root cause localization.
[0093] Figure 2 It is a schematic diagram of the implementation scenario of an embodiment disclosed in this specification. Figure 2 In it, for the target operation with an error, nodes are selected level by level along the root cause troubleshooting tree. Specifically, for any level, first, the corresponding decision rules are respectively extracted from the candidate nodes (shown by gray circles) therein, then the current parameter values of the target parameters involved in each decision rule are respectively determined, and finally, each decision rule and each current parameter are input into the large language model to obtain the node selected at this level. After reaching the leaf level, the error root cause of the target operation is determined based at least on the nodes selected at the leaf level.
[0094] Figure 3 It shows a flowchart of the root cause localization method according to an embodiment of this specification. This method can be executed by any device, equipment, platform, or equipment cluster with computing and processing capabilities. As Figure 3 shown, this method may include the following steps:
[0095] Step S302, obtain the root cause troubleshooting tree corresponding to the target operation with an error.
[0096] Among them, the target operation here may include, but is not limited to, machine deployment operation, Pod startup operation, or remote procedure call (RPC) operation, etc.
[0097] The root cause troubleshooting tree corresponding to the target operation may include nodes at multiple levels. Any node therein represents an error root cause, and any node has a node description, which adopts a mixed manner of natural language and formal language. Specifically, the decision rule for the error root cause in the node description can be described in natural language, and the decision rule is based on the target parameter for determination.
[0098] In addition to the decision rule, the above node description may further include the collection tool for the target parameter involved in the decision rule. More specifically, the relevant introduction of the collection tool can be described in formal language.
[0099] Of course, in practice, the above node description may also include solutions to the root cause of the error, etc. Among them, the solution can also be described in natural language.
[0100] In one example, the root cause troubleshooting tree may be as Figure 1 shown.
[0101] Step S304, select nodes level by level along the root cause troubleshooting tree.
[0102] That is, starting from the topmost level (i.e., the root level) of the root cause troubleshooting tree, in the order from top to bottom, select nodes for each level until reaching the bottommost level (i.e., the leaf level).
[0103] Taking any i-th level as an example, the node selection for the i-th level includes: Step S3042, extract the corresponding decision rules from each candidate node of the i-th level; Step S3044, determine the current parameter values of each target parameter involved in each decision rule respectively; Step S3046, input each decision rule and each current parameter value into a pre-trained large language model, and let it select the nodes that satisfy the corresponding decision rules.
[0104] First, regarding Step S3042, in one embodiment, the candidate nodes of the i-th level can be determined according to the nodes selected in the (i - 1)-th level. For example, in the i-th level, the nodes connected to the nodes selected in the (i - 1)-th level by edges are determined as candidate nodes. That is, the direct children nodes of the nodes selected in the i-th level are determined as the candidate nodes of the (i - 1)-th level.
[0105] Of course, in practice, the sibling nodes of the direct children nodes of the nodes selected in the i-th level can also be used as candidate nodes, and this specification does not limit this.
[0106] It should be understood that the value of i should be greater than 1. When i equals 1, the i-th level is the first level, and there is only one root node in the first level, so the root node can be directly used as the node selected in the first level.
[0107] Taking the node description of each node in the root cause troubleshooting tree as a YAML file as an example, the decision rules can be read from the YAML files corresponding to each candidate node. Taking the node description of a certain candidate node as the aforementioned YAML file as an example, the decision rule read is:
[0108] "Please observe the events finally generated by the following pod scheduling:
[0109] ${podEventInfo}
[0110] If the event is not empty and an error of mirror pull failure occurs, such as "Error: ErrImagePull", then it meets the requirements of this node.
[0111] It should be understood that the target parameter involved in this determination rule is: "podEventInfo".
[0112] Next, regarding step S3044, as mentioned above, the node description can also include the collection tools for the target parameters involved in the determination rule. When the collection tools are also included in the node description, the corresponding collection tools for each candidate node can be called respectively to collect the current parameter values of each target parameter involved in each determination rule.
[0113] It should be understood that when the node description only includes the relevant introduction of the collection tool, the corresponding collection tool can be determined and called according to this relevant introduction (such as the tool name).
[0114] Among them, the input parameters of the above collection tool can include the object identifier of the operation object of the target operation. For example, when the target operation is a machine deployment operation, the input parameters of the collection tool can include the IP address of the machine; when the target operation is a Pod start operation, the input parameters of the collection tool can include the name of the Pod; when the target operation is a remote procedure call (RPC) operation, the input parameters of the collection tool can include the address of the remote machine, etc.
[0115] Taking the aforementioned YAML file as an example, based on the name of the Pod, the collection tool: podEventTool can be called to obtain the scheduling events of this Pod.
[0116] Of course, in practice, the current parameter values of each target parameter involved in each determination rule can also be obtained according to user input, which is not limited in this specification.
[0117] Finally, regarding step S3046, in one embodiment, the current parameter values can be filled into each determination rule respectively, that is, the current parameter values of each target parameter involved in each determination rule are used to replace each target parameter correspondingly. Then, based on each determination rule and its corresponding candidate node after filling in the current parameter values respectively, a prompt word p1 is constructed. The prompt word p1 indicates to select the candidate node for which the corresponding determination rule is satisfied from each candidate node. The prompt word p1 is input into the large language model to obtain the node selected at the i-th level.
[0118] In another embodiment, parameter identifiers of each target parameter can be added respectively corresponding to each current parameter value. Based on each determination rule and its corresponding candidate nodes, and each current parameter value after adding parameter identifiers, a prompt p2 is constructed, and this prompt p2 indicates to select, based on each current parameter value, candidate nodes for which the corresponding determination rule is satisfied from each candidate node. The prompt p2 is input into the large language model to obtain the nodes selected at the i-th level.
[0119] Among them, the parameter identifiers of the above target parameters can be read from the relevant introduction of the acquisition tool in the node description. Taking the above node description as an example, the parameter identifier corresponding to the current parameter value of the target parameter acquired by the acquisition tool is: "podEventInfo".
[0120] It should be understood that in this another embodiment, before the input of each determination rule into the large language model, the target parameters involved therein are not replaced, but directly input into the large language model. Thus, the link of parameter value replacement can be omitted, thereby improving the efficiency of node selection.
[0121] In an example, the above prompt p2 can be as follows:
[0122] Now there is a set of rules. Please judge according to the real-time information which node the reason for the current machine deployment failure most likely belongs to, and only return the node name without returning other information.
[0123] The nodes are as follows:
[0124] ```
[0125] Node name: Node B
[0126] Acquisition tool: [tool1]
[0127] Determination rule: When the information obtained by tool1 is...., the rule of this node is satisfied.
[0128] ---
[0129] Node name: Node C
[0130] Acquisition tool: [tool2]
[0131] Determination rule: When the information obtained by tool2 is..., the rule of this node is satisfied.
[0132] ---
[0133] The following are the actual parameter values:
[0134] tool1.ID1: .....
[0136] tool2.ID2: ......
[0138] ```
[0139] Among them, tool1.ID1 is the parameter identifier output by the acquisition tool corresponding to node B. Or rather, tool1.ID1 is the parameter identifier of the target parameter that the acquisition tool corresponding to node B is responsible for acquiring. tool2.ID2 is the parameter identifier output by the acquisition tool corresponding to node C. Or rather, tool2.ID2 is the parameter identifier of the target parameter that the acquisition tool corresponding to node C is responsible for acquiring.
[0140] It should be understood that after inputting the above prompt word p2 into the large language model, the model will output the nodes in the i-th layer for which the judgment rules are satisfied, and thus the nodes selected in the i-th layer are obtained.
[0141] Figure 4 Fig. 1 shows one of the schematic diagrams of the node selection method in an example of this specification. Figure 4 In it, assuming that node selection is performed for the second layer, then the judgment rules corresponding to node B and node C can be extracted first, and the current parameter values of the target parameters involved in the two judgment rules can be determined respectively. Then, based on the node name of node B, the judgment rule corresponding to node B, the current parameter value of the target parameter involved in this judgment rule, and the node name of node C, the judgment rule corresponding to node C, and the current parameter value of the target parameter involved in this judgment rule, a prompt word is constructed and input into the large language model. Assuming that the large language model outputs the node name of node B, then node B (shown by a gray circle) can be used as the node selected in the second layer.
[0142] Next, node selection can be performed for the third layer, that is, the judgment rules corresponding to node D and node E are extracted first, and the current parameter values of the target parameters involved in the two judgment rules are determined respectively. Then, based on the node name of node D, the judgment rule corresponding to node D, the current parameter value of the target parameter involved in this judgment rule, and the node name of node E, the judgment rule corresponding to node E, and the current parameter value of the target parameter involved in this judgment rule, a prompt word is constructed and input into the large language model. Assuming that the large language model outputs the node name of node D, then node D (shown by a gray circle) can be used as the node selected in the third layer. For details, please refer to Figure 5 as shown.
[0143] Since the third layer is the bottom layer, the node selection process ends.
[0144] As can be seen from the above, the core idea of this solution is to use a large language model to perform multiple node selections, where each node selection corresponds to a level of the root cause investigation tree. The advantages of doing this are as follows:
[0145] 1. The selection of nodes is entrusted to the large language model rather than code logic. Thus, the decision rules can be described in natural language and have a certain generalization ability, and the maintenance cost is also relatively low.
[0146] 2. Since the large language model only performs node selection, its output can be well controlled.
[0147] 3. Since the input of the large language model for each node selection only includes each candidate node and its related information (including decision rules and current parameter values, etc.), the input length of the model can be effectively controlled, so that it will not expand due to the increase in troubleshooting experience.
[0148] 4. Since the large language model only outputs node names, and the time to call the model is strongly related to the output length, there is no need to worry about the time-consuming problem caused by multiple calls to the model.
[0149] Step S306, determine the error root cause of the target operation based at least on the nodes selected at the leaf level.
[0150] For example, the error root cause corresponding to the node selected at the leaf level can be directly determined as the error root cause of the target operation.
[0151] Of course, in practice, in order to improve the clarity and coherence of the error root cause of the target operation, a target prompt word can also be constructed based on the error root causes corresponding to the nodes selected at each level. The target prompt word indicates summarizing the error root causes corresponding to each node. Then, the target prompt word is input into the large language model to obtain the error root cause of the target operation.
[0152] Among them, the large language model here can be the same as or different from the large language model used for node selection above.
[0153] Additionally, the above target prompt word can also include one or more of the following: the current parameter value of the target parameter involved in the decision rule corresponding to the node selected at the target level; the solution corresponding to the node selected at the target level.
[0154] In a specific embodiment, the above target level is the leaf level.
[0155] In other embodiments, the above target level can also include other levels other than the leaf level, and this specification does not make any limitations in this regard.
[0156] For Figure 4 and Figure 5Taking the shown root cause troubleshooting tree as an example, the corresponding root cause summarization process can be as follows Figure 6 as shown. Figure 6 In this case, the root cause 1 corresponding to the selected node A at the first level, the root cause 1.1 corresponding to the selected node B at the second level, and the root cause 1.1.1 corresponding to the selected node D at the third level can be input into the large language model to make it summarize, so as to obtain the error root cause of the target operation.
[0157] In addition, the current parameter value of the target parameter involved in the determination rule corresponding to node D and the solution corresponding to node D can also be input into the large language model for the large language model to refer to when summarizing the content.
[0158] It should be noted that in this solution, using the large language model to perform the above summarization can have the following advantages:
[0159] 1. It can facilitate users to understand specific error information. For example, in the case of an RPC call failure, users usually need to know the specific reason for the failure, such as which service interface was called and what the problem was.
[0160] 2. It can facilitate users to know the reasoning process of the conclusion, which can avoid users repeating the troubleshooting that has already been carried out, thereby saving time.
[0161] 3. It can facilitate users to quickly find the place where the positioning process went wrong. For example, if a user finds that the error root cause corresponding to the selected node at the leaf level is incorrect (this may be caused by an incorrect selection of the node at a certain level), the user can quickly find the error place based on the summary output by the large language model and continue to troubleshoot downward from the error place.
[0162] In summary, the root cause positioning method provided in the embodiments of this specification can utilize the capabilities of the large language model to determine the error root cause of the target operation where an error occurs, thereby improving the root cause positioning efficiency. In addition, in this solution, the large language model only selects hierarchical nodes and does not need to learn the entire root cause positioning logic, which can save costs and is easy to iterate and update.
[0163] Corresponding to the above root cause positioning method, an embodiment of this specification also provides a root cause positioning device, as shown in Figure 7 shown, and this device may include:
[0164] An acquisition unit 702, configured to acquire a root cause troubleshooting tree corresponding to a target operation where an error occurs, where a single node represents an error root cause, and its node description adopts a mixed manner of natural language and formal language, including a determination rule described in natural language and based on target parameters;
[0165] The selection unit 704 is used to select nodes level by level along the root cause investigation tree;
[0166] The selection unit 704 includes:
[0167] The extraction sub-module 7042 is used to extract the corresponding determination rules from each candidate node at the current level;
[0168] The determination sub-module 7044 is used to respectively determine the current parameter values of each target parameter involved in each determination rule;
[0169] The input sub-module 7046 is used to input each determination rule and each current parameter value into the large language model, so that it selects the nodes that meet the corresponding determination rules;
[0170] The determination unit 706 is used to determine the error root cause of the target operation at least according to the nodes selected at the leaf level.
[0171] In one embodiment, the above node description further includes a collection tool for obtaining the target parameter;
[0172] The determination sub-module 7044 is specifically used for:
[0173] Respectively call the collection tools corresponding to each candidate node to collect and obtain each current parameter value.
[0174] In one embodiment, the above candidate nodes are the nodes in the current level that are connected by edges to the nodes selected in the previous level.
[0175] In one embodiment, the input sub-module 7046 is specifically used for:
[0176] Respectively fill each current parameter value into each determination rule;
[0177] Based on each determination rule filled with each current parameter value and its corresponding candidate nodes, construct a first prompt word, which indicates to select the candidate nodes that meet the corresponding determination rules from the candidate nodes;
[0178] Input the first prompt word into the large language model to obtain the nodes selected at the current level.
[0179] In another embodiment, the input sub-module 7046 is specifically used for:
[0180] Respectively add the parameter identifiers of each target parameter to each current parameter value;
[0181] Based on each determination rule and its corresponding candidate nodes, and each current parameter value after adding the parameter identifier, construct a second prompt word, which indicates to select the candidate nodes that meet the corresponding determination rules from the candidate nodes based on each current parameter value;
[0182] Input the second prompt into the large language model to obtain the selected node at the current level.
[0183] In one embodiment, the determination unit 706 is specifically configured to:
[0184] Construct a target prompt based on the root causes of errors corresponding to the selected nodes at each level, where the target prompt indicates a summary of the root causes of errors corresponding to each node;
[0185] Input the target prompt into the large language model to obtain the root cause of the error of the target operation.
[0186] In one embodiment, the above node description further includes a solution;
[0187] The above target prompt further includes one or more of the following:
[0188] The current parameter value of the target parameter involved in the determination rule corresponding to the selected node at the target level;
[0189] The solution corresponding to the selected node at the target level.
[0190] In one embodiment,
[0191] The above target operation is a machine deployment operation, and the above target parameters include log information, environment information on the machine, or configuration information of the machine; or,
[0192] The above target operation is a Pod startup operation, and the above target parameters include the scheduling event of the Pod; or,
[0193] The above target operation is a remote procedure call (RPC) operation, and the above target parameters include an error code.
[0194] In one embodiment, the above node description is recorded as a file in any of the following data formats: YAML format, JSON format, and XML format, etc.
[0195] The functions of the functional modules of the device in the above embodiments of this specification can be implemented by the steps of the above method embodiments. Therefore, the specific working process of the device provided in an embodiment of this specification will not be repeated here.
[0196] The root cause location device provided in an embodiment of this specification can improve the accuracy of error root cause location.
[0197] According to an embodiment of another aspect, there is also provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method described in conjunction with Figure 3 ...
[0198] According to an embodiment of still another aspect, there is also provided a computing device including a memory and a processor, where executable code is stored in the memory, and when the processor executes the executable code, it implements the combination with Figure 3 the method described.
[0199] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the medium or device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for the relevant content.
[0200] The specific embodiments of this specification are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be executed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0201] The specific implementation manners described above further elaborate on the purpose, technical solutions, and beneficial effects of this specification. It should be understood that the above are only the specific implementation manners of this specification and are not used to limit the protection scope of this specification. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of this specification should be included within the protection scope of this specification.
Claims
1. A root cause localization method, comprising: Obtaining a root cause investigation tree corresponding to a target operation with an error, where a single node represents a root cause of the error, and the node description uses a mixed manner of natural language and formal language, including a decision rule described in natural language and determined based on target parameters; Selecting nodes level by level along the root cause investigation tree, including: respectively extracting the corresponding decision rules from each candidate node at the current level; respectively determining the current parameter values of each target parameter involved in each decision rule; inputting the decision rules and the current parameter values into a large language model to select a node for which the corresponding decision rule is satisfied; Determining the root cause of the error of the target operation at least based on the nodes selected at the leaf level.
2. The method according to claim 1, wherein, The node description further includes a collection tool for obtaining the target parameters; The step of respectively determining the current parameter values of each target parameter involved in each decision rule includes: Respectively calling the collection tools corresponding to each candidate node to collect and obtain the current parameter values.
3. The method according to claim 1, wherein Each candidate node is a node in the current level that is connected to the node selected in the previous level by an edge.
4. The method according to claim 1, wherein, The step of inputting the decision rules and the current parameter values into the large language model includes: Respectively filling in the current parameter values into the decision rules; Constructing a first prompt word based on the decision rules filled with the current parameter values and their corresponding candidate nodes, where the first prompt word indicates selecting a candidate node for which the corresponding decision rule is satisfied from the candidate nodes; Inputting the first prompt word into the large language model to obtain the node selected at the current level.
5. The method according to claim 1, wherein The step of inputting the decision rules and the current parameter values into the large language model includes: Respectively adding parameter identifiers of each target parameter to the current parameter values; Constructing a second prompt word based on the decision rules and their corresponding candidate nodes, and the current parameter values with added parameter identifiers, where the second prompt word indicates selecting a candidate node for which the corresponding decision rule is satisfied from the candidate nodes based on the current parameter values; Inputting the second prompt word into the large language model to obtain the node selected at the current level.
6. The method according to claim 1, wherein The step of determining the root cause of the error of the target operation includes: Constructing a target prompt word based on the root causes of the errors corresponding to the nodes selected at each level, where the target prompt word indicates summarizing the root causes of the errors corresponding to the nodes; Inputting the target prompt word into the large language model to obtain the root cause of the error of the target operation.
7. The method according to claim 6, wherein, The node description further includes a solution; The target prompt word further includes one or more of the following: The current parameter values of the target parameters involved in the decision rules corresponding to the nodes selected at the target level; The solutions corresponding to the nodes selected at the target level.
8. The method according to claim 1, wherein, The target operation is a machine deployment operation, and the target parameters include log information, environment information on the machine, or configuration information of the machine; or, The target operation is a Pod startup operation, and the target parameters include scheduling events of the Pod; Or, The target operation is a remote procedure call (RPC) operation, and the target parameters include an error code.
9. The method according to claim 1, wherein The node description record is a file in any of the following data formats: YAML format, JSON format, and XML format.
10. A root cause location device, comprising: An acquisition unit configured to acquire a root cause troubleshooting tree corresponding to a target operation where an error occurs. A single node therein represents an error root cause, and the node description uses a mixed manner of natural language and formal language, including a determination rule described in natural language and determined based on target parameters. A selection unit configured to perform node selection level by level along the root cause troubleshooting tree. The selection unit includes: An extraction sub-module configured to extract corresponding determination rules from each candidate node at the current level. A determination sub-module configured to respectively determine the current parameter values of each target parameter involved in each determination rule. An input sub-module configured to input the determination rules and the current parameter values into a large language model to cause it to select a node where the corresponding determination rule is satisfied. A determination unit configured to determine the error root cause of the target operation at least based on the nodes selected at the leaf level.
11. A computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed on a computer, the computer is caused to execute the method according to any one of claims 1-9.
12. A computing device, comprising a memory and a processor, wherein, An executable code is stored in the memory, and when the processor executes the executable code, the method according to any one of claims 1-9 is implemented.