Fault diagnosis method, device, equipment, medium and product
By acquiring and traversing the call chain of the big data platform, and automatically performing fault diagnosis, the problem of low fault diagnosis efficiency in the existing technology is solved, and fast and accurate fault location and business recovery are achieved.
Patent Information
- Application Number
- CN202211616584.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-15
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-12-15
AI Technical Summary
The existing technology relies on manual inspection in the fault diagnosis of big data platforms, resulting in long diagnosis time and low efficiency, and relying on the experience and judgment of operation and maintenance personnel.
By obtaining the call chain of the object to be arranged, the call chain includes path nodes and processor nodes, traversing each node according to the preset traversal rules, obtaining multiple fault diagnosis results, avoiding human judgment and improving the accuracy of diagnosis.
It has achieved rapid positioning of business abnormal causes, improved fault diagnosis efficiency, reduced intervention of operation and maintenance personnel, and improved the rapid recovery ability of business.
Smart Images

Figure CN116016109B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of information security technology, and particularly relates to a fault diagnosis method, apparatus, device, medium and product. Background Art
[0002] With the development of mobile Internet services, the amount of data managed by big data platforms in various industries is increasing, the scale of clusters is getting larger, and the business processes are becoming more and more complex. Whether it is the Internet industry or the traditional information technology industry, when designing a big data platform, components and middleware based on different versions, different open source parties, and different types are constructed. At the same time, the big data business applications based on them involve multiple aspects and are complex.
[0003] To ensure the normal operation of each business, it is necessary to quickly locate the root cause of the failure in a timely manner when an exception occurs in the business. At present, manual inspection is adopted, and fault analysis is carried out relying on expert experience. After the operation and maintenance personnel receive the business process failure information, the general business fault diagnosis process is as follows: 1) Check the reason for the error from the failure log and determine the components or business nodes involved in the error. 2) Log in to each host to view the relevant logs of the components and determine whether there are any abnormalities in each component. If not, spread the inspection to the previous dependent process nodes, and the information viewed for each business node includes cluster resource status, network quality, user permissions, user operations, etc. 3) After collecting as much operation information of relevant component services as possible, the operation and maintenance personnel rely on operation and maintenance experience to infer the root cause of the problem.
[0004] However, using the method of manual detection, the fault diagnosis time is long and it is relatively dependent on the experience judgment of operation and maintenance personnel. During the fault diagnosis process, it is necessary to log in to multiple server nodes separately, consult the logs of multiple components, and finally still need to rely on expert experience for human judgment. That is to say, the current diagnosis method has low diagnosis efficiency. Summary of the Invention
[0005] The embodiments of this application provide a fault diagnosis method, apparatus, device, medium and product, which can improve the fault diagnosis efficiency and is beneficial to the rapid recovery of the business.
[0006] In a first aspect, the embodiments of this application provide a fault diagnosis method, which includes:
[0007] Obtain the call chain of the object to be sorted, where the call chain includes at least one path node and at least one processor node, the path node is a parent node or a child node, the processor node is a leaf node, the object to be sorted includes multiple sub-objects, and each processor node is used to perform fault diagnosis on one of the multiple sub-objects;
[0008] Traverse each node in the call chain according to the preset traversal rule of the call chain to obtain multiple fault diagnosis results.
[0009] In a second aspect, an embodiment of the present application provides a fault diagnosis device, which includes:
[0010] A first acquisition module, configured to acquire a call chain of an object to be arranged, where the call chain includes at least one path node and at least one processor node, the path node is a parent node or a child node, the processor node is a leaf node, the object to be arranged includes multiple sub-objects, and each processor node is used to perform fault diagnosis on one of the multiple sub-objects;
[0011] A second acquisition module, configured to traverse each node in the call chain according to the preset traversal rule of the call chain to obtain multiple fault diagnosis results.
[0012] In a third aspect, an embodiment of the present application provides an electronic device, which includes: a processor and a memory storing computer program instructions;
[0013] When the processor executes the computer program instructions, the method described in the first aspect is implemented.
[0014] In a fourth aspect, an embodiment of the present application provides a computer storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described in the first aspect is implemented.
[0015] In a fifth aspect, an embodiment of the present application provides a computer program product, and when the instructions in the computer program product are executed by a processor of an electronic device, the electronic device is enabled to execute the method described in the first aspect.
[0016] The fault diagnosis method, device, equipment, medium and product of the embodiments of the present application obtain the call chain of the object to be arranged, where the call chain includes at least one path node and at least one processor node, the path node can be a parent node or a child node, the processor node is a leaf node, and each processor node performs fault diagnosis on one of the multiple sub-objects. According to different node types, the call chain adopts different call strategies, which avoids the repeated development of the call chain. According to the preset traversal rule of the call chain, each node in the call chain is traversed, and multiple fault diagnosis results can be obtained after traversal, without the need for manual judgment of the fault results based on experience, which increases the accuracy of judgment. And the traversal rule of the call chain of this method can traverse each node in the call chain, quickly locate the cause of business anomalies, and improve the fault diagnosis efficiency. Description of the Drawings
[0017] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 It is a schematic flowchart of a fault diagnosis method provided by an embodiment of the present application;
[0019] Figure 2 It is a schematic structural diagram of a service call chain in a fault diagnosis method provided by an embodiment of the present application;
[0020] Figure 3 It is a schematic structural diagram of a capability call chain in a fault diagnosis method provided by an embodiment of the present application;
[0021] Figure 4 It is a schematic flowchart of a call chain fault traceability process in a fault diagnosis method provided by an embodiment of the present application;
[0022] Figure 5 It is a schematic structural diagram of the traceability logic of a fault diagnosis method provided by an embodiment of the present application;
[0023] Figure 6 It is a functional module diagram of a fault diagnosis method provided by an embodiment of the present application;
[0024] Figure 7 It is a schematic structural diagram of a fault diagnosis device provided by an embodiment of the present application;
[0025] Figure 8 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Specific Embodiments
[0026] The following will describe in detail the features and exemplary embodiments of various aspects of the present application. To make the objectives, technical solutions, and advantages of the present application clearer, the following further describes the present application in detail with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without some of these specific details. The following description of the embodiments is only intended to provide a better understanding of the present application by showing examples of the present application.
[0027] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0028] To solve the problems of the prior art, embodiments of the present application provide a fault diagnosis method, apparatus, device, medium and product. The following will, with reference to the accompanying drawings, describe in detail the fault diagnosis method provided by the embodiments of the present application through specific embodiments and their application scenarios.
[0029] Figure 1 It is a schematic flowchart of the fault diagnosis method provided by an embodiment of the present application. As Figure 1 shown, the fault diagnosis method provided by the embodiments of the present application may include steps S110 - S120, where:
[0030] S110. Obtain the call chain of the object to be sorted. The call chain includes at least one path node and at least one processor node. The path node is a parent node or a child node, and the processor node is a leaf node. The object to be sorted includes multiple sub-objects, and each processor node is used to perform fault diagnosis on one of the multiple sub-objects; all leaf nodes are processor nodes.
[0031] S120. Traverse each node in the call chain according to a preset traversal rule of the call chain to obtain multiple fault diagnosis results.
[0032] Thus, by obtaining the call chain of the object to be sorted, where the call chain includes at least one path node and at least one processor node, the path node can be a parent node or a child node, the processor node is a leaf node, and each processor node will perform fault diagnosis on one of the multiple sub-objects. According to different node types, the call chain will adopt different call strategies, thus avoiding duplicate development of the call chain. According to the preset traversal rule of the call chain, each node in the call chain is traversed. After traversal, multiple fault diagnosis results can be obtained, eliminating the need for manual judgment of fault results based on experience, increasing the accuracy of judgment, and the traversal rule of the call chain of this method can traverse each node in the call chain, quickly locate the cause of business anomalies, and improve the efficiency of fault diagnosis.
[0033] The specific implementation methods of the above steps are introduced below.
[0034] In an embodiment of the present application, in S110, the object to be sorted can be a big data platform, including multiple sub-objects. The sub-objects can be devices, components, modules, etc. in the big data platform. Among them, the call chain includes at least one path node and at least one processor node. The path node can be either a parent node or a child node. It should be noted that the path node cannot be a leaf node, and the processor node is a leaf node.
[0035] The call chain is a storage structure similar to the file system in the Linux operating system. The call chain can be created in a custom manner. By combining the characteristics of internal multi-node types and multi-types of call chains with the operation and maintenance knowledge graph, new business scenarios can be quickly matched. The common call chain execution logic starts from the root node of the call chain. It should be noted that the root node of the call chain is also a path node. The call chain also includes reference nodes. The parent node of the reference node is a node in at least one path node. The child nodes of the reference node can be path nodes and processor nodes obtained from other call chains, or processor nodes obtained from other call chains.
[0036] The call chain development and management module is used for the development of the call chain. For the rapid development of business scenarios, according to the characteristics that some components or services are universal during the platform operation and maintenance process, two methods are designed respectively when constructing the call chain: business call chain and capability call chain. For example, the capability call chain can reference the business call chain of other objects to be sorted through reference nodes. At this time, the business call chain of other objects to be sorted also becomes a capability call chain.
[0037] Such as Figure 2 and Figure 3 shown, Figure 2 is a schematic structural diagram of the business call chain in the fault diagnosis method provided by an embodiment of the present application, Figure 3 is a schematic structural diagram of the capability call chain in the fault diagnosis method provided by an embodiment of the present application.
[0038] The business call chain is used to define the chain for performing analysis and positioning when a business exception occurs.
[0039] The capability call chain is used to define the analysis capability chain shared by the business and is referenced in the business call chain.
[0040] Exemplarily, a path node is used to logically layer nodes, similar to a directory in a file system. A path node must have child nodes, and the child nodes can be either path nodes or processor nodes; a processor node consists of a processor program (the processor program is a section of traceability diagnosis logic) and the custom traceability parameters of the node (the traceability parameters are represented in key-value pair format). A processor node has no child nodes, similar to a file in a file system. Each processor node can perform fault diagnosis on one of multiple sub-objects, where multiple processor nodes can diagnose different faults of the same sub-object, or one processor node can diagnose multiple different faults of one sub-object.
[0041] The following steps are also included before this step:
[0042] Construct path nodes and processor nodes. The path nodes are used to logically layer the nodes in the call chain, and the processor nodes include a diagnostic program for troubleshooting sub-objects.
[0043] In the above, in addition to path nodes and processor nodes, the call chain also includes reference nodes. The development of traceability nodes is carried out using a traceability node development module. Among them, there are three types of traceability nodes: path nodes, processor nodes, and reference nodes, which can increase the applicable scenarios of node functions. A reference node is used to point to a capability call chain. The parent node of the reference node is a node in at least one path node, and the child nodes of the reference node include target nodes, where the target nodes include path nodes and processor nodes obtained from other call chains, or processor nodes obtained from other call chains, which is not limited here.
[0044] Exemplarily, as Figure 2 and Figure 3 shown, if Figure 2 is a certain business call chain A, there are a call chain root node, path nodes, and processor nodes in the business call chain A; in the case where there are no reference nodes, Figure 3 is the business call chain B. At this time, there are also a call chain root node, path nodes, and processor nodes in the business call chain B. If it is found that the business call chain B can reuse the call chain of the business call chain A, so the business call chain A can be directly reused through a reference node, which can achieve the purpose of quickly developing the business call chain. At this time, the business call chain B also becomes a capability call chain, and at this time, there are a call chain root node, path nodes, processor nodes, and reference nodes in the business call chain B.
[0045] In the front-end interface, various nodes can be added to the left side of the interface by creating a new call chain. The hierarchical nodes are divided and assigned according to the relationships set by the business, and finally the left call chain is formed. For the development of each node, naming, uploading the script file executed when tracing or pathfinding for this node, and the dependent data source can be carried out. Compared with other path nodes and reference nodes, the development of processor nodes is more complex. The processor node information accessor writes the processor program into the distributed file system, and at the same time stores the resource metadata of the processor program (such as the main class of the processor program, the custom parameters used by the processor program, the data source used by the processor, the network resource location of the processor program) into the database.
[0046] The processor program loader can, according to the resource metadata of the processor program, use the class loader to load the processor program into the memory according to the network resource location and the fully qualified class name of the processor program, and instantiate the processor program. Considering that the processor program package is a compressed (Java Archive, JAR) package containing other dependencies, to prevent dependency conflicts, the processor program class loader provides a class loader for each JAR package at the network resource location, so as to achieve dependency isolation between different processor programs. At the same time, to ensure that the classes that have been loaded do not need to be reloaded when called multiple times in a short period, the processor program class loader will cache a certain number of class loaders corresponding to the network resource locations, and use the Least Recently Used (LRU) algorithm to clear the class loaders that exceed the quantity range.
[0047] At the same time, to ensure that the modified processor program takes effect immediately, when the modified processor program is re-uploaded to the processor module, the processor program loader will remove its corresponding class loader from the cache. When instantiating the processor program object, the processor program loader first obtains the class loader from the cache, and then uses the class loader to obtain the processor program object through reflection. If it cannot be obtained, a class loader is instantiated and added to the cache.
[0048] In an embodiment of the present application, in S120, according to the preset traversal rules of the call chain, each node in the call chain is traversed to obtain multiple fault diagnosis results, including the following three steps:
[0049] The first step: For each first path node, sort the child nodes of the first path node according to the traversal rules to obtain a first order, and the first path node is any one of at least one path node.
[0050] The traversal rules of the preset call chain include starting from the root node of the call chain. When there are multiple parent nodes under the root node, there is a sequential order among the multiple parent nodes. During execution, after traversing the previous parent node and all the processor nodes under the previous parent node in accordance with this sequential order, the next parent node and all the processor nodes under the next parent node are traversed.
[0051] Exemplarily, the first path node has multiple child nodes. The first path may include path node a, path node b, etc. For example, sort the multiple child nodes under path node a and execute the first child node first to obtain a first order. The first path node is any one of at least one path node, such as Figure 4 shown Figure 4 is a schematic diagram of the call chain fault tracing process in a fault diagnosis method provided by an embodiment of the present application. The thick solid lines in the figure represent multiple child nodes in the tracing process, and the dotted lines represent the context information of the tracing process and the processor node parameters. Figure 4 The 1-10 in it represents the process of call chain fault tracing. Starting from the root node, first execute the first path node, such as path node a. Path node a has multiple child nodes. First execute the first child node, that is, processor node a, and then execute the second child node, that is, processor node b. Traverse the child nodes of the first path node according to the traversal rules, and traverse them in sequence according to the execution principle and the sequential numbers.
[0052] Second step: Traverse the child nodes of the first path node in sequence according to the first order until the traversed child node is the first processor node, then trigger the first processor node to execute the diagnostic program of the first processor node to obtain the first fault diagnosis result.
[0053] The first processor node is all the processor nodes under the first path, such as Figure 4 shown, the first processor node may include processor node a, processor node b, etc.
[0054] Exemplarily, as Figure 5 shown Figure 5It is the traceability logic structure diagram of a fault diagnosis method provided by an embodiment of the present application, including the call chain traceability logic process, the call chain analysis traceability engine logic process, and the logic process of the in-node processing logic engine. They are executed in sequence according to the process. According to the first order mentioned above, the child nodes of the first path node are traversed in sequence. The traversed child nodes are the first processor nodes. Processor node a has no child nodes, and processor node b also has no child nodes, and both contain a processor. At this time, the first processor node is triggered to execute the diagnostic program of the first processor node. After the execution process of processor node a ends, the subsequent sibling node of processor node a, that is, processor node b, is executed. After all the child nodes in path node a have been executed, the first fault diagnosis results of all the child nodes are obtained. The first fault diagnosis result includes the fault level. The severity of a fault with a higher fault level is greater than that of a fault with a lower fault level. The fault levels are normal, warning, and severe, which are represented by 0, 1, and 2 respectively. If the fault level is severe, it is considered that the cause of the service failure has been found, and the traversal process ends; if the rating is normal or warning, the traceability traversal of the call chain will continue.
[0055] This step may specifically include the following steps:
[0056] Transfer the preset context information to the processor program loader to obtain the processor object of the first processor node. The context information includes the identifier of the first processor node and the identifier of the first sub-object used by the first processor node for fault diagnosis;
[0057] Transfer the parameter list of the first processor node to the processor object. The processor object executes the diagnostic program of the first processor node and records the obtained first fault diagnosis result in the context information. The parameter list includes the data obtained from the first sub-object and the reference value.
[0058] In the above, the context information includes the identifier of the first processor node, such as processor node a, processor node b, etc., and also includes the identifier of the first sub-object used by the first processor node for fault diagnosis, such as device a, device b, component a, component b, etc. The processor node includes a processor, which will pass the traceability process context information to the processor program loader during execution. The call chain executor will pass the processor node information contained in the first processor node to the processor program loader, obtain a processor object of the first processor node, and then pass the first processor node custom parameters and alarm information to the processor object. The processor object executes the diagnostic program. After the diagnostic function in the diagnostic program is executed, a series of diagnostic indicators and the severity level of these diagnostic indicators will be output, and the fault diagnosis results will be recorded in the context information. In addition, the processor management module used in the embodiment of the present application can be freely defined according to the actual business logic, and business stakeholders can deeply participate in the customization of the traceability logic. The processor program in the call chain is pluggable and supports hot loading. New or modified processor programs do not need to be redeployed or restarted. The program logic can be replaced by re-uploading the call chain program in the processor management module. The traceability logic will be used the next time the traceability process is triggered. By pre-defining the processor cluster and importing the task chain lineage data, the call chain object corresponding to the task chain can be automatically generated without manual intervention to build the call chain object.
[0059] In addition, when the processor program is executed, the parameter list of the first processor node is passed to the processor object. The parameter list includes data such as abnormal data and reference values such as reference thresholds obtained from the first sub-object. The diagnostic program of the first processor node is executed by the processing object, and the diagnostic program executes the function. After the execution is completed, the analyzed indicator items, indicator values and the abnormal level of the indicator are recorded in the traceability process context information, and the maximum abnormal level is updated to the node state of the node. At this point, the execution process of processor node a ends. After the execution process of processor node a ends, the subsequent brother node of node a is executed, that is, processor node b. Similarly, the context information and the parameter list on processor node b are passed to the diagnostic program in processor node b, and the function is executed to obtain the indicator items, indicator values and the abnormal level of the indicator. At this point, the execution process of processor node b ends, and the first fault diagnosis result of the first processor node is obtained.
[0060] This step is followed by the following steps:
[0061] When the first fault diagnosis result of the first processor node includes a fault level, and the fault level is greater than a preset level, the traversal process is terminated.
[0062] Exemplarily, starting from the root node of the call chain, first determine whether the current node is a path node or a processor node. If it is a processor node, the call chain executor will pass the processor node information contained in the processor node to the processor program loader to obtain a processor object, and then pass the custom parameters and alarm information of the processor node to the processor object, and the processor object will execute the diagnostic function. After the diagnostic function is executed, a series of diagnostic metrics and the severity levels of these diagnostic metrics will be output. This level is called the fault level, and the fault level values are normal, warning, and severe, which are represented by 0, 1, and 2 respectively. If the fault level is greater than the preset level and is rated as severe, it is considered that the reason for the task failure has been found, and the tracing process ends.
[0063] Step 3: Determine the fault diagnosis result of the first path node according to the first fault diagnosis results of all the first processor nodes under the first path node.
[0064] Exemplarily, for example, the call chain engine passes the tracing process context information to the root node of the call chain. First, the first path node a is executed. After all the child nodes in the path node a are executed, the node status of the node with the highest fault level among all the child nodes is obtained and updated to the node status of the path node a. Similarly, after all the child nodes in the path node b are executed, the node status of the node with the highest fault level among all the child nodes is taken and updated to the node status of the path node b. The maximum fault levels of the path node a and the path node b are updated to the first path node status to determine the fault diagnosis result of the first path node.
[0065] This step may specifically include the following steps:
[0066] Take the maximum fault level among the first fault diagnosis results of all the first processor nodes of the first path node as the fault diagnosis result of the first path node.
[0067] In the above, for the first path node, the call chain executor sorts its child nodes according to the sequence labels of the child nodes, and then executes each child node in sequence. When all the child nodes are executed, the maximum fault level of all the child nodes is taken and updated to the first path node status. And the execution result of the processor function will be stored in the system cache first until the tracing process ends and the execution result is stored in the database.
[0068] Exemplarily, when all the child nodes of the call chain root node are executed, the node status of the node with the maximum fault level of the child nodes is updated to the root node. When the execution process of the call chain root node ends, the call chain process ends, and the devices, components, and modules in the big data platform are diagnosed according to the obtained node statuses of multiple nodes.
[0069] In addition, the call chain traceability process can achieve visual full-link analysis and provide visual full-link traceability analysis results. First, the call chain tree structure diagram defined by the call chain development management module is used as the background display, and then the analysis results of the traceability processors configured for each leaf node are displayed on each leaf node. Users can view the traceability analysis results of all links in the full business process in one layout.
[0070] To describe the full-scenario and full-link fault analysis method of the fault diagnosis method provided by the embodiments of the present application in a more comprehensive and detailed manner, it can be used for full-scenario and full-link fault analysis of complex and diverse big data services.
[0071] As Figure 6 shown, Figure 6 The functional module diagram of the fault diagnosis method provided by an embodiment of the present application mainly includes a call chain traceability driver module, a traceability node development module, a call chain development management module, a data source management module, and an operation and maintenance knowledge base.
[0072] In an embodiment of the present application, the business traceability trigger node is used to trigger the call chain traceability driver module during business fault troubleshooting;
[0073] The call chain traceability driver module is used to trigger the call chain development management module after receiving the trigger task from the business traceability trigger node, and drive the call chain object to automatically troubleshoot problems, without the need to manually check components or task information one by one according to task lineage;
[0074] The call chain development management module is used to form a call chain by connecting business relationships in combination with traceability nodes, including a business call chain and a capability call chain;
[0075] The operation and maintenance knowledge base includes storing various resources such as hosts, components, platforms, services, etc. and their ownership relationships, various corresponding metrics, log data, and common general exception event information using a knowledge graph database. The operation and maintenance knowledge base converts the networking topology architecture characteristics of components and services involved in the big data platform business into an entity relationship diagram structure, and associates the metrics, alarms, and logs of all-level running objects of the business scenario running instances, and extracts the metadata of the criterion data. Among them, the running objects include hardware, OS, tools, etc., and are also converted into an entity relationship diagram structure.
[0076] Through the flag of the entity relationship, the abstraction and association of all-level running entities of the business are realized, and the dynamic association of the call chain and operation and maintenance data is carried out. Store the relationship mapping of the networking architecture, components of existing tools, and service composition; solidify the entity set of the entity relationship diagram data structure using the graph database entity set to realize the reuse of the tool service-level call chain construction module.
[0077] The data source management module is used to connect the traceability node development module and the operation and maintenance knowledge base. The traceability node development module needs to associate relevant data during development and use.
[0078] The traceability node development module is used to develop various traceability nodes, including path nodes, processor nodes, and reference nodes.
[0079] The traceability node development module, call chain development management module, data source management module, operation and maintenance knowledge base module, call chain traceability drive engine, and traceability result display module included in the embodiments of the present application can meet the requirements of rapid development and multi-scenario analysis by designing three types of nodes in the call chain in the traceability node development module and forming two call chains through the call chain development management module to connect the business relationships of the traceability nodes. At the same time, a call chain fault diagnosis method is designed to make the fault traceability analysis more accurate and reasonable.
[0080] Figure 7 The structure diagram of the fault diagnosis device provided by the embodiments of the present application is shown. As Figure 7 shown, the fault diagnosis device 700 includes:
[0081] The first acquisition module 701 is used to acquire the call chain of the object to be sorted. The call chain includes at least one path node and at least one processor node. The path node is a parent node or a child node, and the processor node is a leaf node. The object to be sorted includes multiple sub-objects, and each processor node is used to perform fault diagnosis on one of the multiple sub-objects.
[0082] The second acquisition module 702 is used to traverse each node in the call chain according to the preset traversal rule of the call chain to obtain multiple fault diagnosis results.
[0083] Therefore, by acquiring the call chain of the object to be sorted, where the call chain includes at least one path node and at least one processor node, the path node can be a parent node or a child node, the processor node is a leaf node, and each processor node will perform fault diagnosis on one of the multiple sub-objects. According to different node types, the call chain will adopt different call strategies, which avoids the repeated development of the call chain. According to the preset traversal rule of the call chain, each node in the call chain is traversed. After traversal, multiple fault diagnosis results can be obtained, and there is no need to manually judge the fault results based on experience, which increases the accuracy of judgment. And the traversal rule of the call chain of this method can traverse each node in the call chain, quickly locate the cause of business anomalies, and improve the efficiency of fault troubleshooting.
[0084] In an embodiment of the present application, the call chain further includes a reference node. The parent node of the reference node is a node in at least one path node, and the child nodes of the reference node include a target node. The target node includes path nodes and processor nodes obtained from other call chains, or processor nodes obtained from other call chains.
[0085] In an embodiment of the present application, the call chain includes a service call chain set according to the service of the object to be sorted, or the call chain includes a capability call chain that reuses the service call chain of other objects to be sorted.
[0086] In an embodiment of the present application, the second acquisition module 702 includes:
[0087] An acquisition sub-module, configured to, for each first path node, sort the child nodes of the first path node according to a traversal rule to obtain a first order. The first path node is any one of at least one path node;
[0088] A trigger sub-module, configured to sequentially traverse the child nodes of the first path node in the first order until the traversed child node is a first processor node, and then trigger the first processor node to execute the diagnostic program of the first processor node to obtain a first fault diagnosis result;
[0089] A determination sub-module, configured to determine the fault diagnosis result of the first path node according to the first fault diagnosis results of all first processor nodes under the first path node.
[0090] In an embodiment of the present application, the first fault diagnosis result includes a fault level, and the severity of a fault with a larger fault level is greater than the severity of a fault with a smaller fault level; the determination sub-module includes:
[0091] Taking the maximum fault level among the first fault diagnosis results of all first processor nodes of the first path node as the fault diagnosis result of the first path node.
[0092] In an embodiment of the present application, the trigger sub-module includes:
[0093] Transmitting preset context information to a processor program loader to obtain a processor object of the first processor node. The context information includes the identifier of the first processor node and the identifier of the first sub-object used by the first processor node for fault diagnosis;
[0094] Transmitting the parameter list of the first processor node to the processor object, and having the processing object execute the diagnostic program of the first processor node and record the obtained first fault diagnosis result in the context information. The parameter list includes data obtained from the first sub-object and reference values.
[0095] In one embodiment of the present application, the apparatus 700 further includes:
[0096] An end module, configured to end the traversal process when the first fault diagnosis result of the first processor node includes a fault level and the fault level is greater than a preset level.
[0097] In one embodiment of the present application, the apparatus 700 further includes:
[0098] A construction module, configured to construct path nodes and processor nodes, where the path nodes are used to perform logical layering on the nodes in the call chain, and the processor nodes include diagnostic programs for troubleshooting sub-objects.
[0099] The fault diagnosis apparatus 700 provided by the embodiments of the present application can implement each process implemented by the foregoing embodiments of the fault diagnosis method. To avoid repetition, it will not be elaborated here.
[0100] Figure 8 The hardware structure diagram of the fault diagnosis method provided by the embodiments of the present application is shown.
[0101] The electronic device may include a processor 801 and a memory 802 storing computer program instructions.
[0102] Specifically, the foregoing processor 801 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0103] The memory 802 may include a mass storage for data or instructions. By way of example and not limitation, the memory 802 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In a suitable case, the memory 802 may include a removable or non-removable (or fixed) medium. In a suitable case, the memory 802 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 802 is a non-volatile solid state memory.
[0104] The memory may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage media device, an optical storage media device, a flash memory device, an electrical, optical, or other physical / tangible memory storage device. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to the first aspect or the second aspect of the present disclosure.
[0105] The processor 801 realizes any one of the fault diagnosis methods in the above embodiments by reading and executing the computer program instructions stored in the memory 802.
[0106] In one example, the electronic device 800 may further include a communication interface 803 and a bus 810. Among them, as Figure 8 shown, the processor 801, the memory 802, and the communication interface 803 are connected through the bus 810 to complete communication with each other.
[0107] The communication interface 803 is mainly used to implement communication between various modules, devices, units, and / or devices in the embodiments of the present application.
[0108] The bus 810 includes hardware, software, or both, and couples the components of the fault diagnosis method or verification device to each other. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses or a combination of two or more of these. In a suitable case, the bus 810 may include one or more buses. Although the embodiments of the present application describe and illustrate specific buses, the present application contemplates any suitable bus or interconnect.
[0109] In addition, in combination with the fault diagnosis method in the above embodiments, the embodiments of the present application may provide a computer storage medium to implement. Computer program instructions are stored on the computer storage medium; when the computer program instructions are executed by a processor, any one of the fault diagnosis methods in the above embodiments is realized.
[0110] It should be clear that the present application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present application.
[0111] The functional blocks shown in the above-described structural block diagrams can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave over a transmission medium or a communication link. A "machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical discs, hard disks, fiber optic media, radio frequency (RF) links, and so on. The code segment can be downloaded via a computer network such as the Internet, an intranet, and so on.
[0112] It should also be noted that the exemplary embodiments mentioned in the present application describe some methods or systems based on a series of steps or devices. However, the present application is not limited to the order of the above steps, that is, the steps can be executed in the order mentioned in the embodiments, or different from the order in the embodiments, or several steps can be executed simultaneously.
[0113] Aspects of the present disclosure have been described above with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block in the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices to produce a machine such that the instructions executed by the processor of the computer or other programmable data processing devices enable the implementation of the functions / actions specified in one or more blocks of the flowcharts and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It can also be understood that each block in the block diagrams and / or flowcharts, and the combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by dedicated hardware that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0114] As described above, this is only the specific implementation manner of the present application. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein. It should be understood that the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present application.
Claims
1. A fault diagnosis method, characterized in that, the method includes: obtaining a call chain of an object to be troubleshot, the call chain including at least one path node and at least one processor node, the path node being a parent node or a child node, the processor node being a leaf node, the object to be troubleshot including multiple sub-objects, and each processor node being used to perform a fault diagnosis on one of the multiple sub-objects; traversing each node in the call chain according to a preset traversal rule of the call chain to obtain multiple fault diagnosis results; the traversing each node in the call chain according to a preset traversal rule of the call chain to obtain multiple fault diagnosis results includes: for each first path node, sorting the child nodes of the first path node according to the traversal rule to obtain a first order, the first path node being any one of the at least one path node; sequentially traversing the child nodes of the first path node according to the first order until the traversed child node is a first processor node, then triggering the first processor node to execute the diagnostic program of the first processor node to obtain a first fault diagnosis result; determining the fault diagnosis result of the first path node according to the first fault diagnosis results of all the first processor nodes under the first path node.
2. The method according to claim 1, characterized in that, the call chain further includes a reference node, the parent node of the reference node being a node among the at least one path node, the child nodes of the reference node including a target node, and the target node including path nodes and processor nodes obtained from other call chains, or processor nodes obtained from other call chains.
3. The method according to claim 1 or 2, characterized in that, the call chain includes a service call chain set according to the service of the object to be troubleshot, or the call chain includes a capability call chain that reuses the service call chain of other objects to be troubleshot.
4. The method according to claim 1, characterized in that, the first fault diagnosis result includes a fault level, and the severity of a fault with a larger fault level is greater than the severity of a fault with a smaller fault level; the determining the fault diagnosis result of the first path node according to the first fault diagnosis results of all the first processor nodes under the first path node includes: taking the maximum fault level among the first fault diagnosis results of all the first processor nodes of the first path node as the fault diagnosis result of the first path node.
5. The method according to claim 1, characterized in that, triggering the first processor node to execute the diagnostic program of the first processor node to obtain a first fault diagnosis result includes: transmitting preset context information to a processor program loader to obtain a processor object of the first processor node, the context information including the identifier of the first processor node and the identifier of the first sub-object used by the first processor node for fault diagnosis. Pass the parameter list of the first processor node to the processor object, execute the diagnostic program of the first processor node by the processor object, and record the obtained first fault diagnosis result in the context information. The parameter list includes data obtained from the first sub-object and reference values.
6. The method according to claim 1, wherein, after obtaining the first fault diagnosis result, the method further includes: ending the traversal process when the first fault diagnosis result of the first processor node includes a fault level and the fault level is greater than a preset level.
7. The method according to claim 1, wherein, the traversal rule includes: starting from the root node of the call chain for traversal. If there are multiple parent nodes under the root node and there is a sequence order among the multiple parent nodes, after traversing the previous parent node and all sub-processor nodes under the previous parent node in accordance with the sequence order, then traverse the next parent node and all processor nodes under the next parent node.
8. The method according to claim 1, wherein, before obtaining the call chain of the object to be sorted, the method further includes: constructing the path node and the processor node. The path node is used for logically hierarchical partitioning of the nodes in the call chain, and the processor node includes a diagnostic program for troubleshooting sub-objects.
9. A fault diagnosis device, wherein, the device includes: a first acquisition module, configured to acquire a call chain of an object to be sorted. The call chain includes at least one path node and at least one processor node. The path node is a parent node or a sub-node, and the processor node is a leaf node. The object to be sorted includes multiple sub-objects, and each processor node is used for fault diagnosis of one of the multiple sub-objects; a second acquisition module, configured to traverse each node in the call chain according to a preset traversal rule of the call chain to obtain multiple fault diagnosis results. The traversing each node in the call chain according to a preset traversal rule of the call chain to obtain multiple fault diagnosis results includes: for each first path node, sorting the sub-nodes of the first path node according to the traversal rule to obtain a first order, where the first path node is any one of the at least one path node; traversing the sub-nodes of the first path node in sequence according to the first order until the traversed sub-node is a first processor node, then triggering the first processor node to execute the diagnostic program of the first processor node to obtain a first fault diagnosis result; determining the fault diagnosis result of the first path node according to the first fault diagnosis results of all first processor nodes under the first path node.
10. An electronic device, wherein, the electronic device includes: a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps of the fault diagnosis method according to any one of claims 1-8 are implemented.
11. A computer-readable storage medium, characterized in that, computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are executed by a processor, the method described in any one of claims 1-8 is implemented.
12. A computer program product, characterized in that, when the instructions in the computer program product are executed by a processor of an electronic device, the electronic device is caused to execute the method described in any one of claims 1-8.
Citation Information
Patent Citations
Service node fault positioning method, call chain generation method and server
CN112737800A
Fault diagnosis method and device and electronic equipment
CN114330138A