Network fault diagnosis method, electronic device, storage medium, and program product

By calling the executable diagnostic interface of potential cause nodes and the pre-trained language model in the fault knowledge graph, the problems of strong subjectivity and insufficient generality in knowledge extraction of the fault knowledge graph are solved, and efficient and accurate network fault diagnosis is achieved.

WO2026021106A1PCT designated stage Publication Date: 2026-01-29ZTE CORP

Patent Information

Application Number
PCT/CN2025/103520
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-22
Filing Date
2025-06-25
Publication Date
2026-01-29

Smart Images

  • Figure CN2025103520_29012026_PF_FP_ABST
    Figure CN2025103520_29012026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of network fault diagnosis, and discloses a network fault diagnosis method, an electronic device, a storage medium, and a program product. The method comprises: acquiring fault diagnosis trigger information, the fault diagnosis trigger information comprising a resource name and a fault identifier; determining resource instance information corresponding to the resource name, a target node matching the fault identifier in a fault knowledge graph, and at least one potential cause node having a causal relationship with the target node; and invoking an executable diagnosis interface corresponding to each potential cause node to execute target resource instance information corresponding to each potential cause node, and when it is verified that a potential cause corresponding to the potential cause node is established, and the node type of the potential cause node is a preset root cause node, acquiring a fault diagnosis result corresponding to the potential cause node. In this way, the accuracy and credibility of fault diagnosis can be improved, and complex and accurate fault diagnosis tasks can be implemented.
Need to check novelty before this filing date? Find Prior Art

Description

Network fault diagnosis method, electronic device, storage medium and program product

[0001] Cross-reference to related applications

[0002] The present application claims priority to the Chinese patent application No. 202410979481.9, filed on July 22, 2024, and entitled "Network fault diagnosis method, electronic device, storage medium and program product", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0003] Embodiments of the present application relate to the technical field of network fault diagnosis, in particular to a network fault diagnosis method, an electronic device, a storage medium and a program product. BACKGROUND

[0004] With the increasing scale of the bearing network, the structure of the network is becoming more and more complex. When the network fails, a lot of manpower needs to be invested for fault diagnosis, the operation and maintenance cost is very high, and the fault diagnosis period is long, and the fault positioning efficiency is low.

[0005] At present, network fault diagnosis can be based on a knowledge graph, that is, fault operation and maintenance related knowledge is constructed into a corresponding fault knowledge graph, and then fault diagnosis and reasoning are performed according to the constructed fault knowledge graph. However, the fault knowledge graph usually faces problems such as field boundary limitation, limited data size, and sparse data, and its knowledge completeness is insufficient, which affects the accuracy and reliability of the fault diagnosis result. SUMMARY

[0006] Embodiments of the present application provide a network fault diagnosis method, an electronic device, a storage medium and a program product.

[0007] In a first aspect, the embodiments of the present application provide a network fault diagnosis method, comprising: obtaining fault diagnosis trigger information, wherein the fault diagnosis trigger information comprises a resource name and a fault identifier; determining resource instance information corresponding to the resource name, and a target node in a fault knowledge graph that matches the fault identifier and at least one potential cause node having a causal relationship with the target node; executing target resource instance information corresponding to each of the potential cause nodes in the resource instance information by calling an executable diagnosis interface corresponding to each of the potential cause nodes; and obtaining a fault diagnosis result corresponding to the potential cause node in a case where the potential cause corresponding to the potential cause node is verified to be established, and the node type of the potential cause node is a preset root cause node.

[0008] In a second aspect, an electronic device is provided, which includes a processor and a memory. The memory stores programs or instructions executable on the processor. When the programs or instructions are executed by the processor, the steps of the method according to the first aspect are implemented.

[0009] In a third aspect, a computer readable storage medium is provided, which stores programs or instructions. When the programs or instructions are executed by a processor, the steps of the method according to the first aspect are implemented.

[0010] In a fourth aspect, a computer program product is provided, which includes a computer program stored on a non-transitory computer readable storage medium. The computer program includes program instructions, which, when executed by a computer, cause the computer to perform the steps of the method according to the first aspect.

[0011] It should be understood that the general description above and the detailed description below are only exemplary and explanatory, and are not limiting of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0012] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and serve to explain the principles of the present application, together with the description.

[0013] FIG. 1 shows a flowchart of a network fault diagnosis method according to an embodiment of the present application;

[0014] FIG. 2 shows another flowchart of a network fault diagnosis method according to an embodiment of the present application;

[0015] FIG. 3 shows still another flowchart of a network fault diagnosis method according to an embodiment of the present application;

[0016] FIG. 4 shows a flowchart of a fault knowledge graph generation method according to an embodiment of the present application;

[0017] FIG. 5 shows another flowchart of a fault knowledge graph generation method according to an embodiment of the present application;

[0018] FIG. 6 shows an example of a fault knowledge graph according to an embodiment of the present application;

[0019] FIG. 7 shows still another flowchart of a network fault diagnosis method according to an embodiment of the present application;

[0020] FIG. 8 shows a structural block diagram of an online fault diagnosis reasoning system according to an embodiment of the present application;

[0021] FIG. 9 shows a structural block diagram of an offline knowledge mining production system according to an embodiment of the present application;

[0022] FIG. 10 shows a structural schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0023] The exemplary embodiments will be described in detail herein with reference to the attached drawings. In the following description, unless otherwise indicated, like numbers in the attached drawings refer to the same or similar elements. The embodiments described in the following exemplary embodiments are not meant to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the present application as detailed in the appended claims.

[0024] Network fault diagnosis based on a knowledge graph is to construct fault operation and maintenance related knowledge into a corresponding fault knowledge graph, and then to perform fault diagnosis and reasoning according to the constructed fault knowledge graph. The effect of implementing this scheme depends on the completeness of the constructed fault knowledge graph. This approach at least has the following problems:

[0025] 1) There are great difficulties in constructing a fault knowledge graph. Knowledge extraction of a fault knowledge graph mainly depends on manual work, which is highly subjective, and the knowledge extraction is difficult and the quality is difficult to guarantee;

[0026] 2) The knowledge of the fault knowledge graph is not universal. The fault knowledge graph has strong industry attributes and domain specificity, and lacks universality and generalization ability;

[0027] 3) The knowledge of the fault knowledge graph is not complete. The fault knowledge graph usually faces the problems of limited domain boundary, limited data size, and sparsity in data, resulting in insufficient knowledge completeness;

[0028] 4) Semantic understanding and natural language processing are difficult. The knowledge graph still lacks effective technical means when facing problems such as semantic ambiguity in natural language, context understanding, and language common sense reasoning.

[0029] In view of the problems existing in the network fault diagnosis process, the embodiments of the present application provide a network fault diagnosis method. When the fault diagnosis trigger information is acquired, the method executes the target resource instance information corresponding to each potential cause node in the fault knowledge graph based on the executable diagnosis interface corresponding to each potential cause node in the fault knowledge graph, and then locates the root cause node causing the fault and acquires the corresponding fault diagnosis result, so as to improve the accuracy and reliability of fault diagnosis.

[0030] Please refer to FIG. 1, which shows a flowchart of a network fault diagnosis method provided by an embodiment of the present application. The execution subject of the method can be a terminal device or a server, where the terminal device can be a device such as a personal computer, or a mobile terminal device such as a mobile phone or a tablet computer, and the terminal device can be a terminal device used by a user. The server can be a standalone server, or a server cluster composed of multiple servers, and the server can be a background server of a certain service, or a background server of a certain platform or application (for example, a network management system, a fault monitoring and management platform, a network performance monitoring tool, etc.). The embodiment of the present application is mainly used in the application scenario of network fault diagnosis, and the specific network can be any kind of network, including but not limited to the following networks, such as a packet transport network (PTN), a slicing packet network (SPN), an IP radio access network (IPRAN), an optical transport network (OTN), a time synchronization network, etc. If a network fails, the network fault diagnosis method 100 provided by the embodiment of the present application can be used for fault diagnosis. In the embodiment of the present application, the execution subject is taken as an example of a server. For the case of a terminal device, the relevant content described below can be processed, and details are not described herein. As shown in the figure, the network fault diagnosis method 100 can include the following steps:

[0031] S101: Obtain fault diagnosis trigger information, where the fault diagnosis trigger information includes a resource name and a fault identifier.

[0032] In an example embodiment, when a network fault or anomaly is detected, fault diagnosis trigger information of the network fault or anomaly is acquired, which can be an abnormal event (e.g., an alarm, a state anomaly, a performance anomaly, a log anomaly, etc.) monitored in real time, or information input or fed back by a user, and the fault diagnosis trigger information includes a resource name and a fault identifier. For example, in the field of PTN network fault, the resource name can include, for example, a Layer 2 Virtual Private Network (L2VPN), a Time Division Multiplexing (TDM), a pseudo-wire, a tunnel, a Traffic Management System (TMS), a link, a port, or a network element name, etc., and the fault identifier can be a fault name or an event name, which can include, for example, an L2VPN service interruption, an L2VPN service packet loss, an Ethernet physical port Loss of Signal (LOS) alarm, etc. When the fault diagnosis trigger information is acquired, fault information identification is performed on the fault diagnosis trigger information to identify the resource name and the fault identifier in the fault diagnosis trigger information, and then subsequent network fault diagnosis is performed according to the resource name and the fault identifier.

[0033] S102: Determine resource instance information corresponding to the resource name, and a target node in the fault knowledge graph that matches the fault identifier, and at least one potential cause node having a causal relationship with the target node.

[0034] The potential cause node can include an event node, a fault node, and a root cause node.

[0035] In an example embodiment, a fault knowledge graph including a plurality of potential cause nodes and their associated relationships can be constructed in advance. When the resource name and the fault identifier included in the fault diagnosis trigger information are identified, resource instance information corresponding to the resource name is determined, which can be a resource instance corresponding to the fault identifier and resource instances dependent thereon. According to the identified fault identifier such as a fault name or an event name, an entity node corresponding to the fault knowledge graph is matched. In another example embodiment, a SELECT or MATCH query statement provided by a graph database can be used to achieve accurate matching of the target node by setting the fault name or the event name as a query filtering condition. If the target node of the fault knowledge graph is matched, at least one potential cause node having a causal relationship with the target node is acquired. If the entity node of the fault knowledge graph cannot be matched, a fault diagnosis knowledge question and answer can be given based on a pre-trained language model, and the diagnosis process is ended.

[0036] S103: In a case that the potential cause node corresponding to the target resource instance information corresponding to each of the potential cause nodes in the resource instance information is verified to be true, and the node type of the potential cause node is a preset root cause node, a fault diagnosis result corresponding to the potential cause node is obtained by calling the executable diagnosis interface corresponding to each of the potential cause nodes.

[0037] In an example embodiment, the potential cause node is associated with an executable diagnosis interface. By calling the executable diagnosis interface corresponding to each potential cause node, the target resource instance information corresponding to each potential cause node in the resource instance information can be executed, and whether the potential cause corresponding to each potential cause node is true can be verified. In a case that the potential cause corresponding to the potential cause node is verified to be true, and the node type of the potential cause node is a preset root cause node, a fault diagnosis result corresponding to the potential cause node is obtained. The fault diagnosis result can be a specific cause of the fault, or a solution or repair suggestion is provided.

[0038] Through the above steps, when the fault diagnosis trigger information is obtained, the target resource instance information corresponding to the potential cause node is executed based on the executable diagnosis interface corresponding to each potential cause node in the fault knowledge graph, and then the root cause node causing the fault is investigated and the corresponding fault diagnosis result is obtained. This method does not completely rely on the pre-constructed fault knowledge graph. By executing the target resource instance information corresponding to each potential cause node, whether the potential cause corresponding to the potential cause node is true is verified, so that the accuracy and reliability of fault diagnosis can be improved, and complex and precise fault diagnosis tasks can be implemented.

[0039] In a possible implementation manner, the above-mentioned fault diagnosis trigger information can be a passive diagnosis scene corresponding to user dialogue interaction. The identification of the fault diagnosis trigger information is completed by means of the language service capability of the pre-trained language model through multiple rounds of dialogue. That is, in S101, the fault diagnosis trigger information is obtained, including:

[0040] In response to a trigger operation of fault diagnosis of a user terminal, at least one question sentence is sent to the user terminal. The fault diagnosis trigger information is obtained according to the at least one question sentence and an answer sentence returned by the user terminal in response to the at least one question sentence.

[0041] In an embodiment of the present application, in response to a triggering operation of fault diagnosis by the user terminal, at least one question sentence can be sent to the user terminal, for example, for a fault, the question format is: what is the fault name of the following fault phenomenon? The fault phenomenon is described as follows: …, and the corresponding fault name is taken as the answer; for an event, the question format is: what is the event name of the following event description? The event description is as follows: …, and the corresponding event name is taken as the answer. According to the at least one question sentence and the answer sentence returned by the user terminal in response to the at least one question sentence, such as the fault name or the fault phenomenon, the fault diagnosis triggering information is obtained. The fault phenomenon or the event description input by the user on the user terminal can also be obtained after the at least one question sentence is sent to the user terminal. In an exemplary embodiment, the question sentence and the answer sentence form a question and answer knowledge pair, the formed question and answer knowledge pair is vectorized, common vectorization models include M3E-base, BERT, GPT, Sentence-BERT, etc., and the vectorized content is stored in a vector database, common vector databases include Faiss, Annoy, Milvus and Pinecone, etc. When identifying, the fault name or the event name can be quickly queried through the fault phenomenon or the event description input by the user on the user terminal, so as to obtain the fault diagnosis triggering information.

[0042] In another possible implementation, the fault diagnosis triggering message described above can be an active diagnosis scenario corresponding to abnormal monitoring, that is, in S101 described above, the fault diagnosis triggering information is obtained, including:

[0043] monitoring an abnormal event of the target system during operation; and obtaining the fault diagnosis triggering information according to the abnormal event.

[0044] In an embodiment of the present application, the abnormal event (including alarm, state anomaly, performance anomaly, log anomaly, etc.) of the target system during operation is monitored in real time, and the fault diagnosis is automatically triggered to obtain the fault diagnosis triggering information.

[0045] In a possible implementation, after the target resource instance corresponding to each of the potential cause nodes in the resource instance information is executed by calling the executable diagnosis interface corresponding to each of the potential cause nodes in S103 described above, the method further includes:

[0046] In a case where the potential cause corresponding to the potential cause node is verified to be true and the node type of the potential cause node is not a preset root cause node, an associated cause node having a cause-effect relationship with the potential cause node in the fault knowledge graph is acquired; target resource instance information corresponding to the associated cause node is subjected to fault diagnosis; in a case where the potential cause corresponding to the associated cause node is verified to be true and the node type of the associated cause node is the root cause node, a fault diagnosis result corresponding to the associated cause node is acquired.

[0047] In an example embodiment, if the potential cause corresponding to the potential cause node is true but the potential cause node is not a root cause node, a possible potential cause leading to the potential cause node is further investigated, that is, an associated cause node having a cause-effect relationship with the potential cause node is acquired, and the target resource instance information corresponding to the associated cause node is executed by calling an executable diagnosis interface corresponding to the associated cause node; in a case where the potential cause corresponding to the associated cause node is verified to be true and the node type of the associated cause node is the root cause node, a fault diagnosis result corresponding to the associated cause node is acquired.

[0048] In a possible implementation, after the execution of the target resource instance information corresponding to each potential cause node in the resource instance information by calling the executable diagnosis interface corresponding to each potential cause node in S103, the method further includes:

[0049] In a case where the potential cause corresponding to the potential cause node is verified to be false, the potential cause node is excluded.

[0050] In the embodiment, the target resource instance information corresponding to each potential cause node in the resource instance information is executed by calling the executable diagnosis interface corresponding to each potential cause node; in a case where the potential cause corresponding to the potential cause node is verified to be false, the potential cause node is excluded, and the next potential cause node is further investigated.

[0051] In a possible implementation, after the resource instance information corresponding to the resource name and the target node in the fault knowledge graph matching the fault identifier and the at least one potential cause node having a cause-effect relationship with the target node are determined, the method further includes:

[0052] For any potential cause node in the at least one potential cause node, in a case where the any potential cause node does not have a corresponding executable diagnosis interface, a diagnosis rule of the any potential cause node is displayed on a preset visual interface; a fault diagnosis result corresponding to the any potential cause node input by the visual interface is acquired.

[0053] In the embodiments of the present application, when performing fault diagnosis, for any one of the at least one potential cause node, if there is no corresponding executable diagnosis interface (i.e., the potential cause node does not have automatic diagnosis capability) in the any one potential cause node, the diagnosis rule of the potential cause node is displayed on the preset visual interface, a user performs manual diagnosis according to the diagnosis rule, and after the manual diagnosis, the fault diagnosis result corresponding to the potential cause node input by the visual interface is obtained. Considering that this way will block the current network diagnosis process, the potential cause node that needs manual diagnosis can also be recorded first, and after all potential cause nodes are investigated, the potential cause node that needs manual diagnosis is displayed on the visual interface for manual diagnosis.

[0054] In a possible implementation, the method further includes:

[0055] In a case where the node type of the any one potential cause node is the root cause node, root cause information corresponding to the any one potential cause node is stored, and the root cause information is displayed on the visual interface.

[0056] In the embodiments of the present application, in a case where there is no corresponding executable diagnosis interface (i.e., the potential cause node does not have automatic diagnosis capability) in the any one potential cause node, if the node type of the potential cause node is a root cause node, root cause information corresponding to the potential cause node is stored, and the root cause information is displayed on the visual interface.

[0057] In a case where the potential cause node is determined to be a root cause node, the fault diagnosis result corresponding to the potential cause node is obtained.

[0058] In an example embodiment, as shown in FIG. 2, the network fault diagnosis method can include the following steps:

[0059] Step 201: Based on the matched target node and the corresponding resource instance information, all possible potential cause nodes (including event nodes, fault nodes, and root cause nodes) are investigated in sequence.

[0060] Step 202: An existing diagnosis result is obtained.

[0061] Step 203: Whether the potential cause corresponding to the potential cause node has been diagnosed (i.e., whether the potential cause is established) is determined according to the existing diagnosis result. If not, step 204 is entered; otherwise, step 208 is entered.

[0062] Step 204: determining whether the potential cause node has automatic diagnosis capability (i.e., whether the potential cause node has a corresponding executable diagnosis interface), if yes, entering step 205; otherwise, entering step 206;

[0063] Step 205: calling the executable diagnosis interface (diagnosis API) corresponding to the potential cause node to execute the target resource instance information corresponding to each potential cause node in the resource instance information, wherein, the step 205 includes obtaining and filling the diagnosis API parameters by means of the pre-trained language model, and obtaining the diagnosis result by executing the diagnosis API;

[0064] Step 206: outputting the diagnosis rule and performing manual diagnosis;

[0065] Step 207: verifying whether the potential cause node is diagnosed (i.e., verifying whether the potential cause corresponding to the potential cause node is established), if yes, entering step 208; otherwise, entering step 210.

[0066] Step 208: determining whether the potential cause node is a root cause node, if yes, entering step 208; otherwise, returning to step 201;

[0067] Step 209: recording the diagnosis result (i.e., the above-mentioned obtaining the fault diagnosis result corresponding to the potential cause node);

[0068] Step 210: determining whether all potential cause nodes have been investigated, if yes, ending the process; otherwise, returning to step 201.

[0069] In another exemplary embodiment, as shown in FIG. 3, the above-mentioned network fault diagnosis method further comprises:

[0070] Step 211: after step 204, if the potential cause node does not have automatic diagnosis capability (i.e., the potential cause node does not have a corresponding executable diagnosis interface), determining whether the potential cause node is a root cause node, if yes, entering step 212; otherwise, returning to step 201;

[0071] Step 212: recording the root cause information, and returning to step 201.

[0072] In a possible implementation manner, as shown in FIG. 4, in S101, before the fault diagnosis trigger information is obtained, further comprising:

[0073] S1011: cleaning the obtained at least one operation and maintenance document to obtain a target document in a preset format.

[0074] In the embodiments of the present application, different formats of operation and maintenance documents (for example, asset information, fault handling, alarm handling, fault case set, inspection knowledge, etc.) are obtained, at least one obtained operation and maintenance document is cleaned to obtain a target document in a preset format, for example, the operation and maintenance document is cleaned into a markdown format file, so that the first pre-trained language model can better understand and extract knowledge from the operation and maintenance document. In specific applications, the cleaning function can be automatically completed by a special cleaning tool, and documents in file formats such as pdf, doc / docx, html, etc. can be cleaned. After cleaning, the effective content such as document chapter title, chapter complete content and table information can be retained to ensure the effectiveness of the target document content in the knowledge extraction process and improve the knowledge extraction efficiency.

[0075] S1012: Determine the knowledge extraction prompt word of the first pre-trained language model according to the target document, the preset entity type and the relationship type.

[0076] Among them, the knowledge extraction prompt can be automatically generated according to the fault knowledge graph, and the knowledge extraction prompt is input into the first pre-trained language model to realize knowledge extraction of the target document; or for each target document, the knowledge extraction prompt is defined according to the knowledge content contained in the target document and the fault knowledge graph, and the knowledge extraction prompt is input into the first pre-trained language model to realize knowledge extraction of the target document.

[0077] It should be noted that the knowledge extraction prompt needs to cover all entity types and attributes, as well as edge types and attributes defined in the fault knowledge graph. The definition of the knowledge extraction prompt depends on the document content to be extracted and the entity type and relationship type to be extracted.

[0078] In an exemplary embodiment, taking the fault diagnosis scene of the PTN network as an example, the format of the knowledge extraction prompt is roughly described as follows, and the specific implementation needs to be adjusted according to the knowledge extraction document content:

[0079] You are a PTN network operation and maintenance expert, please extract knowledge from the given text according to the given knowledge extraction example and its interpretation. The knowledge extraction example and its interpretation are as follows:

[0080] Fault_schema_description={

[0081] 'name':' fault name, fault name should reflect "fault location + fault characteristics" as much as possible. Return in string form',

[0082] 'type': 'fault type, classify the fault for easy management and maintenance. Current fault classification includes: specification, capacity, performance, limit, and state. Returned as a string.',

[0083] 'code': 'fault code, corresponding to the defined fault uniform code. If not defined, it can be empty. Returned as a string',

[0084] 'description': 'fault phenomenon. Describe the phenomenon after the fault occurs, which needs to be described from the user's perspective. Returned as a string',

[0085] 'ConfirmRule': 'diagnosis rule, describes the judgment rule for confirming and diagnosing whether the fault exists. Returned as a string',

[0086] 'ReasonList':'reason list, describes the list of possible reason names that cause the fault. Returned as a list'

[0087] }

[0088] RootCause_schema_description={

[0089] 'RootCauseName': 'root cause name, root cause name needs to be globally unique, otherwise, it is considered as the same kind of root cause. Different root causes need to be separated. The same kind of root cause needs to correspond to clear diagnosis rules and solutions, otherwise, it cannot be considered as a root cause. Returned as a string',

[0090] 'RootCauseType': 'root cause type, classify the root cause for easy management and maintenance. Includes link, configuration, hardware, software, and external. Returned as a string.',

[0091] 'RootCauseDescription': 'root cause description. From the user's perspective, describe the root cause in detail to facilitate user understanding of the root cause. Returned as a string',

[0092] 'ConfirmRule': 'diagnosis rule, is the judgment rule for confirming and diagnosing whether the root cause exists. Returned as a string',

[0093] 'SolutionName':'solution name, solution name needs to be globally unique, otherwise, it is considered as the same kind of solution. Returned as a string',

[0094] 'SolutionDescription': 'Solution description. Detailed description of the solution, including scenarios of manual intervention or automatic repair. Returned as a string'

[0095] }、

[0096] …

[0097] The knowledge extraction result is output in json format defined by the knowledge extraction schema. If the information cannot be found, the corresponding value in the json text is empty.

[0098] The text is as follows:

[0099] …

[0100] S1013: input the target document and the knowledge extraction prompt into the first pre-trained language model, so that the first pre-trained language model extracts a plurality of entity nodes, attribute information of each entity node, and entity relationships between the entity nodes in the target document according to the knowledge extraction prompt, the entity nodes including a potential cause node.

[0101] In the embodiments of the present application, the target document and the knowledge extraction prompt are input into the first pre-trained language model, the first pre-trained language model loads the pre-processed markdown format target document in sequence according to the knowledge extraction prompt, performs text segmentation according to the title chapter, inputs each sub-title chapter content as text, and performs knowledge extraction; meanwhile, the content of knowledge extraction is processed, and a specified file format is output, including but not limited to csv, json and the like. According to the entity type and the relationship type, the corresponding csv files are output, the file name is Vertex_Entity Type Name.csv or Edge_Relationship Type Name.csv, such as Vertex_Fault.csv or Edge_dependOn.csv. For the output entity type csv file, each row represents an entity instance, and each column corresponds to an attribute of the entity type. If no extraction is performed, it is empty. For the output relationship type csv file, each row represents a relationship instance, with fixed Out and In columns, Out indicating the direction entity instance name, In indicating the in-direction entity instance name, and each attribute corresponding to a column if there are other attributes. In specific implementation, the large model service mainly provides large model semantic understanding and language generation capability, and can deploy pre-trained language large models such as Qwen1.5-110B-Chat, Meta-Llama-3-70B-Instruct, etc. or adopt third-party deployed pre-trained language large models such as ChatGPT3.5, ChatGPT4, etc.

[0102] According to the entity name and entity attribute information, the same entity instances of the same entity type extracted from different documents are de-duplicated to remove duplicate entity instances. At the same time, the relationship instances extracted from different documents are knowledge-aligned according to the corresponding out- and in- entity instance names or attributes (such as alarm code) information to establish and supplement the entity relationships between different knowledge. Finally, the knowledge extraction results after knowledge de-duplication and alignment are stored in a graph database (such as OrientDB graph database). Thus, multiple entity nodes in the target document, attribute information of each entity node, and entity relationships between each entity node are extracted, wherein the entity nodes include the above-mentioned potential cause nodes.

[0103] In a specific application, since the adoption rate of knowledge extraction using a pre-trained language model cannot reach 100%, the fault knowledge graph content stored in the graph database can also be manually audited, corrected, and supplemented and improved, and the extracted diagnosis rules can be converted into an executable method. In an exemplary embodiment, a knowledge graph visualization and correction tool can be provided to facilitate manual auditing, correction, and supplementation.

[0104] Optionally, the tool can provide an operation interface, and the user can set query conditions according to the confirmation status, entity type, or relationship type, query entity instances or relationship instance records that meet the conditions, and support presentation in a table manner and a graph visualization manner (graphically presenting relationships between entities); the user can perform correction operations on the table or the graph based on the queried entity instances or relationship instances, including correcting inaccurate or unclear content, supplementing incomplete content, deleting redundant or erroneous content, updating the confirmation status to confirmed after completion, and canceling the confirmation after an error operation. Optionally, the tool can also automatically persistently record the correction content, the confirmation status, the confirmation time, and the confirmation personnel information. At the same time, the tool can also support exporting all entity instances and relationship instances to an excel table format (including but not limited to xls, xlsx, csv, etc.) for batch correction and correction, and support full import and incremental import into the graph data for updating after completion. Full import will clear all records and regenerate records, supporting deletion operations. Incremental import will not delete existing records, update records, add new records, and not support deletion records.

[0105] S1014: According to the fault diagnosis rules included in the attribute information of each entity node, determine the executable diagnosis interface corresponding to each entity node, and based on the multiple entity nodes, the attribute information and executable diagnosis interface of each entity node, and the entity relationships between each entity node, generate the fault knowledge graph.

[0106] In the embodiments of the present application, in order to realize automatic diagnosis and enhance the reasoning effect of fault diagnosis, according to the fault diagnosis rules included in the attribute information of each entity node, the executable diagnosis interface (diagnosis API) corresponding to each entity node is determined. In an exemplary embodiment, for the fault diagnosis rules that can be automatically executed, the pre-trained language model is used to realize the diagnosis API that can be called by the program, and the executable diagnosis interface corresponding to each entity node in the fault knowledge graph is updated. Thus, based on the multiple entity nodes, the attribute information and the executable diagnosis interface of each entity node, and the entity relationship between each entity node, the fault knowledge graph is generated.

[0107] In a possible implementation, after generating the fault knowledge graph, a knowledge graph management function can also be provided at the user end, including knowledge graph content viewing and knowledge graph updating functions. In the embodiments of the present application, when using the knowledge graph content viewing function, the query conditions can be set, including setting entity types, entity names or relationship types, etc. The query results are presented in two ways, table and graph visualization. The table mode is convenient for viewing entity instances and attribute information, and the graph visualization mode is convenient for viewing relationship instances, that is, various relationships between different entity instances. At the same time, the fault knowledge graph updating function is provided, and the operation interface can be operated by the user to update the latest fault knowledge graph into the online fault diagnosis reasoning system. The update operation supports full update and incremental update. Full update will delete all existing fault knowledge graph content, and then generate a graph based on the latest fault knowledge graph information. Incremental update will not delete the existing fault knowledge graph information, but will update the existing graph information and supplement the new graph information.

[0108] In an exemplary embodiment, as shown in FIG. 5, the fault knowledge graph generation method includes the following steps:

[0109] Step 501: Operation and maintenance document cleaning: cleaning the obtained at least one operation and maintenance document to obtain a target document in a preset format;

[0110] Step 502: Define knowledge extraction prompt (i.e., knowledge extraction prompt word): according to the target document, the preset entity type and relationship type, determine the knowledge extraction prompt word;

[0111] Step 503: Document knowledge extraction based on a large model (i.e., a first pre-trained language model);

[0112] Step 504: Large model extraction result deduplication and alignment processing;

[0113] Step 505: Fault knowledge graph manual review, correction and supplement;

[0114] Step 506: implement the fault diagnosis rule as an executable diagnosis API (i.e., an executable diagnosis interface).

[0115] In a possible implementation, the entity types include resources, events, faults, root causes, and solutions, and the relationship types include dependency, possession, cause, and solution.

[0116] In the embodiments of the present application, as shown in FIG. 6, the fault knowledge graph includes entity types such as Resource (resource), Event (event), Fault (fault), RootCause (root cause), and Solution (solution), and edge types such as dependOn (dependent on), has (has), causedBy (caused by), and solvedBy (solved by). In addition, entity types and relationship types can be extended according to business needs. For each entity type, general attributes such as applicable network type (optional, fill in default general, applicable to all network types, support multiple network types, separated by “,” or “&” symbols), applicable device type (optional, fill in default general, applicable to all device types, support multiple device types, separated by “,” or “&” symbols), knowledge source (optional), knowledge creation time (optional), knowledge confirmation status (optional, True indicates confirmed, False indicates not confirmed, default not confirmed), knowledge confirmation time (optional), knowledge confirmation personnel (optional), and the like are included, and the entity type also includes its own special attributes. For each relationship type, it can be decided according to business needs whether to strictly constrain the out-node type and the in-node type, or to add attributes. The following will be described respectively:

[0117] 1) Resource (resource): an abstraction of network management resource objects. Different network types (such as PTN, SPN, IPRAN, OTN network, etc.) may have different resource types, and there are business hierarchy and dependency relationships between resources. In addition to general attributes, the attributes of the resource also include resource name (mandatory), resource type (mandatory), and resource description (optional). Among them, the resource name can be used for interface display, and if internationalization is required, corresponding Chinese, English, or other language descriptions need to be provided, and the resource name corresponding to the same entity type needs to be globally unique; the resource type is consistent with the network resource object type definition; the resource description provides more detailed resource information. At the same time, attributes can also be extended according to business needs.

[0118] 2) Event (event): is the abstraction of various alarms, state abnormalities, performance abnormalities, log abnormalities and the like generated by the system, which describes an abnormality of the system. In addition to the general attributes, the attributes of the event also include event name (mandatory), event type (mandatory), event code (optional), event description (mandatory), diagnosis rule (i.e. the network diagnosis rule described above, mandatory), diagnosis API (i.e. the executable diagnosis interface described above, optional) and the like. Among them, the event name can be used for interface display, if internationalization is required, corresponding Chinese, English or other language descriptions need to be provided, and the event name corresponding to the same entity type needs to be globally unique; the event type is a classification of the event, which can include alarms, performance, state, logs and the like, and can be defined or extended by itself according to business needs; the event code is a unified standardized code for each event, such as the code of the alarm event corresponding to the alarm detection point + alarm code, which is convenient for knowledge alignment processing and knowledge management, and can not be filled in; the event description is a detailed description of the event from the user's point of view; the diagnosis rule describes the judgment rule for confirming and diagnosing whether the event exists; the diagnosis API provides an API interface that can be called and executed, which is used to automatically judge whether the event exists. At the same time, attributes can also be extended according to business needs.

[0119] 3) Fault (fault): is the abstraction of system operation abnormalities and problems, which usually accompanies a series of related abnormal events after the fault occurs. In addition to the general attributes, the attributes of the fault also include fault name (mandatory), fault type (mandatory), fault code (optional), fault phenomenon (mandatory), diagnosis rule (mandatory), diagnosis API (optional) and the like. Among them, the fault name can be used for interface display, if internationalization is required, corresponding Chinese, English or other language descriptions need to be provided, and the same entity type needs to be globally unique; the fault type is a classification of the fault, including specification class, capacity class, performance class, overrun class, state class and the like, which can be defined or extended by itself according to business needs; the fault code is a unified standardized code for each fault, which is convenient for knowledge alignment processing and knowledge management, and can not be filled in; the fault phenomenon describes the phenomenon description seen by the user after the fault occurs; the fault diagnosis rule describes the judgment rule for confirming and diagnosing whether the fault exists; the diagnosis API provides an API interface that can be called and executed, which is used to automatically judge whether the fault exists. At the same time, attributes can also be extended according to business needs.

[0120] 4) RootCause (root cause): Root cause refers to the root cause of the fault, which is the information element that fault diagnosis needs to output. In addition to the general attributes, root cause also includes root cause name (mandatory), root cause type (mandatory), root cause code (optional), root cause description (mandatory), diagnosis rule (mandatory), diagnosis API (optional) and other attributes. Among them, the root cause name can be used for interface display, such as supporting internationalization, providing corresponding Chinese, English or other language descriptions, and the root cause information corresponding to the same entity type needs to be globally unique; the root cause type is a classification of the root cause, which is convenient for management and maintenance, including link class, configuration class, hardware class, software class, external class, etc., which can be extended according to business needs; the root cause code is a standardized coding for each root cause, which is convenient for knowledge alignment processing and knowledge management, and can be left blank; the root cause description is a detailed description of the root cause from the user's perspective, which is convenient for the user to understand the root cause; the diagnosis rule describes the judgment rule for confirming and diagnosing whether the root cause exists; the diagnosis API provides an API interface that can be called and executed, which is used to automatically judge whether the root cause exists. At the same time, attributes can also be extended according to business needs.

[0121] 5) Solution (solution): Solution is a description of how to solve the fault after finding the root cause. Usually, one root cause corresponds to one solution, and there are cases where different root causes correspond to the same solution. The case where the same root cause corresponds to multiple solutions is less common. In addition to the general attributes, the solution also includes solution name (mandatory), solution type (optional), solution code (optional), solution description (mandatory), repair rule (mandatory), repair API (optional) and other attributes. Among them, the solution name can be used for interface display, such as supporting internationalization, providing corresponding Chinese, English or other language descriptions, and the same entity type needs to be globally unique; the solution type is a classification of the solution, which is convenient for management and maintenance, including configuration class, link class, hardware class, software class, external class, etc., which can be defined and extended according to business needs; the solution code is a standardized coding for each solution, which is convenient for knowledge alignment processing and knowledge management, and can be left blank; the solution description is a detailed description of the solution from the user's perspective, which is convenient for the user to understand the solution; the repair rule describes a method for the user to repair the fault after finding the cause of the fault; the repair API provides an API interface that can be called and executed, which is used to automatically repair the fault. At the same time, attributes can also be extended according to business needs.

[0122] 6) dependOn (depend on): used to describe the dependency relationship between resources, limiting the node type and the entry node type to be Resource (resource) entity type. The direction of the edge represents the dependency relationship, when A->B, it means that resource A depends on resource B, and when resource B is abnormal, it will affect the state of resource A.

[0123] 7) has: Used to describe the relationship between resources and events, failures, and root causes, indicating that a certain resource type has a corresponding event, failure, or root cause.

[0124] 8) Caused By: This describes the propagation relationship between events, events and faults, faults and faults, events and root causes, and faults and root causes. This relationship type also includes a Priority attribute, which corresponds to the investigation order, probability, or investigation priority. Generally, the higher the probability or the higher the investigation priority, the higher the priority should be investigated. If left blank, it means that the priorities are the same, and the investigation order is randomly generated.

[0125] 9) sovledBy: This describes the relationship between a root cause and a solution. Usually, one root cause corresponds to one solution. Different root causes may correspond to the same solution. It is rare for the same root cause to correspond to multiple solutions.

[0126] In one possible implementation, after obtaining the fault diagnosis result corresponding to the potential cause node in S103 above, the method further includes:

[0127] Obtain the fault diagnosis results of each potential cause node and the corresponding fault solutions; sort each potential cause node according to a preset fault investigation order or fault investigation priority, and display the fault diagnosis results and fault solutions corresponding to a preset number of potential cause nodes that are sorted first.

[0128] In this embodiment, the fault diagnosis results of each potential cause node and the corresponding fault solutions are obtained. All recorded fault diagnosis results are summarized and displayed to the user, and the root cause and solution of the fault are output. For solutions with automatic repair APIs, the user is prompted whether to perform automatic repair actions. For solutions requiring manual intervention, specific handling suggestions are given. If there are too many potential cause nodes, the fault diagnosis results and fault solutions corresponding to a preset number N of potential cause nodes with the highest probability can be given according to the Priority attribute of the causeBy edge, based on the fault investigation order or fault investigation priority.

[0129] In one possible implementation, after obtaining the fault diagnosis trigger information, step S103 above further includes:

[0130] In a case where it is determined that the target resource node matching the fault identification is not included in the fault knowledge graph, a first diagnostic knowledge answer corresponding to the fault diagnosis trigger information is determined based on a second pre-trained language model and the fault knowledge graph.

[0131] In the embodiments of the present application, in a case where it is determined that the target resource node matching the fault identification is not included in the fault knowledge graph (i.e., no target node in the fault knowledge graph is matched), a first diagnostic knowledge answer corresponding to the fault diagnosis trigger information is determined based on a second pre-trained language model and the fault knowledge graph.

[0132] In a possible implementation, the network fault diagnosis method described above further includes:

[0133] In a case where the fault identification included in the fault diagnosis trigger information is not recognized, a second diagnostic knowledge answer corresponding to the fault diagnosis trigger information is determined based on a third pre-trained language model.

[0134] In the embodiments of the present application, if the fault identification included in the fault diagnosis trigger information is not recognized, a second diagnostic knowledge answer corresponding to the fault diagnosis trigger information is determined based on a third pre-trained language model.

[0135] In an exemplary embodiment, as shown in FIG. 7, the network fault diagnosis method can include the following steps:

[0136] Step 701: User dialogue interaction / system anomaly monitoring, obtaining fault diagnosis trigger information;

[0137] Step 702: Recognizing the fault identification in the fault diagnosis trigger information;

[0138] Step 703: Determining whether the fault identification is recognized successfully, if yes, proceeding to step 704; otherwise, proceeding to step 705;

[0139] Step 704: Matching the target node in the fault knowledge graph;

[0140] Step 705: Fault knowledge question answering based on a large model (i.e., the second pre-trained language model described above);

[0141] Step 706: Determining whether the target node in the fault knowledge graph is matched, if yes, proceeding to step 707; otherwise, proceeding to step 705;

[0142] Step 707: Obtaining resource instance information corresponding to the resource name in the fault diagnosis trigger information;

[0143] Step 708: Determining whether the resource instance information is obtained, if yes, proceeding to step 709; otherwise, proceeding to step 710;

[0144] Step 709: fault diagnosis reasoning based on a large model (i.e., the first pre-trained language model described above) and a fault knowledge graph;

[0145] Step 710: fault knowledge question answering based on a large model (i.e., the third pre-trained language model described above) and a fault knowledge graph;

[0146] Step 711: summarizing and outputting fault diagnosis results.

[0147] The embodiment of the present application provides a network fault diagnosis method, obtains fault diagnosis trigger information, the fault diagnosis trigger information includes a resource name and a fault identifier; determines resource instance information corresponding to the resource name, and a target node in a fault knowledge graph matched with the fault identifier and at least one potential cause node having a causal relationship with the target node; by calling an executable diagnosis interface corresponding to each potential cause node, executing target resource instance information corresponding to each potential cause node in the resource instance information, in the case that the potential cause corresponding to the potential cause node is established and the node type of the potential cause node is a preset root cause node, obtaining a fault diagnosis result corresponding to the potential cause node. In this way, when the fault diagnosis trigger information is obtained, the pre-constructed fault knowledge graph can not be completely relied on, the target resource instance information is executed through the executable diagnosis interface corresponding to each potential cause node to verify whether the potential cause corresponding to the potential cause node is established, so that the accuracy and reliability of fault diagnosis can be improved, and complex and accurate fault diagnosis tasks can be realized.

[0148] In addition, by passive triggering of user dialogue interaction and / or active triggering of system abnormal monitoring, the fault diagnosis trigger information is obtained, which can ensure that as many network faults as possible are found and processed, and the stability and reliability of network operation are ensured; meanwhile, the embodiment of the present application provides entity types and relationship types of the fault knowledge graph, which can effectively represent the knowledge in the network fault field and provide strong support for fault diagnosis reasoning, which is conducive to further improving the accuracy of fault diagnosis and improving the efficiency of network fault processing.

[0149] FIG. 8 shows a structural block diagram of an online fault diagnosis reasoning system provided by an embodiment of the present application. As shown in the figure, the online fault diagnosis reasoning system 800 includes several key components such as user dialogue interaction 810, system anomaly monitoring 820, fault identification and recognition 830, fault knowledge graph matching 840, resource example information acquisition 850, fault diagnosis reasoning 860, diagnosis result presentation 870, first large model service 880, fault knowledge graph and knowledge graph management 890, etc. The system 800 supports both active diagnosis and passive diagnosis scenarios. The user dialogue interaction corresponds to the passive diagnosis scenario, and the user intent is recognized by means of the large model capability according to the user input, triggering fault diagnosis passively. The system anomaly monitoring corresponds to the active diagnosis scenario, and real-time monitoring of system abnormal events (including alarms, state abnormalities, performance abnormalities, log abnormalities, etc.) automatically triggers fault diagnosis. After triggering fault diagnosis, the fault phenomenon input by the user or the abnormal event monitored based on the system is used to identify the resource name such as the business, link or network element with the fault, and the fault name or event name information. Then, the fault or event name is accurately matched with the fault knowledge graph node, and the resource and dependent resource instance information is acquired. Based on the matched knowledge graph node and the corresponding resource instance, all potential cause nodes are sequentially investigated. The investigation process investigates the possible potential cause nodes based on the fault diagnosis rules and executable diagnosis interfaces corresponding to the potential cause nodes, and records the fault diagnosis results. Finally, the fault diagnosis results are summarized and displayed.

[0150] FIG. 9 shows a structural block diagram of an offline knowledge mining production system provided by an embodiment of the present application. As shown in the figure, the system 900 includes several key components such as document cleaning 910, knowledge graph design 920, knowledge extraction 930, knowledge review 940, executable diagnosis interface implementation 950, second large model service 960, fault knowledge graph, etc. According to the knowledge graph design, the system 900 fully utilizes the semantic understanding and language generation capability of the large model, and can quickly and effectively extract relevant knowledge from the specified operation and maintenance documents (including asset information, inspection knowledge, alarm processing, fault processing, fault case set, etc.) after knowledge deduplication and alignment processing, and then after knowledge manual review, correction and supplement, finally converts the fault diagnosis rules into executable diagnosis interfaces, thereby completing the construction of the fault knowledge graph.

[0151] In this way, by fully utilizing the technical advantages of the large model and the knowledge graph, the complex and precise fault diagnosis task can be realized, the controllability, reliability and explainability of fault diagnosis are improved, and the fault diagnosis can be quickly and accurately completed, and the fault cause is output.

[0152] Fig. 10 shows a structural schematic diagram of an electronic device according to an embodiment of the present application. Referring to Fig. 10, at the hardware level, the electronic device 1000 comprises a processor 1010, and optionally comprises an internal bus 1020, a network interface 1030, and a memory 1040. The memory 1040 can include an internal memory 1041, such as a high-speed Random-Access Memory (RAM), and can also include a non-volatile memory 1042, such as at least one disk memory. Of course, the electronic device 1000 can also include other hardware required by other services.

[0153] The processor 1010, the network interface 1030, and the memory can be connected to each other through the internal bus 1020, which can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one bidirectional arrow is shown in the figure, but it does not mean that there is only one bus or only one type of bus.

[0154] The memory 1040 stores programs. In an exemplary embodiment, the programs can include program codes including computer operation instructions. The memory 1040 can include the internal memory 1041 and the non-volatile memory 1042, and provide instructions and data to the processor 1010.

[0155] The processor 1010 reads the corresponding computer programs from the non-volatile memory 1042 into the internal memory and then runs, and forms a device for positioning a target user at the logical level. The processor 1010 executes the programs stored in the memory, and specifically executes the methods disclosed in Figs. 1-5 and 7, and realizes the functions and advantages of each method described in the foregoing method embodiments, which will not be repeated here.

[0156] The method disclosed in the embodiments of the present application as shown in FIG. 1 to FIG. 5 and FIG. 7 can be applied to the processor 1010 or implemented by the processor 1010. The processor 1010 can be an integrated circuit chip having a signal processing capability. In the implementation process, each step of the above method can be completed by an integrated logic circuit or an instruction in the form of software in the processor 1010. The processor 1010 described above can be a general processor, including a central processing unit (CPU), a network processor (NP), etc.; or a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. Each method, step and logic block diagram disclosed in the embodiments of the present application can be implemented or executed. The general processor can be a microprocessor or any conventional processor. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium in the memory is read by the processor, and the hardware thereof is combined to complete the steps of the above method.

[0157] The computer device can also execute the methods described in the foregoing method embodiments, and achieve the functions and beneficial effects of the methods described in the foregoing method embodiments, which will not be repeated here.

[0158] Of course, in addition to the software implementation, the electronic device of the present application does not exclude other implementation manners, such as logic devices or a combination of software and hardware, etc. That is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or a logic device.

[0159] The embodiments of the present application also propose a computer readable storage medium, which stores one or more programs, and when the one or more programs are executed by an electronic device including a plurality of application programs, the electronic device executes the method disclosed in the embodiments of FIG. 1 to FIG. 5 and FIG. 7 and achieves the functions and beneficial effects of the methods described in the foregoing method embodiments, which will not be repeated here.

[0160] The computer readable storage medium includes a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0161] Further, the embodiment of the present application further provides a computer program product, the computer program product comprising a computer program stored on a non-transitory computer readable storage medium, the computer program comprising program instructions which, when executed by a computer, implement the method disclosed in the embodiments shown in FIG. 1 to FIG. 5 and FIG. 7, and achieve the functions and advantages of each method described in the foregoing method embodiments, which will not be repeated here.

[0162] The embodiment of the present application can be applied to various electronic device cooperation or interconnection scenarios, including: mobile phone and notebook computer / tablet computer cooperation and interconnection; mobile terminal and smart television / display cooperation and interconnection; mobile phone or tablet computer and vehicle-mounted entertainment system cooperation and interconnection; mobile terminal and smart conference system cooperation and interconnection, etc. Thus, the needs of users in smart home, smart office, smart travel and other diversified scenarios are met.

[0163] In summary, the above only describes the preferred embodiments of the present application, and does not limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

[0164] The system, device, module or unit illustrated in the above embodiments can be specifically implemented by a computer chip or entity, or by a product with certain functions. A typical implementation device is a computer. Specifically, the computer may, for example, be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0165] Computer-readable media includes permanent and non-permanent, movable and non-movable media that can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can store information accessible by a computing device. According to the definition herein, computer-readable media does not include transitory media such as modulated data signals and carriers.

[0166] It should also be noted that the terms "comprising", "containing", or any other variant thereof are intended to cover non-exclusive inclusions, so that a process, method, article or apparatus that includes a list of elements does not only include those elements, but also includes other elements not explicitly listed, or inherent to such a process, method, article or apparatus. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article or apparatus that includes the element.

[0167] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiments.

Claims

1. A network fault diagnosis method, comprising: obtaining fault diagnosis trigger information, wherein the fault diagnosis trigger information comprises a resource name and a fault identifier; determining resource instance information corresponding to the resource name, and a target node in a fault knowledge graph that matches the fault identifier and at least one potential cause node that has a causal relationship with the target node; executing target resource instance information corresponding to each of the potential cause nodes in the resource instance information by calling an executable diagnosis interface corresponding to each of the potential cause nodes, and obtaining a fault diagnosis result corresponding to the potential cause node in a case where a potential cause corresponding to the potential cause node is verified to be true and a node type of the potential cause node is a preset root cause node.

2. The method of claim 1, wherein, The obtaining of the fault diagnosis trigger information comprises: sending at least one question sentence to a user terminal in response to a triggering operation of the user terminal for fault diagnosis; obtaining the fault diagnosis trigger information according to the at least one question sentence and an answer sentence returned by the user terminal for the at least one question sentence.

3. The method of claim 1, wherein, The obtaining of the fault diagnosis trigger information comprises: monitoring an abnormal event of a target system during operation; obtaining the fault diagnosis trigger information according to the abnormal event.

4. The method of claim 1, wherein, After the executing of the target resource instance information corresponding to each of the potential cause nodes in the resource instance information by calling the executable diagnosis interface corresponding to each of the potential cause nodes, the method further comprises: obtaining an associated cause node in the fault knowledge graph that has a causal relationship with the potential cause node in a case where the potential cause corresponding to the potential cause node is verified to be not true and the node type of the potential cause node is not the preset root cause node; performing fault diagnosis on target resource instance information corresponding to the associated cause node, and obtaining a fault diagnosis result corresponding to the associated cause node in a case where a potential cause corresponding to the associated cause node is verified to be true and a node type of the associated cause node is the root cause node.

5. The method of claim 1, wherein, After the executing of the target resource instance information corresponding to each of the potential cause nodes in the resource instance information by calling the executable diagnosis interface corresponding to each of the potential cause nodes, the method further comprises: excluding the potential cause node in a case where the potential cause corresponding to the potential cause node is verified to be not true.

6. The method of claim 1, wherein, After the determining of the resource instance information corresponding to the resource name, and the target node in the fault knowledge graph that matches the fault identifier and the at least one potential cause node that has a causal relationship with the target node, the method further comprises: in a case where there is no executable diagnosis interface corresponding to any one of the at least one potential cause node, displaying a diagnosis rule of the any one of the at least one potential cause node on a preset visual interface; obtaining a fault diagnosis result corresponding to the any one of the at least one potential cause node input by the visual interface.

7. The method of claim 6, wherein, The method further comprises: in a case where the node type of the any one of the at least one potential cause node is the root cause node, storing root cause information corresponding to the any one of the at least one potential cause node. displaying the root cause information on the visualization interface.

8. The method of claim 1, wherein, Before the obtaining the fault diagnosis trigger information, further comprising: cleaning the obtained at least one operation and maintenance document to obtain a target document in a preset format; determining a knowledge extraction prompt word of a first pre-trained language model according to the target document, a preset entity type, and a preset relationship type; inputting the target document and the knowledge extraction prompt word into the first pre-trained language model, so that the first pre-trained language model extracts a plurality of entity nodes, attribute information of each of the entity nodes, and entity relationships between the entity nodes from the target document according to the knowledge extraction prompt word, wherein the entity nodes include a potential cause node; determining an executable diagnosis interface corresponding to each of the entity nodes according to a fault diagnosis rule included in the attribute information of each of the entity nodes, and generating the fault knowledge graph based on the plurality of entity nodes, the attribute information and the executable diagnosis interface of each of the entity nodes, and the entity relationships between the entity nodes.

9. The method of claim 8, wherein, The entity types include resources, events, faults, root causes, and solutions, and the relationship types include dependency, possession, cause, and solution.

10. The method according to any one of claims 1 to 9, wherein, After the obtaining the fault diagnosis result corresponding to the potential cause node, further comprising: obtaining a fault diagnosis result of each of the potential cause nodes and a fault solution corresponding to the fault diagnosis result; sorting each of the potential cause nodes according to a preset fault troubleshooting order or a fault troubleshooting priority, and displaying the fault diagnosis result and the fault solution corresponding to a preset number of potential cause nodes sorted in front.

11. The method according to any one of claims 1 to 9, wherein, After the obtaining the fault diagnosis trigger information, further comprising: in a case where it is determined that the fault knowledge graph does not include a target resource node matching the fault identifier, determining a first diagnosis knowledge answer corresponding to the fault diagnosis trigger information based on a second pre-trained language model and the fault knowledge graph.

12. The method according to any one of claims 1 to 9, wherein, Further comprising: in a case where the fault identifier included in the fault diagnosis trigger information is not identified, determining a second diagnosis knowledge answer corresponding to the fault diagnosis trigger information based on a third pre-trained language model.

13. An electronic device, comprising a processor and a memory, the memory storing programs or instructions executable on the processor, the programs or instructions being executed by the processor to implement the steps of the method of any one of claims 1 to 12.

14. A computer-readable storage medium, the computer-readable storage medium storing programs or instructions, the programs or instructions being executed by a processor to implement the steps of the method of any one of claims 1 to 12.

15. A computer program product, the computer program product comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions that, when executed by a computer, cause the computer to perform the steps of the method of any one of claims 1 to 12.

Citation Information

Patent Citations

  • Fault root cause analysis method and device

    CN113259168A

  • Fault reason determination method and device, equipment and storage medium

    CN117459365A

  • Automatic construction of fault-finding trees

    WO2022033988A1

Cited By

  • MES intelligent fault processing method and device, computer equipment and storage medium

    CN121682542A

  • Equipment knowledge management system and method

    CN121722926A

  • A device knowledge management system and method

    CN121722926B

  • WAT parameter abnormity intelligent tracing method, system and device and storage medium

    CN121899622A

  • Method, device and equipment for auditing power transmission line fault report

    CN122347418A