Obstacle removing knowledge graph acquisition method and device, equipment and medium

CN120725115BActive Publication Date: 2026-09-25BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510887847.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2026-09-25
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

[0002]随着技术的发展,智能化运维系统在人们的工作和生活中占据着愈加重要的位置,在运维系统的运行过程中,可能会出现运行故障的情况,从而对系统运行的安全性和稳定性造成影响

Benefits of technology

[0009]应当理解,本部分所描述的内容并非旨在标识本公开的实施例的关键或重要特征,也不用于限制本公开的范围。本公开的其它特征将通过以下的说明书而变得容易理解。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120725115B_ABST
    Figure CN120725115B_ABST
Patent Text Reader

Abstract

The present disclosure provides an obstacle knowledge graph acquisition method, device, equipment and medium, relating to artificial intelligence technology fields such as natural language processing, the method comprising: acquiring a system running fault event, and acquiring a candidate obstacle knowledge graph corresponding to the system; in response to the candidate obstacle knowledge graph being unable to troubleshoot the running fault event, obtaining a first target troubleshooting strategy for the running fault event based on the model capability of a pre-acquired large model; and updating the candidate obstacle knowledge graph based on the first target troubleshooting strategy to obtain an updated target obstacle knowledge graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of large language model technology, and in particular to the field of artificial intelligence technology such as natural language processing. Background Technology

[0002] With the development of technology, intelligent operation and maintenance systems are playing an increasingly important role in people's work and life. During the operation of these systems, malfunctions may occur, affecting their security and stability. Currently, troubleshooting these system faults manually is inefficient. Summary of the Invention

[0003] This disclosure presents a method, apparatus, device, and medium for acquiring a knowledge graph for troubleshooting.

[0004] According to a first aspect of this disclosure, a method for obtaining a troubleshooting knowledge graph is proposed, comprising: obtaining operational failure events of a system and obtaining candidate troubleshooting knowledge graphs corresponding to the system; in response to the inability to troubleshoot the operational failure events through the candidate troubleshooting knowledge graphs, obtaining a first target troubleshooting strategy for the operational failure events based on the model capabilities of a pre-acquired large model; and updating the candidate troubleshooting knowledge graphs based on the first target troubleshooting strategy to obtain an updated target troubleshooting knowledge graph.

[0005] According to a second aspect of this disclosure, a device for acquiring a troubleshooting knowledge graph is proposed, comprising: a first acquisition module, configured to acquire operational failure events of a system and acquire candidate troubleshooting knowledge graphs corresponding to the system; a second acquisition module, configured to, in response to the inability to troubleshoot the operational failure event through the candidate troubleshooting knowledge graph, obtain a first target troubleshooting strategy for the operational failure event based on the model capabilities of a pre-acquired large model; and an update module, configured to update the candidate troubleshooting knowledge graph based on the first target troubleshooting strategy to obtain an updated target troubleshooting knowledge graph.

[0006] According to a third aspect of this disclosure, an electronic device is proposed, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the obstacle removal knowledge graph acquisition method proposed in the first aspect above.

[0007] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the obstacle removal knowledge graph acquisition method proposed in the first aspect above.

[0008] According to the fifth aspect of this disclosure, a computer program product is proposed, comprising a computer program that, when executed by a processor, implements the method for acquiring the troubleshooting knowledge graph proposed in the first aspect above.

[0009] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0010] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0011] Figure 1 This is a flowchart illustrating a method for acquiring a knowledge graph for troubleshooting according to an embodiment of this disclosure;

[0012] Figure 2 This is a flowchart illustrating a method for acquiring a knowledge graph for troubleshooting according to another embodiment of this disclosure.

[0013] Figure 3 This is a flowchart illustrating a method for acquiring a knowledge graph for troubleshooting according to another embodiment of this disclosure.

[0014] Figure 4 This is a schematic diagram of the structure of a knowledge graph acquisition device for obstacle removal according to an embodiment of the present disclosure;

[0015] Figure 5 This is a schematic block diagram of an electronic device according to an embodiment of the present disclosure. Detailed Implementation

[0016] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0017] Data processing is a fundamental aspect of systems engineering and automatic control. Data is a form of expression of facts, concepts, or instructions, which can be processed manually or by automated devices. After data is interpreted and given meaning, it becomes information. Data processing involves the acquisition, storage, retrieval, processing, transformation, and transmission of data. The basic purpose of data processing is to extract and derive valuable and meaningful data from large amounts of potentially chaotic and difficult-to-understand data.

[0018] Natural Language Processing (NLP) is an important field within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Its main applications include machine translation, public opinion monitoring, automatic summarization, opinion extraction, text classification, question answering, text semantic comparison, speech recognition, and Chinese OCR.

[0019] Artificial Intelligence (AI) is a key driving force behind the new round of technological revolution and industrial transformation. It is a new technological science that studies and develops theories, methods, technologies, and application systems to simulate, extend, and expand human intelligence. AI is an important component of the discipline of intelligence; it attempts to understand the essence of intelligence and produce a new kind of intelligent machine capable of reacting in a manner similar to human intelligence. AI is a very broad science, encompassing robotics, speech recognition, image recognition, natural language processing, expert systems, machine learning, computer vision, and more.

[0020] Figure 1 This is a flowchart illustrating a method for acquiring a troubleshooting knowledge graph according to an embodiment of this disclosure, as shown below. Figure 1 As shown, the method includes:

[0021] S101, obtain the system's operational fault events and the corresponding candidate troubleshooting knowledge graph.

[0022] During the daily operation of the system, operational failures may occur. In such cases, relevant fault information about the system's operational failure can be obtained, and the events generated based on this fault information can be identified as operational failure events of the system.

[0023] In this embodiment of the disclosure, the system has a pre-built knowledge graph, which includes information such as historical fault events that occurred in the system within a set historical time range and the troubleshooting strategies used when the historical fault events were troubleshooted. In this scenario, the knowledge graph can be identified as the candidate troubleshooting knowledge graph corresponding to the system.

[0024] S102, in response to the inability to troubleshoot the operational failure event through the candidate troubleshooting knowledge graph, the first target troubleshooting strategy for the operational failure event is obtained based on the model capabilities of the pre-acquired large model.

[0025] In this embodiment of the disclosure, after obtaining the candidate troubleshooting knowledge graph, it is possible to identify from the various troubleshooting strategies already included in the candidate troubleshooting knowledge graph whether there is a strategy that can troubleshoot the operational failure events that have occurred in the current system.

[0026] This can be understood as follows: when it is identified that there is a strategy among the various troubleshooting strategies included in the candidate troubleshooting knowledge graph that can troubleshoot the current operational failure event, it can be determined that the operational failure can be troubleshooted through the current candidate troubleshooting knowledge graph.

[0027] Accordingly, when it is identified that there is no strategy among the various troubleshooting strategies included in the candidate troubleshooting knowledge graph that can troubleshoot the current operational failure event, it can be determined that the current candidate troubleshooting knowledge graph may not be able to troubleshoot the operational failure event.

[0028] In other words, for the candidate troubleshooting knowledge graph, the operational failure events that occur in the current system can be understood as new types of failure events that have not been learned by the candidate troubleshooting knowledge graph. In this scenario, new troubleshooting strategies can be obtained for new types of operational failure events, and troubleshooting can be carried out on the operational failure events based on the new troubleshooting strategies. Among them, the obtained new troubleshooting strategies can be determined as the first target troubleshooting strategies for operational failure events.

[0029] Optionally, a large model capable of generating troubleshooting strategies can be obtained, relevant information of the current running fault event can be input into the large model, and a troubleshooting strategy capable of handling the running fault event, i.e., the first target troubleshooting strategy, can be generated by calling the model capabilities of the large model.

[0030] S103, update the candidate obstacle removal knowledge graph based on the first target obstacle removal strategy to obtain the updated target obstacle removal knowledge graph.

[0031] In this embodiment of the disclosure, the first target obstacle removal strategy is an obstacle removal strategy not included in the candidate obstacle removal knowledge graph. In this scenario, the first target obstacle removal strategy can be updated to the candidate obstacle removal knowledge graph, and the updated candidate obstacle removal knowledge graph can be determined as the updated target obstacle removal knowledge graph.

[0032] This can be understood as follows: the candidate troubleshooting knowledge graph does not contain relevant fault information and corresponding troubleshooting strategy information for operational fault events. After updating the first target troubleshooting strategy to the graph, when other fault events similar to the operational fault event occur, the system can directly obtain the first target troubleshooting strategy from the updated target troubleshooting knowledge graph for troubleshooting, without having to execute a separate task to generate a new troubleshooting strategy.

[0033] Optionally, the first target obstacle removal strategy can be updated to the candidate obstacle removal knowledge graph based on the knowledge graph update method in related technologies, thereby obtaining the updated target obstacle removal knowledge graph.

[0034] The troubleshooting knowledge graph acquisition method proposed in this disclosure acquires system operation failure events and candidate troubleshooting knowledge graphs. When it is identified that an operation failure event cannot be troubleshooted through the candidate troubleshooting knowledge graph, a first target troubleshooting strategy is generated through the model capabilities of a large model, and the first target troubleshooting strategy is updated to the candidate troubleshooting knowledge graph to obtain an updated target troubleshooting knowledge graph. In this disclosure, in scenarios where operation failure events cannot be troubleshooted through the candidate troubleshooting knowledge graph, the model capabilities of the large model are invoked to obtain corresponding troubleshooting strategies for new types of operation failure events, improving the efficiency of obtaining troubleshooting strategies for new types of operation failure events. Updating the candidate troubleshooting knowledge graph based on the first target troubleshooting strategy expands the coverage of system failure types in the troubleshooting knowledge graph, optimizes the information completeness and timeliness of the troubleshooting knowledge graph, and increases the likelihood of obtaining the required troubleshooting strategy through the target troubleshooting knowledge graph in scenarios where fault events are troubleshooted based on the updated target troubleshooting knowledge graph, thereby improving the efficiency of troubleshooting strategy acquisition, improving system troubleshooting efficiency, and optimizing system operation security and stability.

[0035] In the above embodiments, the acquisition of the target troubleshooting knowledge graph and the troubleshooting of operational failure events can be combined with... Figure 2 To understand further, Figure 2 This is a flowchart illustrating a method for acquiring a troubleshooting knowledge graph according to another embodiment of this disclosure, as shown below. Figure 2 As shown, the method includes:

[0036] S201, Obtain system operation failure events.

[0037] Optionally, the operating status data of each system node of the system is obtained, and the target fault determination threshold matrix of the system is obtained;

[0038] In this embodiment of the disclosure, the operating status of each system node in the system that needs to be fault detected can be detected, thereby obtaining relevant data that can characterize the operating status of each system node, which can be used as the operating status data of each system node.

[0039] In this embodiment of the disclosure, the system has a preset fault determination threshold matrix. This can be understood as the ability to determine and identify whether a fault has occurred in the current operating state of each system node by using the fault determination thresholds included in the fault determination threshold matrix.

[0040] It should be noted that the preset fault determination threshold matrix can be dynamically updated as the system runs. In other words, for any fault determination threshold in the matrix, if it is found that the fault determination threshold does not match the actual situation of the corresponding system node as the system runs, the fault determination threshold corresponding to that system node can be updated.

[0041] Optionally, the candidate decision threshold matrix of the system can be obtained.

[0042] In this embodiment of the disclosure, the threshold matrix corresponding to the system that needs to be updated for fault determination can be determined as the candidate determination threshold matrix of the system.

[0043] Optionally, the system can obtain the operation logs of each system node and perform semantic analysis on the operation logs using the model capabilities of a large model to generate a dynamic update strategy for the candidate decision threshold matrix.

[0044] In this embodiment of the disclosure, each system node of the system has a log composed of relevant operating information, which can be identified as the operating log of each system node. In this scenario, for any system node, the operating log of the system node can be analyzed to identify whether the fault judgment threshold corresponding to the system node in the candidate judgment threshold matrix needs to be updated.

[0045] This involves using the model capabilities of a large model to perform semantic analysis on the operation logs of each system node, and then obtaining the status data of each system node when a fault occurs during actual operation based on the semantic analysis results of the operation logs. Specifically, for any system node, if the error between the candidate fault judgment threshold corresponding to the system node in the candidate fault judgment threshold matrix and the status data exceeds the set error range, it can be determined that the candidate fault judgment threshold of the system node needs to be updated.

[0046] In this scenario, based on the model capabilities of the large model and the status data of the system node that needs to be updated when it fails during actual operation, the update failure judgment threshold of the system node can be obtained.

[0047] Furthermore, an overall update strategy for candidate fault determination thresholds is obtained based on the updated fault determination thresholds of each system node, serving as a dynamic update strategy.

[0048] Optionally, the candidate decision thresholds in the candidate decision threshold matrix are updated based on a dynamic update strategy to obtain the updated target fault decision threshold matrix.

[0049] In this embodiment of the disclosure, each updated judgment threshold that needs to be updated in the candidate fault judgment threshold matrix can be obtained from the dynamic update strategy, and the update position of each updated judgment threshold in the candidate fault judgment threshold matrix can be determined.

[0050] Optionally, for any update judgment threshold, the data at the update position of the update judgment threshold can be cleaned to obtain a cleaned empty space, and the update judgment threshold can be filled into the empty space to realize the update of the candidate fault judgment threshold matrix based on the update judgment threshold.

[0051] Furthermore, based on the completion of updating all the update judgment thresholds in the dynamic update strategy, the updated target fault judgment threshold matrix is ​​obtained.

[0052] It should be noted that the threshold update of the candidate fault determination threshold matrix is ​​performed dynamically. That is, as the system runs, when a determination threshold that needs to be updated is found, the update process of the candidate fault determination threshold matrix can be initiated to obtain the updated target fault determination threshold matrix. Furthermore, during the running time after obtaining the target fault determination threshold matrix, if a determination threshold that needs to be updated is found, the target fault determination threshold matrix will be used as the candidate fault determination threshold matrix in the new round of updates for updating.

[0053] Optionally, based on the target fault determination threshold matrix and the operating status data of each system node, the existence of a faulty node in the system can be identified to obtain the operating fault event.

[0054] In this embodiment of the disclosure, the system can determine whether a fault has occurred in each system node by using the target fault determination threshold matrix and the operating status data of each system node.

[0055] Specifically, for any system node, the target fault determination threshold corresponding to the system node is obtained from the target fault determination threshold matrix.

[0056] In this embodiment of the disclosure, the system node has a corresponding node identifier, which can be used to search the target fault determination threshold matrix to determine the determination threshold of the system node from the determination thresholds included in the target fault determination threshold matrix, and use it as the target fault determination threshold of the system node.

[0057] Optionally, in response to the matching of the system node's operating status data with the target fault determination threshold, the system node is identified as a system operating fault node, and the existence of an operating fault node in the system is confirmed. Based on the operating fault node, the system's operating fault event is obtained.

[0058] In this embodiment of the disclosure, for any system node, there are set judgment conditions for the fault determination of the system node. The judgment conditions can be understood as the comparison relationship that the system node’s operating status data and its corresponding target fault determination threshold need to satisfy.

[0059] In other words, when the comparison between the system node's operating status data and its corresponding target fault determination threshold meets the corresponding determination conditions, the operating status data in that scenario can be determined as data that matches the target fault determination threshold.

[0060] In this scenario, it can be determined that the system node may experience operational failures, and it can be identified as an operational failure node in the system.

[0061] Furthermore, based on the events constructed from related technologies, and based on the relevant operational failure information of the operational failure node, a corresponding failure event is generated, namely, an operational failure event in the system.

[0062] It should be noted that after obtaining a system malfunction event, the troubleshooting strategy corresponding to the malfunction event may or may not be obtained through the candidate troubleshooting knowledge graph. The specific details regarding the scenario where a troubleshooting strategy cannot be obtained can be understood in conjunction with the following:

[0063] S202, in response to the identification that there is no reference fault event corresponding to the operational fault event in the candidate troubleshooting knowledge graph, it is determined that the operational fault event cannot be troubleshooted through the candidate troubleshooting knowledge graph.

[0064] In this embodiment of the disclosure, fault event information of operational fault events can be obtained, and information retrieval can be performed on the candidate troubleshooting knowledge graph based on the fault event information.

[0065] Specifically, when the search results indicate that there is no matching reference fault event information for the operational fault event in the candidate troubleshooting knowledge graph, it can be determined that the current operational fault event is a fault event under a fault type that has not been learned by the candidate troubleshooting knowledge graph.

[0066] Furthermore, it can be determined that no strategy for troubleshooting the operational failure event can be obtained through the candidate troubleshooting knowledge graph, and therefore it can be determined that the operational failure event cannot be troubleshooted through the candidate troubleshooting knowledge graph.

[0067] S203, Obtain fault event information of running fault events.

[0068] In this embodiment of the disclosure, the system's operational failure event includes relevant information about the failure, and this information can be identified as the failure event information of the operational failure event.

[0069] The fault event information may include the fault time when the system experiences a failure, the fault node, the current fault type of the fault node, the load information corresponding to the fault node, the hardware status information of the hardware to which the fault node belongs, and the environmental information of the environment in which the fault node is located. It may also include other fault information related to the operational failure, which is not specifically limited here.

[0070] S204 invokes the model capabilities of the large model to generate a set of candidate troubleshooting strategies for running fault events based on fault event information.

[0071] In this embodiment of the disclosure, when it is found that the troubleshooting strategy required for the running failure event cannot be obtained from the candidate troubleshooting knowledge graph, the corresponding large model can be obtained and the failure event information of the running failure event can be input into the large model, and then the troubleshooting strategy available for the running failure event can be generated based on the model capability of the large model.

[0072] Optionally, the large model can output multiple troubleshooting strategies available for operational failure events, and select the strategy that meets the preset troubleshooting conditions from these multiple available troubleshooting strategies as the first target troubleshooting strategy for the operational failure event.

[0073] In this scenario, multiple available troubleshooting strategies output by the large model can be identified as candidate troubleshooting strategies, and the set of multiple candidate troubleshooting strategies can be identified as the set of candidate troubleshooting strategies for running fault events.

[0074] The large model can be a large language model (LLM) or any other large model capable of generating obstacle avoidance strategies; no specific limitation is made here.

[0075] S205, perform feasibility verification and comprehensive evaluation on each candidate troubleshooting strategy in the candidate troubleshooting strategy set, so as to determine the first target troubleshooting strategy for the operational failure event from the candidate troubleshooting strategy set.

[0076] In this embodiment of the disclosure, after obtaining each candidate troubleshooting strategy in the candidate troubleshooting strategy set, it is necessary to compare and analyze each candidate troubleshooting strategy from multiple dimensions, and then select the first target troubleshooting strategy from the candidate troubleshooting strategy set based on the results of the comparison and analysis.

[0077] Optionally, each candidate obstacle removal strategy can be analyzed from a feasibility perspective, wherein a subset of feasible obstacle removal strategies that have passed feasibility verification can be obtained from the set of candidate obstacle removal strategies.

[0078] In this embodiment of the disclosure, the feasibility of each candidate obstacle removal strategy can be verified by the feasibility verification method in the related technology, so as to realize the analysis of candidate obstacle removal strategies under the feasibility dimension.

[0079] As an example, a corresponding feasibility verification model can be constructed based on the causal model in the relevant technology, and the feasibility of each candidate obstacle removal strategy can be verified based on the model algorithm deployed in the feasibility verification model and the corresponding feasibility verification results can be output. Then, the candidate obstacle removal strategies that indicate that they have passed the feasibility verification in the feasibility verification results can be determined as candidate feasible obstacle removal strategies.

[0080] Furthermore, the set of each candidate feasible obstacle removal strategy is determined as a subset of the candidate feasible obstacle removal strategies in the candidate obstacle removal strategy set.

[0081] Optionally, a comprehensive evaluation is performed on each candidate feasible obstacle removal strategy in the subset of candidate feasible obstacle removal strategies to obtain a candidate evaluation score for each candidate feasible obstacle removal strategy.

[0082] In this embodiment of the disclosure, after obtaining a subset of candidate feasible obstacle removal strategies, each candidate feasible obstacle removal strategy can be further comprehensively evaluated and scored to obtain an evaluation score for each candidate feasible obstacle removal strategy. This score can be determined as the candidate evaluation score for each candidate feasible obstacle removal strategy.

[0083] As an example, a corresponding comprehensive evaluation model can be constructed based on the confidence model in related technologies. The evaluation algorithm deployed in the model can be used to process each candidate feasible obstacle removal strategy, thereby outputting the candidate evaluation score of each candidate feasible obstacle removal strategy.

[0084] Optionally, the candidate feasible troubleshooting strategy with the highest score among the candidate evaluation scores is selected as the first target troubleshooting strategy for the operational failure event.

[0085] In this embodiment of the disclosure, each candidate feasible troubleshooting strategy can be screened based on the evaluation scores of each candidate, and the candidate feasible troubleshooting strategy that meets the set screening conditions can be determined. This strategy can be determined as the first target troubleshooting strategy to be used when troubleshooting operational failure events.

[0086] Among these options, the candidate feasible obstacle removal strategy with the highest score among all candidate evaluation scores can be selected as the first target obstacle removal strategy. Alternatively, the candidate evaluation scores can be further analyzed and calculated to obtain the first target obstacle removal strategy from the subset of candidate feasible obstacle removal strategies based on the results of the further analysis and calculation.

[0087] Optionally, troubleshooting of operational failure events can be performed based on the first objective troubleshooting strategy.

[0088] In this embodiment of the disclosure, after obtaining the first target troubleshooting strategy, the system troubleshooting method in the related technology can be used to troubleshoot the operation failure event based on the first target troubleshooting strategy. Specifically, the fault node to which the operation failure event belongs can be repaired based on the first target troubleshooting strategy so that the fault node can be restored to normal operation. Also, the fault node of the operation failure event can be isolated based on the first target troubleshooting strategy, and a backup node corresponding to the fault node can be obtained to replace it so that the system can be restored to normal operation. No specific limitations are made here.

[0089] S206, Update the candidate obstacle removal knowledge graph based on the first target obstacle removal strategy to obtain the updated target obstacle removal knowledge graph.

[0090] In this embodiment of the disclosure, since the troubleshooting of system operation failure events requires a certain degree of timeliness, in this scenario, although the first target troubleshooting strategy obtained from the candidate troubleshooting strategy set output by the large model can achieve the troubleshooting of operation failure events, its troubleshooting effect may still have a certain degree of room for optimization.

[0091] In this scenario, when it is necessary to update the first target obstacle removal strategy to the candidate obstacle removal knowledge graph, the first target obstacle removal strategy can be optimized based on the strategy optimization method in related technologies, and the candidate obstacle removal knowledge graph can be updated based on the optimized obstacle removal strategy, thereby obtaining the updated knowledge graph as the target obstacle removal knowledge graph.

[0092] Optionally, a troubleshooting simulation sandbox is constructed. In the troubleshooting simulation sandbox, simulated fault events that run fault events are constructed, and the troubleshooting strategy of the first target is optimized based on the simulated fault events to obtain the optimized second target troubleshooting strategy.

[0093] In this embodiment of the disclosure, a sandbox corresponding to the optimization of the first target obstacle removal strategy can be constructed based on the simulation sandbox construction algorithm in the related technology, which serves as the obstacle removal simulation sandbox. The obstacle removal simulation sandbox can be constructed based on the system architecture of the system to which the running fault event belongs, or it can be constructed based on the fault environment in which the running fault event occurs. No specific limitation is made here.

[0094] Optionally, simulation events corresponding to operational failure events can be constructed in the obstacle removal simulation sandbox as simulated failure events of operational failure events in the obstacle removal simulation sandbox. Furthermore, simulated obstacle removal processing is performed on the simulated failure events in the obstacle removal simulation sandbox based on the first target obstacle removal strategy.

[0095] Specifically, the simulated fault events can be troubleshooted based on the first target troubleshooting strategy to obtain the simulated troubleshooting results of the first target troubleshooting strategy.

[0096] In this embodiment of the present disclosure, the simulated fault events in the fault simulation sandbox can be processed for fault removal based on the first target fault removal strategy, and the result of the fault removal of the simulated fault events based on the first target fault removal strategy is determined as the simulated fault removal result of the simulated fault events.

[0097] Optionally, in response to the simulated obstacle removal result not meeting the preset obstacle removal conditions, the non-compliant obstacle removal strategy in the first target obstacle removal strategy that caused the simulated obstacle removal result not to meet the obstacle removal conditions is obtained.

[0098] In this embodiment of the disclosure, when the first target troubleshooting strategy troubleshoots a simulated fault event, the simulated troubleshooting result obtained may not meet the preset troubleshooting conditions. It can be understood that when the first target troubleshooting strategy troubleshoots a running fault event, the fault node corresponding to the running fault event may be able to execute various tasks loaded on the node, but it still has not recovered to a completely normal state. In this scenario, it can be determined that the simulated troubleshooting result does not meet the preset troubleshooting conditions.

[0099] In this scenario, the first target obstacle removal strategy can be analyzed based on the parts of the simulated obstacle removal results that do not meet the preset obstacle removal conditions. Then, based on the results of the strategy analysis, the part of the first target obstacle removal strategy that causes the simulated obstacle removal results to not meet the obstacle removal conditions can be identified. This part of the strategy can be determined as the non-compliance obstacle removal strategy in the first target obstacle removal strategy.

[0100] Optionally, the non-compliance troubleshooting strategies in the first target troubleshooting strategy are optimized to obtain a new optimized first target troubleshooting strategy. Then, the process is returned to continue troubleshooting the simulated fault events based on the new first target troubleshooting strategy until the new simulated troubleshooting results of the new first target troubleshooting strategy meet the troubleshooting conditions. The optimization ends, and a second target troubleshooting strategy to be updated to the candidate troubleshooting knowledge graph is obtained.

[0101] In this embodiment of the disclosure, the fault event information of the operation failure event can be further analyzed, and then the failure troubleshooting strategy can be adjusted and optimized according to the analysis results. The optimized strategy is then updated into the first target troubleshooting strategy, thereby realizing the update and optimization of the first target troubleshooting strategy, and the updated and optimized first target troubleshooting strategy is determined as the new first target troubleshooting strategy.

[0102] In this scenario, the simulated fault events in the obstacle removal simulation sandbox can be processed in the next round based on the new first-objective obstacle removal strategy, and the result obtained by the obstacle removal process based on the new first-objective obstacle removal strategy can be determined as the new simulated obstacle removal result of the new first-objective obstacle removal strategy.

[0103] Furthermore, the system continues to evaluate whether the new simulated obstacle removal results meet the preset obstacle removal conditions. If it is determined that the new simulated obstacle removal results still cannot meet the obstacle removal conditions, the system returns to optimizing the new first target obstacle removal strategy and continues to obtain new simulated obstacle removal results corresponding to the optimized obstacle removal strategy until the latest obtained simulated obstacle removal result meets the preset obstacle removal conditions. At this point, the optimization of the first target obstacle removal strategy ends, and the obstacle removal strategy corresponding to the simulated obstacle removal result that meets the obstacle removal conditions is determined as the second target obstacle removal strategy.

[0104] It should be noted that when troubleshooting simulated fault events based on the first target troubleshooting strategy, there is a possibility that the first target troubleshooting strategy does not need to be optimized. In this case, the simulated troubleshooting result of the first target troubleshooting strategy satisfies the troubleshooting condition, and the first target troubleshooting strategy is determined to be the second target troubleshooting strategy.

[0105] In this embodiment of the disclosure, when the simulated troubleshooting result corresponding to the first target troubleshooting strategy meets the preset troubleshooting conditions, it can be determined that after the fault node of the faulty running event is troubleshooted based on the first target troubleshooting strategy, the node recovers to the point where it can execute all the tasks on its load and recovers to a completely normal running state. In this scenario, it can be determined that the first target troubleshooting strategy does not need to be optimized, and thus the first target troubleshooting strategy can be directly determined as the second target troubleshooting strategy that needs to be updated to the candidate troubleshooting knowledge graph.

[0106] Optionally, the candidate obstacle removal knowledge graph is updated based on the second objective obstacle removal strategy to obtain the objective obstacle removal knowledge graph.

[0107] In this embodiment of the disclosure, the simulated troubleshooting result obtained after troubleshooting the operational failure event based on the second target troubleshooting strategy can meet the preset troubleshooting conditions. In this scenario, the second target troubleshooting strategy can be used as the troubleshooting strategy to update the candidate troubleshooting knowledge graph.

[0108] Optionally, obtain the reference event chain associated with the runtime failure event in the candidate troubleshooting knowledge graph.

[0109] In this embodiment of the disclosure, there is a certain degree of correlation between the system nodes running in the system. In this scenario, when any system node fails, it may be due to the failure of other system nodes. In this scenario, there is a certain degree of correlation between the failure event corresponding to the system node and the failure events on other running nodes that caused its failure.

[0110] Optionally, the correlation between the operational failure events occurring in the current system and the various failure events included in the candidate troubleshooting knowledge graph is analyzed, and then some events that are related to the operational failure events are selected from the candidate troubleshooting knowledge graph as reference events associated with the operational failure events in the candidate troubleshooting knowledge graph.

[0111] Furthermore, based on the existing relationships between each reference event in the candidate obstacle removal knowledge graph, a corresponding event chain can be constructed as a reference event chain composed of each reference event.

[0112] Optionally, based on the association between the operational failure event and each reference event in the reference event chain, the event position of the operational failure event in the candidate troubleshooting knowledge graph to which the reference event chain belongs is determined.

[0113] In this embodiment of the disclosure, the operational failure event and each reference event in the reference event chain have a certain degree of correlation. The correlation can be a causal relationship or other types of correlation.

[0114] In this scenario, based on the correlation between each reference event and the operational failure event, the upstream and downstream reference events with the highest correlation among the reference events can be determined. The position between these two reference events is the position where the operational failure event needs to be placed in the reference event chain.

[0115] Furthermore, corresponding empty slots are constructed in the candidate troubleshooting knowledge graph, and these empty slots are determined as the event positions required for the running fault events in the candidate troubleshooting knowledge graph.

[0116] Optionally, based on the event location, the operational failure events are populated into the candidate troubleshooting knowledge graph, and the populated operational failure events in the candidate troubleshooting knowledge graph are associated with the second target troubleshooting strategy to update the candidate troubleshooting knowledge graph and obtain the target troubleshooting knowledge graph.

[0117] In this embodiment of the disclosure, an associated troubleshooting strategy update position can be constructed based on the event position, and the operation failure event can be filled into the event position, and the second target troubleshooting strategy corresponding to the operation failure event can be filled into the troubleshooting strategy update position, thereby realizing the event strategy association between the operation failure event filled into the event position and the second target troubleshooting strategy.

[0118] Furthermore, based on the operational failure events and the filling of the second target troubleshooting strategy, the candidate troubleshooting knowledge graph was updated, resulting in the updated target troubleshooting knowledge graph.

[0119] The troubleshooting knowledge graph acquisition method proposed in this disclosure dynamically updates the candidate fault judgment threshold matrix, improving the accuracy and precision of each target fault judgment threshold in the updated target fault judgment threshold matrix. This, in turn, improves the accuracy and precision of identifying operational fault events in the system. Compared to related technologies that rely on manual fault detection, this method reduces the dependence on manual intervention and improves the efficiency and accuracy of fault detection. In scenarios where the candidate troubleshooting knowledge graph cannot troubleshoot operational fault events, the method utilizes the capabilities of a large model to obtain corresponding troubleshooting strategies for new types of operational fault events, improving the efficiency of obtaining troubleshooting strategies for new types of operational fault events. The method performs sandbox simulation and optimization on the first target troubleshooting strategy to obtain a second target troubleshooting strategy, thus optimizing the troubleshooting strategy corresponding to the operational fault event. Based on the second target troubleshooting strategy, the candidate troubleshooting knowledge graph is updated, improving the strategy quality of the troubleshooting strategies in the target troubleshooting knowledge graph, expanding the coverage of system fault types in the troubleshooting knowledge graph, optimizing the information completeness and timeliness of the troubleshooting knowledge graph, and improving the practicality and applicability of the target troubleshooting knowledge graph.

[0120] In the above embodiments, troubleshooting operational failure events can be combined with... Figure 3 understand, Figure 3 This is a flowchart illustrating a method for acquiring a troubleshooting knowledge graph according to another embodiment of this disclosure, as shown below. Figure 3 As shown, the method includes:

[0121] Optionally, in response to identifying a reference fault event corresponding to an operational fault event in the candidate troubleshooting knowledge graph, it is determined that the operational fault event can be troubleshooted through the candidate troubleshooting knowledge graph.

[0122] In this embodiment of the disclosure, information retrieval can be performed on the candidate troubleshooting knowledge graph based on the fault event information of the operational fault event. When a fault event corresponding to the operational fault event is found in the candidate troubleshooting knowledge graph, it can be determined that the operational fault event can be troubleshooted through the candidate troubleshooting knowledge graph.

[0123] Among them, the fault event can be identified as a reference fault event in the candidate troubleshooting knowledge graph of the running fault event.

[0124] S301, obtain the reference fault event corresponding to the running fault event in the candidate troubleshooting knowledge graph.

[0125] In this embodiment of the disclosure, relevant event identification information of operational failure events can be obtained, and information retrieval can be performed in the candidate troubleshooting knowledge graph based on the event identification information, so as to filter out the identification information that matches the event identification information of the operational failure event from the candidate troubleshooting knowledge graph.

[0126] Among them, the event corresponding to the matching identification information in the candidate troubleshooting knowledge graph can be identified as the reference fault event of the running fault event in the candidate troubleshooting knowledge graph.

[0127] S302, obtain the reference troubleshooting strategy associated with the reference fault event in the candidate troubleshooting knowledge graph, and determine the reference troubleshooting strategy as the third target troubleshooting strategy.

[0128] In this embodiment of the disclosure, if a troubleshooting strategy is associated with a reference fault event in the candidate troubleshooting knowledge graph, the strategy can be determined as the reference troubleshooting strategy associated with the reference fault event.

[0129] The reference troubleshooting strategy can be understood as a troubleshooting strategy that can achieve troubleshooting of reference fault events. In the scenario where the reference fault event is the fault event corresponding to the running fault event, it can be determined that the troubleshooting of the running fault event can be achieved through the reference troubleshooting strategy.

[0130] In this scenario, the reference troubleshooting strategy can be determined as the troubleshooting strategy obtained from the candidate troubleshooting knowledge graph of the running fault event, that is, the third target troubleshooting strategy.

[0131] S303, troubleshooting operational failure events based on the third-objective troubleshooting strategy.

[0132] Optionally, based on the relevant troubleshooting information included in the third-target troubleshooting strategy, troubleshooting can be performed on the faulty nodes in the operational failure event, thereby enabling the system to resume normal operation.

[0133] The process of troubleshooting operational fault events based on the third-target troubleshooting strategy can be understood in conjunction with the process of troubleshooting operational fault events based on the first-target troubleshooting strategy proposed in the above embodiments, and will not be repeated here.

[0134] The troubleshooting knowledge graph acquisition method proposed in this disclosure directly obtains the troubleshooting strategies required for operational failure events through the knowledge graph, eliminating the need for separate troubleshooting strategy generation. This simplifies the troubleshooting strategy acquisition process, improves the efficiency of troubleshooting strategy acquisition, and optimizes the quality of troubleshooting strategies. In scenarios where troubleshooting operational failure events is based on troubleshooting strategies, it optimizes the troubleshooting effect of operational failure events and improves the stability and security of system operation.

[0135] An embodiment of this disclosure also proposes a device for acquiring a troubleshooting knowledge graph. Since the device for acquiring a troubleshooting knowledge graph proposed in this disclosure corresponds to the method for acquiring a troubleshooting knowledge graph proposed in the above embodiments, the implementation methods of the above-mentioned methods for acquiring a troubleshooting knowledge graph are also applicable to the device for acquiring a troubleshooting knowledge graph proposed in this disclosure. It will not be described in detail in the following embodiments.

[0136] Figure 4 This is a schematic diagram of the structure of a knowledge graph acquisition device for obstacle removal according to an embodiment of this disclosure, as shown below. Figure 4 As shown, the obstacle removal knowledge graph acquisition device 400 includes a first acquisition module 410, a second acquisition module 430, and an update module 430, wherein:

[0137] The first acquisition module 410 is used to acquire system operation failure events and acquire the corresponding candidate troubleshooting knowledge graph of the system.

[0138] The second acquisition module 420 is used to obtain the first target troubleshooting strategy for the running failure event based on the model capabilities of the pre-acquired large model in response to the inability to troubleshoot the running failure event through the candidate troubleshooting knowledge graph.

[0139] The update module 430 is used to update the candidate obstacle removal knowledge graph based on the first target obstacle removal strategy to obtain the updated target obstacle removal knowledge graph.

[0140] In one embodiment of this disclosure, the second acquisition module 420 is further configured to: acquire fault event information of the running fault event; invoke the model capabilities of the large model to generate a set of candidate troubleshooting strategies for the running fault event based on the fault event information; and perform feasibility verification and comprehensive evaluation on each candidate troubleshooting strategy in the set of candidate troubleshooting strategies to determine the first target troubleshooting strategy for the running fault event from the set of candidate troubleshooting strategies.

[0141] In one embodiment of this disclosure, the second acquisition module 420 is further configured to: acquire a subset of candidate feasible troubleshooting strategies that have passed feasibility verification from the candidate troubleshooting strategy set; comprehensively evaluate each candidate feasible troubleshooting strategy in the subset of candidate feasible troubleshooting strategies to obtain a candidate evaluation score for each candidate feasible troubleshooting strategy; and acquire the candidate feasible troubleshooting strategy with the highest score from each candidate evaluation score as the first target troubleshooting strategy for the operational failure event.

[0142] In one embodiment of this disclosure, the apparatus further includes a troubleshooting module for: troubleshooting operational failure events based on a first target troubleshooting strategy.

[0143] In one embodiment of this disclosure, the update module 430 is further configured to: construct a troubleshooting simulation sandbox; construct simulated fault events of running fault events in the troubleshooting simulation sandbox, and optimize the first target troubleshooting strategy based on the simulated fault events to obtain a second target troubleshooting strategy to be updated to the candidate troubleshooting knowledge graph; update the candidate troubleshooting knowledge graph based on the second target troubleshooting strategy to obtain the target troubleshooting knowledge graph.

[0144] In one embodiment of this disclosure, the update module 430 is further configured to: troubleshoot simulated fault events based on a first target troubleshooting strategy to obtain simulated troubleshooting results of the first target troubleshooting strategy; in response to the simulated troubleshooting results not meeting preset troubleshooting conditions, obtain the non-compliance troubleshooting strategies in the first target troubleshooting strategy that cause the simulated troubleshooting results not to meet the troubleshooting conditions; optimize the non-compliance troubleshooting strategies in the first target troubleshooting strategy to obtain an optimized new first target troubleshooting strategy, and return to continue troubleshooting simulated fault events based on the new first target troubleshooting strategy until the new simulated troubleshooting results of the new first target troubleshooting strategy meet the troubleshooting conditions, end the optimization, and obtain an optimized second target troubleshooting strategy.

[0145] In one embodiment of this disclosure, the update module 430 is further configured to: determine the first target obstacle removal strategy as the second target obstacle removal strategy in response to the simulated obstacle removal result of the first target obstacle removal strategy meeting the obstacle removal conditions.

[0146] In one embodiment of this disclosure, the update module 430 is further configured to: obtain the reference event chain associated with the running failure event in the candidate troubleshooting knowledge graph; obtain the event position of the running failure event in the candidate troubleshooting knowledge graph to which the reference event chain belongs based on the association relationship between the running failure event and each reference event in the reference event chain; fill the running failure event into the candidate troubleshooting knowledge graph based on the event position, and associate the filled running failure event in the candidate troubleshooting knowledge graph with the second target troubleshooting strategy to update the candidate troubleshooting knowledge graph and obtain the target troubleshooting knowledge graph.

[0147] In one embodiment of this disclosure, the apparatus further includes a determination module, configured to: determine that the operational fault event can be troubleshooted through the candidate troubleshooting knowledge graph in response to identifying a reference fault event corresponding to the operational fault event in the candidate troubleshooting knowledge graph; and determine that the operational fault event cannot be troubleshooted through the candidate troubleshooting knowledge graph in response to identifying a reference fault event corresponding to the operational fault event in the candidate troubleshooting knowledge graph.

[0148] In one embodiment of this disclosure, the determination module is further configured to: obtain a reference fault event corresponding to the operational fault event in the candidate troubleshooting knowledge graph; obtain a reference troubleshooting strategy associated with the reference fault event in the candidate troubleshooting knowledge graph, and determine the reference troubleshooting strategy as the third target troubleshooting strategy; and troubleshoot the operational fault event based on the third target troubleshooting strategy.

[0149] In one embodiment of this disclosure, the first acquisition module 410 is further configured to: acquire the operating status data of each system node of the system, and acquire the target fault determination threshold matrix of the system; and identify whether there is an operating fault node in the system based on the target fault determination threshold matrix and the operating status data of each system node, so as to obtain the operating fault event.

[0150] In one embodiment of this disclosure, the first acquisition module 410 is further configured to: acquire the candidate decision threshold matrix of the system; acquire the operation logs of each system node of the system, and perform semantic analysis on the operation logs through the model capabilities of the large model to generate a dynamic update strategy for the candidate decision threshold matrix; update each candidate decision threshold in the candidate decision threshold matrix based on the dynamic update strategy to obtain the updated target fault decision threshold matrix.

[0151] In one embodiment of this disclosure, the first acquisition module 410 is further configured to: for any system node, acquire the target fault judgment threshold corresponding to the system node from the target fault judgment threshold matrix; in response to the matching of the system node's operating status data with the target fault judgment threshold, determine the system node as a system operating fault node, and determine that an operating fault node has been identified in the system; and obtain the system operating fault event based on the operating fault node.

[0152] It should be noted that the foregoing explanation of the method for obtaining the obstacle removal knowledge graph also applies to the obstacle removal knowledge graph acquisition device of this embodiment, and will not be repeated here.

[0153] The troubleshooting knowledge graph acquisition device disclosed herein acquires system operational fault events and candidate troubleshooting knowledge graphs. When it is identified that an operational fault event cannot be troubleshooted through the candidate troubleshooting knowledge graphs, a first target troubleshooting strategy is generated using the model capabilities of a large model, and the first target troubleshooting strategy is updated in the candidate troubleshooting knowledge graph to obtain an updated target troubleshooting knowledge graph. In this disclosure, in scenarios where operational fault events cannot be troubleshooted through the candidate troubleshooting knowledge graphs, the model capabilities of the large model are invoked to obtain corresponding troubleshooting strategies for new types of operational fault events, improving the efficiency of obtaining troubleshooting strategies for new types of operational fault events. Updating the candidate troubleshooting knowledge graph based on the first target troubleshooting strategy expands the coverage of system fault types in the troubleshooting knowledge graph, optimizes the information completeness and timeliness of the troubleshooting knowledge graph, and increases the likelihood of obtaining the required troubleshooting strategy through the target troubleshooting knowledge graph in scenarios where fault events are troubleshooted based on the updated target troubleshooting knowledge graph, thereby improving the efficiency of troubleshooting strategy acquisition, and ultimately improving the system's troubleshooting efficiency and optimizing the system's operational security and stability.

[0154] According to embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0155] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0156] like Figure 5 As shown, device 500 includes a computing unit 501, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 502 or a computer program loaded from storage unit 508 into random access memory (RAM) 503. RAM 503 may also store various programs and data required for the operation of device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. Input / output (I / O) interface 505 is also connected to bus 504.

[0157] Multiple components in device 500 are connected to I / O interface 505, including: input unit 506, such as keyboard, mouse, etc.; output unit 506, such as various types of monitors, speakers, etc.; storage unit 508, such as disk, optical disk, etc.; and communication unit 509, such as network card, modem, wireless transceiver, etc. Communication unit 509 allows device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0158] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as the troubleshooting knowledge graph acquisition method and / or the troubleshooting knowledge graph acquisition method. For example, in some embodiments, the troubleshooting knowledge graph acquisition method and / or the troubleshooting knowledge graph acquisition method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by computing unit 501, it can perform the troubleshooting knowledge graph acquisition method and / or one or more steps of the troubleshooting knowledge graph acquisition method described above. Alternatively, in other embodiments, computing unit 501 can be configured to perform the troubleshooting knowledge graph acquisition method and / or the troubleshooting knowledge graph acquisition method by any other suitable means (e.g., by means of firmware).

[0159] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0160] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0161] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0162] To initiate interaction with a user account, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user account; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user account can submit input to the computer. Other types of devices can also be used to initiate interaction with the user account; for example, feedback submitted to the user account can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user account can be received in any form (including voice input, speech input, or tactile input).

[0163] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user account computer with a graphical user interface or web browser through which a user account can interact with the implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0164] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0165] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0166] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for acquiring a troubleshooting knowledge graph, wherein, The method includes: Obtain system operation failure events and obtain the corresponding candidate troubleshooting knowledge graph for the system; In response to the inability to troubleshoot the operational failure event through the candidate troubleshooting knowledge graph, a first target troubleshooting strategy for the operational failure event is obtained based on the model capabilities of the pre-acquired large model. The candidate obstacle removal knowledge graph is updated based on the first target obstacle removal strategy to obtain the updated target obstacle removal knowledge graph. The step of updating the candidate obstacle removal knowledge graph based on the first target obstacle removal strategy to obtain the updated target obstacle removal knowledge graph includes: Construct a sandbox for obstacle removal simulation; In the troubleshooting simulation sandbox, simulated fault events of the running fault events are constructed, and the first target troubleshooting strategy is optimized based on the simulated fault events to obtain a second target troubleshooting strategy to be updated to the candidate troubleshooting knowledge graph. Obtain the reference event chain associated with the operational failure event in the candidate troubleshooting knowledge graph. The reference event chain is constructed based on the existing association relationships of each reference event in the candidate troubleshooting knowledge graph. Based on the correlation between the operational failure event and each reference event in the reference event chain, the event position of the operational failure event in the candidate troubleshooting knowledge graph to which the reference event chain belongs is obtained; Based on the event location, the operational failure event is filled into the candidate troubleshooting knowledge graph, and the operational failure event filled into the candidate troubleshooting knowledge graph is associated with the second target troubleshooting strategy to update the candidate troubleshooting knowledge graph and obtain the target troubleshooting knowledge graph.

2. The method according to claim 1, wherein, In response to the inability to troubleshoot the operational failure event through the candidate troubleshooting knowledge graph, a first target troubleshooting strategy for the operational failure event is obtained based on the model capabilities of the pre-acquired large model, including: Obtain the fault event information of the aforementioned operational fault event; The model capabilities of the large model are invoked, and a set of candidate troubleshooting strategies for the operational failure event is generated based on the failure event information; Feasibility verification and comprehensive evaluation are performed on each candidate troubleshooting strategy in the candidate troubleshooting strategy set to determine the first target troubleshooting strategy for the operational failure event from the candidate troubleshooting strategy set.

3. The method according to claim 2, wherein, The step of verifying the feasibility and comprehensively evaluating each candidate troubleshooting strategy in the candidate troubleshooting strategy set to determine the first target troubleshooting strategy for the operational failure event from the candidate troubleshooting strategy set includes: Obtain a subset of feasible candidate obstacle removal strategies that have passed feasibility verification from the set of candidate obstacle removal strategies; A comprehensive evaluation is performed on each candidate feasible obstacle removal strategy in the subset of candidate feasible obstacle removal strategies to obtain a candidate evaluation score for each candidate feasible obstacle removal strategy. The candidate feasible troubleshooting strategy with the highest score among all candidate evaluation scores is selected as the first target troubleshooting strategy for the operational failure event.

4. The method according to any one of claims 1-3, wherein, The method further includes: Troubleshooting is performed on the operational failure events based on the first target troubleshooting strategy.

5. The method according to claim 1, wherein, In the obstacle-clearing simulation sandbox, simulated failure events of the operational failure events are constructed, and the first target obstacle-clearing strategy is optimized based on the simulated failure events to obtain a second target obstacle-clearing strategy to be updated to the candidate obstacle-clearing knowledge graph, including: Based on the first target troubleshooting strategy, the simulated fault event is troubleshooted to obtain the simulated troubleshooting result of the first target troubleshooting strategy; In response to the simulated obstacle removal result not meeting the preset obstacle removal conditions, the non-compliance obstacle removal strategy in the first target obstacle removal strategy that caused the simulated obstacle removal result not to meet the obstacle removal conditions is obtained; The non-compliance troubleshooting strategies in the first target troubleshooting strategy are optimized to obtain a new optimized first target troubleshooting strategy. Then, the process is returned to continue troubleshooting simulated fault events based on the new first target troubleshooting strategy until the new simulated troubleshooting results of the new first target troubleshooting strategy meet the troubleshooting conditions. The optimization ends, and the optimized second target troubleshooting strategy is obtained.

6. The method according to claim 5, wherein, The method further includes: In response to the simulated obstacle removal result of the first target obstacle removal strategy satisfying the obstacle removal condition, the first target obstacle removal strategy is determined to be the second target obstacle removal strategy.

7. The method according to claim 1, wherein, After acquiring the system's operational failure events and the corresponding candidate troubleshooting knowledge graph, the method further includes: In response to the identification of a reference fault event corresponding to the operational fault event in the candidate troubleshooting knowledge graph, it is determined that the operational fault event can be troubleshooted through the candidate troubleshooting knowledge graph. In response to the identification that the reference fault event corresponding to the operational fault event does not exist in the candidate troubleshooting knowledge graph, it is determined that the operational fault event cannot be troubleshooted through the candidate troubleshooting knowledge graph.

8. The method according to claim 7, wherein, The step of responding to the identification of a reference fault event corresponding to the operational fault event in the candidate troubleshooting knowledge graph, and determining that the operational fault event can be troubleshooted through the candidate troubleshooting knowledge graph, includes: Obtain the reference fault event corresponding to the operational fault event in the candidate troubleshooting knowledge graph; Obtain the reference troubleshooting strategy associated with the reference fault event in the candidate troubleshooting knowledge graph, and determine the reference troubleshooting strategy as the third target troubleshooting strategy; The operational failure event is troubleshooted based on the third objective troubleshooting strategy.

9. The method according to claim 1, wherein, The acquisition of system malfunction events includes: Obtain the operating status data of each system node of the system, and obtain the target fault determination threshold matrix of the system; Based on the target fault determination threshold matrix and the operating status data of each system node, identify whether there are any operating fault nodes in the system, so as to obtain the operating fault event.

10. The method according to claim 9, wherein, The step of obtaining the target fault determination threshold matrix of the system includes: Obtain the candidate decision threshold matrix of the system; The system obtains the operation logs of each system node and performs semantic analysis on the operation logs using the model capabilities of the large model to generate a dynamic update strategy for the candidate decision threshold matrix. The candidate decision thresholds in the candidate decision threshold matrix are updated based on the dynamic update strategy to obtain the updated target fault decision threshold matrix.

11. The method according to claim 10, wherein, The step of identifying whether there are operational fault nodes in the system based on the target fault determination threshold matrix and the operating status data of each system node, in order to obtain the operational fault event, includes: For any system node, obtain the target fault determination threshold corresponding to the system node from the target fault determination threshold matrix; In response to the system node's operating status data matching the target fault determination threshold, the system node is identified as the system's operating fault node, and the existence of the operating fault node in the system is confirmed. Based on the operational failure node, the operational failure event of the system is obtained.

12. A device for acquiring a knowledge graph for obstacle removal, wherein, The device includes: The first acquisition module is used to acquire system operation failure events and acquire the corresponding candidate troubleshooting knowledge graph of the system. The second acquisition module is used to obtain a first target troubleshooting strategy for the operational failure event based on the model capabilities of the pre-acquired large model in response to the inability to troubleshoot the operational failure event through the candidate troubleshooting knowledge graph. The update module is used to update the candidate obstacle removal knowledge graph based on the first target obstacle removal strategy to obtain the updated target obstacle removal knowledge graph. Specifically, the update module is used for: Construct a sandbox for obstacle removal simulation; In the troubleshooting simulation sandbox, simulated fault events of the running fault events are constructed, and the first target troubleshooting strategy is optimized based on the simulated fault events to obtain a second target troubleshooting strategy to be updated to the candidate troubleshooting knowledge graph. Obtain the reference event chain associated with the operational failure event in the candidate troubleshooting knowledge graph. The reference event chain is constructed based on the existing association relationships of each reference event in the candidate troubleshooting knowledge graph. Based on the correlation between the operational failure event and each reference event in the reference event chain, the event position of the operational failure event in the candidate troubleshooting knowledge graph to which the reference event chain belongs is obtained; Based on the event location, the operational failure event is filled into the candidate troubleshooting knowledge graph, and the operational failure event filled into the candidate troubleshooting knowledge graph is associated with the second target troubleshooting strategy to update the candidate troubleshooting knowledge graph and obtain the target troubleshooting knowledge graph.

13. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-11.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-11.

15. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-11.

Citation Information

Patent Citations

  • Method and system for access safety detection and isolation of virtualized user

    CN102902920A

  • Distribution network fault recovery method and system based on artificial intelligence

    CN117458432A

  • Fault diagnosis method, electronic equipment and storage medium

    CN119865418A

  • Numerical control system fault diagnosis method and system based on knowledge injection

    CN120122611A