Fault event processing method and device, equipment, medium and product
By generating and updating knowledge graphs in the financial system, the problem of low efficiency in manually reviewing fault events is solved, achieving high efficiency and accuracy in fault analysis and improving the stability and reliability of the system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-24
- Publication Date
- 2026-04-14
AI Technical Summary
In industries with stringent requirements for system reliability, such as the financial system, existing technologies rely on manual review of failure events, resulting in low efficiency in failure analysis and difficulty in quickly reusing historical experience.
By generating a knowledge graph based on historical failure events, failure solutions are dynamically updated, and the knowledge graph can be used to quickly provide accurate solutions when similar failures occur.
It improves the efficiency of fault location and resolution, shortens fault repair time, enhances system stability and reliability, and reduces losses caused by faults.
Smart Images

Figure CN121860009A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of big data, and in particular to a method, apparatus, equipment, medium and product for handling fault events. Background Technology
[0002] In industries with stringent reliability requirements, such as the financial system, business continuity is crucial for maintaining normal operations and protecting customer rights. Real-time analysis of failure events and the accumulation of experience are core aspects of ensuring business continuity. If a system failure occurs and the cause is not analyzed and addressed promptly and accurately, it may lead to business interruption and losses.
[0003] To address the complex needs of failure time analysis, the industry currently primarily employs an analytical approach that combines multi-source heterogeneous data integration with manual review. Data is collected from multiple channels, and then technical personnel, leveraging their expertise and experience, manually tracing the event timeline, filtering key information from massive amounts of data, analyzing the root causes of the failures, and writing detailed review reports.
[0004] However, multi-source heterogeneous data obtained from various channels is usually in unstructured or semi-structured form, lacking unified semantic relationships. This makes it difficult to comprehensively and accurately uncover the potential correlations behind the data, thus affecting the accuracy and efficiency of fault analysis. The fault review process relies heavily on manual operation, resulting in low efficiency. Furthermore, fault analysis results are typically stored in static document form, which cannot be parsed or correlated by machines, making it difficult to quickly reuse historical experience in subsequent events. Summary of the Invention
[0005] This application provides a fault event handling method, apparatus, equipment, medium, and product to solve the problem of low efficiency in fault analysis in the prior art, which relies on manual handling and review of fault events.
[0006] In a first aspect, this application provides a fault event handling method, including:
[0007] Based on the first fault information and the first knowledge graph of the target event, a first fault solution for the target event is determined, wherein the first knowledge graph is generated based on all historical fault events of the banking system;
[0008] The first knowledge graph is updated based on the first fault solution to obtain the second knowledge graph;
[0009] When the target event causes a second fault, a second fault solution is determined based on the second knowledge graph, wherein the second fault is of the same type as the first fault.
[0010] Secondly, this application provides a fault event handling apparatus, comprising:
[0011] The determination module is used to determine a first fault solution for the target event based on the first fault information of the target event and a first knowledge graph, wherein the first knowledge graph is generated based on all historical fault events of the banking system;
[0012] The processing module is used to update the first knowledge graph based on the first fault solution to obtain a second knowledge graph;
[0013] The determination module is used to determine a second fault solution for the second fault based on the second knowledge graph when the target event causes the second fault, wherein the second fault is of the same type as the first fault.
[0014] Thirdly, this application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0015] The memory stores computer-executed instructions;
[0016] The processor executes computer execution instructions stored in the memory to implement the fault event handling method as described in the first aspect and various possible implementations of the first aspect.
[0017] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions thereon, which, when executed by a processor, are used to implement the fault event handling method as described in the first aspect and various possible implementations of the first aspect.
[0018] Fifthly, this application provides a program product including a computer program that, when executed by a processor, implements the fault event handling method described above.
[0019] The fault event handling method, apparatus, equipment, medium, and product provided in this application obtain first fault information of a fault event and call a first knowledge graph generated based on historical fault events. After comparison and matching, a first fault solution is determined. After resolving the fault, new experience is integrated into the first knowledge graph for updating to obtain a second knowledge graph. When the same type of fault occurs again, a more accurate second fault solution is quickly determined based on the second knowledge graph. This method, through dynamic updating of the knowledge graph, achieves continuous accumulation and optimization of experience in handling similar faults, improves the efficiency of locating the root cause of faults and formulating solutions, effectively shortens fault repair time, improves the stability and reliability of bank system operation, and reduces losses caused by faults. Attached Figure Description
[0020] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0021] Figure 1 A flowchart illustrating a fault event handling method provided in this application embodiment. Figure 1 ;
[0022] Figure 2 A flowchart illustrating a fault event handling method provided in this application embodiment. Figure 2 ;
[0023] Figure 3 A flowchart illustrating a fault event handling method provided in this application embodiment. Figure 3 ;
[0024] Figure 4 A flowchart illustrating a fault event handling method provided in this application embodiment. Figure 4 ;
[0025] Figure 5 A flowchart illustrating a fault event handling method provided in this application embodiment. Figure 5 ;
[0026] Figure 6 This application provides a schematic diagram of the structure of a fault event handling device;
[0027] Figure 7 This is a schematic diagram of the structure of an electronic device provided in this application.
[0028] The accompanying drawings have illustrated specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to specific embodiments. Detailed Implementation
[0029] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0030] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with relevant laws, regulations, and standards, necessary confidentiality measures have been taken, they do not violate public order and good morals, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0031] Furthermore, the technical solution involved in this application, which involves big data analysis of user information (including but not limited to personal biometrics, identity data, consumption data, asset data, electronic terminal operation data, etc.) and the use of artificial intelligence technology for automated decision-making, and makes decisions that have a significant impact on personal rights based on the results of automated decision-making, provides users with corresponding operation entry points for users to choose to agree to or reject the results of automated decision-making; if the user chooses to reject, the process will proceed to the expert decision-making process.
[0032] It should be noted that the fault event handling methods, devices, equipment, media and products provided in this application can be used in the field of big data, or in any field other than big data. The application fields of the fault event handling methods, devices, equipment, media and products in this application are not limited.
[0033] In industries with stringent reliability requirements, such as the financial system, business continuity is crucial for maintaining normal operations and protecting customer rights. Real-time analysis of failure events and the reuse of learned knowledge are core elements in ensuring business continuity. If a system failure occurs and the cause cannot be analyzed and resolved promptly and accurately, it may lead to business interruption and losses.
[0034] To address the complex needs of failure time analysis, the industry currently primarily employs an analytical approach that combines data integration with manual review. Data is collected from multiple sources, and then technical personnel, leveraging their expertise and experience, manually tracing the event timeline, filtering key information from massive amounts of data, analyzing the root causes of the failures, and writing detailed review reports.
[0035] However, as system complexity increases, technicians need to simultaneously process multi-source heterogeneous data from multiple channels. This data is typically unstructured or semi-structured, lacking unified semantic relationships, making it difficult to comprehensively and accurately uncover potential correlations within the data, thus impacting the accuracy and efficiency of fault analysis. Fault review processes heavily rely on manual operation, resulting in low efficiency. Furthermore, fault analysis results are usually stored as static documents, which cannot be parsed or correlated by machines, making it difficult to quickly reuse historical experience in subsequent events.
[0036] The fault event handling method provided in this application transforms the analysis results of fault events into structured knowledge. By extracting this structured knowledge, key entities and relationships from the fault events are populated into a graph database to generate an indicator graph. Upon the occurrence of a new event, the method automatically retrieves relevant knowledge from the knowledge graph, improving the efficiency and accuracy of event analysis and fault solutions. Furthermore, the knowledge graph can be dynamically updated based on the fault handling solution for the current event. This allows the knowledge graph to not only support accurate analysis of new events but also proactively associate fault handling experience, enabling the dynamic accumulation, verification, and reuse of operational knowledge.
[0037] This application can be applied to IT operations and maintenance scenarios with high reliability requirements, such as finance, e-commerce, and cloud computing. For example, when a service experiences a database connection pool exhaustion failure, this application can generate a structured debriefing report by aggregating monitoring metrics, logs, chat logs, and other data in real time. The root cause (e.g., configuration error) and solutions (e.g., adjusting the connection pool limit) are dynamically updated to a knowledge graph. When a similar event occurs again, the system can automatically retrieve historical context and recommend validated solutions, significantly shortening the repair time.
[0038] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.
[0039] Figure 1 A flowchart illustrating a fault event handling method provided in this application embodiment. Figure 1 .like Figure 1 As shown, the fault event handling method provided in this embodiment includes:
[0040] S101: Determine the first fault solution for the target event based on the first fault information and the first knowledge graph of the target event.
[0041] The first knowledge graph is generated based on all historical failure events of the banking system. It includes details such as the causes, solutions, and failure modes of these historical failure events. The target event refers to the event currently causing a failure in the banking system. The first failure information refers to information such as the cause of the failure corresponding to the target event.
[0042] Understandably, during the operation of a banking system, when a fault event (i.e., the target event) occurs, the fault event processing system can collect the initial fault information. This initial fault information includes various data at the time of the fault, such as the specific manifestation of the fault, the location of the fault, the faulty equipment component, and abnormal parameters. This information helps the system analyze the cause of the fault event.
[0043] Simultaneously, the fault event handling system can access a pre-built first knowledge graph. This first knowledge graph, presented as a structured knowledge network, clearly displays the relationships between various fault types, their causes, and corresponding solutions. The fault event handling system can meticulously compare and intelligently match the collected first fault information with the knowledge in the first knowledge graph, filtering from the vast knowledge graph to generate the most suitable first fault solution for the current fault situation, providing clear guidance for subsequent fault repair work.
[0044] S102: Update the first knowledge graph based on the first fault solution to obtain the second knowledge graph.
[0045] Understandably, after handling the current fault event according to the first fault solution, the fault event handling system can comprehensively review the entire fault handling process, analyzing and determining the specific procedures and execution effects of the first fault solution during implementation. Specific analysis content may include whether the repair was successful, the speed of fault repair, the operational stability of the equipment after repair, and whether the repair caused any other potential problems.
[0046] Based on these comprehensive analyses and summaries, the system can perform targeted updates to the first knowledge graph, adding new fault handling experience and knowledge to it, optimizing its structure and content, and thus obtaining a more complete and richer second knowledge graph, providing more accurate knowledge support for handling similar faults in the future.
[0047] S103: When a second fault occurs due to a target event, determine a second fault solution based on the second knowledge graph.
[0048] The second fault is of the same type as the first fault.
[0049] Understandably, when a second fault occurs during the subsequent operation of the target event, and the system determines that the second fault belongs to the same fault type as the first fault, the fault event handling system can generate a new fault solution based on the second knowledge graph.
[0050] The system can collect information about the second fault again. Although the fault type is the same, the fault information may differ due to factors such as the equipment operating environment. Collecting the first fault information and then generating a fault solution can make the fault solution more accurate and realistic.
[0051] By combining the second fault information with the first knowledge graph and leveraging the richer and more accurate knowledge in the second knowledge graph, a second fault solution can be quickly determined through analysis. This second fault solution is optimized and adjusted based on the implementation experience of the first fault solution and the actual situation of the second fault, enabling a more accurate and efficient resolution of the second fault, ensuring the banking system returns to normal operation as soon as possible, and minimizing losses and impacts caused by the fault.
[0052] The fault event handling method provided in this embodiment determines a first fault solution for a target event based on first fault information and a first knowledge graph. The first knowledge graph is generated based on all historical fault events of the banking system. The first knowledge graph is then updated based on the first fault solution to obtain a second knowledge graph. When a second fault occurs in the target event, a second fault solution is determined based on the second knowledge graph. The second fault has the same fault type as the first fault. This method, through dynamic updates to the knowledge graph, enables continuous accumulation and optimization of experience in handling similar faults, improving the efficiency of locating fault causes and formulating solutions, effectively shortening fault repair time, and enhancing the stability and reliability of the banking system.
[0053] Figure 2 A flowchart illustrating a fault event handling method provided in this application embodiment. Figure 2 .like Figure 2 As shown, in Figure 1 Based on the examples, the process of generating knowledge graphs is described in detail, including:
[0054] S201: Obtain multi-source heterogeneous operation and maintenance data of historical fault events.
[0055] As is understandable, multi-source heterogeneous operation and maintenance data refers to operation and maintenance data from different sources and in different formats, such as monitoring systems, log platforms, alarm tools, and collaboration tools. Operation and maintenance data includes, but is not limited to: time-series metrics (such as memory usage, processors), event data, log text, alarm information, processing records, configuration change records, and unstructured chat logs.
[0056] After acquiring heterogeneous operation and maintenance data from multiple sources, the data can be preprocessed, including data cleaning, format conversion, and normalization.
[0057] S202: Perform timestamp standardization and semantic association processing on multi-source heterogeneous operation and maintenance data to generate a unified event stream.
[0058] Understandably, timestamp standardization refers to aligning time information in heterogeneous data using a unified time base to ensure the consistency of the event stream's time sequence. Semantic association refers to establishing semantic relationships between unstructured text (such as chat logs) and structured data (such as metrics) through natural language processing techniques (such as named entity recognition and relation extraction). A unified event stream refers to a multimodal data set ordered by time and semantically associated, including event timelines, key metrics, log summaries, chat content, etc.
[0059] The specific implementation process is as follows: First, the timestamps of multi-source data are standardized, converting different time formats into a unified standard time format to ensure comparability and consistency of events across time dimensions. Second, the semantic information describing events in each data source is analyzed in depth. Through techniques such as natural language processing or large speech models, the inherent semantic connections between events are mined, linking related events and removing redundant and conflicting information. Finally, the multi-source data, after time standardization and semantic association processing, is integrated into a logically coherent, temporally ordered, and semantically clear unified event stream.
[0060] S203: Extracting structured knowledge from a unified event flow.
[0061] Structured knowledge includes entity nodes and entity relationship edges. Specifically, it refers to entity nodes and entity relationship edges extracted through a preset schema. Entity nodes can include, for example, events, services, hosts, root causes, solutions, and deployments. Entity relationship edges represent the relationships between entities, and can include, for example, effects, caused by, dependent on, and resolved by.
[0062] Understandably, a fault event handling system can be based on a defined schema, such as a root cause schema (used to standardize the structure and information related to the root cause of an event) and a solution schema (used to specify the methods and steps for solving a specific problem), to parse and extract a unified event flow.
[0063] When extracting entity nodes, the system can identify and extract various entities closely related to the business scenario, such as services and root causes, based on the schema. When extracting entity relationship edges, the system can analyze the logical connections between entities in the unified event flow and extract relationship edges that can express the association between entities according to the relationship types defined in the schema (such as relationship edges describing causal relationships like "caused by..."), thereby obtaining structured knowledge.
[0064] S204: Assign state attributes to entity nodes and entity relationship edges in structured knowledge.
[0065] In this context, state attributes refer to metadata used to identify the lifecycle state of nodes or edges in a knowledge graph. State attributes can be any of the following: unverified, manually confirmed, verified and valid, or verified and invalid.
[0066] For example, "unverified" means that the knowledge (specifically, the causal or resolution relationship between nodes) has not been manually or automatically verified. "Manually verified" means that it has been verified and labeled by technical personnel. "Verified valid" means that the knowledge has been automatically verified and the corresponding knowledge is valid. "Verified invalid" means that the knowledge has been automatically verified and the corresponding knowledge is invalid. Valid knowledge means that the problem can be solved by following the knowledge. Conversely, invalid knowledge means that the problem cannot be solved by following the knowledge, and that the verification of the knowledge has some additional effects, which can be viewed later.
[0067] Understandably, in the process of persisting structured knowledge to a graph database, the system can assign this state attribute to each entity node and relation edge, and the reliability and timeliness of the knowledge graph can be managed through the state attribute.
[0068] Optionally, the system receives user requests for corrections to structured knowledge through a manual calibration interface, and updates entity nodes, entity relationship edges, and corresponding state attributes in the first knowledge graph based on the correction requests.
[0069] Understandably, the system also includes a manual calibration interface as an interactive module. This interface serves as a channel for interaction between the system and the user, and can receive correction requests from the user in real time regarding the extracted structured knowledge and corresponding state attributes.
[0070] Users can correct the structured knowledge output by the system. Correction principles include: some entity nodes contain incorrect or incomplete information; entity nodes do not conform to actual business scenarios; the association logic expressed by entity relationship edges is biased or inaccurate; state attribute annotations are incorrect; and the state attributes of nodes have changed after manual verification. Users can submit specific correction requests to the system through the manual calibration interface for any of these situations.
[0071] After receiving a user's correction request, the system can activate its internal processing mechanism to locate the extracted entity nodes and entity relationship edges based on the content that needs to be modified as specified in the request, and make detailed updates and modifications according to the user's requirements, thereby ensuring the accuracy of structured knowledge.
[0072] S205: Sort entity nodes and entity relationship edges according to their status attributes to determine the priority sorting result.
[0073] As is understandable, priority ranking refers to sorting nodes or edges in a knowledge graph based on state attributes (such as "human-confirmed" taking precedence over "unverified"), ensuring that highly reliable knowledge is retrieved first. After state attributes are assigned, the system can prioritize nodes and edges in the knowledge graph according to these attributes. For example, nodes with the state "human-confirmed" or "verified and valid" have higher priority than "unverified" nodes during retrieval. When a new event occurs, the system can prioritize retrieving high-priority knowledge, ensuring the high reliability of the context information input to the model when generating fault solutions.
[0074] S206: Generate the first knowledge graph based on priority ranking results and structured knowledge.
[0075] Understandably, after obtaining relevant information on all historical failure events, the entity nodes, entity relationship edges, and corresponding state attributes from the structured knowledge, along with their priority ranking, can be stored in a graph database to obtain the first knowledge graph. A graph database is a database that supports efficient storage and querying of graph-structured data. Storing structured knowledge in a graph database enables semantic association and dynamic updates of knowledge.
[0076] The fault event handling method provided in this embodiment unifies unstructured text (such as chat logs) and structured data (such as metrics) into a machine-processable event stream format through timestamp standardization and semantic association technologies. Aggregating multi-source heterogeneous data eliminates data silos. Storing structured knowledge through a graph database enables semantic association and dynamic updates of knowledge, making operational knowledge a reusable knowledge asset.
[0077] Figure 3 A flowchart illustrating a fault event handling method provided in this application embodiment. Figure 3 .like Figure 3 As shown, in Figure 1 Based on the embodiments, the process of generating the first fault solution is described in detail, including:
[0078] S301: Based on the first fault information and the first knowledge graph, determine the verified solution corresponding to the first fault information.
[0079] The verified solution is used to indicate the verified historical solutions in the first knowledge graph.
[0080] Understandably, the first fault information is a collection of relevant data on the various characteristics and phenomena exhibited when a fault occurs. The first knowledge graph is a knowledge system constructed based on historical fault information, which connects various fault-related knowledge elements in the form of a graph.
[0081] Retrieving verified solutions that match the first fault information from the first knowledge graph can be done using similarity calculation. Similarity calculation involves comparing and analyzing each key element of the first fault information with the knowledge represented by entity nodes and entity relationship edges in the first knowledge graph. By determining the degree of similarity between them, it is possible to identify which parts of the knowledge graph are most closely related to the current fault.
[0082] In determining a solution, relevant knowledge needs to be identified based on the state attributes marked in the knowledge graph. State attributes serve as a filter; during this process, solutions marked as manually processed or verified as effective should be given priority. These solutions represent knowledge obtained through actual operation by professionals or automated system testing, and thus possess a certain degree of reliability and practicality.
[0083] By taking into account both similarity and state attributes, we can more accurately identify verified solutions that match the first fault information from the first knowledge graph, thus providing support for fault resolution.
[0084] S302: Based on the verified solution and the first fault information, generate the first fault solution for the target event.
[0085] Understandably, once a validated solution is determined, a prompt report can be generated based on it. This prompt report can summarize the key steps and operational points of the validated solution. The purpose of generating the prompt report is to clearly present the main content of the validated solution, providing accurate and comprehensive input for the large language model.
[0086] Furthermore, integrating this alert report with the initial fault information creates a complete and comprehensive input dataset. This dataset contains the actual situation of the fault event and existing effective handling experience for that fault, providing rich contextual information for large language models.
[0087] Then, the integrated dataset is input into the large language model. The large language model is an advanced language processing tool based on deep learning technology, possessing powerful language understanding and generation capabilities. It can perform in-depth analysis and understanding of the input dataset, uncovering its underlying logic and relationships. It can combine the processing procedures from validated solutions with the specific features of the first fault information, utilizing its vast learned knowledge and language patterns to generate a first fault solution for the target event.
[0088] The fault event handling method provided in this embodiment filters out verified solutions corresponding to the first fault information from the first knowledge graph through similarity calculation and state attributes. This process can accurately locate reliable and effective handling experience that has been tested in practice, improving the efficiency and accuracy of the preliminary preparation for fault resolution. Then, a prompt report is generated based on the verified solution, and the prompt report and the first fault information are input into a large language model to generate a first fault solution. This process, with the help of the language understanding and generation capabilities of the large language model, fully integrates existing experience and the specific situation of the fault to generate a comprehensive and practical solution, which can effectively shorten the time for fault investigation and resolution, reduce the impact of the fault on the system operation, and improve the overall stability and reliability of the system operation.
[0089] Figure 4 A flowchart illustrating a fault event handling method provided in this application embodiment. Figure 4 .like Figure 4 As shown, in Figure 1 Based on the embodiments, the process of updating the first knowledge graph is described in detail, including:
[0090] S401: After implementing the first fault solution, determine the implementation result corresponding to the first fault solution.
[0091] Understandably, after implementing the first fault solution for the target event, the implementation results of the fault solution can be evaluated and added to the first knowledge graph to complete the update process of the first knowledge graph.
[0092] The fault event handling system can collect all kinds of data generated during the execution of the first fault solution. This data can include a variety of information from the moment the solution is started to the end of the execution period, such as changes in equipment operating parameters, system log records, network communication status, etc.
[0093] Through in-depth mining and analysis of this massive amount of data, the system can clearly understand the impact of the first fault solution on the target event during actual operation.
[0094] S402: If the implementation results indicate that the fault of the target event has been resolved, determine that the verification result of the first fault solution is verified as effective.
[0095] Understandably, if the implementation results obtained after data analysis indicate that the fault in the target event has been successfully resolved, the system can determine the verification result of the first fault solution as valid based on the pre-set rules and logical processes.
[0096] The proven effectiveness indicates that this first fault solution is feasible and effective in resolving the current target event fault. From a practical application perspective, this fault solution not only accurately locates the root cause of the target event fault, but also effectively eliminates the fault, restoring the target event to normal operation. This fault solution can provide effective successful experience for encountering similar faults in the future, helping to solve problems quickly and efficiently.
[0097] S403: If the implementation results indicate that the fault of the target event has not been resolved, determine that the verification result of the first fault solution is invalid.
[0098] Understandably, if the implementation results obtained after data analysis by the system indicate that the fault of the target event still exists and has not been successfully resolved, then the system can determine the verification result of the first fault solution as invalid based on the pre-set rules and logical flow.
[0099] The ineffectiveness of the solution indicates that the first fault solution has shortcomings and limitations in handling the current target event fault. This could be due to a deviation in the fault localization process, failing to accurately identify the true cause of the target event fault; or it could be due to problems in the design and implementation of the solution, making it unable to effectively eliminate the fault.
[0100] Ineffective solutions to faults can provide feedback for optimization. By analyzing these ineffective solutions and identifying their problems, the model can be improved. When encountering similar faults in the future, the model can avoid those steps or optimize the solution, prompting it to generate more reasonable and effective solutions, thereby increasing the success rate of fault resolution.
[0101] S404: Add the first fault solution and verification results to the first knowledge graph to obtain the second knowledge graph.
[0102] Understandably, after verifying the first fault solution, the system can add the first fault solution and its corresponding verification result to the existing first knowledge graph, thereby updating and improving the first knowledge graph.
[0103] The first knowledge graph, as a crucial carrier for storing and managing fault solutions and related knowledge, significantly impacts the efficiency and quality of subsequent fault handling by the system, influencing its richness and accuracy. Adding first fault solutions and their verification results to the graph further expands and enriches its content.
[0104] The newly added solutions can provide new practical experience for knowledge graphs. Validated solutions can serve as successful cases for future reference. When encountering similar faults, the system can quickly retrieve the solution from the knowledge graph and apply it, improving the efficiency of fault resolution. Validated ineffective solutions can provide negative examples for optimizing subsequent solutions, avoiding similar errors.
[0105] The second knowledge graph, obtained after data updates, has a more complete knowledge structure and can provide a more comprehensive, accurate, and reliable reference for future fault handling.
[0106] Figure 5 A flowchart illustrating a fault event handling method provided in this application embodiment. Figure 5 .like Figure 5 As shown, in Figure 1 Based on the implementation examples, the process of early warning based on knowledge graphs is described in detail, including:
[0107] S501: Determine the attribute information of each entity node in the third knowledge graph.
[0108] The third knowledge graph is the knowledge graph in the current state, and its attribute information includes multiple dimensions.
[0109] Understandably, the fault event handling system also includes a multi-factor anomaly assessment model. This model can comprehensively analyze multiple dimensions of the root cause nodes of each fault in the knowledge graph. These multiple dimensions include, but are not limited to, their historical frequency of occurrence, recent activity, the severity of the business impact associated with them, and the breadth of their downstream dependencies in the topology network.
[0110] This model can be triggered according to a preset cycle or by specific events, such as system changes or new code deployments. When the multi-factor anomaly assessment model is triggered, the current knowledge graph becomes the third knowledge graph, which can be used to assess the faults of entity nodes in the knowledge graph.
[0111] S502: Weight the attribute information from multiple dimensions to obtain the anomaly score for each entity node.
[0112] Understandably, by weighting the attribute information from multiple dimensions, a dynamically changing anomaly score can be output for each potential fault source. The formula for calculating the anomaly score is as follows:
[0113]
[0114] Where F is the occurrence frequency, To indicate relevance, B represents the weight of the largest business impact associated with the graph, and D represents the dependency of downstream services. This is an adjustable coefficient.
[0115] S503: Determine the warning node based on the anomaly score.
[0116] Understandably, when the anomaly score of any entity node exceeds the preset security threshold, that entity node becomes a warning node, meaning that the warning node has a high probability and high impact of potential risk points, and further prevention is needed.
[0117] S504: Based on a third knowledge graph, identify at least one solution that has a relationship with the warning node and whose state attribute is verified to be valid.
[0118] Understandably, once a high-risk warning node is identified, the warning event processing system can automatically query a third-party knowledge graph. From this graph, it can accurately retrieve solutions associated with the warning node that have historically been flagged as proven effective. This ensures that recommended preventative measures are based on past successes and can effectively address the potential failures of the warning node.
[0119] S505: Generate preventative work orders based on the attribute information of the solution and the early warning node.
[0120] Among them, preventative work orders are used to remind users to adjust and correct entity nodes.
[0121] Understandably, after obtaining a solution, the solution along with the multi-dimensional attribute information of the early warning node can be input into a large language model to generate a structured, executable preventative maintenance work order. This work order can be directly pushed to the task queue of technicians or integrated with an external task management system to remind relevant personnel to handle the early warning node in a timely manner and prevent failures from occurring.
[0122] The fault event handling method provided in this embodiment identifies high-risk items by real-time monitoring of the knowledge graph and periodically calculating the anomaly scores of nodes. It automatically retrieves "verified effective" solutions associated with high-risk items, generates preventative work orders, proactively avoids risks, and reduces the occurrence of recurring faults, thereby lowering the risk of business interruption.
[0123] Figure 6 This is a schematic diagram of a fault event handling device provided in this application. Figure 6 As shown, this application provides a fault event handling apparatus 600, which includes:
[0124] The determination module 601 is used to determine the first fault solution for the target event based on the first fault information and the first knowledge graph of the target event. The first knowledge graph is generated based on all historical fault events of the banking system.
[0125] Processing module 602 is used to update the first knowledge graph based on the first fault solution to obtain the second knowledge graph;
[0126] The determination module 601 is used to determine a second fault solution for the second fault based on a second knowledge graph when the target event causes a second fault, wherein the second fault has the same fault type as the first fault.
[0127] Optionally, the device may also include: an acquisition module 603;
[0128] Module 603 is used to acquire multi-source heterogeneous operation and maintenance data of historical fault events;
[0129] The processing module 602 is also used to perform timestamp standardization and semantic association processing on multi-source heterogeneous operation and maintenance data to generate a unified event stream;
[0130] The processing module 602 is also used to extract structured knowledge from the unified event stream, which includes entity nodes and entity relationship edges.
[0131] The determination module 601 is also used to persist structured knowledge to a graph database and generate the first knowledge graph.
[0132] Optionally, the processing module 602 is specifically used to assign state attributes to entity nodes and entity relationship edges in structured knowledge. The state attributes include any one of the following: unverified, manually confirmed, verified and valid, and verified and invalid.
[0133] Module 601 is specifically used to prioritize entity nodes and entity relationship edges according to their status attributes and determine the priority ranking result.
[0134] Module 601 is defined as being used to generate the first knowledge graph based on priority sorting results and structured knowledge.
[0135] Optionally, the acquisition module 603 is also used to receive user requests for correction of structured knowledge through a manual calibration interface;
[0136] The processing module 602 is also used to update the entity nodes, entity relationship edges and corresponding state attributes in the first knowledge graph according to the correction request.
[0137] Optionally, the determining module 601 is specifically used to determine the verified solution corresponding to the first fault information based on the first fault information and the first knowledge graph. The verified solution is used to indicate the verified historical solutions in the first knowledge graph.
[0138] The determination module 601 is specifically used to generate a first fault solution for the target event based on the verified solution and the first fault information.
[0139] Optionally, module 601 is specifically used to determine the implementation result corresponding to the first fault solution after implementing the first fault solution;
[0140] The determination module 601 is specifically used to determine that the verification result of the first fault solution is verified as effective when the implementation result indicates that the fault of the target event has been resolved.
[0141] The determination module 601 is specifically used to determine that the verification result of the first fault solution is invalid when the implementation result indicates that the fault of the target event has not been resolved.
[0142] The processing module 602 is specifically used to add the first fault solution and the verification result to the first knowledge graph to obtain the second knowledge graph.
[0143] Optionally, the determining module 601 is also used to determine the attribute information of each entity node in the third knowledge graph, which is the knowledge graph in the current state, and the attribute information includes multiple dimensions;
[0144] The processing module 602 is also used to perform weighted processing on attribute information of multiple dimensions to obtain an anomaly score for each entity node;
[0145] The determination module 601 is also used to determine the early warning node based on the anomaly score;
[0146] The determination module 601 is also used to determine, based on the third knowledge graph, at least one solution that has a relationship with the warning node and whose state attribute is verified to be valid;
[0147] The determination module 601 is also used to generate preventive work orders based on the attribute information of the solution and the early warning node. The preventive work orders are used to remind users to adjust and correct the entity nodes.
[0148] The fault event handling device provided in this application embodiment is similar in principle and technical effect to the implementation of each part of the aforementioned fault event handling method, and will not be described again here.
[0149] Figure 7 This is a schematic diagram of the structure of an electronic device provided in this application. Figure 7 This application provides an electronic device 700, which includes a receiver 701, a transmitter 702, a processor 703, and a memory 704.
[0150] Receiver 701 is used to receive commands and data;
[0151] Transmitter 702 is used to send commands and data;
[0152] Memory 704 is used to store instructions executed by the computer;
[0153] The processor 703 is used to execute computer execution instructions stored in the memory 704 to implement the various steps of the fault event handling method in the above embodiments. For details, please refer to the relevant descriptions in the foregoing embodiments of the fault event handling method.
[0154] Optionally, the memory 704 can be either standalone or integrated with the processor 703.
[0155] When the memory 704 is set up independently, the electronic device also includes a bus for connecting the memory 704 and the processor 703.
[0156] The implementation principle and technical effects of the electronic device provided in this embodiment can be found in the foregoing embodiments, and will not be repeated here.
[0157] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the method of any of the foregoing embodiments.
[0158] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the method of any of the foregoing embodiments.
[0159] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.
[0160] It should be further noted that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0161] It should be understood that the above-described device embodiments are merely illustrative, and the device of this application can also be implemented in other ways. For example, the division of units / modules in the above embodiments is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units, modules, or components may be combined, or integrated into another system, or some features may be ignored or not executed.
[0162] Furthermore, unless otherwise specified, the functional units / modules in the various embodiments of this application can be integrated into one unit / module, or each unit / module can exist physically separately, or two or more units / modules can be integrated together. The integrated units / modules described above can be implemented in hardware or in the form of software program modules.
[0163] When integrated units / modules are implemented in hardware, the hardware can be digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, the processor can be any suitable hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC, etc. Unless otherwise specified, the storage unit can be any suitable magnetic or magneto-optical storage medium, such as Resistive Random Access Memory (RRAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Enhanced Dynamic Random Access Memory (EDRAM), High-Bandwidth Memory (HBM), Hybrid Memory Cube (HMC), etc.
[0164] If the integrated unit / module is implemented as a software program module and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0165] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0166] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0167] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A fault event handling method, characterized in that, include: Based on the first fault information and the first knowledge graph of the target event, a first fault solution for the target event is determined, wherein the first knowledge graph is generated based on all historical fault events of the banking system; The first knowledge graph is updated based on the first fault solution to obtain the second knowledge graph; When the target event causes a second fault, a second fault solution is determined based on the second knowledge graph, wherein the second fault is of the same type as the first fault.
2. The method according to claim 1, characterized in that, The method further includes: Acquire multi-source heterogeneous operation and maintenance data of historical failure events; The multi-source heterogeneous operation and maintenance data are subjected to timestamp standardization and semantic association processing to generate a unified event stream; Structured knowledge is extracted from the unified event stream, and the structured knowledge includes entity nodes and entity relationship edges; The structured knowledge is persisted to a graph database to generate the first knowledge graph.
3. The method according to claim 2, characterized in that, The step of persisting the structured knowledge to a graph database and generating a first knowledge graph includes: Assign state attributes to the entity nodes and entity relationship edges in the structured knowledge, wherein the state attributes include any one of unverified, manually confirmed, verified and valid, and verified and invalid. Based on the state attributes, the entity nodes and entity relationship edges are prioritized and sorted to determine the priority sorting result; Based on the priority ranking result and the structured knowledge, the first knowledge graph is generated.
4. The method according to claim 3, characterized in that, After assigning state attributes to the entity nodes and entity relationship edges in the structured knowledge, the method further includes: The system receives user requests for corrections to the structured knowledge through a manual calibration interface. Update the entity nodes, entity relationship edges, and corresponding state attributes in the first knowledge graph according to the correction request.
5. The method according to claim 1, characterized in that, The step of determining the first fault solution for the target event based on the first fault information and the first knowledge graph of the target event includes: Based on the first fault information and the first knowledge graph, a verified solution corresponding to the first fault information is determined. The verified solution is used to indicate a verified historical solution in the first knowledge graph. Based on the verified solution and the first fault information, a first fault solution for the target event is generated.
6. The method according to claim 1, characterized in that, The step of updating the first knowledge graph based on the first fault solution to obtain a second knowledge graph includes: After implementing the first fault solution, determine the implementation result corresponding to the first fault solution; If the implementation results indicate that the fault of the target event has been resolved, the verification result of the first fault solution is determined to be verified as effective; If the implementation results indicate that the fault of the target event has not been resolved, the verification result of the first fault solution is determined to be invalid. The first fault solution and the verification result are added to the first knowledge graph to obtain the second knowledge graph.
7. The method according to claim 1, characterized in that, The method further includes: Determine the attribute information of each entity node in the third knowledge graph, which is the knowledge graph in the current state, and the attribute information includes multiple dimensions; The attribute information from multiple dimensions is weighted to obtain an anomaly score for each entity node; Based on the anomaly score, the warning node is determined; Based on the third knowledge graph, at least one solution is identified that has an association with the warning node and whose status attribute is verified to be valid. Based on the solution and the attribute information of the early warning node, a preventive work order is generated. The preventive work order is used to remind the user to adjust and correct the entity node.
8. A fault event handling device, characterized in that, include: The determination module is used to determine a first fault solution for the target event based on the first fault information of the target event and a first knowledge graph, wherein the first knowledge graph is generated based on all historical fault events of the banking system; The processing module is used to update the first knowledge graph based on the first fault solution to obtain a second knowledge graph; The determination module is used to determine a second fault solution for the second fault based on the second knowledge graph when the target event causes the second fault, wherein the second fault is of the same type as the first fault.
9. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1 to 7.
11. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method of any one of claims 1 to 7.