Methods and devices for locating the root cause of incident failures

By utilizing the inspection methods and similarity analysis of primary and secondary event fault root cause lists in banking business systems, the root cause of faults can be quickly located and a remediation solution can be provided. This solves the problem of long event fault location time in large-scale multi-module software systems and improves the accuracy of location and the efficiency of remediation.

CN119515520BActive Publication Date: 2025-10-28INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311024638.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-15
Publication Date
2025-10-28
Estimated Expiration
2043-08-15

AI Technical Summary

Technical Problem

The fault location and repair of large-scale, multi-module software systems suffer from problems such as long location time and inappropriate temporary alternative measures, leading to frequent incidents.

Method used

By acquiring pending fault events during the banking business process, the system uses the inspection methods in the primary event fault root cause list to determine the matching primary event fault root cause, and then matches it in the secondary event fault root cause list before sending it to the user. Upon receiving the user's fault management instruction, the system retrieves and executes the corresponding management solution information and scripts. When a primary event fault root cause cannot be matched, the system determines the management solution by analyzing historical events for similarity.

Benefits of technology

It enables rapid location of the root cause of event failures, improves the accuracy of location and the efficiency of governance, and provides targeted governance solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119515520B_ABST
    Figure CN119515520B_ABST
Patent Text Reader

Abstract

This invention discloses a method and apparatus for locating the root cause of an event failure, applicable to the field of artificial intelligence technology. The method includes: acquiring pending failure events during the banking business process; determining the primary event failure root cause matching the pending failure event based on the checking methods associated with each primary event failure root cause in the primary event failure root cause list; the checking methods associated with the primary event failure root causes are used to verify whether the corresponding primary event failure root cause has occurred; determining the secondary event failure root cause list under the primary event failure root cause based on the primary event failure root cause matching the pending failure event; and when a secondary event failure root cause matching the pending failure event exists in the secondary event failure root cause list, sending the secondary event failure root cause matching the pending failure event to the user. This invention achieves rapid location of event failure root causes, improves the accuracy of event failure root cause location, and enhances governance efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method and apparatus for locating the root cause of event failures. Background Technology

[0002] This section is intended to provide background or context for the embodiments of the invention set forth in the claims. The description herein is not an admission that it is prior art simply because it is included in this section.

[0003] With the continuous updates of new technologies and the rapid demand for new solutions, large-scale, multi-module software systems need to keep up with technological changes, policy updates, and market changes, leading to frequent incidents. The multi-layered architecture, distributed applications, fragmented data, and frequent version iterations result in long waiting times for incident location and repair, with temporary alternatives mostly involving closing transactions or rolling back versions.

[0004] In summary, there is an urgent need for a method to locate the root cause of event failures in order to solve the above problems. Summary of the Invention

[0005] This invention provides a method for locating the root cause of an event fault, thereby improving the accuracy and efficiency of locating the root cause of an event fault. The method includes:

[0006] Obtain pending fault events during the banking transaction process;

[0007] Based on the checking methods associated with each primary event fault root cause in the primary event fault root cause list, determine the primary event fault root cause that matches the fault event to be processed; the checking methods associated with the primary event fault root causes are used to verify whether the corresponding primary event fault root cause has occurred.

[0008] Based on the primary event fault root cause that matches the fault event to be processed, determine the list of secondary event fault root causes under the primary event fault root cause;

[0009] If a secondary event fault root cause that matches the pending fault event exists in the secondary event fault root cause list, the secondary event fault root cause that matches the pending fault event will be sent to the user.

[0010] Upon receiving a fault management instruction from a user, the system retrieves corresponding management solution information based on the root cause of the secondary event that matches the fault event to be processed; the management solution information includes a management script; and the management script is executed.

[0011] When there is no primary event fault root cause that matches the fault event to be processed, determine the event description corresponding to the fault event to be processed;

[0012] Based on the event descriptions corresponding to the fault events to be processed, determine the similarity between the fault events to be processed and each historical event;

[0013] Identify the governance scheme information corresponding to the historical event with the highest similarity; execute the governance script contained in the governance scheme information corresponding to the historical event with the highest similarity.

[0014] This invention also provides an event fault root cause localization device to improve the accuracy and efficiency of event fault root cause localization. The device includes:

[0015] The matching module is used to acquire pending fault events during the banking business processing; determine the primary event fault root source that matches the pending fault event based on the check methods associated with each primary event fault root source in the primary event fault root source list; the check methods associated with the primary event fault root sources are used to verify whether the corresponding primary event fault root source has occurred; determine the list of secondary event fault root sources under the primary event fault root source based on the primary event fault root source that matches the pending fault event; when there is a secondary event fault root source that matches the pending fault event in the list of secondary event fault root sources, send the secondary event fault root source that matches the pending fault event to the user;

[0016] The governance module, upon receiving a user's fault governance instruction, retrieves corresponding governance scheme information based on the secondary event root cause matching the fault event to be processed; the governance scheme information includes a governance script; the governance script is executed; if no primary event root cause matching the fault event to be processed exists, the event description corresponding to the fault event to be processed is determined; based on the event description corresponding to the fault event to be processed, the similarity between the fault event to be processed and each historical event is determined; the governance scheme information corresponding to the historical event with the highest similarity is determined; and the governance script contained in the governance scheme information corresponding to the historical event with the highest similarity is executed.

[0017] This invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the above-described method for locating the root cause of an event failure.

[0018] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for locating the root cause of event failures.

[0019] This invention also provides a computer program product, which includes a computer program that, when executed by a processor, implements the above-described event fault root cause localization method.

[0020] In this embodiment of the invention, pending fault events during the banking business processing are obtained; based on the checking methods associated with each primary event fault root source in the primary event fault root source list, a primary event fault root source matching the pending fault event is determined; the checking methods associated with the primary event fault root sources are used to verify whether the corresponding primary event fault root source has occurred; based on the primary event fault root source matching the pending fault event, a list of secondary event fault root sources under the primary event fault root source is determined; when there is a secondary event fault root source matching the pending fault event in the secondary event fault root source list, the secondary event fault root source matching the pending fault event is... The process involves sending the root cause of an event fault to the user. Upon receiving the user's fault management instruction, the system retrieves corresponding management solution information based on the secondary event fault root causes matching the fault event to be processed. This management solution information includes a management script. The management script is then executed. If no primary event fault root cause matches the fault event to be processed, the system determines the event description corresponding to the fault event. Based on the event description, the system determines the similarity between the fault event to be processed and various historical events. The system then determines the management solution information corresponding to the historical event with the highest similarity. Finally, the system executes the management script contained in the management solution information corresponding to the historical event with the highest similarity. Compared to existing technologies, this method, after determining the primary event fault root cause matching the fault event to be processed based on the checking method associated with each primary event fault root cause in the primary event fault root cause list, continues to match the fault event to be processed with secondary event fault root causes. This achieves rapid location of the event fault root cause, improves the accuracy of event fault root cause location, and provides a management solution, thus improving management efficiency. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:

[0022] Figure 1 This is a flowchart illustrating the method for locating the root cause of an event failure provided by the present invention.

[0023] Figure 2 This is a flowchart illustrating the method for locating the root cause of an event failure provided by the present invention.

[0024] Figure 3 This is a flowchart illustrating the method for locating the root cause of an event failure provided by the present invention.

[0025] Figure 4 This is a flowchart illustrating the method for locating the root cause of an event failure provided by the present invention.

[0026] Figure 5 This is a schematic diagram of the structure of the event fault root cause localization device provided by the present invention. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.

[0028] Figure 1 This is a flowchart illustrating a method for locating the root cause of an event failure provided in an embodiment of the present invention, as shown below. Figure 1 As shown, the method includes:

[0029] Step 101: Obtain pending fault events during the banking business processing.

[0030] Step 102: Based on the inspection method associated with each primary event fault root source in the primary event fault root source list, determine the primary event fault root source that matches the fault event to be processed.

[0031] It should be noted that the method of checking the root cause association of a Level 1 event is used to verify whether the corresponding Level 1 event root cause has occurred.

[0032] For example, based on the pending fault event, retrieve its primary event fault root source list, poll the primary event root sources q1, q2...qn, and check the matching degree. For example, if q1 is "day-to-day switching failure", the database is queried through the check method "sql statement" associated with q1 to find that the current date is T-1 (should be T), and the conclusion of day-to-day switching failure is drawn. The primary event fault root source matching of q1 is successful.

[0033] If the database receives the date T when checking the date q1, then the q1 check fails. It can be assumed that the pending fault event is not caused by q1. The matching fails, and the process is repeated until the traversal of qi(1~n) ends.

[0034] Step 103: Based on the primary event fault root source that matches the fault event to be processed, determine the list of secondary event fault root sources under the primary event fault root source.

[0035] Step 104: If a secondary event fault root cause that matches the fault event to be processed exists in the secondary event fault root cause list, send the secondary event fault root cause that matches the fault event to be processed to the user.

[0036] For example, the system polls the list of secondary event root causes to check for a match. If check q1 passes, it then matches the corresponding secondary root cause list q1-1, q1-2, q1-i...q1-n. The secondary root cause list can be empty, meaning that q1 is both the apparent root cause and the deepest root cause. If none of the secondary root causes match after traversing all of them, it means that the apparent cause is known, but the specific deepest cause is unknown, requiring manual intervention. For example, if q1 is "date unsuccessful" and q1-1 is "batch not completed," then it needs to check if q1-1 meets the criteria. Through the "check script" associated with q1-1, a database query is performed to find that the batch execution status is true, indicating that there are batches that have not been completed, and the q1-1 check result is passed.

[0037] For example, if q1 is "Communication area error, cardtype type mismatch," and q1-1 is "Downstream communication area modification," then checking if q1-1 matches requires a more complex checking script, including the following steps: 1. Pull the error-reporting service from the logs; 2. Pull the downstream service from the production rule base; 3. Search the software requirements to confirm whether the downstream service has modification steps; 4. If the software requirements match and there is a modification, then the q1-1 check passes. If no modification requirement is found, then the q1-1 check fails. Here, a semantic similarity of 80% between the software requirements is required for the check to pass.

[0038] The results of the matching q1-1 check are recorded and returned to the front end when the root cause of the problem is returned.

[0039] Step 105: After receiving the user's fault management instruction, retrieve the corresponding management solution information based on the secondary event fault root cause that matches the fault event to be processed.

[0040] It should be noted that the governance solution information includes the governance script; execute the governance script;

[0041] Step 106: If there is no primary event fault root cause that matches the fault event to be processed, determine the event description corresponding to the fault event to be processed.

[0042] Step 107: Determine the similarity between the fault event to be processed and each historical event based on the event description corresponding to the fault event to be processed.

[0043] Step 108: Determine the governance scheme information corresponding to the historical event with the highest similarity; execute the governance script contained in the governance scheme information corresponding to the historical event with the highest similarity.

[0044] The above solution determines the primary event fault root cause that matches the fault event to be processed based on the inspection method associated with each primary event fault root cause in the primary event fault root cause list. Then, it continues to match the fault event to be processed with secondary event fault root causes to achieve rapid location of event fault root causes, improve the accuracy of event fault root cause location, and provide a governance solution to improve governance efficiency.

[0045] In this embodiment of the invention, the event description includes event name, event ID, event level, occurrence time, resolution time, event phenomenon, application modules involved, scope of impact, primary event fault root cause, secondary event fault root cause, primary inspection script, secondary inspection script, secondary governance plan, secondary governance script, primary governance plan, primary governance script, etc.

[0046] Among them, the resolution time, root cause of the incident, inspection script, governance plan, and governance script were initially blank and will be supplemented by the user after final confirmation.

[0047] In one possible implementation, the inspection method confirms whether the root cause of the event failure is met by defining SQL statements, retrieving logs, and screening the environment, and then returns a Y / N indication when determining whether the root cause of the event failure is successfully matched.

[0048] The governance solution information includes governance scripts that execute database, container, and network requests and changes via shell to eliminate the root cause of the event failure. If both a primary and secondary event failure root cause are matched, the secondary governance solution and script will be output first. If no secondary event failure root cause is matched, the primary governance solution information and script will be output.

[0049] In this embodiment of the invention, after obtaining pending fault events during the banking business processing, a list of primary event fault root causes is generated and displayed on the front end.

[0050] In this embodiment of the invention, the event analysis system is divided into a front-end management system and a back-end retrieval system. The front-end management system receives standardized event descriptions from users, performs routine text segmentation, intent recognition, and similar word replacement, and then transmits the optimized event descriptions to the back-end retrieval system to generate event root causes.

[0051] The background retrieval system mainly includes the following components: S01, S02, S03, S04, and S05.

[0052] S01: Pull the event library to form an event rule library.

[0053] The event database contains event name, event ID, event level, occurrence time, resolution time, event symptoms, involved application modules, scope of impact, primary event root cause, secondary event root cause, primary inspection script, secondary inspection script, primary governance solution, primary governance script, secondary governance solution, and secondary governance script. Based on this, further summaries and additions are made to form the "Event Rule Database," which includes events, the number of primary event root causes, and the number of secondary event root causes.

[0054] S02: Retrieve business data to form a business rule base.

[0055] Enterprises streamline their business architecture, breaking it down into product models, process models, and entity models. The product model describes service conditions and processing rules, the process model includes standardized processing logic, and the entity model includes rules for how the enterprise stores data. It describes the business rules of the software product; for example, opening a bank account at a branch requires users to provide their ID card, mobile phone number, address, etc. After the teller successfully verifies the materials, they enter the information, open the account, and provide the card to the user.

[0056] By storing relevant product models, process models, and entity models, we can provide business processes and rules for normal and abnormal transactions.

[0057] S03: Retrieve change records to form a change rule base.

[0058] Changes refer to adjustments made to the system environment or business rules in the production environment. They are generally not reflected in software requirements or business rules and need to be recorded and analyzed independently. System environment changes include disaster recovery failover, database migration, network changes, container replication, and scaling up or down CPU and memory.

[0059] A change rule base typically includes: date, module, background, execution content, and business impact.

[0060] S04: Retrieve software requirement records to form a software requirement rule base.

[0061] The software requirements record the incremental changes in each version, documenting version changes in a time series. Analyzing the sequence of changes can provide firsthand information on the occurrence of events. The software requirement change record includes an overall overview (describing the change background), logical processing (change implementation flow), system menus and entry points (related transaction sections), and technical and business checkpoints (related verification points).

[0062] The software requirement record is based on the causal reasoning formed by SCM, and records the software requirement rule base.

[0063] S05: Pull the full production log library to form a module call rule library.

[0064] The logic between production operation services is stored in the database. The module call rule base generally includes the following information: timestamp, globally unique tracking ID, module ID, parent module ID, and request / return message.

[0065] Based on the tracking ID and timestamp, a call chain can be formed from service A to service B to service C. This helps to locate the call relationships between module services and ultimately pinpoint the deepest source of the event.

[0066] S06: Pull the code repository and table structure to form a code rule base.

[0067] To achieve a correspondence between the codebase and its semantics, it is necessary to analyze the code and comments line by line to form a code rule base.

[0068] In step 102 of this embodiment of the invention, the primary event fault root source that matches the fault event to be processed is determined according to the inspection method associated with each primary event fault root source in the primary event fault root source list. The step flow is as follows: Figure 2 As shown, the details are as follows:

[0069] Step 201: Sort the multiple first-level event fault root sources according to the occurrence frequency of each first-level event fault root source in the first-level event fault root source list. Then, sort the event root sources from high to low according to the occurrence frequency of 100, 20...1, and check the qi matching degree according to the sorting.

[0070] Step 202: According to the sorting results, the fault events to be processed are matched with the root causes of each first-level event fault in turn until a match is successful, so as to obtain the root causes of the first-level events that match the fault events to be processed.

[0071] In this embodiment of the invention, when the current first-level event fault root cause association check is performed and the check passes, it is determined that the current first-level event fault root cause and the fault event to be processed are successfully matched.

[0072] In this embodiment of the invention, after determining the list of secondary event root causes under the primary event root cause based on the primary event root cause matching the fault event to be processed, the process flow is as follows: Figure 3 As shown, the details are as follows:

[0073] Step 301: When the list of secondary event fault root causes is empty, send the primary event fault root cause that matches the fault event to be processed to the user.

[0074] Step 302: After receiving the user's fault management instruction, retrieve the corresponding management solution information based on the root cause of the primary event fault that matches the fault event to be processed, and execute the management script contained in the management solution information.

[0075] In this embodiment of the invention, the system automatically saves event snapshots, which include event descriptions and recommended root causes. If maintenance personnel confirm that the event was indeed caused by a recommended root cause after a certain hour or day, they return to the system and retrieve the event snapshot based on the date and keywords.

[0076] In one possible implementation, the user clicks "Generate Event Governance Solution" in the snapshot, and the system retrieves the governance solution and governance script based on the event root cause in the snapshot and outputs them to the front end for the operation and maintenance personnel to refer to.

[0077] In this embodiment of the invention, the specific governance solutions, described in Chinese, include, but are not limited to, executing SQL statements, rolling back the container version, increasing the container's CPU and memory, increasing the number of replicas, and closing the program.

[0078] Users can click "Script Edit" to query and edit governance scripts. The scripts are historical scripts of historical events. After automatically correcting the application's database environment, container registration environment, and network environment, users can click "Submit" to make the changes effective in the production environment.

[0079] In step 107 of this embodiment of the invention, the similarity between the fault event to be processed and each historical event is determined based on the event description corresponding to the fault event to be processed. The process flow is as follows: Figure 4 As shown, the details are as follows:

[0080] Step 401: Use the word2vec model to determine the event description vector of the fault event to be processed based on the event description corresponding to the fault event to be processed.

[0081] Step 402: Determine the similarity between the fault event to be processed and each historical event based on the event description vector of the fault event to be processed.

[0082] For a newly input event 'a', extract the event description, use word2vec to obtain the event description vector, and calculate the similarity r using the following formula.<a,b> Calculate the similarity of 'a' with other historical event questions.

[0083] For example:

[0084] r<a,b> =cov(a, b)

[0085] In this embodiment of the invention, the smaller the r value, the higher the similarity. By comparing the r values, the optimal matching historical event can be found.

[0086] In the above scheme, for cases where the event description is matched but the root cause of the event failure is not matched, word2vec is used to match the root cause of historical events, find the optimal matching historical event, and use its secondary governance scheme to display the event governance scheme and governance script.

[0087] In this embodiment of the invention, users can supplement the input of event root causes, governance solutions, and governance scripts, and this information is sent back to the event rule base for continuous storage.

[0088] Event snapshots are categorized into "Case Filed" and "Case Closed" statuses based on information completeness. Specifically, when retrieving the corresponding governance plan information based on the secondary event root cause matching the pending fault event, the snapshot is initialized to the "Case Filed" status. After the event information is supplemented, the snapshot is updated to the "Case Closed" status.

[0089] For cases that have been "filed," alert emails will be continuously sent to the user to facilitate timely tracking and processing, and to close the case as "filed" as soon as possible.

[0090] Events that have been "filed" will have their complexity calculated based on the time of case closure, the related modules, the event level, and manual scoring. These events will be classified as high-complexity events for learning, production, and modification purposes.

[0091] This invention also provides an interface mapping device, as described in the following embodiments. This device is as follows... Figure 5 As shown, the device includes:

[0092] The matching module 501 is used to obtain pending fault events during the banking business processing; determine the primary event fault root source that matches the pending fault event based on the check method associated with each primary event fault root source in the primary event fault root source list; the check method associated with the primary event fault root source is used to verify whether the corresponding primary event fault root source has occurred; determine the secondary event fault root source list under the primary event fault root source based on the primary event fault root source that matches the pending fault event; when there is a secondary event fault root source that matches the pending fault event in the secondary event fault root source list, send the secondary event fault root source that matches the pending fault event to the user;

[0093] The governance module 502 is used to, upon receiving a user's fault governance instruction, retrieve corresponding governance scheme information based on the secondary event root cause matching the fault event to be processed; the governance scheme information includes a governance script; execute the governance script; when no primary event root cause matching the fault event to be processed exists, determine the event description corresponding to the fault event to be processed; based on the event description corresponding to the fault event to be processed, determine the similarity between the fault event to be processed and each historical event; determine the governance scheme information corresponding to the historical event with the highest similarity; and execute the governance script contained in the governance scheme information corresponding to the historical event with the highest similarity.

[0094] In this embodiment of the invention, the matching module 501 is specifically used for:

[0095] Sort multiple first-level event fault sources according to the occurrence frequency of each first-level event fault source in the first-level event fault source list.

[0096] Based on the sorting results, the fault events to be processed are matched sequentially with the root causes of each primary event until a match is found, thus obtaining the primary event root causes that match the fault events to be processed.

[0097] In this embodiment of the invention, the matching module 501 is specifically used for:

[0098] If the current first-level event fault root cause association check is performed and the check passes, it is determined that the current first-level event fault root cause and the fault event to be processed are successfully matched.

[0099] In this embodiment of the invention, the matching module 501 is further configured to:

[0100] After determining the list of secondary event root causes under the primary event root cause based on the primary event root cause that matches the fault event to be processed, the primary event root cause that matches the fault event to be processed is sent to the user when the list of secondary event root causes is empty.

[0101] Upon receiving a fault management instruction from a user, the system retrieves the corresponding management solution information based on the root cause of the primary event fault that matches the fault event to be processed, and executes the management script contained in the management solution information.

[0102] In this embodiment of the invention, the governance module 502 is specifically used for:

[0103] The word2vec model is used to determine the event description vector of the fault event to be processed based on the event description corresponding to the fault event to be processed;

[0104] The similarity between the fault event to be processed and each historical event is determined based on the event description vector of the fault event to be processed.

[0105] Since the principle by which this device solves problems is similar to that of the event fault root cause localization method, the implementation of this device can be referred to the implementation of the event fault root cause localization method, and the repeated parts will not be repeated.

[0106] This invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the above-described method for locating the root cause of an event failure.

[0107] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for locating the root cause of event failures.

[0108] This invention also provides a computer program product, which includes a computer program that, when executed by a processor, implements the above-described event fault root cause localization method.

[0109] In this embodiment of the invention, pending fault events during the banking business processing are obtained; based on the checking methods associated with each primary event fault root source in the primary event fault root source list, a primary event fault root source matching the pending fault event is determined; the checking methods associated with the primary event fault root sources are used to verify whether the corresponding primary event fault root source has occurred; based on the primary event fault root source matching the pending fault event, a list of secondary event fault root sources under the primary event fault root source is determined; when there is a secondary event fault root source matching the pending fault event in the secondary event fault root source list, the secondary event fault root source matching the pending fault event is... The process involves sending the root cause of an event fault to the user. Upon receiving the user's fault management instruction, the system retrieves corresponding management solution information based on the secondary event fault root causes matching the fault event to be processed. This management solution information includes a management script. The management script is then executed. If no primary event fault root cause matches the fault event to be processed, the system determines the event description corresponding to the fault event. Based on the event description, the system determines the similarity between the fault event to be processed and various historical events. The system then determines the management solution information corresponding to the historical event with the highest similarity. Finally, the system executes the management script contained in the management solution information corresponding to the historical event with the highest similarity. Compared to existing technologies, this method, after determining the primary event fault root cause matching the fault event to be processed based on the checking method associated with each primary event fault root cause in the primary event fault root cause list, continues to match the fault event to be processed with secondary event fault root causes. This achieves rapid location of the event fault root cause, improves the accuracy of event fault root cause location, and provides a management solution, thus improving management efficiency.

[0110] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0111] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0112] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0113] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0114] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for locating the root cause of an event failure, characterized in that, include: Obtain pending fault events during the banking transaction process; Based on the checking methods associated with each primary event fault root cause in the primary event fault root cause list, determine the primary event fault root cause that matches the fault event to be processed; the checking methods associated with the primary event fault root causes are used to verify whether the corresponding primary event fault root cause has occurred. Based on the primary event fault root cause that matches the fault event to be processed, determine the list of secondary event fault root causes under the primary event fault root cause; If a secondary event fault root cause that matches the pending fault event exists in the secondary event fault root cause list, the secondary event fault root cause that matches the pending fault event will be sent to the user. After receiving the user's fault management instruction, the corresponding management solution information is retrieved based on the root cause of the secondary event fault that matches the fault event to be processed; The governance solution information includes the governance script; execute the governance script; When there is no primary event fault root cause that matches the fault event to be processed, determine the event description corresponding to the fault event to be processed; Based on the event descriptions corresponding to the fault events to be processed, determine the similarity between the fault events to be processed and each historical event; Identify the governance scheme information corresponding to the historical event with the highest similarity; execute the governance script contained in the governance scheme information corresponding to the historical event with the highest similarity.

2. The method for locating the root cause of an event failure as described in claim 1, characterized in that, Based on the checking methods associated with each primary event fault root cause in the primary event fault root cause list, determine the primary event fault root causes that match the fault event to be processed, including: Sort multiple first-level event fault sources according to the occurrence frequency of each first-level event fault source in the first-level event fault source list. Based on the sorting results, the fault events to be processed are matched sequentially with the root causes of each primary event until a match is found, thus obtaining the primary event root causes that match the fault events to be processed.

3. The event fault root cause localization method as described in claim 2, characterized in that, Based on the sorting results, the fault events to be processed are matched sequentially with the root causes of each primary event until a match is found, including: If the current first-level event fault root cause association check is performed and the check passes, it is determined that the current first-level event fault root cause and the fault event to be processed are successfully matched.

4. The method for locating the root cause of an event failure as described in claim 1, characterized in that, After determining the list of secondary event root causes under the primary event root causes based on the primary event root causes that match the fault events to be processed, the following is also included: When the list of secondary event fault root causes is empty, the primary event fault root cause that matches the fault event to be processed will be sent to the user.

5. The event fault root cause localization method as described in claim 4, characterized in that, When the list of secondary event fault root causes is empty, after sending the primary event fault root cause that matches the fault event to be processed to the user, the following steps are also included: Upon receiving a fault management instruction from a user, the system retrieves the corresponding management solution information based on the root cause of the primary event fault that matches the fault event to be processed, and executes the management script contained in the management solution information.

6. The method for locating the root cause of an event failure as described in claim 1, characterized in that, Based on the event description corresponding to the fault event to be processed, determine the similarity between the fault event to be processed and each historical event, including: Determine the event description vector of the fault event to be processed based on the event description corresponding to the fault event to be processed; The similarity between the fault event to be processed and each historical event is determined based on the event description vector of the fault event to be processed.

7. The method for locating the root cause of an event failure as described in claim 6, characterized in that, The event description vector of the fault event to be processed is determined based on the event description corresponding to the fault event to be processed, including: The word2vec model is used to determine the event description vector of the fault event to be processed based on the event description corresponding to the fault event to be processed.

8. A device for locating the root cause of an event failure, characterized in that, include: The matching module is used to obtain pending fault events in the banking business process; based on the checking methods associated with each primary event fault root source in the primary event fault root source list, it determines the primary event fault root source that matches the pending fault event; the checking methods associated with the primary event fault root source are used to verify whether the corresponding primary event fault root source has occurred. Based on the primary event fault root cause that matches the fault event to be processed, determine the list of secondary event fault root causes under the primary event fault root cause; If a secondary event fault root cause that matches the pending fault event exists in the secondary event fault root cause list, the secondary event fault root cause that matches the pending fault event will be sent to the user. The governance module is used to retrieve corresponding governance solution information based on the root cause of the secondary event that matches the fault event to be processed after receiving the user's fault governance instruction; The governance solution information includes the governance script; execute the governance script; When there is no primary event fault root cause that matches the fault event to be processed, determine the event description corresponding to the fault event to be processed; Based on the event descriptions corresponding to the fault events to be processed, determine the similarity between the fault events to be processed and each historical event; Identify the governance scheme information corresponding to the most similar historical events; The governance scripts are included in the governance scheme information corresponding to the historical event with the highest similarity.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and device for fault locating

    CN103001811A

  • Fault positioning method and device and storage medium

    CN111930547A