Event root cause analysis method, computer equipment and storage medium
By building a question-and-answer guided incident root cause analysis model, the problem of relying on human experience in existing technologies is solved, a more standardized and efficient incident root cause analysis is achieved, and the analysis quality and enterprise management level are improved.
Patent Information
- Application Number
- CN202510715719.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-16
AI Technical Summary
Existing root cause analysis methods for incidents rely on the knowledge and experience of analysts, resulting in uneven analysis quality, difficulty in ensuring comprehensiveness and depth, and difficulty in quickly matching the actual needs of the enterprise.
Build an incident root cause analysis model based on question-and-answer guidance. By guiding event review, failure point identification and question-and-answer guidance, combined with enterprise workflow and equipment function design, it provides a logical failure point cause map and corrective action guide, reducing dependence on personnel knowledge and experience.
It improves the standardization and efficiency of root cause analysis of incidents, ensures the quality of analysis, reduces dependence on personnel experience, and improves the company's management level and production performance.
Smart Images

Figure CN120653779A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of experience feedback, and in particular to an event root cause analysis method, computer equipment, and storage medium. Background Art
[0002] Root cause analysis is a crucial step in the feedback process. The nuclear power industry typically conducts root cause analysis on critical status report events, defined as Level A or B. The goal is to identify the underlying causes behind the events through investigation and develop appropriate corrective actions to eliminate them, thereby effectively preventing similar incidents from recurring. Therefore, root cause analysis is crucial to the reliable operation of equipment and systems, the continuous improvement of employee behavior and organizational management, and the sustained improvement of corporate performance.
[0003] In general practice, the root cause analysis process typically includes planning, information collection and interviews, failure point analysis, root cause determination, corrective action development, and report preparation. Root cause analysts may selectively employ methods such as event factor diagrams, variation analysis, barrier analysis, task analysis, or fault tree analysis to conduct root cause analysis.
[0004] While mastering the aforementioned methods is crucial for root cause analysts to identify failure points during incidents, they often lack standardized and detailed guidance, as well as systematic and logical supporting documentation to facilitate the analysis process. This is particularly true when root cause analysts lack sufficient knowledge and experience, which can lead to a lack of logic, comprehensiveness, and depth in the failure cause analysis process. For example, if a worker incorrectly operated a valve, causing an incident, this failure could be related to factors such as the "work environment," "physical or mental state," "tools, maintenance aids, and personal protective equipment," "human-machine interface," and "procedures." Inexperienced root cause analysts may overlook less obvious factors like "physical or mental state," "work environment," and "human-machine interface" during their investigations, focusing instead on the fact that the "procedures" clearly state that the worker should have opened the correct valve, leading to corrective actions such as penalties, assessments, and retraining. However, the root cause of similar failures could also be due to poor lighting at the workplace, resulting in misreading of valve labels, or human-machine interface design issues where two valves appear identical but don't correspond closely to their labels, leading to confusion among workers. Consequently, insufficient knowledge and experience make it difficult for root cause analysts to ensure comprehensiveness, accuracy, and depth in their investigations and analysis. These factors ultimately lead to varying degrees of success in the quality of root cause analyses conducted across various companies.
[0005] Because mastering incident root cause analysis methods and acquiring and digesting knowledge and experience from historical events are difficult and require long-term participation in incident investigation and analysis, it is difficult to solve this problem in the short term through training. As a result, the number of personnel capable of completing high-quality incident root cause analysis work has always been unable to match the actual needs of the company's incident root cause analysis work. Therefore, it is necessary to study a more guided incident investigation and cause analysis method to reduce the heavy reliance of incident root cause analysis work on personnel knowledge and skills, ensure the improvement and maintenance of the company's incident root cause analysis quality level, further improve the efficiency and effectiveness of the company's experience feedback work, and continuously improve the company's organizational management level and production performance. Summary of the Invention
[0006] Based on this, it is necessary to provide an incident root cause analysis method, computer equipment and storage medium to address the problem that the quality of existing incident root cause analysis is seriously dependent on the knowledge and experience level of incident root cause analysts. By constructing a comprehensive and in-depth failure point cause map and adopting a step-by-step guidance form of questions and answers, incident root cause analysts are guided to conduct a comprehensive and in-depth investigation and cause analysis of incidents occurring in the enterprise, which greatly reduces the serious dependence of incident root cause analysis quality on incident root cause analysts and their experience level.
[0007] To achieve the above object, the present invention provides a method for analyzing the root cause of an event, comprising the following steps:
[0008] Step 1: Guide the incident review and improve the incident process;
[0009] Step 2: Identification of the guide failure point;
[0010] Step 3: Based on the incident root cause analysis model, guide the incident root cause analysis through questions and answers;
[0011] Step 4: Provide corrective action guidance corresponding to the root cause of the incident.
[0012] As one of the feasible methods, step one is to guide the incident review and improve the incident process, which includes the following steps:
[0013] Guide the root cause analysts of the incident to conduct a preliminary investigation into the actual behavior of the personnel and the actual action of the equipment in the incident, combining the enterprise workflow, management requirements and equipment function design related to the incident, and conduct a complete review of the incident to improve the process of the incident.
[0014] As one possible approach, step 2, guiding the identification of failure points, includes the following steps:
[0015] Guide the root cause analysts of the incident to compare the actual behavior of the personnel and the actual actions of the equipment reflected in the incident with the behavior of the personnel and the actions of the equipment required by the enterprise, determine the deviations in personnel behavior and equipment action that contribute to the incident, and identify them as the failure points of the incident.
[0016] As one of the possible implementation methods, human behavior deviation includes incorrectly performing actions that should be performed and not performing actions that should be performed; equipment action deviation includes not performing actions according to designed functions and refusing to perform actions according to designed functions.
[0017] As one of the possible implementation approaches, the incident root cause analysis model is constructed by the following method:
[0018] By summarizing and generalizing the main problems and historical data of root cause analysis of incidents in existing incident reports in the nuclear power industry, investigating and understanding and referring to advanced root cause analysis practices at home and abroad, integrating the root cause analysis work process management and equipment life cycle management processes of the domestic nuclear power industry and other advanced industries, and after multiple expert desktop talks, we constructed a root cause analysis model and its corresponding root cause database.
[0019] As one of the feasible ways, multiple failure point types are set in the event root cause analysis model; under each failure point type, all cause categories that may cause the occurrence of this type of failure point are set; each cause category is set with a complete cause node map according to the logical relationship; the cause node map logically expands the cause nodes layer by layer, and the bottom cause node is connected to other cause categories until it goes deep into the cause category of "management effectiveness"; each failure point type, cause category and cause node is set with guiding questions to guide the event root cause analyst to conduct event root cause analysis.
[0020] As one of the possible ways to achieve this, logical relationships include cause-effect relationships and general-specific relationships.
[0021] As one possible approach, step three is to guide the root cause analysis of the incident through questions and answers based on the incident root cause analysis model, which includes the following steps:
[0022] Step 301: The event root cause analysis model provides guiding questions corresponding to the failure point type; the event root cause analyst identifies the failure point type by answering the guiding questions corresponding to the failure point type one by one;
[0023] When the event root cause analyst answers "yes" to the guiding question corresponding to the failure point type, the event root cause analysis model continues to provide guiding questions corresponding to the cause categories that may cause the failure point of this type. The event root cause analyst answers the guiding questions corresponding to the cause categories that may cause the failure point one by one, and then the process proceeds to step 302.
[0024] When the event root cause analyst answers "no" to the guidance question corresponding to the failure point type, the event root cause analysis model stops providing guidance questions corresponding to the cause category that may have caused the occurrence of this type of failure point;
[0025] The guiding question corresponding to the cause category that may cause the occurrence of this type of failure point is whether this cause category is the cause of the occurrence of this type of failure point;
[0026] Step 302: When the event root cause analyst answers "yes" to the guiding question corresponding to the cause category that may cause the occurrence of this type of failure point, the event root cause analysis model continues to provide guiding questions corresponding to the cause nodes under this cause category. The event root cause analyst answers the guiding questions corresponding to the cause nodes under this cause category one by one to identify whether the cause node is the cause of the occurrence of this type of failure point, and then proceeds to step 303.
[0027] When the event root cause analyst answers "no" to the guidance question corresponding to the cause category that may cause the occurrence of this type of failure point, the event root cause analysis model stops providing guidance questions corresponding to the cause node under this cause category;
[0028] The guiding question corresponding to the cause node under this cause category is whether the cause node is the cause of the occurrence of this type of failure point;
[0029] Step 303: When the event root cause analyst answers "yes" to the guiding question corresponding to the cause node under the cause category, the event root cause analysis model continues to provide guiding questions corresponding to the cause node under the cause node;
[0030] When the event root cause analyst answers "no" to the guidance question corresponding to the cause node under the cause category, the event root cause analysis model stops providing guidance questions corresponding to the cause nodes under the cause node;
[0031] When the event root cause analyst answers "yes" to the guiding question corresponding to the lowest-level cause node under the cause category, the event root cause analysis model continues to provide guiding questions corresponding to the cause categories subsequent to the cause node, and the process returns to step 302;
[0032] When the event root cause analyst answers "no" to the guidance question corresponding to the lowest-level cause node under the cause category, the event root cause analysis model stops providing guidance questions corresponding to the cause categories subsequent to the cause node;
[0033] The event root cause analysis model takes the lowest-level cause nodes under each cause category that causes each failure point as the root cause of the event and sorts them by correlation.
[0034] As one of the feasible ways, the event root cause analysis model provides the deepest cause category that may cause each type of failure point to be "management effectiveness"; if the enterprise believes that it is not necessary to analyze the causes of all failure points to reach the cause category of "management effectiveness", the event root cause analysis model requires the event root cause analyst to provide corresponding reasons before the cause node that the enterprise believes is sufficient to avoid the occurrence of similar failure points can be reflected as the root cause or contributing cause of the incident.
[0035] As one possible approach, step 4 provides corrective action guidance corresponding to the root cause of the incident, including the following steps:
[0036] The incident root cause analysis model forms a corrective action guide corresponding to the root cause of the incident based on the corrective action data of historical outstanding incidents and good practices in the industry for reference and signature by incident root cause analysts.
[0037] In order to achieve the above objectives, in a second aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the above-mentioned event root cause analysis method when executing the computer-readable instructions.
[0038] In order to achieve the above objectives, in a third aspect, the present invention provides a computer-readable storage medium having computer-readable instructions stored thereon, which implement the steps of the above-mentioned event root cause analysis method when executed.
[0039] Beneficial technical effects of the present invention:
[0040] The event root cause analysis method, computer equipment and storage medium of the present invention, combined with historical event cause analysis data and the experience of experts in the field of experience feedback, construct a professional and comprehensive question-and-answer-guided event root cause analysis model and its corresponding event root cause database, sort out and summarize the guiding question data corresponding to the event type, cause category and cause node in the event root cause analysis model, construct the logical relationship between the event type, cause category and cause node, and at the same time develop a corrective action guide corresponding to the event root cause for event root cause analysis personnel to borrow and refer to, transforming event root cause analysis from a difficult and demanding knowledge-based task into a guiding rule-based task, which can guide enterprises to carry out event root cause analysis in a more standardized and detailed manner, greatly reduce the difficulty of event root cause analysis, improve the efficiency and effectiveness of event root cause analysis work, and ensure the quality of event root cause analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 A flowchart of an embodiment of the event root cause analysis method of the present invention;
[0042] Figure 2 A logical diagram of one embodiment of an incident root cause analysis model. DETAILED DESCRIPTION
[0043] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs; the terms used in the specification of the application are only for the purpose of describing specific embodiments and are not intended to limit this application; the term "include" and any variations thereof in the specification and claims of this application and the above-mentioned figure descriptions are intended to cover non-exclusive inclusions.
[0044] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0045] In previous incident root cause analysis practices, incident root cause analysis work was heavily dependent on the knowledge and experience of incident root cause analysts. Since the number of incident root cause analysts with rich knowledge and experience could not always match the number of enterprise incidents in a short period of time, the quality of incident root cause analysis was uneven.
[0046] The present invention constructs an event root cause analysis model based on question-and-answer guidance, sorts out and summarizes the guiding question data corresponding to the failure point type, cause category and cause node in the event root cause analysis model, constructs the logical relationship between the failure point type, cause category and cause node, and transforms the event root cause analysis from a difficult and demanding knowledge-based task into a guiding rule-based task. It can guide enterprises to carry out event root cause analysis in a more standardized and detailed manner, greatly reduce the difficulty of event root cause analysis, improve the efficiency and effectiveness of event root cause analysis work, and ensure the quality of event root cause analysis.
[0047] refer to Figure 1 The present invention provides a method for analyzing the root cause of an event, comprising the following steps:
[0048] Step 1: Guide the incident review and improve the incident process;
[0049] Step 2: Identification of the guide failure point;
[0050] Step 3: Based on the incident root cause analysis model, guide the incident root cause analysis through questions and answers;
[0051] Step 4: Provide corrective action guidance corresponding to the root cause of the incident.
[0052] In the present invention, as one of the possible implementation methods, step 1, guiding the event review and improving the event process, includes the following steps:
[0053] Guide the root cause analysts of the incident to conduct a preliminary investigation into the actual behavior of the personnel and the actual action of the equipment in the incident, combining the enterprise workflow, management requirements and equipment function design related to the incident, and conduct a complete review of the incident to improve the process of the incident.
[0054] In the present invention, as one of the possible implementation methods, step 2, guiding the identification of the failure point, includes the following steps:
[0055] Guide the root cause analysts to compare the actual behaviors of the personnel and equipment reflected in the incident with the behaviors of the personnel and equipment required by the enterprise, determine the deviations in personnel behavior and equipment operation that contributed to the incident, and identify them as the failure points of the incident;
[0056] Personnel behavior deviations include incorrectly performing actions that should be performed and not performing actions that should be performed; equipment action deviations include not performing actions according to designed functions and refusing to perform actions according to designed functions.
[0057] In the present invention, as one of the possible implementations, the event root cause analysis model is constructed by the following method:
[0058] By summarizing and analyzing the main problems and historical data of root cause analysis in existing incident reports in the nuclear power industry, investigating and understanding advanced root cause analysis practices at home and abroad, integrating the root cause analysis work process management and equipment life cycle management processes of the domestic nuclear power industry and other advanced industries, and through multiple expert desktop discussions, a professional and comprehensive root cause analysis model and its corresponding root cause database were constructed.
[0059] The event root cause analysis model sets multiple failure point types; under each failure point type, all possible cause categories that may cause that type of failure point are set; each cause category sets a complete cause node map based on logical relationships, including causal relationships and total-score relationships; the cause node map logically expands the cause nodes layer by layer, with the bottom-level cause node connected to other cause categories until it reaches the cause category of "management effectiveness"; each failure point type, cause category, and cause node is designed with guiding questions to guide the event root cause analyst in the event root cause analysis;
[0060] Step 3: Based on the incident root cause analysis model, guide the incident root cause analysis through questions and answers, including the following steps:
[0061] Step 301: The event root cause analysis model provides guidance questions corresponding to the failure point type; the guidance questions corresponding to the failure point type include whether the failure point type is "personnel performance", whether the failure point type is "equipment", and whether the failure point type is "external environment";
[0062] The incident root cause analyst identifies the failure point type by answering the corresponding guiding questions one by one;
[0063] When the event root cause analyst answers "yes" to the guiding question corresponding to the failure point type, the event root cause analysis model continues to provide guiding questions corresponding to the cause categories that may cause the occurrence of this type of failure point. Since there may be multiple cause categories that may cause the occurrence of this type of failure point, the event root cause analyst needs to answer the guiding questions corresponding to the cause categories that may cause the occurrence of the failure point one by one, and then proceeds to step 302;
[0064] When the event root cause analyst answers "no" to the guidance question corresponding to the failure point type, the event root cause analysis model stops providing guidance questions corresponding to the cause category that may have caused the occurrence of this type of failure point;
[0065] The guiding question corresponding to the cause category that may cause the occurrence of this type of failure point is whether this cause category is the cause of the occurrence of this type of failure point;
[0066] Step 302: When the event root cause analyst answers "yes" to the guiding question corresponding to the cause category that may cause the occurrence of this type of failure point, the event root cause analysis model continues to provide guiding questions corresponding to the cause nodes under this cause category. Since there are multiple cause nodes under this cause category, the event root cause analyst needs to answer the guiding questions corresponding to the cause nodes under this cause category one by one to identify whether the cause node is the cause of the occurrence of this type of failure point, and then proceeds to step 303;
[0067] When the event root cause analyst answers "no" to the guidance question corresponding to the cause category that may cause the occurrence of this type of failure point, the event root cause analysis model stops providing guidance questions corresponding to the cause node under this cause category;
[0068] The guiding question corresponding to the cause node under this cause category is whether the cause node is the cause of the occurrence of this type of failure point;
[0069] Step 303: When the event root cause analyst answers "yes" to the guiding question corresponding to the cause node under the cause category, the event root cause analysis model continues to provide guiding questions corresponding to the cause node under the cause node;
[0070] When the event root cause analyst answers "no" to the guidance question corresponding to the cause node under the cause category, the event root cause analysis model stops providing guidance questions corresponding to the cause nodes under the cause node;
[0071] When the event root cause analyst answers "yes" to the guiding question corresponding to the lowest-level cause node under the cause category, the event root cause analysis model continues to provide guiding questions corresponding to the cause categories subsequent to the cause node, and the process returns to step 302;
[0072] When the event root cause analyst answers "no" to the guidance question corresponding to the lowest-level cause node under the cause category, the event root cause analysis model stops providing guidance questions corresponding to the cause categories subsequent to the cause node;
[0073] The event root cause analysis model takes the lowest-level cause nodes under each cause category that causes each failure point as the root cause of the event and sorts them by correlation.
[0074] In the present invention, as one of the possible implementation methods, the failure point types are divided into "personnel performance", "equipment" and "external environment";
[0075] Among them, the nine cause categories that may cause or contribute to the occurrence of "personnel performance" type failure points include "procedures", "human-machine interface", "communication", "working environment", "tools, maintenance auxiliary facilities and personal protective equipment", "physical or mental state", "knowledge and skills", "work process management" and "management effectiveness";
[0076] The eight cause categories that may cause or contribute to the occurrence of "equipment" type failure points include "defects left over from installation during the construction phase", "equipment changes", "equipment damaged during storage or incorrect issuance", "improper equipment management", "defects left over from commissioning", "maintenance", "operation", and "deficiencies in quality, performance (mechanical, electrical, instrumentation and control, etc.) and system layout of equipment before it arrives on site".
[0077] The two categories of causes that may cause or contribute to the occurrence of "external environment" type failure points include "natural environment" and "human events";
[0078] In summary, there are 19 cause categories that may cause failure points. A total of 683 cause nodes are constructed in the 19 cause node maps corresponding to the 19 cause categories. Cause categories corresponding to the same failure point type can contribute to the occurrence of the failure point individually or collectively, and cause nodes corresponding to the same cause category can also contribute to the occurrence of the failure point individually or collectively.
[0079] All of the aforementioned failure point types, cause categories, and cause nodes are derived through logical deduction, drawing on management processes and requirements across various areas of the enterprise, as well as feedback from historical incidents. Furthermore, the logic and integrity of each failure point type, cause category, and cause node have been fully verified through numerous real-world case studies.
[0080] Although the cause node map constructed by the event root cause analysis model is already perfect enough, it is still considered that such an event root cause analysis model may not be able to fully guide event root cause analysts who lack knowledge and experience, and cannot quickly let them know how to conduct supplementary investigations to determine whether the corresponding cause node has contributed to the occurrence of the event. Therefore, the present invention designs guiding questions for each cause node of the event root cause analysis model through a lot of sorting, summarizing and induction. These guiding questions are actually the explanations and definitions of the corresponding cause nodes. When the event root cause analyst cannot accurately answer these guiding questions, he must conduct a supplementary investigation on the details of the event in order to obtain more detailed facts to confirm these cause nodes. In this way, the guiding nature of the event root cause analysis model is fully reflected, so that the event root cause analyst can use the event root cause analysis model to conveniently and quickly carry out cause analysis of the event.
[0081] During the construction of the incident root cause analysis model, not only was the completeness and guidance of individual cause categories considered, but also the logical relationships between cause categories. For example, if a worker operated the wrong valve during work due to failure to follow procedures, the corresponding cause category would be "Procedures." Further investigation revealed that the failure to follow procedures was due to the work preparer not including the correct procedure documents in the work package, resulting in a lack of available procedures. The corresponding cause category for this failure is "Work Process Management." Furthermore, the lowest-level cause category contributing to the incident can be further refined to "Management Effectiveness," a cause category related to the enterprise management system. Therefore, the incident root cause analysis model logically connects additional cause categories at the lowest level of the 19 cause categories, ultimately reaching the "Management Effectiveness" cause category. This approach further improves the logic of the incident root cause analysis model and greatly ensures its maintainability. With the subsequent changes in enterprise management models and technological innovations, the event root cause analysis model can be upgraded after sufficient verification. By simply adjusting the cause node map logic of a single cause category, all cause node maps involving this cause category can be uniformly adjusted, avoiding the possibility of errors caused by maintenance. According to statistics, through this method of connecting cause categories and cause nodes, the cause node map constructed by the event root cause analysis model already contains 163,587 cause nodes (including duplicate nodes), and duplicate cause nodes do not need to have their guidance issues maintained separately. After sufficient case verification, its completeness is sufficient to cover the root cause analysis logic of the vast majority of enterprise events that currently occur.
[0082] Following the above steps, the root cause analyst, guided by the questions and answers designed into the root cause analysis model, gradually investigates and analyzes the root cause of the incident. The root cause analysis model's construction scheme sets the deepest possible cause category for the incident as "management effectiveness." However, because the depth of root cause analysis is affected by resources and progress, the root cause analysis model does not impose a mandatory requirement. If, during the ongoing in-depth process, the root cause analyst determines that the cause node under investigation is sufficient to represent the root cause of the incident, correcting it can prevent the recurrence of most similar incidents. Subsequent analysis can be discontinued and the analysis can be terminated.
[0083] For example, if there is an error in the procedure of a certain work, which causes errors when personnel perform tasks and ultimately leads to an incident, the incident root cause analyst believes that modifying the procedure and training the personnel who perform the task can prevent similar incidents from happening again, and can terminate the analysis, even if the incident root cause analysis model still provides the next level of cause nodes that lead to "errors in the procedure", and the next level of cause nodes include but are not limited to "lack of procedure development specifications or incomplete procedure development specifications" and "failure to update procedures according to upstream information in a timely manner."
[0084] The event root cause analysis model is equivalent to providing a map of all possible routes from the starting point (the occurrence of the event) to the end point (the root cause of the event), and guiding questions are set at each fork in the road leading to the end point. The event root cause analyst only needs to conduct additional investigations according to the guidance of the questions, and continue to move forward after completing the answers to reach the end point.
[0085] In the present invention, as one of the feasible ways, the event root cause analysis model provides cause nodes of sufficient depth and reaches the cause category of "management effectiveness"; if the enterprise believes that it is not necessary to conduct cause analysis of all failure points to reach the cause category of "management effectiveness", the event root cause analysis model requires the event root cause analyst to provide reasonable and sufficient reasons before the cause node that the enterprise believes is sufficient to avoid the occurrence of similar failure points can be reflected as the root cause or contributing cause of the event.
[0086] In the present invention, as one possible implementation method, step 4, providing a corrective action guide corresponding to the root cause of the incident, includes the following steps:
[0087] The incident root cause analysis model forms a corrective action guide corresponding to the root cause of the incident based on the corrective action data of historical outstanding incidents and good practices in the industry for reference and signature by incident root cause analysts, thereby improving the effectiveness of the corrective actions corresponding to the root cause of the incident.
[0088] As an implementation of the above method, the present invention provides a computer device, which corresponds to the above event root cause analysis method.
[0089] The computer device described in the present invention includes a memory, a processor, and a network interface that are communicatively connected to each other via a system bus. It should be noted that the present invention only illustrates a computer device having a memory, a processor, and a network interface, but it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead. Among them, those skilled in the art will understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to a microprocessor, an application-specific integrated circuit, a programmable gate array, a digital processor, an embedded device, etc.
[0090] The computer device may be a desktop computer, notebook computer, PDA, cloud server, etc. The computer device may interact with the user via a keyboard, mouse, remote control, touchpad, or voice control device.
[0091] The memory includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory, random access memory, static random access memory, read-only memory, electrically erasable programmable read-only memory, programmable read-only memory, magnetic memory, magnetic disk, optical disk, etc. The memory can be an internal storage unit of the computer device, such as the computer device's hard disk or memory, or an external storage device of the computer device, such as a plug-in hard disk, smart memory card, secure digital card, flash memory card, etc. equipped with the computer device. Of course, the memory can also include both the internal storage unit of the computer device and its external storage device. In the present invention, the memory is generally used to store the operating system and various application software installed on the computer device, such as the computer-readable instructions of the aforementioned event root cause analysis method. In addition, the memory can also be used to temporarily store various types of data that have been output or are about to be output.
[0092] In some embodiments, the processor may be a central processing unit, a controller, a microcontroller, a microprocessor, or other data processing chip. The processor is typically used to control the overall operation of the computer device. In the present invention, the processor is used to execute computer-readable instructions stored in the memory or process data, such as computer-readable instructions for executing the aforementioned event root cause analysis method.
[0093] The network interface may include a wireless network interface or a wired network interface, which is generally used to establish a communication connection between the computer device and other electronic devices.
[0094] As an implementation of the above method, the present invention provides a computer-readable storage medium, which corresponds to the above event root cause analysis method.
[0095] The computer-readable storage medium of the present invention stores computer-readable instructions, and the computer-readable instructions can be executed by at least one processor to enable the at least one processor to perform the steps of the above-mentioned event root cause analysis method.
[0096] The present invention will be further described in detail below with reference to specific embodiments.
[0097] This embodiment provides a method for analyzing the root cause of an event, including the following steps:
[0098] Step 1: Guide the incident review and improve the incident process;
[0099] Step 2: Identification of the guide failure point;
[0100] Step 3: Based on the incident root cause analysis model, guide the incident root cause analysis through questions and answers;
[0101] Step 4: Provide corrective action guidance corresponding to the root cause of the incident.
[0102] In this embodiment, the event is an event at a nuclear power plant, and the initial information of the event at the nuclear power plant is:
[0103] On XX / XX / XXXX, the primary circuit pressure of Unit 1 of a nuclear power plant was 14.5 MPa, and the primary circuit temperature was 292°C;
[0104] During a leak inspection inside the containment, operators discovered water on the floor of room 11209. They also discovered water mist in pipe L703, one inch upstream of test valve 1-CVS-V703 in the mixed bed connection pipeline, confirming a leak.
[0105] The operating personnel bypassed the CVS purification unit through the purification loop bypass valve 1-CVS-V062, isolated the CVS resin bed of the chemical volume control system, opened the mixed bed connecting pipeline test valve 1-CVS-V703 to relieve pressure, and stopped the leakage; the main control initiated work order 22037568 to deal with the leakage of the upstream pipeline of the mixed bed connecting pipeline test valve 1-CVS-V703; when the maintenance personnel executed work order 22037568, they found that the pipeline leakage point was 1.5mm long and 0.2mm wide, and there were traces of grinding on the leaking pipeline. It is suspected that the personnel were accidentally injured when cutting nearby pipelines due to inadequate protective measures.
[0106] In this embodiment, step one, guiding the event review and improving the event process, includes the following steps:
[0107] Guide the root cause analysts of the incident to conduct a preliminary investigation into the actual behavior of the personnel and the actual action of the equipment in the incident, combining the enterprise workflow, management requirements and equipment function design related to the incident, and conduct a complete review of the incident to improve the process of the incident.
[0108] Specifically, the root cause analysts of the incident conducted a preliminary investigation of the work orders near the pipeline and confirmed that the pipeline near the pipeline had executed work order 22007083 during the 103 overhaul; work order 22007083 carried out the optimization and transformation of the chemical volume control system CVS pipeline of Unit 1, and the adjacent pipeline valve 70mm away from the damaged pipeline was cut and polished and a new valve was welded; the root cause analysts of the incident used this to conduct a detailed investigation of the implementation process of work order 22007083 and reviewed the detailed process of the entire incident.
[0109] The detailed process of the incident is as follows:
[0110] On XX / XX / XXXX, the preparation engineer completed work order 22007083 according to the design change document, but did not include any reminders regarding accidental damage to equipment or pipelines in the work document package;
[0111] On XX / XX / XXXX, the spare parts required for the change arrived, and work order 22007083 was advanced to the work order application start status;
[0112] On XX / XX / XXXX, the person in charge of the work noticed the nearby pipeline during the on-site confirmation and used fire blankets, three-proof cloth and other protective items, but did not set up physical cutting protection on the pipeline;
[0113] On XX / XX / XXXX, the work group first performed welding work on the newly added pneumatic isolation valve CVS-PL-V069 from the chemical volume control system CVS to the normal waste heat removal system RNS shell, and the isolation valve CVS-PL-V079 from the chemical volume control system CVS to the waste heat removal pump inlet in accordance with the upgraded design documents. During the cutting and grinding of the on-site pipeline, a nearby pipeline was accidentally cut;
[0114] On XX / XX / XXXX, the equipment QC personnel conducted on-site foreign body prevention witnessing before valve welding and found no abnormalities;
[0115] On XX / XX / XXXX, the welding personnel performed valve welding work;
[0116] On XX / XX / XXXX, the on-site pipeline non-destructive inspection was carried out and the inspection was qualified;
[0117] On XX / XX / XXXX, after replacing the valve, the person in charge of the work cleaned and inspected the site and discovered a small amount of damage from the grinding disc on the upstream pipeline of the test valve 1-CVS-V703 of the mixed bed connecting pipeline nearby. However, the inspection results were not reported;
[0118] On XX / XX / XXXX, the operating personnel discovered that the upstream pipeline of the mixed bed connecting pipeline test valve 1-CVS-V703 was leaking due to damage.
[0119] At this point, the process of the event has been completed through the technical solution of step one.
[0120] In this embodiment, step two, guiding the identification of failure points, includes the following steps: guiding the root cause analyst of the event to compare the actual behavior of the personnel and the actual actions of the equipment reflected in the event with the behavior of the personnel and the actions of the equipment required by the enterprise, determining the deviations in personnel behavior and equipment action that contributed to the event, and identifying them as the failure points of the event.
[0121] Specifically, after investigating the nuclear power plant's management procedures, the root cause analysts learned that maintenance work preparation engineers should fully identify risks in the work process when preparing work orders, that work supervisors should identify risks on site and take preventive measures when confirming the on-site work environment, that workers should avoid damaging other equipment during implementation, and that any abnormalities discovered during work should be reported promptly. By comparing these regulations with the actual process of the incident and evaluating whether the corresponding deviations contributed to the consequences of the incident, the failure points of the above incident were determined as follows:
[0122] 1. The preparation engineer completed work order 22007083 according to the design change document, but did not include any reminders about accidental damage to equipment or pipelines in the work document package;
[0123] 2. During on-site confirmation, the person in charge of the work noticed the nearby pipelines and used fire blankets, three-proof cloth and other protective items, but did not set up physical cutting protection on the pipelines;
[0124] 3. The working group first carried out welding work on the newly added pneumatic isolation valve CVS-PL-V069 from the chemical volume control system CVS to the normal waste heat removal system RNS shell, and the isolation valve CVS-PL-V079 from the chemical volume control system CVS to the waste heat removal pump inlet. When cutting and grinding the on-site pipeline, a nearby pipeline was accidentally cut;
[0125] 4. After the work leader replaced the valve, he cleaned and inspected the site and found that the upstream pipeline of the nearby mixed bed connecting pipeline test valve 1-CVS-V703 had a small amount of damage caused by the grinding disc, but did not report the inspection results.
[0126] In this embodiment, step three, based on the event root cause analysis model, guides the event root cause analysis through questions and answers, including the following steps:
[0127] The event root cause analyst determines and classifies the failure point based on the guiding questions corresponding to the failure point type, and then conducts a supplementary investigation on the failure point according to the guidance of the questions set in the event root cause analysis model. After the event root cause analyst obtains the factual evidence that proves the corresponding cause node, he or she continues to conduct further supplementary investigation guided by the questions in the next level cause node in the event root cause analysis model until the cause node with sufficient depth is analyzed;
[0128] The event root cause analysis model provides cause nodes of sufficient depth and up to the cause category of "management effectiveness". If the enterprise believes that it is not necessary to conduct such a deep cause analysis for all failure points, the event root cause analysis model requires the event root cause analyst to provide reasonable and sufficient reasons before the cause node that the enterprise believes is sufficient to avoid the occurrence of similar failure points can be reflected as the root cause or contributing cause of the incident.
[0129] Taking the third failure point in the above incident as an example, "the working group first performed the welding work of the newly added pneumatic isolation valve CVS-PL-V069 from the chemical volume control system CVS to the normal waste heat removal system RNS shell and the isolation valve CVS-PL-V079 from the chemical volume control system CVS to the waste heat removal pump inlet according to the upgraded design documents. When cutting and grinding the on-site pipeline, a nearby pipeline was accidentally cut." The root cause analysis model of the incident determined that the failure point type was "personnel performance" through question and answer guidance. The possible cause categories of this type of failure point include "human-machine interface", "procedures", "communication" and "working environment". The corresponding guiding questions are:
[0130] Human-machine interface: Did the staff fail to use or use the wrong tools, maintenance aids, and personal protective equipment, or use tools, maintenance aids, and personal protective equipment that did not meet the work requirements, causing the problem? --- According to the survey, answer "No"
[0131] Procedures: Are procedures incomplete, inaccurate, inapplicable, or not used? --- "No" based on survey answer
[0132] Communication: Is communication effectiveness insufficient? --- According to the survey, answer "No"
[0133] Work Environment: Are the working environment conditions not what the workers expected or uncomfortable? --- Answer "yes" according to the survey.
[0134] After interviewing the staff involved and inspecting the site, it was determined that the operation had nothing to do with the human-machine interface or communication, and that the correct pipes were cut according to the correct procedures. However, due to the small working space, the angle grinder accidentally touched other surrounding pipes during the cutting process, causing damage. Therefore, the answer to the guiding question about the "working environment" was "yes", ruling out the possibility of other types of causes.
[0135] After the incident root cause analyst selects "Work Environment," the incident root cause analysis model provides the next level of cause nodes, including "Vibration / Shaking," "Noise," "Ambient Temperature," and "Buildings / Structures and Workspaces." The guiding questions are:
[0136] 1. Are the steps too high, causing workers to slip or trip? --- According to the survey, the answer is "No"
[0137] 2. Have any buildings or structures caused workers to hit or trip over, resulting in falls or injuries? For example, are there pipes or cable trays crossing the walkway at ankle height? --- "No" according to the survey.
[0138] 3. Is the workspace too cramped, causing the staff member to be unable to complete the task, requiring replacement and ultimately delaying the task? --- According to the survey, answer "No"
[0139] 4. Is the workspace too narrow, making it difficult for workers to avoid accidentally touching other equipment while performing their tasks, causing problems? --- According to the survey, answer "yes"
[0140] 5. Is there insufficient workspace, requiring personnel to work on ladders or platforms or to adopt unusual body positions, causing physical discomfort and leading to delays or errors in the performance of tasks? --- Answer "No" based on the survey
[0141] Based on the guiding questions above, the root cause analyst answers "yes" to question 4, thus determining the relevance of the cause node "Buildings / Structures and Workspaces." After the root cause analyst determines "Buildings / Structures and Workspaces," the root cause analysis model provides the following next-level cause node and its guiding questions:
[0142] Work Process Management: 1. Were adverse factors in the work environment identified by the operating unit, but not identified or considered during the work process management phase of the task, leading to problems? For example, were the methods and procedures for operating in a confined workspace not clearly defined during pre-job meetings? Were there no staff rotations or rest periods? --- "Yes" based on the survey answer.
[0143] Inadequate human factors engineering design: 1. Are there deficiencies in human factors engineering design that result in the failure to consider the characteristics of the operating environment during the installation of new equipment or systems, leading to adverse factors that could have been eliminated? --- According to the survey, the answer is "No";
[0144] Deficiencies in identifying risks in the work environment: 1. Did the company fail to assess the relevant work environment, or did this result in deficiencies, leading to unidentified or inadequate identification of hazardous factors in the work environment? --- Based on the survey, answer "No." 2. Did the company fail to take necessary measures to minimize the risks associated with the hazardous factors in the work environment? For example, failing to wrap brackets protruding from the wall with anti-collision pads? --- Based on the survey, answer "No." 3. Did the company fail to promptly identify and implement effective measures for changes in hazardous factors in the work environment? For example, technological upgrades to equipment resulted in a narrower workspace or protruding building components, but the company failed to implement anti-collision measures? --- Based on the survey, answer "No."
[0145] The root cause analysts of the incident conducted additional investigations according to the above cause nodes and guiding questions, and ruled out the two cause nodes of "inadequate human factors engineering design" and "inadequate identification of operating environment risks", and confirmed the cause node of "work process management". In the root cause analysis model of the incident, "work process management" is an independent cause category following the cause node of "buildings / structures and workspaces". In the subsequent levels of this cause category, multiple levels of cause nodes are still set up. The root cause analysts of the incident can conduct additional investigations and cause analysis layer by layer in the form of the above question and answer guidance until they analyze the cause node that the nuclear power plant believes is sufficient to avoid the recurrence of similar failure points, and reflect it as the root cause or contributing cause of the incident.
[0146] Since the problem guidance process and the technical solutions for supplementary investigation / cause analysis in the event root cause analysis model are repetitive implementations, they will not be repeated here. The root cause of the event analyzed in the above embodiment is "the risk of accidentally cutting other pipelines due to the narrow working space not being fully identified during work preparation." The logical relationship of the cause nodes corresponding to the analysis results is: working environment > buildings / structures and workspace > inadequate work process management > work preparation > inadequate preparation of work file packages > insufficient work risk analysis.
[0147] In the actual incident report corresponding to this embodiment, the root cause analyst identified the narrow workspace as the root cause of the incident, without further analyzing the management failures in the work preparation process. Furthermore, the original report compared the power plant's personnel behavioral standards with the staff's mistaken cutting of other pipelines, arguing that poor personnel behavioral standards were the root cause of the incident and instituting corrective actions related to training and learning. The aforementioned cause analysis process suffered from issues such as a lack of evidence, an imperfect logical deduction process, and insufficient depth in the cause analysis. As can be seen from the above embodiment, the root cause analysis model can provide root cause analysts with sufficiently standardized cause analysis guidance and a sufficiently complete logical chain, thereby improving the quality of root cause analysis of incidents.
[0148] In this embodiment, step four, providing a corrective action guide corresponding to the root cause of the incident, includes the following steps: the incident root cause analysis model forms a corrective action guide corresponding to the root cause of the incident based on the corrective action data in historical outstanding incidents and good practices in the industry for reference and signature by incident root cause analysts, thereby improving the effectiveness of the corrective actions corresponding to the root cause of the incident.
[0149] In this example, the root cause analysis results of step 3 are "Work environment > Buildings / Structures and Workspaces > Inadequate Work Process Management > Work Preparation > Inadequate Work Document Package Preparation > Inadequate Work Risk Analysis." The root cause analysis model provides the following corrective action guidelines for the root cause analysis results of step 3:
[0150] 1. If the work involved in this incident is not frequently performed and the person preparing the work lacks sufficient time and energy to fully prepare, you may consider developing a common work risk analysis template / checklist for frequently performed work and reviewing these templates / checklists to ensure their accuracy and completeness. This will allow the person preparing the work more time to fully identify the operational risks of infrequently performed work.
[0151] 2. If the person preparing the work for this incident used a general work risk analysis template / checklist, you may consider adding / modifying the risk types reflected in this incident to the template / checklist to prompt the person preparing the work to fully identify them;
[0152] 3. If the experience feedback information required by the job preparer for this incident is insufficient, you may consider including this incident in the company's experience feedback data and consider linking this experience feedback information with the work order system to provide effective support information when the job preparer prepares similar jobs in the future;
[0153] 4........
[0154] In previous similar incidents involving insufficient risk analysis during the work preparation phase, some root cause analysts often used corrective actions such as enhanced training and learning to increase the knowledge and experience of those preparing for the work to prevent similar incidents from occurring. However, this has been ineffective. In this example, root cause analysts simulated and developed corrective actions based on the corrective action guidelines for the corresponding cause nodes. After expert discussion and evaluation, the following corrective actions were found to be effective in preventing the recurrence of similar incidents:
[0155] 1. Add the risk category of accidental damage to pipes and equipment within xx mm during cutting and grinding operations to the power plant's existing standard risk analysis checklist, set physical protection requirements for equipment and pipes that are relatively close, and set self-inspection and external inspection requirements after the work is completed;
[0156] 2. Conduct publicity and training for those who prepare work. When preparing work orders for cutting and grinding, a risk analysis of accidental damage to equipment should be conducted and the experience feedback database should be entered to facilitate subsequent work order preparation and retrieval.
[0157] The above-described embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
Claims
1. A method for analyzing the root cause of an incident, characterized in that: The following steps are involved: Step 1: Guide the incident review and improve the incident process; Step 2: Identification of the guide failure point; Step 3: Based on the incident root cause analysis model, guide the incident root cause analysis through questions and answers; Step 4: Provide corrective action guidance corresponding to the root cause of the incident.
2. The event root cause analysis method according to claim 1, characterized in that: Step 1: Guide the incident review and improve the incident process, including the following steps: Guide the root cause analysts of the incident to conduct a preliminary investigation into the actual behavior of the personnel and the actual action of the equipment in the incident, combining the enterprise workflow, management requirements and equipment function design related to the incident, and conduct a complete review of the incident to improve the process of the incident.
3. The event root cause analysis method according to claim 1, characterized in that: Step 2, guiding the identification of failure points, includes the following steps: Guide the root cause analysts of the incident to compare the actual behavior of the personnel and the actual actions of the equipment reflected in the incident with the behavior of the personnel and the actions of the equipment required by the enterprise, determine the deviations in personnel behavior and equipment action that contribute to the incident, and identify them as the failure points of the incident.
4. The event root cause analysis method according to claim 3, characterized in that: Personnel behavior deviations include incorrectly performing actions that should be performed and not performing actions that should be performed; equipment action deviations include not performing actions according to designed functions and refusing to perform actions according to designed functions.
5. The event root cause analysis method according to claim 1, characterized in that: The incident root cause analysis model is constructed using the following methods: By summarizing and generalizing the main problems and historical data of root cause analysis of incidents in existing incident reports in the nuclear power industry, investigating and understanding and referring to advanced root cause analysis practices at home and abroad, integrating the root cause analysis work process management and equipment life cycle management processes of the domestic nuclear power industry and other advanced industries, and after multiple expert desktop talks, we constructed a root cause analysis model and its corresponding root cause database.
6. The event root cause analysis method according to claim 1, characterized in that: In the incident root cause analysis model, multiple failure point types are set; under each failure point type, all possible cause categories that may cause the occurrence of that type of failure point are set; each cause category is set up with a complete cause node map based on logical relationships; the cause node map logically expands the cause nodes layer by layer, with the lowest cause node connected to other cause categories until it reaches the cause category of "management effectiveness"; each failure point type, cause category, and cause node is set with guiding questions to guide the incident root cause analyst in the incident root cause analysis.
7. The event root cause analysis method according to claim 1, characterized in that: Step 3: Based on the incident root cause analysis model, guide the incident root cause analysis through questions and answers, including the following steps: Step 301: The event root cause analysis model provides guiding questions corresponding to the failure point type; the event root cause analyst identifies the failure point type by answering the guiding questions corresponding to the failure point type one by one; When the event root cause analyst answers "yes" to the guiding question corresponding to the failure point type, the event root cause analysis model continues to provide guiding questions corresponding to the categories of causes that may have caused the failure point of this type. The event root cause analyst answers the guiding questions corresponding to the categories of causes that may have caused the failure point one by one, and then the process proceeds to step 302. When the event root cause analyst answers "no" to the guidance question corresponding to the failure point type, the event root cause analysis model stops providing guidance questions corresponding to the cause category that may have caused the occurrence of this type of failure point; The guiding question corresponding to the cause category that may cause the occurrence of this type of failure point is whether this cause category is the cause of the occurrence of this type of failure point; Step 302: When the event root cause analyst answers "yes" to the guiding question corresponding to the cause category that may cause the occurrence of this type of failure point, the event root cause analysis model continues to provide guiding questions corresponding to the cause nodes under this cause category. The event root cause analyst answers the guiding questions corresponding to the cause nodes under this cause category one by one to identify whether the cause node is the cause of the occurrence of this type of failure point, and then proceeds to step 303. If the event root cause analyst answers "no" to the guidance question corresponding to the cause category that may cause the occurrence of this type of failure point, the event root cause analysis model stops providing guidance questions corresponding to the cause node under this cause category; The guiding question corresponding to the cause node under this cause category is whether the cause node is the cause of the occurrence of this type of failure point; Step 303: When the event root cause analyst answers "yes" to the guiding question corresponding to the cause node under the cause category, the event root cause analysis model continues to provide guiding questions corresponding to the cause node under the cause node; If the event root cause analyst answers "no" to the guidance question corresponding to the cause node under the cause category, the event root cause analysis model stops providing guidance questions corresponding to the cause nodes under the cause node; When the event root cause analyst answers "yes" to the guiding question corresponding to the lowest-level cause node under the cause category, the event root cause analysis model continues to provide guiding questions corresponding to the cause categories subsequent to the cause node, and the process returns to step 302; If the event root cause analyst answers "no" to the guidance question corresponding to the lowest-level cause node under the cause category, the event root cause analysis model stops providing guidance questions corresponding to the cause categories subsequent to the cause node; If the enterprise believes that it is not necessary to analyze the causes of all failure points to reach the cause category of "management effectiveness", the incident root cause analysis model requires the incident root cause analyst to provide corresponding reasons before the cause node that the enterprise believes is sufficient to prevent similar failure points from occurring can be reflected as the root cause or contributing cause of the incident; The event root cause analysis model takes the lowest-level cause nodes under each cause category that causes each failure point as the root cause of the event and sorts them by correlation.
8. The event root cause analysis method according to claim 1, characterized in that: Step 4: Provide corrective action guidance corresponding to the root cause of the incident, including the following steps: The incident root cause analysis model forms a corrective action guide corresponding to the root cause of the incident based on the corrective action data of historical outstanding incidents and good practices in the industry for reference and signature by incident root cause analysts.
9. A computer device comprising a memory and a processor, wherein the memory stores computer-readable instructions, wherein: When the processor executes the computer-readable instructions, the steps of the event root cause analysis method according to any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium having computer-readable instructions stored thereon, characterized in that: When the computer-readable instructions are executed, the steps of the event root cause analysis method according to any one of claims 1 to 8 are implemented.