Root cause analysis method and device, computer equipment, storage medium and program product
By obtaining event data of change events associated with fault events in the information technology system, determining the correlation level and conducting root cause analysis, the problem of low accuracy of root cause analysis in the prior art is solved, and more efficient fault diagnosis and system stability are achieved.
Patent Information
- Application Number
- CN202510387102.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-01
AI Technical Summary
The existing root cause analysis methods have poor accuracy in information technology systems, especially when service changes frequently, it is difficult to accurately identify the root cause of the failure.
By obtaining event data for the change event associated with the fault event, determining the level of association between the change event and the fault event, and performing root cause analysis based on the target event, diagnosing whether the change event is the root cause of the fault.
Improve the accuracy of root cause analysis, help quickly restore the normal operation of the system, improve system stability, and optimize subsequent change management.
Smart Images

Figure CN120234176A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical fields of data processing technology and text generation technology, and particularly relates to a root cause analysis method, apparatus, computer device, storage medium, and program product. Background Art
[0002] In an information technology system that provides various services to meet specific user needs, an integrated technology system is incorporated, which includes several interrelated elements, such as computer hardware (servers, terminal devices, etc.), software (operating systems, application programs, etc.), network communication devices (routers, switches, etc.), data, and relevant personnel, etc. Its core lies in using information technology to achieve functions such as information collection, storage, processing, transmission, and application, so as to support the business operations and decision-making management of an organization or enterprise. When a system fails, it is often necessary to automatically attribute and identify the root cause of the system failure. For example, by setting thresholds, pattern recognition, or machine learning algorithms, etc., to detect abnormal changes in data and timely discover problems such as failures or performance degradation in the service platform. For example, when the processor usage rate of a server suddenly rises above a certain threshold and lasts for a period of time, it may be determined that a failure has occurred.
[0003] However, in related root cause analysis solutions, the system failures or abnormal situations caused by frequent service changes are often ignored. For some complex failures involving multiple system components and intertwined multiple technical factors, automatic root cause analysis may require more powerful algorithms and more complex models to accurately find the root cause, otherwise inaccurate conclusions may be drawn, resulting in poor accuracy of root cause analysis. Summary of the Invention
[0004] In view of this, the present disclosure provides a root cause analysis method, apparatus, computer device, storage medium, and program product to solve the problem of poor accuracy of root cause analysis in an information technology system.
[0005] In a first aspect, the present disclosure provides a root cause analysis method, and the method includes:
[0006] When a failure event is detected, obtaining a change event associated with the failure event in the information technology system, where the change event is used to indicate an event generated by updating a service in the information technology system;
[0007] Obtaining event data of the change event, where the event data is used to indicate record data generated by a specific operation or state change in the information technology system corresponding to the change event;
[0008] Based on the event data, determining the association level between the change event and the failure event;
[0009] Determine a target event in a change event according to the association level, and perform root cause analysis on the fault event based on the target event to obtain an analysis result.
[0010] In a second aspect, the present disclosure provides a root cause analysis device, which includes:
[0011] A detection module, configured to obtain a change event associated with a fault event in an information technology system when detecting the fault event, where the change event is used to indicate an event generated by updating a service in the information technology system;
[0012] An acquisition module, configured to acquire event data of the change event, where the event data is used to indicate record data generated by a specific operation or state change in the information technology system corresponding to the change event;
[0013] A determination module, configured to determine the association level between the change event and the fault event based on the event data;
[0014] An analysis module, configured to determine a target event in the change event according to the association level, and perform root cause analysis on the fault event based on the target event to obtain an analysis result.
[0015] In a third aspect, the present disclosure provides a computer device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the root cause analysis method according to the first aspect or any corresponding embodiment thereof.
[0016] In a fourth aspect, the present disclosure provides a computer-readable storage medium, on which computer instructions are stored. The computer instructions are used to cause a computer to execute the root cause analysis method according to the first aspect or any corresponding embodiment thereof.
[0017] In a fifth aspect, the present invention provides a computer program product, including computer instructions, which are used to cause a computer to execute the root cause analysis method according to the first aspect or any corresponding embodiment thereof.
[0018] In an embodiment of the present disclosure, in response to a detected fault event, a change event associated with the fault event may be obtained, where the change event is used to indicate an event generated by updating a service in an information technology system. Then, event data of the change event may be obtained, where the event data is used to indicate recorded data generated by a specific operation or a state change corresponding to the change event. Next, based on the event data, the association level between the change event and the fault event may be determined, and a target event may be determined from the change events according to the association level, so as to perform root cause analysis on the fault event based on the target event to obtain an analysis result. Thus, in the process of root cause analysis, change events that may sometimes cause system failures or abnormal conditions are analyzed to diagnose whether the change event is the root cause of the fault, improving the accuracy of root cause analysis, which is crucial for quickly restoring the normal operation of the system, enhancing system stability, and optimizing subsequent change management, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0020] Figure 1 is a schematic flowchart of a root cause analysis method according to an embodiment of the present disclosure;
[0021] Figure 2 is a schematic flowchart of another root cause analysis method according to an embodiment of the present disclosure;
[0022] Figure 3 is a schematic flowchart of yet another root cause analysis method according to an embodiment of the present disclosure;
[0023] Figure 4 is a schematic diagram of a topology structure display interface according to an embodiment of the present disclosure;
[0024] Figure 5 is a schematic diagram of the display of a target event according to an embodiment of the present disclosure;
[0025] Figure 6 is a schematic flowchart of still another root cause analysis method according to an embodiment of the present disclosure;
[0026] Figure 7 is an architecture diagram of a root cause analysis device according to an embodiment of the present disclosure;
[0027] Figure 8 is a schematic hardware structure diagram of a computer device according to an embodiment of the present disclosure. Specific Embodiments
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.
[0029] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure shall be informed to the user and the user's authorization shall be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0030] For example, when responding to receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application program, server, or storage medium that performs the operations of the technical solutions of the present disclosure based on the prompt message.
[0031] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0032] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0033] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of the corresponding laws, regulations, and related provisions.
[0034] Combined with the application scenarios on which the execution of the root cause analysis method depends, the application scenarios are described herein.
[0035] In an information technology system that provides various services to meet specific user needs, an integrated technical system is incorporated, which includes several interrelated elements, such as computer hardware (servers, terminal devices, etc.), software (operating systems, application programs, etc.), network communication devices (routers, switches, etc.), data, and relevant personnel, etc. Its core lies in using information technology to achieve functions such as information collection, storage, processing, transmission, and application, so as to support the business operations and decision-making management of an organization or enterprise. When a system fails, it is often necessary to automatically attribute and identify the root cause of the system failure. For example, by setting thresholds, pattern recognition, or machine learning algorithms, etc., to detect abnormal changes in data and timely discover problems such as failures or performance degradation in the information technology system. For example, when the processor usage rate of a server suddenly rises above a certain threshold and lasts for a period of time, it may be judged that a failure has occurred.
[0036] However, in related root cause analysis solutions, the system failures or abnormal situations caused by frequent service changes are often ignored. For some complex failures involving multiple system components and intertwined technical factors, automatic root cause analysis may require more powerful algorithms and more complex models to accurately find the root cause. Otherwise, inaccurate conclusions may be drawn, resulting in poor accuracy of root cause analysis.
[0037] Based on this, the embodiments of the present disclosure provide a root cause analysis method, which can, in response to a detected fault event, obtain a change event associated with the fault event, where the change event is used to indicate an event generated by updating a service. Then, it can obtain the event data of the change event, where the event data is used to indicate the recorded data generated by a specific operation or state change corresponding to the change event. Next, based on the event data, it can determine the association level between the change event and the fault event, and determine a target event in the change event according to the association level, so as to perform root cause analysis on the fault event based on the target event to obtain an analysis result. Thus, in the process of root cause analysis, the change events that may sometimes cause system failures or abnormal situations are analyzed to diagnose whether the change event is the root cause of the failure, improving the accuracy of root cause analysis, which is crucial for quickly restoring the normal operation of the system, enhancing system stability, and optimizing subsequent change management, etc.
[0038] According to the embodiments of the present disclosure, an embodiment of a root cause analysis method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0039] According to an embodiment of the present disclosure, an embodiment of a root cause analysis method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0040] In this embodiment, a root cause analysis method is provided, which can be used in the above-mentioned information technology systems, such as supply management systems, online information technology systems, etc. Figure 1 It is a flowchart of the root cause analysis method according to an embodiment of the present disclosure, as Figure 1 shown, the process includes the following steps:
[0041] Step S101, in response to a detected fault event, obtain a change event associated with the fault event, where the change event is used to indicate an event generated by updating a service in the information technology system.
[0042] In the embodiment of the present disclosure, the fault event can be a software and hardware fault detected in the information system, such as increased latency, network disconnection, file corruption, etc. In addition, according to the dependency relationship between services in the information technology system, a change event associated with the fault event can be determined.
[0043] Specifically, the dependency relationship can include data dependency, function dependency, performance dependency, etc. After determining the fault event, a change event associated with the fault event can be searched based on the service corresponding to the fault event. For example, if the service corresponding to the fault event is Service 1, the change event 1 of Service 2 depends on the change event 2 of Service 3, the change event 2 of Service 3 depends on the change event 3 corresponding to Service 1, and Service 1 causes the fault event "increased latency". Then the change events corresponding to the fault event include change event 1 - change event 3.
[0044] Step S102, obtain the event data of the change event, where the event data is used to indicate the recorded data generated by a specific operation or state change corresponding to the change event.
[0045] In the embodiment of the present disclosure, the event data of the change event can include service-related data of the service corresponding to the change event and event-related data corresponding to the change event. For example, the service-related data can include service performance indicators, dependency relationships between services, root cause nodes for characterizing fault events, etc. The event-related data can include event records of the occurrence of the change event, such as event types (such as configuration changes, code deployments, etc.), risk warnings during the process of releasing the change event (such as the number of rollbacks, the number of warnings, etc.), the change time sequence of the change event (such as change frequency, change interval, etc.), metric data (such as service response time, error rate, etc.).
[0046] Step S103: Determine the association level between the change event and the fault event based on the event data.
[0047] In an embodiment of the present disclosure, data features of the event data can be extracted based on a preset network model, and compared with the data features of the fault event to obtain the association degree therebetween, and the association level can be determined according to the association degree.
[0048] For example, the association degree can be a percentage value, and the association level can be a level preset for each association degree range. For example, an association degree of 0%-20% corresponds to the first level, an association degree of 20%-40% corresponds to the second level, etc. It should be understood that the specific implementation manner of determining the association level is subject to being able to be implemented, and the present disclosure does not make specific limitations thereon.
[0049] Step S104: Determine the target event in the change event according to the association level, and perform root cause analysis on the fault event based on the target event to obtain an analysis result.
[0050] In an embodiment of the present disclosure, the target event can be a change event whose association level meets a preset association condition, and the target event that meets the preset association condition should be a change event that has a strong association with the occurrence of the fault event.
[0051] For example, if the preset association condition is the target level, the change event with the association level being the target level can be determined as the target event. For example, the association level includes the first to fifth levels, where the higher the level, the higher the association degree between the change event and the fault event. Then, the second level can be determined as the preset association condition, that is, the change event of the fifth level is determined as the target event.
[0052] For another example, if the association level includes the above-mentioned association degree, the preset association condition can be used to indicate taking the k update times with the highest association degree as the target events.
[0053] After determining the target event, analysis can be performed based on the target event through a root cause analysis model to obtain the fault cause that leads to the fault event, and the information of these target events can be highlighted to facilitate the operation and maintenance personnel to quickly understand and pay attention to the target event.
[0054] As described above, in the embodiments of the present disclosure, in response to a detected fault event, a change event associated with the fault event can be obtained, where the change event is used to indicate an event generated by updating a service. Then, the event data of the change event can be obtained, where the event data is used to indicate the recorded data generated by a specific operation or a state change corresponding to the change event. Next, based on the event data, the association level between the change event and the fault event can be determined, and a target event can be determined from the change events according to the association level, so as to perform root cause analysis on the fault event based on the target event to obtain an analysis result. Thus, during the root cause analysis process, change events that may sometimes cause system failures or abnormal conditions can be analyzed to diagnose whether the change event is the root cause of the fault, improving the accuracy of root cause analysis, which is crucial for quickly restoring the normal operation of the system, enhancing system stability, and optimizing subsequent change management, etc.
[0055] In this embodiment, another root cause analysis method is provided, which can be used in the above-mentioned information technology systems, such as a supply management system, an online information technology system, etc. Figure 2 It is a flowchart of the root cause analysis method according to the embodiments of the present disclosure, as Figure 2 shown, and the process includes the following steps:
[0056] Step S201, in response to a detected fault event, obtain a change event associated with the fault event, where the change event is used to indicate an event generated by updating a service. For details, please refer to Figure 1 Step S101 of the embodiment shown, which will not be elaborated here.
[0057] Step S202, obtain the event data of the change event, where the event data is used to indicate the recorded data generated by a specific operation or a state change corresponding to the change event.
[0058] Specifically, the above step S202 includes:
[0059] Step S2021, determine the target service corresponding to the change event.
[0060] Step S2022, obtain the detection data of the preset performance indicators of the target service.
[0061] In the embodiments of the present disclosure, the preset performance indicators of the target service can be the key performance indicators of the service, and can include resource utilization rate, key performance indicator latency, QPS (Queries Per Second), error rate, processor utilization rate, memory occupancy rate, etc. The detection data can be the detection data obtained by monitoring these preset performance indicators.
[0062] Step S2023, obtain the dependency relationship between the change event and the fault event, where the dependency relationship is used to indicate the causal relationship between the change event and the fault event.
[0063] In the embodiments of the present disclosure, the dependency relationship can be determined based on the dependency relationship between the target service and the service corresponding to the change event. Here, the dependency relationship between services can include: data dependency, functional dependency, performance dependency, etc.
[0064] For example, data dependency can include input-output dependency, data sharing dependency, etc. Among them, input-output dependency is used to indicate that the output of one service may be the input of another service, and data sharing dependency is used to indicate that the data processing results of some services will affect the operation logic of other services. Functional dependency can include business process dependency, interface call dependency, etc. Among them, business process dependency can be used to represent the dependency between the functions of each service in the job process, such as the approval process, and interface call dependency can be used to indicate the interaction and function call between services through application programming interfaces. Performance dependency can include response time dependency, load capacity dependency, etc. For example, in a distributed system, the microservice architecture enables each service to be deployed on different servers. Response time dependency is used to indicate that if the response time of a certain basic service (such as an authentication service) is slow, then other business services (such as user information query service) that depend on it will also be affected. Load capacity dependency is used to indicate the dependency relationship of load capacity existing between services such as search service and cache service.
[0065] Step S2024, obtain the time record corresponding to the fault event.
[0066] Step S2025, obtain the release record corresponding to the release of the change event. The event data includes detection data, dependency relationship, time record, and release record.
[0067] In the embodiments of the present disclosure, the time record corresponding to the fault time can be used to indicate the occurrence time of the fault event and the change event associated with the fault event. The release record can include the event type of the change event (such as configuration change, code deployment, etc.), risk warnings during the process of releasing the change event (such as the number of rollbacks, the number of warnings, etc.), the change timing of the change event (such as change frequency, change interval, etc.), metric data (such as the response time, error rate, etc. of the service), etc. It should be understood that the release record is used to determine the weight of the change event when calculating the correlation degree between the change event and the fault event.
[0068] Step S203, based on the event data, determine the association level between the change event and the fault event. For details, please refer to Figure 1 Step S103 of the illustrated embodiment, which will not be elaborated here.
[0069] Step S204, determine a target event in the change event according to the association level, and perform root cause analysis on the fault event based on the target event to obtain an analysis result. For details, please refer to Figure 1 step S104 of the embodiment shown, which will not be elaborated here.
[0070] In the embodiment of the present disclosure, various dimensions of event data corresponding to the change event can be obtained. For example, detection data, dependency relationships, time records, and release records, so that the calculation of the association degree from multiple dimensions realizes a more comprehensive and accurate diagnosis of the change event.
[0071] In this embodiment, another root cause analysis method is provided, which can be used for the above-mentioned information technology systems, such as supply management systems, online information technology systems, etc. Figure 3 is a flowchart of the root cause analysis method according to the embodiment of the present disclosure, as Figure 3 shown, the process includes the following steps:
[0072] Step S301, in response to the detected fault event, obtain a change event associated with the fault event, where the change event is used to indicate an event generated by updating a service in the information technology system. For details, please refer to Figure 1 step S101 of the embodiment shown, which will not be elaborated here.
[0073] Step S302, obtain the event data of the change event, where the event data is used to indicate the record data generated by a specific operation or state change corresponding to the change event. For details, please refer to Figure 1 step S102 of the embodiment shown, which will not be elaborated here.
[0074] Step S303, based on the event data, determine the association level between the change event and the fault event.
[0075] Specifically, the above step S303 includes:
[0076] Step S3031, obtain the preset scoring dimensions corresponding to the change event.
[0077] Step S3032, score the change event based on the event data corresponding to the preset scoring dimensions to obtain the association degree between the change event and the fault event under the preset scoring dimensions.
[0078] Step S3033, determine the association level based on the association degrees corresponding to each preset scoring dimension.
[0079] In the embodiment of the present disclosure, the preset scoring dimension can be a scoring dimension preset for the event data of each dimension, and the preset scoring dimension can score the association degree between the event data and the fault event to obtain the association degree.
[0080] Specifically, if the dimensions of the event data include: the detection data corresponding to the change event, the dependency between the change event and the fault event, and the time record of the change event, step S3032, scoring the change event based on the event data corresponding to the preset scoring dimension to obtain the correlation degree between the change event and the fault event under the preset scoring dimension, including:
[0081] Step a1, obtain the detection data of the preset performance indicators corresponding to the preset scoring dimension, and determine the first sub-correlation degree of the change event based on the proportion of abnormal data in the detection data.
[0082] In the embodiments of the present disclosure, the first sub-correlation degree is used to indicate the correlation degree between the service corresponding to the change event and the fault event. Here, the preset scoring dimension may be the characteristic of the proportion of abnormal indicators in the total indicators, and this characteristic of the proportion of abnormal indicators in the total indicators can be used to indicate the proportion of the abnormal data corresponding to the abnormal indicators in the detection data corresponding to all indicators. The specific calculation method can be expressed as \(\text{anomaly KPI Score}=\frac{\text{Number ofabnormal KPIs}}{\text{Total number of KPIs}}\). It should be understood that a service with more abnormal instances will obtain a higher score, which means that the service is more likely to be involved in defective changes and has a stronger correlation with the fault event.
[0083] In addition, the preset scoring dimension can also be used to indicate the characteristic of the proportion of abnormal instances in the total instances in the service corresponding to the change event. Similarly, this characteristic of the proportion of abnormal instances in the total instances can be used to indicate the proportion of abnormal instances in the total instances in the service. The specific calculation method can be expressed as \(\text{anomaly KPI Score}=\frac{\text{Numberof abnormal instances}}{\text{Total number of service instances}}\). It should be understood that a service with more abnormal instances will obtain a higher score, which means that the service is more likely to be involved in defective changes and has a stronger correlation with the fault event.
[0084] Step a2, determine the distance value between the change event and the fault event based on the dependency corresponding to the preset scoring dimension, and determine the second sub-correlation degree corresponding to the change event based on the distance value.
[0085] In the embodiments of the present disclosure, the second sub - correlation degree can be used to indicate the relevance between the service change corresponding to the change event and its dependent services. Given that the dependent services are more likely to be defective services, the root cause of potential faults can be deeply explored through the scoring of the second sub - correlation degree.
[0086] Specifically, when calculating the second sub - correlation degree, the preset scoring dimension can be used to indicate the distance feature between the service of the change event and the root cause node in the service topology corresponding to the dependency relationship, where the root cause node is the fault cause node. Specifically, when calculating the above - mentioned second sub - correlation degree, the dependency graph corresponding to the dependency relationship can be obtained, and according to the position of the service corresponding to the change event in the dependency graph, the above - mentioned distance value can be determined, and the dependency score can be calculated based on this distance value, so as to determine the second sub - correlation degree based on this dependency score.
[0087] Here, the transitive closure of the dependency graph can be used to calculate the dependency score of the service, and this closure can effectively capture all direct and indirect dependency relationships existing between services. The specific calculation method can be expressed as: \(\text{DistanceScore}=\frac{\text{Distance Score}_{\text{max}}}{\text{Distance Score}_{\text{max}}+\text{Tier}(S,S_{\text{RCA}})}\), where the services closer to the target service corresponding to the fault event will be assigned higher scores, while the services farther away will get relatively lower scores.
[0088] Step a3: Determine the occurrence time of the change event based on the time record, and determine the third sub - correlation degree corresponding to the change event based on the occurrence time.
[0089] In the embodiments of the present disclosure, when calculating the second sub - correlation degree, the preset scoring dimension can be used to indicate the distance feature between the key node time of the change event and the fault time, so that the third sub - correlation degree can fully consider the time factor of the change. By analyzing the occurrence time of the change and the fault, the potential impact of the change on the fault can be effectively evaluated, and the relevance between the change and the fault can be measured from the time dimension.
[0090] Specifically, the calculation method of the third sub - correlation degree can be expressed as: \(\text{Time Score}=\frac{1}{1 + \text{Time Difference}}\), where the time difference refers to the time interval from the implementation of the change to the occurrence of the fault. The smaller the time difference, the higher the time score, indicating a greater relevance between the change and the fault.
[0091] Step a4: Determine the correlation degree according to the first sub - correlation degree, the second sub - correlation degree and the third sub - correlation degree.
[0092] In the embodiments of the present disclosure, the first sub - correlation degree, the second sub - correlation degree, and the third sub - correlation degree can be weighted and summed to obtain the correlation score corresponding to the change event, and the correlation level can be determined according to the correlation score. The specific process is as follows:
[0093] (1) Obtain the change weights corresponding to the first sub - correlation degree, the second sub - correlation degree, and the third sub - correlation degree.
[0094] (2) Based on the change weights, perform weighted summation on the first sub - correlation degree, the second sub - correlation degree, and the third sub - correlation degree to obtain the correlation score, and determine the correlation level according to the correlation score.
[0095] In the embodiments of the present disclosure, through a preset comprehensive evaluation algorithm (such as weighted average, etc.), the evaluation results of the multi - dimensional correlation degrees can be integrated to obtain the comprehensive impact degree score (score) of the change event on RCA (Root Cause Analysis), and based on this score, it can be determined whether the change event is the root cause of the failure event and the level of its impact, providing strong support for subsequent operation and maintenance decisions.
[0096] Specifically, the correlation score y i =β0 + β1x i1 +β2x i2 +…+β p x ip +∈ i , where y i is the dependent variable of the score of the i - th change event. x ij is the j - th independent variable of the i - th change event, that is, the sub - correlation degree. β0 is the intercept term, β1, β2,…,β p are the slope coefficients, representing the weight coefficients of the current sub - correlation degree, and ∈ i is the error term of the i - th observation.
[0097] It should be understood that the above - mentioned correlation score can be calculated through a scoring model, where the above - mentioned β i and ∈ i are trainable parameters. Specifically, the weights β i corresponding to the sub - correlation degrees of each feature can be determined based on the feature importance method. The specific way of determining the weights is subject to being able to be implemented, and the present disclosure does not make specific limitations on this.
[0098] Step S304: Determine the target event among the change events according to the correlation level, and perform root cause analysis on the failure event based on the target event to obtain the analysis result. For details, please refer to step S104 of the embodiment shown in Figure 1 and will not be elaborated here.
[0099] In the embodiments of the present disclosure, a preset scoring dimension corresponding to a change event can be obtained, and the change event can be scored based on the event data corresponding to the preset scoring dimension, so as to obtain the correlation degree between the change event and the fault event under the preset scoring dimension. Thus, the limitation of single-dimensional evaluation in the traditional method is overcome, multi-dimensional diagnosis of change events is realized, which can more effectively help operation and maintenance personnel locate the root cause of faults, reduce the fault troubleshooting time, improve the stability and reliability of the system, and at the same time optimize the change management process and reduce the risk of faults caused by changes.
[0100] In some alternative embodiments, the above Figure 1 corresponding embodiments further include:
[0101] Step S11: Sort the change events based on the correlation level between the change events and the fault events to obtain a sorting result.
[0102] Step S12: Determine the change events at preset positions in the sorting result, and use the determined change events as target events.
[0103] In the embodiments of the present disclosure, when sorting the target events, ascending or descending sorting can be used to obtain the sorting result. Here, the change events with the top k correlation levels can be selected from the sorting result as the target change events, for example, k = 3.
[0104] Specifically, the above step S11 of sorting the change events based on the correlation level to obtain the sorting result includes:
[0105] Step b1: Obtain the event type of the change event, where the event type is used to indicate the release path of the change event.
[0106] Step b2: Determine the sorting weight of the change event based on the event type.
[0107] Step b3: Sort the change events based on the sorting weight and the correlation level to obtain the sorting result.
[0108] In the embodiments of the present disclosure, considering that different types of events have different degrees of importance in terms of their impact on the system, corresponding weight values are assigned to various change events. When diagnosing change events, combining the weights of event types can more reasonably measure the contribution degree of the events related to the change to the overall system fault possibility, making the diagnosis result more in line with the actual situation.
[0109] For example, event types can include TCE (Toutiao Cloud Engine) events and TCC (Toutiao Config Center) events. Here, the TCE cloud engine can provide users with a fast and efficient service deployment solution, focusing on service lifecycle management such as creation, upgrade, and rollback, and aiming to achieve highly available and elastically scalable container services. TCC can provide a configuration management solution that combines a platform and an SDK package for business parties, including functions such as configuration management, version management, multi-region and multi-environment support, permission management, and gray release.
[0110] When setting the corresponding sorting weights for the above TCE events and TCC events, sorting weights can be set for the events corresponding to the action (specific operation steps of each service during transaction processing), env (environment variables or relevant information about the configured environment), status (status during transaction execution), and stage (different stages of transaction processing) of the TCE events and TCC events respectively.
[0111] For example, in TCE events, the weight of the upgrade cluster event corresponding to action can be 1, the weight of the upgrade sidecar event can be 1, the weight of the create ab experiment event can be 1, the weight of the delete cluster event can be 1, the weight of the delete service event can be 1, the weight of the update cluster information event can be 0.6, the weight of the auto-scaling event can be 0.8, the weight of the update service information event can be 0.6, and the weight of the cancel work order event can be 0.3. Additionally, the weight of the online event corresponding to env can be 1, and the weight of the unknown can be 0.5. The weight corresponding to the status of the event being running can be 1, and the weight corresponding to the status of the event being finished can be 0.9, etc. The weight corresponding to the stage of the event being full traffic can be 1, and the weight corresponding to the stage of the event being single data center can be 0.9, etc.
[0112] Again, for example, in TCC events, the weight of the event of modifying TCC configuration through the platform corresponding to action can be 1, and the weight of the event of modifying TCC configuration through the model interface can be 1. The weight of the online event corresponding to env can be 1, and the weight of the unknown can be 0.5. The weight corresponding to the status of the event being running can be 1, and the weight corresponding to the status of the event being finished can be 1, etc. The weight corresponding to the stage of the event being full traffic can be 1, and the weight corresponding to the stage of the event being small traffic can be 0.6, etc.
[0113] In the embodiments of the present disclosure, sorting weights can be set for change events based on the event type, so that when diagnosing change events, combined with the weights of the event type, the contribution degree of the change-related events to the overall system failure possibility can be more reasonably measured, making the diagnosis result more in line with the actual situation.
[0114] In some alternative embodiments, the above Figure 1 corresponding embodiment further includes:
[0115] After obtaining the analysis result, based on the dependency relationship, a topological structure corresponding to the change event and the failure event is generated and the topological structure is displayed.
[0116] In the embodiments of the present disclosure, as Figure 4 shown in the schematic diagram of the topological structure display interface, among which, the failure event is the increase in the target service latency, the root cause of this failure event is the event change, and the services having a dependency relationship with the target service include Service 1, Service 2, Service 3, Service 4, Service 5, and Service 6.
[0117] In the display interface as Figure 4 shown, the suspected root cause determined based on the analysis result, the impact surface of this failure in the information technology system, the associated change events corresponding to the failure event, and the associated alarm section can also be displayed. The administrator can trigger the corresponding section to view the information to be viewed.
[0118] In the embodiments of the present disclosure, the analysis result can be displayed through the display interface, so as to provide clear navigation and operation guidance for the administrator. For example, a graphical interface is used to display the monitoring data and the diagnosis result, thereby reducing the learning cost of the user.
[0119] In some alternative embodiments, after obtaining the analysis result, the above Figure 1 corresponding embodiment further includes:
[0120] Step S31: Obtain the time information corresponding to the target event.
[0121] Step S32: Display the target event based on the time information.
[0122] In the embodiments of the present disclosure, the time information may include the time axis of the target event. Here, as Figure 5 shown in the schematic diagram of the target event display, among which, the target event can be displayed based on the time axis of the target event. For example, the event type of the target event is displayed.
[0123] It should be understood that the target event display interface may also include the details of the target event. For example, the specific event name, the occurrence time of the failure event, etc. The management personnel can configure the content displayed on the display page according to specific usage requirements.
[0124] In the embodiments of the present disclosure, the target event can be displayed through the display interface, thereby highlighting the information of the key change events that lead to the failure event, facilitating the operation and maintenance personnel to quickly understand and pay attention to the key change events, and reducing the learning cost of the operation and maintenance personnel.
[0125] In this embodiment, another root cause analysis method is provided, which can be used for the above-mentioned information technology systems, such as supply management systems, online information technology systems, etc. Figure 6 It is a flowchart of the root cause analysis method according to the embodiments of the present disclosure, as Figure 6 shown, the process includes the following steps:
[0126] Step S601, collect the event data corresponding to the change events associated with the failure event in the information technology system.
[0127] Step S602, perform feature calculation on the event data using the features of each dimension to obtain the first sub-correlation degree, the second sub-correlation degree, and the third sub-correlation degree.
[0128] Step S603, integrate the first sub-correlation degree, the second sub-correlation degree, and the third sub-correlation degree through a preset fusion feature model to obtain the correlation degree.
[0129] Step S604, determine the correlation level of the change event according to the correlation degree.
[0130] Step S605, sort the change events according to the correlation level to obtain the sorting result.
[0131] Step S606, output the change events whose correlation level exceeds the preset threshold, and display the target events whose correlation level is among the top k in the sorting result.
[0132] Step S607, perform page display on the target events.
[0133] In the embodiments of the present disclosure, the specific implementation manners of steps S601 - S607 are as described above Figure 1 corresponding to the embodiments described herein, and will not be elaborated herein.
[0134] In summary, in the embodiments of the present disclosure, in response to a detected fault event, a change event associated with the fault event can be obtained, where the change event is used to indicate an event generated by updating a service. Then, the event data of the change event can be obtained, where the event data is used to indicate the recorded data generated by a specific operation or state change corresponding to the change event. Next, based on the event data, the association level between the change event and the fault event can be determined, and a target event can be determined from the change events according to the association level, so as to perform root cause analysis on the fault event based on the target event to obtain an analysis result. Thus, in the process of root cause analysis, the change events that may sometimes cause system failures or abnormal conditions can be analyzed to diagnose whether the change event is the root cause of the fault, improving the accuracy of root cause analysis, which is crucial for quickly restoring the normal operation of the system, enhancing system stability, and optimizing subsequent change management, etc.
[0135] In this embodiment, a root cause analysis device is further provided. The device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that realizes a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0136] This embodiment provides a root cause analysis device, as Figure 7 shown, including:
[0137] A detection module 701, configured to obtain a change event associated with a detected fault event, where the change event is used to indicate an event generated by updating a service;
[0138] An acquisition module 702, configured to obtain the event data of the change event, where the event data is used to indicate the recorded data generated by a specific operation or state change corresponding to the change event;
[0139] A determination module 703, configured to determine the association level between the change event and the fault event based on the event data;
[0140] An analysis module 704, configured to determine a target event from the change events according to the association level, and perform root cause analysis on the fault event based on the target event to obtain an analysis result.
[0141] In some alternative implementation manners, the acquisition module 702 is further configured to:
[0142] Determine the target service corresponding to the change event;
[0143] Obtain the detection data of the preset performance indicators of the target service;
[0144] Obtain the dependency relationship between the change event and the fault event, where the dependency relationship is used to indicate the causal relationship between the change event and the fault event;
[0145] Obtain the time record corresponding to the fault event;
[0146] Obtain the release record corresponding to the release of the change event. The event data includes detection data, dependency relationship, time record, and release record.
[0147] In some alternative embodiments, the determining module 703 is further configured to:
[0148] Obtain the preset scoring dimension corresponding to the change event;
[0149] Score the change event based on the event data corresponding to the preset scoring dimension to obtain the correlation degree between the change event and the fault event under the preset scoring dimension;
[0150] Determine the correlation level based on the correlation degrees corresponding to each preset scoring dimension.
[0151] In some alternative embodiments, the event data includes: the detection data corresponding to the change event, the dependency relationship between the change event and the fault event, and the time record of the change event; the determining module 703 is further configured to:
[0152] Obtain the detection data of the preset performance index corresponding to the preset scoring dimension, and determine the first sub-correlation degree of the change event based on the proportion of abnormal data in the detection data;
[0153] Determine the distance value between the change event and the fault event based on the dependency relationship corresponding to the preset scoring dimension, and determine the second sub-correlation degree corresponding to the change event based on the distance value;
[0154] Determine the occurrence time of the change event based on the time record, and determine the third sub-correlation degree corresponding to the change event based on the occurrence time;
[0155] Determine the correlation degree according to the first sub-correlation degree, the second sub-correlation degree, and the third sub-correlation degree.
[0156] In some alternative embodiments, the determining module 703 is further configured to:
[0157] Obtain the change weights corresponding to the first sub-correlation degree, the second sub-correlation degree, and the third sub-correlation degree;
[0158] Perform weighted summation on the first sub-correlation degree, the second sub-correlation degree, and the third sub-correlation degree based on the change weights to obtain the correlation score, and determine the correlation level according to the correlation score.
[0159] In some alternative embodiments, the analysis module 704 is further configured to:
[0160] Sort the change events based on the association level between the change events and the fault events to obtain a sorting result;
[0161] Determine the change events at the preset positions in the sorting result, and use the determined change events as target events.
[0162] In some alternative embodiments, the analysis module 704 is further configured to:
[0163] Obtain the event type of the change event, where the event type is used to indicate the release path of the change event;
[0164] Based on the event type, determine the sorting weight of the change event;
[0165] Sort the change events based on the sorting weight and the association level to obtain a sorting result.
[0166] In some alternative embodiments, the event data includes: the dependency relationship between the change event and the fault event; the apparatus is further configured to:
[0167] After obtaining the analysis result, generate a topology corresponding to the change event and the fault event based on the dependency relationship, and display the topology.
[0168] In some alternative embodiments, the apparatus is further configured to:
[0169] After obtaining the analysis result, obtain the time information corresponding to the target event;
[0170] Based on the time information, display the target event.
[0171] The further function descriptions of the above-mentioned various modules and units are the same as those in the corresponding above-mentioned embodiments, and will not be elaborated here.
[0172] The root cause analysis apparatus in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0173] This embodiment of the present disclosure also provides a computer device having the above-mentioned Figure 7 root cause analysis apparatus as shown.
[0174] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of a computer device provided by an alternative embodiment of the present disclosure, as shown in Figure 8As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting the components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if needed, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (such as an array of servers, a set of blade servers, or a multi-processor system). Figure 8 In [the figure], a processor 10 is taken as an example.
[0175] The processor 10 can be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 can further include a hardware chip. The above-mentioned hardware chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above-mentioned programmable logic device can be a complex programmable logic device, a field-programmable gate array, a generic array logic, or any combination thereof.
[0176] Among them, the memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiments.
[0177] The memory 20 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device, etc. In addition, the memory 20 can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 can optionally include a memory remotely set relative to the processor 10, and these remote memories can be connected to the computer device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0178] The memory 20 can include a volatile memory, such as a random access memory; the memory can also include a non-volatile memory, such as a flash memory, a hard disk, or a solid-state drive; the memory 20 can also include a combination of the above types of memories.
[0179] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or a communication network.
[0180] Embodiments of the present disclosure also provide a computer-readable storage medium. The methods according to the embodiments of the present disclosure can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code that is originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the methods described herein can be processed by such software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the methods shown in the above embodiments are implemented.
[0181] A part of the present invention can be applied as a computer program product, for example, computer program instructions, which when executed by a computer, can call or provide the methods and / or technical solutions according to the present invention through the operation of the computer. Those skilled in the art should be able to understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible by the computer.
[0182] Although the embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A root cause analysis method, characterized in that: The method comprises: In response to a detected fault event, obtaining a change event associated with the fault event, wherein the change event is used to indicate an event generated by updating a service; Acquire event data of the change event, wherein the event data is used to indicate record data generated by a specific operation or state change corresponding to the change event; Based on the event data, determining a correlation level between the change event and the fault event; A target event is determined in the change event according to the correlation level, and a root cause analysis is performed on the fault event based on the target event to obtain an analysis result.
2. The method according to claim 1, characterized in that The acquiring the event data of the change event includes: Determine a target service corresponding to the change event; Obtaining detection data of preset performance indicators of the target service; Acquire a dependency relationship between the change event and the fault event, wherein the dependency relationship is used to indicate a causal relationship between the change event and the fault event; Obtaining a time record corresponding to the fault event; A publishing record corresponding to when the change event is published is obtained, wherein the event data includes the detection data, the dependency relationship, the time record, and the publishing record.
3. The method according to claim 1, characterized in that The determining, based on the event data, a correlation level between the change event and the fault event includes: Obtaining a preset scoring dimension corresponding to the change event; Scoring the change event based on the event data corresponding to the preset scoring dimension to obtain the correlation between the change event and the fault event under the preset scoring dimension; The association level is determined based on the association degree corresponding to each of the preset scoring dimensions.
4. The method according to claim 3, characterized in that The event data includes: detection data corresponding to the change event, a dependency relationship between the change event and the fault event, and a time record of the change event; Scoring the change event based on the event data corresponding to the preset scoring dimension to obtain the correlation between the change event and the fault event under the preset scoring dimension includes: Acquire detection data of a preset performance indicator corresponding to the preset scoring dimension, and determine a first sub-correlation of the change event based on a proportion of abnormal data in the detection data; Based on the dependency relationship corresponding to the preset scoring dimension, determining a distance value between the change event and the fault event, and determining a second sub-correlation degree corresponding to the change event based on the distance value; Determining an occurrence time of the change event based on the time record, and determining a third sub-relevance corresponding to the change event based on the occurrence time; The degree of association is determined according to the first sub-degree of association, the second sub-degree of association, and the third sub-degree of association.
5. The method according to claim 4, characterized in that The determining the association level based on the association degree includes: Obtaining change weights corresponding to the first sub-relevance, the second sub-relevance, and the third sub-relevance; The first sub-degree of association, the second sub-degree of association and the third sub-degree of association are weightedly summed based on the change weight to obtain a correlation score, and the correlation level is determined according to the correlation score.
6. The method according to claim 1, characterized in that The determining a target event in the change event according to the association level includes: sorting the change events based on the correlation levels between the change events and the fault events to obtain a sorting result; A change event at a preset position in the sorting result is determined, and the determined change event is used as the target event.
7. The method according to claim 6, characterized in that The step of sorting the target events based on the correlation levels of the change events to obtain a sorting result includes: Acquire the event type of the change event, wherein the event type is used to indicate a publishing channel of the change event; Based on the event type, determining a ranking weight of the change event; The change events are sorted based on the sorting weights and the association levels to obtain a sorting result.
8. The method according to claim 1, characterized in that The event data includes: a dependency relationship between the change event and the fault event; The method further comprises: After the analysis result is obtained, a topology structure corresponding to the change event and the fault event is generated based on the dependency relationship, and the topology structure is displayed.
9. The method according to claim 8, characterized in that The method further comprises: After obtaining the analysis result, obtaining the time information corresponding to the target event; Based on the time information, the target event is displayed.
10. A root cause analysis device, characterized in that: The device comprises: A detection module, configured to obtain, in response to a detected fault event, a change event associated with the fault event, wherein the change event is used to indicate an event generated by updating a service; An acquisition module, used to acquire event data of the change event, wherein the event data is used to indicate record data generated by a specific operation or state change corresponding to the change event; a determination module, configured to determine a correlation level between the change event and the fault event based on the event data; The analysis module is used to determine a target event in the change event according to the correlation level, and perform a root cause analysis on the fault event based on the target event to obtain an analysis result.
11. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the root cause analysis method according to any one of claims 1 to 9 by executing the computer instructions.
12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the root cause analysis method according to any one of claims 1 to 9.
13. A computer program product, characterized in that The method comprises computer instructions for causing a computer to execute the root cause analysis method according to any one of claims 1 to 9.
Citation Information
Cited By
Fault evaluation method and device, electronic equipment and storage medium
CN121664635A