Fault evaluation method and device, electronic equipment and storage medium
By monitoring service failures in the IT operations and maintenance environment, and performing aggregated correlation analysis on the acquired failure records and change information, the system can automatically locate and determine the cause of failures, solving the problem of low efficiency in fault location caused by service changes and achieving efficient and accurate fault analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-03-13
AI Technical Summary
In IT operations and maintenance environments, fault location and cause determination caused by service changes are inefficient, and existing technologies rely on human experience, making it difficult to efficiently link change information and fault data.
When a service failure is monitored, fault record information and change information in the online order are obtained, aggregated correlation analysis is performed to determine the relationship between changes, and the importance of the related objects is assessed. Finally, a fault evaluation result is generated to automatically locate and determine the cause of the failure.
It improves the accuracy and efficiency of fault location and fault cause determination, realizes automated fault analysis, and reduces reliance on human experience.
Smart Images

Figure CN121664635A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of operation and maintenance engineering technology, and in particular to a fault evaluation method, device, electronic device and storage medium. Background Technology
[0002] In the current IT operations and maintenance environment, service changes and releases can sometimes lead to service failures, affecting system stability and availability.
[0003] Currently, fault diagnosis tools are used to detect faults in the operation and maintenance environment. However, because change data and fault data are at different stages of the software lifecycle, the correlation between change information and fault data can only rely on manual experience. This leads to low efficiency in fault location and cause determination. Summary of the Invention
[0004] This invention provides a fault evaluation method, apparatus, electronic device, and storage medium, which enables automatic analysis of service faults and improves the accuracy and efficiency of fault location and fault cause determination.
[0005] In a first aspect, embodiments of the present invention provide a fault evaluation method, the method comprising:
[0006] When a service failure is detected, the failure record information is determined, and the change information is obtained by accessing the online order. The failure record information and the change information are aggregated and correlated to obtain at least two change correlation relationships.
[0007] For each type of change relationship, determine the corresponding relationship evaluation based on the related objects in the change relationship.
[0008] The fault evaluation results are determined based on the associated evaluations corresponding to all changes and then displayed.
[0009] The fault evaluation method provided in this invention not only requires fault record information of service faults, but also obtains change information related to the time point of the service fault recorded in the online order. It then performs correlation analysis on the fault record information and change information, automatically aggregating and correlating change information during the software release phase with fault record information during the software operation phase. For each change correlation relationship, the correlation between the fault record information and its corresponding change information can be analyzed to determine the importance of this correlation in affecting the occurrence of the fault, thus obtaining a correlation evaluation. This enables accurate correlation analysis after automatically aggregating fault conditions and change information. Integrating the correlation evaluations corresponding to all change correlation relationships to determine the fault evaluation result not only shows the degree of impact of the current change on the service fault, but also shows the ranking of the degree of impact of different changes on the service fault. Therefore, after showing the fault evaluation result to staff, staff can accurately and quickly determine the cause of the service fault and achieve rapid fault location based on the fault evaluation result. This solves the problem of low efficiency in locating and determining the cause of faults by relying solely on human experience, achieving automatic analysis of service faults and improving the accuracy and efficiency of fault location and cause determination.
[0010] Secondly, embodiments of the present invention also provide a fault evaluation device, the device comprising:
[0011] The analysis module is used to determine the fault record information of the service failure when a service failure is detected, and to access the online order to obtain the change information. It then performs aggregated correlation analysis on the fault record information and the change information to obtain at least two change correlation relationships.
[0012] The determination module is used to determine the association evaluation corresponding to each type of change association based on the associated objects in the change association.
[0013] The display module is used to determine the fault evaluation result based on the associated evaluations corresponding to all changes and to display the fault evaluation result.
[0014] Thirdly, embodiments of the present invention also provide an electronic device, the electronic device comprising:
[0015] At least one processor; and
[0016] A memory that is communicatively connected to at least one processor; wherein,
[0017] The memory stores a computer program that can be executed by at least one processor, such that the at least one processor is able to perform the fault evaluation method of any embodiment of the present invention.
[0018] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing computer instructions that are used to cause a processor to execute and implement the fault evaluation method of any embodiment of the present invention.
[0019] Fifthly, embodiments of the present invention also provide a computer program product, including a computer program that, when executed by a processor, implements the fault evaluation method of any embodiment of the present invention.
[0020] It should be noted that the aforementioned computer instructions may be stored, in whole or in part, on a computer-readable storage medium. This computer-readable storage medium may be packaged together with the processor of the fault evaluation device, or it may be packaged separately from the processor of the fault evaluation device; this application does not impose any limitations on this.
[0021] The descriptions of the second, third, fourth, and fifth aspects in this application can be referred to the detailed description of the first aspect; and the beneficial effects of the descriptions of the second, third, fourth, and fifth aspects can be referred to the analysis of the beneficial effects of the first aspect, which will not be repeated here.
[0022] In this application, the name of the aforementioned fault evaluation device does not limit the equipment or functional module itself. In actual implementation, these devices or functional modules may appear under other names. As long as the function of each device or functional module is similar to that of this application, it falls within the scope of the claims of this application and its equivalents.
[0023] These or other aspects of this application will become more readily apparent in the following description. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 This is a flowchart illustrating a fault evaluation method provided in an embodiment of the present invention.
[0026] Figure 2 A flowchart illustrating another fault evaluation method provided in an embodiment of the present invention;
[0027] Figure 3 This is a schematic diagram of the structure of a fault evaluation device provided in an embodiment of the present invention;
[0028] Figure 4This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0029] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.
[0030] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0031] The terms “initial” and “target” in the specification and drawings of this application are used to distinguish different objects or to distinguish different treatments of the same object, rather than to describe a specific order of objects.
[0032] Furthermore, the terms "comprising" and "having," and any variations thereof, used in the description of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the steps or units listed, but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to such process, method, product, or apparatus.
[0033] Before discussing the exemplary embodiments in more detail, it should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of these operations can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations can be rearranged. The process can be terminated when its operation is completed, but may also have additional steps not included in the figures. The process can correspond to a method, function, procedure, subroutine, subroutine, etc. Moreover, without conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0034] It should be noted that in the embodiments of this application, the words "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of the words "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0035] In the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0036] Figure 1 This is a flowchart illustrating a fault evaluation method provided in an embodiment of the present invention. This embodiment is applicable to situations where service faults detected in an operation and maintenance environment are analyzed. The method can be executed by a fault evaluation device, which can be implemented in hardware and / or software and can be configured in an electronic device. In this embodiment, the electronic device can be a computer or server belonging to the operation and maintenance personnel. (Continue referring to...) Figure 1 This embodiment specifically includes the following steps:
[0037] S101. When a service failure is detected, determine the failure record information of the service failure, access the online order to obtain the change information, and perform aggregated correlation analysis on the failure record information and the change information to obtain at least two change correlation relationships.
[0038] Service failure is used to characterize the state in which a software service fails to provide functionality according to its design specifications, service level agreement, or user expectations. In this embodiment, service failure refers to an anomaly or malfunction detected by the electronic device. Fault log information is used to characterize fault-related information at the time the fault occurred. A deployment order is a formal document or electronic record used to record, approve, track, and control the migration of a software change from the development environment to the production environment. In this embodiment, the deployment order can be a change order, recording relevant information for all change operations. Change information refers to the relevant change information of the software with service failure at the time the failure occurred. Aggregate correlation analysis combines the two types of information for correlation analysis.
[0039] Specifically, during the operation of the software service, fault diagnosis tools are continuously used to detect runtime anomalies. Therefore, when the fault diagnosis tool detects an anomaly, a service failure exists. At this time, the electronic device can automatically record the time point of the service failure detection and specific fault information, forming a fault log. Simultaneously, the electronic device can access the deployment order to obtain change information related to the time point of the service failure. Furthermore, the fault log information and related content in the change information can be aggregated to obtain at least one aggregated set of information. By performing correlation analysis on each aggregated set of information, at least two change correlation relationships can be obtained.
[0040] In this embodiment, not only is it necessary to obtain the fault record information of the service failure, but also to obtain the change information related to the time point of the service failure recorded in the online order. Then, the fault record information and change information are correlated and analyzed. This realizes the automatic aggregation and correlation analysis of change information in the software release stage and fault record information in the software operation stage, providing a basis for the correlation between information for accurate and rapid fault location and determination of fault causes.
[0041] S102. For each type of change relationship, determine the corresponding relationship evaluation based on the related objects in the change relationship.
[0042] In this embodiment, different change associations are used to characterize the correlation between fault record information and change information across different analytical dimensions. The associated objects represent fault record information and change information that are related within a change association. The association evaluation characterizes the importance of a change association at the time of the service failure.
[0043] Specifically, after obtaining at least two types of change relationships as described above, for each change relationship, the correlation evaluation corresponding to that change relationship can be determined based on the associated objects, i.e., a certain fault record information and its corresponding change information. For example, if the fault occurrence time in the fault record information corresponds to the change time, the correlation between the fault occurrence time and the change time in this change relationship can be analyzed to determine the importance of this correlation in the final impact on the occurrence of the fault, i.e., to determine the correlation evaluation.
[0044] In this embodiment, for each type of change association, the correlation between the fault record information and its corresponding change information in the change association can be analyzed, thereby determining the importance of the correlation to the occurrence of the fault and obtaining the association evaluation. This enables the automatic aggregation of fault conditions and change information for accurate association analysis, providing a basis for accurately determining the cause of the fault in the future.
[0045] S103. Determine the fault evaluation results based on the associated evaluations corresponding to all changes and display the fault evaluation results.
[0046] The fault evaluation results are a ranking of the degree of influence among the specific causes that led to the service failure.
[0047] Specifically, after determining the associated evaluation for each change relationship, all associated evaluations can be integrated, such as using weighted summation, analytic hierarchy process (AHP), and rule engines, to determine the fault evaluation result. Furthermore, after determining the fault evaluation result, it can be displayed to staff, allowing them to accurately and quickly determine the cause of the service failure and achieve rapid fault localization.
[0048] In this embodiment, the associated evaluations corresponding to all changes are integrated to determine the fault evaluation results. This not only shows the degree of impact of the current change on the service fault, but also shows the ranking of the degree of impact of different changes on the service fault. Therefore, after showing the fault evaluation results to the staff, the staff can accurately and quickly determine the cause of the service fault and achieve rapid fault location based on the fault evaluation results. This solves the problem of low efficiency in locating and determining the cause of faults by relying solely on human experience. It realizes automatic analysis of service faults and improves the accuracy and efficiency of fault location and cause determination.
[0049] The fault evaluation method provided in this invention not only requires fault record information of service faults, but also obtains change information related to the time point of the service fault recorded in the online order. It then performs correlation analysis on the fault record information and change information, automatically aggregating and correlating change information during the software release phase with fault record information during the software operation phase. For each change correlation relationship, the correlation between the fault record information and its corresponding change information can be analyzed to determine the importance of this correlation in affecting the occurrence of the fault, thus obtaining a correlation evaluation. This enables accurate correlation analysis after automatically aggregating fault conditions and change information. Integrating the correlation evaluations corresponding to all change correlation relationships to determine the fault evaluation result not only shows the degree of impact of the current change on the service fault, but also shows the ranking of the degree of impact of different changes on the service fault. Therefore, after showing the fault evaluation result to staff, staff can accurately and quickly determine the cause of the service fault and achieve rapid fault location based on the fault evaluation result. This solves the problem of low efficiency in locating and determining the cause of faults by relying solely on human experience, achieving automatic analysis of service faults and improving the accuracy and efficiency of fault location and cause determination.
[0050] Figure 2 This is a flowchart illustrating another fault evaluation method provided by an embodiment of the present invention. This embodiment, based on the above embodiment, specifies the steps of obtaining at least two change associations, determining the association evaluation corresponding to the change associations, and determining the fault evaluation result. In this embodiment, the method may include:
[0051] S201. When a service failure is detected, determine the failure record information of the service failure and access the online order to obtain change information.
[0052] Specifically, monitoring tools or log analysis can be used to detect service failures. When a service failure is detected, the failure record information can be determined, such as: the failure identification time point, i.e., the specific time when the failure occurred; failure information, such as error logs, exception types, etc.; and the number of failures, such as the number of times the failure occurred or the scope of impact. Simultaneously, deployment records can be accessed, and change information related to the failure identification time point can be obtained from the deployment records, such as: change time, i.e., the specific time when the change occurred; change coordinates, such as the code repository location (file path or line number in the code repository); change platform, such as changing the deployment platform or environment (production environment, testing environment, etc.); change batch, such as the batch to which the change belongs or the release batch; and change context, such as the change description, reason for the change, purpose of the change, and the modules involved.
[0053] S202. Perform topological distance analysis on fault record information and change information to obtain topological association relationship.
[0054] In this embodiment, the aggregation correlation analysis includes topological distance analysis, time series analysis, and change attribute analysis. Topological distance analysis is used to analyze the topological relationship between changes and faulty services.
[0055] Specifically, since topology distance analysis helps to understand the location of changes in the system architecture and their impact on faulty services, topology distance analysis can be performed using fault service location (such as fault service identifier, fault instance location, fault level), fault propagation characteristics (single point of failure / multi-point failure, temporal propagation order, spatial / topology propagation mode, etc.) and fault dependency chain information (real-time call chain data and service dependency graph snapshots) in fault log information, and change target identifier (such as service / application identifier, i.e., service name, application unique identifier; deployment example identifier; cluster / environment identifier), change dependency relationship (such as direct upstream dependency; direct downstream dependency; data / infrastructure dependency, etc.), and change effective topology layer (such as application layer, middleware layer, infrastructure layer, or data layer) in change information to obtain topology relationships.
[0056] For example, the topological distance analysis of the above information can be performed according to the following steps to obtain the topological association relationship:
[0057] (1) Standardize the identification to ensure that change and fault information use the same service naming and identification specifications; and find the related change information from the change information based on the fault occurrence time (i.e., the fault identification time point); at the same time, the system topology map at the fault identification time point can be obtained. (2) Map the "change target identifier" in the change information to one or more nodes on the topology map; map the "fault service location" in the fault record information to nodes on the topology map; in graph theory, find all possible paths from the "change node" to the "fault node", such as the direct path can be change service → fault service, or the indirect path can be change service → service B → service C → fault service. (3) Calculate the topology relationship metric, that is, calculate the minimum number of hops between two nodes in the topology map and use it as the shortest path distance; then, identify the main path or core path of traffic to obtain the target propagation path; then, find the key hub service on the path and calculate the hierarchical difference between change and fault. Finally, the topology relationship after topology distance analysis can be obtained.
[0058] Alternatively, in a simpler approach, topological distance analysis can be performed on the fault record information and change information to obtain the topological association relationship, including:
[0059] (a) Determine the fault location information based on the service identifier and instance information in the fault record information.
[0060] The service identifier is a logically unique identity information defining a software service; in this embodiment, the service identifier is used to characterize a software service that has failed. Instance information refers to the specific, physical, or virtual computing unit that hosts the service; in this embodiment, the service identifier is used to characterize a specific server, container, or process that has failed.
[0061] Specifically, the fault can be accurately located based on the service identifier and instance information in the fault record information.
[0062] (ii) Perform correlation analysis on the deployment location information and fault location information in the change information to obtain the node location correlation.
[0063] Specifically, since deployment location information can indicate which node the service starts from, and fault location information can indicate where the entire service process is interrupted (ended), the "location description" in deployment location information and fault location information can be converted into node identifiers in the topology graph. Then, by using the node identifier in deployment location information as the starting point and the node identifier in fault location information as the ending point, "path planning" can be performed to obtain the node location association.
[0064] (iii) Determine the topological association relationship based on the node location association and the dependency relationship of each node in the node location association.
[0065] Specifically, since each node's corresponding service / instance has dependencies, such as the change dependencies mentioned above, the dependencies between nodes can be determined directly based on these dependencies. Then, the node location associations and the dependencies between nodes can be integrated to obtain the topological associations.
[0066] In this embodiment, by performing topological distance analysis on fault record information and change information, a topological relationship is obtained, which provides a more intuitive topological relationship basis for accurately and quickly finding the cause of the fault based on the topological relationship.
[0067] S203. Perform time series analysis on fault record information and change information to obtain time correlation.
[0068] Time series analysis is used to analyze the correlation between changes and service failures over time.
[0069] Specifically, since time correlation analysis helps determine whether changes were made before the failure occurred and to determine the length of the time interval, time series analysis can be performed on the failure timeline (such as the first failure event, the failure detection / alarm event, the peak time of failure impact, the failure start recovery time, and the failure complete recovery time) in the failure record information, the time pattern and characteristics of the failure (such as determining whether the failure is instantaneous, intermittent, or continuous; time series data of failure indicators, such as error rate, latency, traffic, etc.), the scope and background of the failure (such as the list of affected services / instances and the business load at the time of the failure), and the change timestamps in the change information (such as the change start time, change effective time, change end / completion time, and time type description), the time attributes and context of the change (such as change type, change strategy and batch, and change duration), and the change unique identifier and scope (such as the change unique identifier / number and change target) to obtain time correlation relationships.
[0070] For example, the above information can be subjected to time series analysis to obtain the time correlation relationship by following the steps below:
[0071] (1) Establish a unified timeline and convert all timestamps to the same time zone, defining analysis windows such as look-ahead and look-back windows. (2) Calculate key time metrics, such as subtracting the change's effective time from the first occurrence of the fault to determine the time difference. (3) By performing correlation analysis on the time information of the faults and the time information of the corresponding changes, determine multiple correlation indicators, such as the temporal proximity between the change's effective time and the fault occurrence event, the Spearman rank correlation between the effective time series of multiple changes and the occurrence time series of multiple faults, the cross-correlation between the continuous time series of fault indicators (such as error rate) and the change event as a binary time series, and the effective time of each change and the time difference from that time to the occurrence of the fault, thus obtaining the survival analysis correlation. After obtaining the above correlation indicators, the time relationship can be determined based on these indicators.
[0072] Alternatively, in a simpler approach, time-series analysis can be performed on fault record information and change information to obtain time correlations, including:
[0073] (i) Perform time correlation analysis on the fault occurrence time in the fault record information and the change time in the change information to obtain the time correlation relationship.
[0074] Specifically, the time difference can be obtained by directly calculating the difference between the time of the failure and the time of the change. Alternatively, one approach is to determine the time difference based on a time threshold (determined from the time difference of failures caused by changes based on historical data). If the time difference is less than the time threshold, the change and failure are considered to be strongly time-related. Another approach is to calculate the time difference using statistical tests or correlation coefficients, again using historical data in the calculation process to ultimately determine the time correlation.
[0075] S204. Perform change attribute analysis on fault record information and change information to obtain the relationship between change attributes.
[0076] Among them, change attribute analysis is used to analyze the type of change and its relevance to the failure service.
[0077] Specifically, since change level analysis helps to distinguish different types of changes and their potential impacts, change attribute analysis can be performed on abnormal reports and abnormal types in fault log information, as well as change types in change information, to obtain the relationship between change attributes.
[0078] For example, change attribute analysis is performed on fault record information and change information to obtain the change attribute association relationship, including:
[0079] (i) Determine the impact indicators of the change information based on the change type in the change information.
[0080] The change types include general changes, important changes, and urgent changes. Change impact metrics are used to characterize the degree to which a change of this type is more likely to cause service outages.
[0081] Specifically, since different types of changes have varying degrees of importance and thus different probabilities of causing service outages, assessing the severity of different change types—such as release changes, configuration changes, and experimental changes—can begin by determining the change impact index based on the change type in the change information. For example, an urgent change will have a higher impact index, while a general change will have a lower impact index.
[0082] (ii) Determine the actual impact based on the abnormal reports and abnormal types in the fault record information.
[0083] Specifically, the anomaly report and anomaly type can determine the actual situation of this service failure. For example, the scope of the anomaly can be determined based on the anomaly report, which serves as the scope of impact; the depth of impact can be determined based on different anomaly types; and the actual degree of impact can be determined by multiplying the scope of impact by the depth of impact.
[0084] (iii) Determine the relationship between change attributes based on the change impact indicators and the actual degree of impact.
[0085] Specifically, the change impact index can be multiplied by the actual degree of impact to obtain the change attribute correlation.
[0086] In this embodiment, by analyzing fault record information and change information at different levels, at least two change correlation relationships are obtained. This enables correlation analysis of fault record information and change information from different levels, achieving a comprehensive and accurate determination of the relationship between faults and changes, and providing a more accurate and comprehensive basis for determining the cause of service failures.
[0087] It is worth noting that in this embodiment, S202, S203 and S204 are parallel steps, which can be executed simultaneously or in stages.
[0088] S205. For each type of change association, determine the association gap index based on the first associated object corresponding to the change information and the second associated object corresponding to the fault record information in the change association.
[0089] Specifically, for each type of change association, the association gap index between the two can be determined based on the first associated object corresponding to the change information and the second associated object corresponding to the fault record information in the change association.
[0090] For example, if the change association is a topological association, then the first associated object is the change node, the second associated object is the faulty service node, and the association gap index is the topological distance between the change node and the faulty service node; if the change association is a time association, then the first associated object is the change time, the second associated object is the fault occurrence time, and the association gap index is the time difference between the change time and the fault occurrence time; if the change association is a change attribute association, then the first associated object is the change type, the second associated object is the fault type, and the association gap index is the level corresponding to each change type and fault type.
[0091] S206. Determine the correlation evaluation corresponding to the change in correlation relationship based on the correlation gap index and the correlation threshold corresponding to the correlation gap index.
[0092] In this embodiment, the relevance threshold is determined based on experience or experiments / trials during historical processes. In this embodiment, the correlation evaluation corresponding to topological relationships measures the topological relationship between changes and faulty services; the closer the topological distance, the higher the score, indicating a greater impact of the change on the faulty service. The correlation evaluation corresponding to temporal relationships measures the correlation between the change time and the fault occurrence time; the smaller the time difference, the higher the score, indicating a stronger correlation between the change and the fault. The correlation evaluation corresponding to change attribute relationships measures the type and scope of the change; the higher the change level (e.g., releasing a change), the higher the score, indicating a greater impact of the change on the system.
[0093] Specifically, each correlation gap indicator has its own correlation threshold. Therefore, the correlation evaluation corresponding to the change in correlation relationship can be obtained by subtracting the corresponding correlation threshold from the correlation gap indicator.
[0094] For example, the correlation evaluation can be calculated as follows:
[0095] If the relationship is topologically related, the distance difference is calculated by subtracting a pre-set topological distance threshold from the correlation gap index of "topological distance between the changed node and the faulty service node". This distance difference is then used as the base of the Euler number e to obtain e^(distance difference). Finally, the correlation evaluation corresponding to the topological relationship is obtained by using the constant 1 - e^(distance difference). The topological distance threshold is used to determine whether the topological distance between the changed node and the faulty service is sufficiently close. If the relationship is temporally related, the time difference is calculated by subtracting a pre-set time correlation threshold from the correlation gap index of "time difference between the change time and the fault occurrence time". The correlation evaluation corresponding to the temporally related relationship is then calculated according to the above steps. The time correlation threshold is used to determine whether the topological distance between the changed node and the faulty service is sufficiently close. If the relationship is based on change attributes, the level difference is determined by using the levels corresponding to the change type and the fault type, such as the difference in change levels and the change level threshold. The correlation evaluation corresponding to the temporally related relationship is then calculated according to the above steps. The change level threshold is used to determine whether the change type and its impact scope are sufficiently important.
[0096] In this embodiment, the corresponding correlation evaluation is calculated for each change relationship. This can accurately present the impact of faults and changes on service failures in the form of mathematical correlation evaluation, which can provide a mathematical basis for subsequent staff to intuitively and accurately understand the impact of faults and changes.
[0097] S207. For each changed relationship, determine the single evaluation result corresponding to the changed relationship based on the allowed weight index and the relationship evaluation.
[0098] In this embodiment, the allowed weighting indicators are used to characterize the importance of the impact of each change relationship on the service failure. For example, the allowed weighting indicator corresponding to the topological relationship is used to measure the importance of the topological relationship between the change and the failed service; the allowed weighting indicator corresponding to the time relationship is used to measure the importance of the correlation between the change time and the failure occurrence time; and the allowed weighting indicator corresponding to the change attribute relationship is used to measure the importance of the change type and the scope of impact.
[0099] Specifically, for each change relationship, the allowed weight index corresponding to the change relationship can be multiplied by the change relationship to obtain a single evaluation result for that change relationship.
[0100] S208. Sum the single evaluation results corresponding to all changes and relationships to obtain the fault evaluation result.
[0101] Specifically, the fault evaluation result is obtained by summing the individual evaluation results corresponding to all changes and relationships. That is, the fault evaluation result is a comprehensive score.
[0102] S209. Display the fault evaluation results.
[0103] Specifically, after obtaining the fault evaluation results, one approach is to directly display the results to staff. Another approach involves using a pre-trained model to analyze the fault record information, change information, at least two change relationships, the corresponding evaluations for each change relationship, and the fault evaluation results. The model then generates ranking results, such as "Release Change Causes Service Failure," "Configuration Change," and "Experimental Change." Here, "Release Change Causes Service Failure" indicates that the release change is the primary cause of the service failure, and it is ranked first; "Configuration Change" indicates that the configuration change may be one of the causes of the service failure; and "Experimental Change" indicates that the experimental change may be one of the causes of the service failure.
[0104] In this embodiment, the fault evaluation results are displayed to the operations and maintenance developers. This intuitive display helps them quickly locate and resolve fault issues.
[0105] Figure 3 This is a schematic diagram of the structure of a fault evaluation device provided in an embodiment of the present invention, as shown below. Figure 3 As shown, the device includes:
[0106] The analysis module 301 is used to determine the fault record information of the service failure when a service failure is detected, and to access the online order to obtain change information. It then performs aggregated correlation analysis on the fault record information and change information to obtain at least two change correlation relationships.
[0107] The determination module 302 is used to determine the association evaluation corresponding to the change association relationship based on the associated objects in the change association relationship for each type of change association relationship.
[0108] The display module 303 is used to determine the fault evaluation result based on the associated evaluation corresponding to all changes and display the fault evaluation result.
[0109] Based on the above embodiments, the aggregation correlation analysis includes topological distance analysis, time series analysis, and change attribute analysis; the analysis module 301 is specifically used for:
[0110] Topological distance analysis is performed on fault record information and change information to obtain topological correlation; time series analysis is performed on fault record information and change information to obtain time correlation; change attribute analysis is performed on fault record information and change information to obtain change attribute correlation.
[0111] Based on the above embodiments, change attribute analysis is performed on fault record information and change information to obtain the change attribute association relationship. The analysis module 301 is specifically used for:
[0112] Determine the impact indicators of the change information based on the change type in the change information; determine the actual impact level based on the anomaly reports and anomaly types in the fault record information; and determine the relationship between the change attributes based on the change impact indicators and the actual impact level.
[0113] Based on the above embodiments, topological distance analysis is performed on fault record information and change information to obtain topological association relationships. The analysis module 301 is specifically used for:
[0114] The fault location information is determined based on the service identifier and instance information in the fault record information; the deployment location information in the change information is correlated with the fault location information to obtain the node location correlation; the topology correlation is determined based on the node location correlation and the dependency relationship of each node in the node location correlation.
[0115] Based on the above embodiments, time series analysis is performed on fault record information and change information to obtain time correlation. The analysis module 301 is specifically used for:
[0116] A time correlation analysis was performed on the fault occurrence time in the fault record information and the change time in the change information to obtain the time correlation relationship.
[0117] Based on the above embodiments, the association evaluation corresponding to the changed association relationship is determined according to the associated objects in the changed association relationship. The determining module 302 is specifically used for:
[0118] The association gap index is determined based on the first associated object corresponding to the change information and the second associated object corresponding to the fault record information in the change association relationship; the association evaluation corresponding to the change association relationship is determined based on the association gap index and the correlation threshold corresponding to the association gap index.
[0119] Based on the above embodiments, the fault evaluation result is determined according to the association evaluation corresponding to all changed relationships. The display module 303 is specifically used for:
[0120] For each change relationship, determine the single evaluation result corresponding to the change relationship based on the allowable weight index and the relationship evaluation; sum the single evaluation results corresponding to all change relationships to obtain the fault evaluation result.
[0121] The fault evaluation device provided in the embodiments of the present invention can execute the fault evaluation method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0122] It is worth noting that in the embodiments of the above-mentioned fault evaluation device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of the present invention.
[0123] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Figure 4 A block diagram is shown of an exemplary electronic device 11 suitable for implementing embodiments of the present invention. Figure 4 The electronic device 11 shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.
[0124] like Figure 4 As shown, the electronic device 11 is represented in the form of a general-purpose computing electronic device. The components of the electronic device 11 may include, but are not limited to: one or more processors or processing units 16, system memory 28, and bus 18 connecting different system components (including system memory 28 and processing unit 16).
[0125] Bus 18 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. For example, these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.
[0126] Electronic device 11 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by electronic device 11, including volatile and non-volatile media, removable and non-removable media.
[0127] System memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Electronic device 11 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be used to read and write non-removable, non-volatile magnetic media (… Figure 4 Not shown; usually referred to as a "hard drive"). Although Figure 4As not shown, disk drives for reading and writing to removable non-volatile disks (e.g., "floppy disks") and optical disc drives for reading and writing to removable non-volatile optical discs (e.g., CD-ROMs, DVD-ROMs, or other optical media) may be provided. In these cases, each drive may be connected to bus 18 via one or more data media interfaces. System memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present invention.
[0128] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in system memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 42 typically perform the functions and / or methods described in the embodiments of the present invention.
[0129] Electronic device 11 can also communicate with one or more external devices 14 (e.g., keyboard, pointing device, display 24, etc.), and with one or more devices that enable a user to interact with electronic device 11, and / or with any device that enables electronic device 11 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed via input / output (I / O) interface 22. Furthermore, electronic device 11 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 20. Figure 4 As shown, network adapter 20 communicates with other modules of electronic device 11 via bus 18. It should be understood that, although... Figure 4 As not shown, other hardware and / or software modules may be used in conjunction with electronic device 11, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0130] The processing unit 16 executes various functional applications and page displays by running programs stored in the system memory 28, such as implementing the fault evaluation method provided in this embodiment. Of course, those skilled in the art will understand that the processor can also implement the technical solutions of the fault evaluation method provided in any embodiment of this invention.
[0131] This invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements, for example, the fault evaluation method provided in this invention.
[0132] The computer storage medium of this invention can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0133] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.
[0134] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0135] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the fault evaluation method provided in any embodiment of this invention.
[0136] Computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0137] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computing device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.
[0138] Furthermore, the acquisition, storage, use, and processing of data in the technical solution of this invention all comply with the relevant provisions of national laws and regulations.
[0139] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.
Claims
1. A fault evaluation method, characterized in that, The method includes: When a service failure is detected, the failure record information of the service failure is determined, and the change information is obtained by accessing the online order. The failure record information and the change information are aggregated and correlated to obtain at least two change correlation relationships. For each type of change association, the association evaluation corresponding to the change association is determined based on the associated objects in the change association; The fault evaluation results are determined based on the associated evaluations corresponding to all changes in the relationship, and the fault evaluation results are displayed.
2. The method according to claim 1, characterized in that, The aggregation correlation analysis includes topological distance analysis, time series analysis, and change attribute analysis; the aggregation correlation analysis of the fault record information and the change information yields at least two change correlation relationships, including: Perform topological distance analysis on the fault record information and the change information to obtain the topological association relationship; Time series analysis was performed on the fault record information and the change information to obtain the time correlation relationship; The fault record information and the change information are analyzed for change attributes to obtain the correlation relationship of change attributes.
3. The method according to claim 2, characterized in that, The step of performing change attribute analysis on the fault record information and the change information to obtain the change attribute association relationship includes: Determine the impact index of the change information based on the change type in the change information; The actual impact is determined based on the anomaly reports and anomaly types in the fault record information; The relationship between the change attributes is determined based on the change impact indicators and the actual degree of impact.
4. The method according to claim 2, characterized in that, The step of performing topological distance analysis on the fault record information and the change information to obtain the topological association relationship includes: The fault location information is determined based on the service identifier and instance information in the fault record information; The deployment location information in the change information and the fault location information are correlated to obtain the node location association; The topological association is determined based on the node location association and the dependency relationship between each node in the node location association.
5. The method according to claim 2, characterized in that, The step of performing time series analysis on the fault record information and the change information to obtain the time correlation includes: A time correlation analysis is performed between the fault occurrence time in the fault record information and the change time in the change information to obtain the time correlation relationship.
6. The method according to claim 1, characterized in that, The step of determining the association evaluation corresponding to the changed association relationship based on the associated objects in the changed association relationship includes: The association gap index is determined based on the first associated object corresponding to the change information and the second associated object corresponding to the fault record information in the change association relationship. The correlation evaluation corresponding to the change in the correlation relationship is determined based on the correlation gap index and the correlation threshold corresponding to the correlation gap index.
7. The method according to claim 1, characterized in that, The step of determining the fault evaluation result based on the association evaluation corresponding to all changes includes: For each of the aforementioned changes and relationships, a single evaluation result corresponding to the change and relationship is determined based on the allowed weight index corresponding to the change and relationship evaluation. The fault evaluation result is obtained by summing the individual evaluation results corresponding to all changes and relationships.
8. A fault evaluation device, characterized in that, The device includes: The analysis module is used to determine the fault record information of the service failure when a service failure is detected, and to access the online order to obtain change information. The fault record information and the change information are aggregated and correlated to obtain at least two change correlation relationships. The determination module is used to determine the association evaluation corresponding to each of the aforementioned change associations based on the associated objects in the change association; The display module is used to determine the fault evaluation result based on the associated evaluation corresponding to all changes and display the fault evaluation result.
9. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the fault evaluation method as described in any one of claims 1 to 7.
10. A readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the fault evaluation method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Fault prediction method and device, electronic equipment and storage medium
CN111858120A
Knowledge graph-based operation and maintenance change influence assessment method and system
CN119829389A
Root cause analysis method and device, computer equipment, storage medium and program product
CN120234176A
Fault information processing method and related device
WO2016095716A1