Change event recommendation method and device, storage medium and program product

By introducing correlation indicators between system operation status data and location results, the association between change events and system anomalies is quantified, solving the problem of low accuracy in change event recommendations in existing technologies and achieving more efficient fault diagnosis and repair.

CN120909822APending Publication Date: 2025-11-07ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510999256.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-11-07

Smart Images

  • Figure CN120909822A_ABST
    Figure CN120909822A_ABST
Patent Text Reader

Abstract

One or more embodiments of the invention provide a change event recommendation method and device, a storage medium and a program product. The method comprises the following steps: determining a plurality of change events based on an alarm event, wherein the alarm event is triggered after detecting that a system is abnormal; determining a correlation index of each change event in the plurality of change events, the correlation index of each change event at least comprising a system operation state data correlation index, the system operation state data correlation index is used for representing the closeness degree of a time node when the system operation state data is suddenly changed and a triggering time node of the change event; and determining a recommendation sequence of the plurality of change events as the change events causing the system abnormality based on the correlation index, and recommending the plurality of change events based on the recommendation sequence.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] One or more embodiments of the present specification relate to the technical field of communication, and in particular, to a change event recommendation method, device, storage medium and program product. BACKGROUND

[0002] Quick positioning and repairing of system exceptions is the key to ensure stable operation of the system. System exceptions can be caused by various factors, among which, change operations (such as code update, configuration modification, database adjustment, etc.) are one of the common reasons. In order to quickly locate and repair system exceptions, it is usually necessary to quickly identify and recommend change events related to the exception to the operation and maintenance personnel to help them quickly focus on the root cause of the problem and repair the exception in time. In related technologies, when recommending change events related to the exception, only the closeness between the occurrence time of the change event and the occurrence time of the alarm event is considered, and the recommendation accuracy is usually low. Therefore, it is necessary to provide a more accurate change event recommendation scheme to help operation and maintenance personnel quickly locate system exceptions. SUMMARY

[0003] Therefore, one or more embodiments of the present specification provide a change event recommendation method, device, storage medium and program product.

[0004] To achieve the above-mentioned purpose, one or more embodiments of the present specification provide technical solutions as follows:

[0005] According to a first aspect of one or more embodiments of the present specification, a change event recommendation method is provided, comprising:

[0006] determining a plurality of change events based on an alarm event, the alarm event being triggered after detecting a system exception;

[0007] determining a relevance indicator of each change event in the plurality of change events, wherein the relevance indicator of each change event at least includes a system running state data relevance indicator, the system running state data relevance indicator being used to represent the closeness between the time node of the mutation of the system running state data and the triggering time node of the change event;

[0008] determining a recommendation order of the plurality of change events as change events causing the system exception based on the relevance indicator, and recommending the plurality of change events based on the recommendation order.

[0009] According to a second aspect of the embodiments of the present specification, an electronic device is provided, comprising:

[0010] a processor;

[0011] a memory for storing processor-executable instructions;

[0012] The processor is configured to implement the method of the first aspect.

[0013] According to a third aspect of the embodiments of the present specification, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program. The computer program is executed by a processor to implement the steps of the method of the first aspect.

[0014] According to a fourth aspect of the embodiments of the present specification, a computer program product is provided, and the computer program product comprises a computer program. The computer program is executed by a processor to implement the steps of the method of the first aspect.

[0015] The technical solutions provided by the embodiments of the present specification can include the following beneficial effects:

[0016] In the embodiments of the present specification, when determining the recommended order of the change events, the correlation indicators of the change events can be determined, the recommended order of the change events as the change events triggering the system anomaly can be determined based on the correlation indicators of the change events, and the change events can be recommended based on the recommended order. In the determination of the correlation indicators, the system running state data correlation indicators are introduced, which can measure the closeness between the time node of the mutation of the system running state data and the time node of the triggering of the change events, and can quantify the influence of the change events on the system running state, so as to more accurately evaluate the correlation degree between the change events and the system anomaly, and more comprehensively and accurately identify and recommend the change events related to the system anomaly to the operation and maintenance personnel. This method significantly improves the accuracy and efficiency of system anomaly troubleshooting, helps operation and maintenance personnel to quickly focus on the root cause of the problem, and improves the efficiency and accuracy of anomaly troubleshooting.

[0017] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and cannot limit the present specification. BRIEF DESCRIPTION OF DRAWINGS

[0018] Figure 1 FIG. 1 is a schematic diagram of an application scenario provided by an exemplary embodiment.

[0019] Figure 2 FIG. 2 is a flowchart of a change event recommendation method provided by an exemplary embodiment.

[0020] Figure 3 FIG. 3 is a schematic diagram of determining the correlation degree between a change event comprising only a single change operation and system running state data provided by an exemplary embodiment.

[0021] Figure 4 FIG. 4 is a schematic diagram of determining the correlation degree between a change event comprising multiple batches of change operations and system running state data provided by an exemplary embodiment.

[0022] Figure 5 FIG. 1 is a schematic diagram of determining an object correlation index according to an example embodiment.

[0023] Figure 6 FIG. 2 is a schematic diagram of an electronic device according to an example embodiment. DETAILED DESCRIPTION

[0024] The example embodiments will be described in detail herein with reference to the attached drawings. The following description is made with reference to the accompanying drawings in which like reference numerals refer to like elements, unless the context of use indicates otherwise. The following description of example embodiments is not representative of all embodiments consistent with one or more aspects of the present specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of one or more embodiments of the present specification as detailed in the appended claims.

[0025] It should be noted that the steps of the methods in other embodiments are not necessarily performed in the order shown and described in the specification. In some other embodiments, the steps of the methods can be more or less than those described in the specification. Furthermore, a single step described in the specification can be broken down into multiple steps in other embodiments, and multiple steps described in the specification can be combined into a single step in other embodiments.

[0026] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the specification are information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.

[0027] Quick positioning and repair of system anomalies is the key to ensure stable operation of the system, and changes have always been one of the most common problem types that cause system anomalies. Changes include code updates, configuration modifications, hardware upgrades, etc. In complex systems, changes can cause unexpected side effects, thus causing anomalies. In order to enable operation and maintenance personnel to focus on the real anomaly point, the concept of unique root cause is introduced in the emergency positioning field, that is, when receiving an alarm event, the anomaly positioning tool analyzes and summarizes the conclusions of each SPI, summarizes the unique anomaly root cause and recommends it to the user, and classifies the positioning conclusions into levels, such as L1 / L2 / L3, wherein L1 shows which application is abnormal, L2 shows which part of the application is abnormal (such as single machine / OB / message, etc.), and L3 shows which change event causes the anomaly. Change event recommendation belongs to L3 conclusion, that is, to determine which change event is most likely to cause the above system anomaly and recommend it to the operation and maintenance personnel, so as to enable the operation and maintenance personnel to quickly focus on the problem root cause and repair the system anomaly in a timely manner.

[0028] In related technologies, when performing change event recommendation, a plurality of change events related to the alarm event are first determined, and then the recommendation order of the plurality of change events is determined based on the proximity of the trigger time of each change event to the trigger time of the alarm event, that is, the change event with the trigger time close to the trigger time of the alarm event is preferentially recommended. However, many times the time point of the alarm is not the start time point of the anomaly, and the anomaly may have occurred long before the alarm event is triggered, so the change event recommendation scheme based on the time point often has low accuracy and it is difficult to accurately focus on the change event that actually causes the anomaly.

[0029] Based on this, the embodiments of the present specification provide a change event recommendation method, which can first determine a plurality of change events based on an alarm event after the system anomaly and the alarm event are triggered, and determine a correlation index that can measure the correlation degree of each change event to the system anomaly. When determining the correlation index, a system running state data correlation index can be introduced, which is used to represent the proximity of the time node at which the system running state data mutates to the trigger time node of the change event, so as to quantify the influence of the change event on the system running state. By introducing the system running state data correlation index, the change event related to the system anomaly can be more comprehensively and accurately identified and recommended to the operation and maintenance personnel, thereby improving the efficiency and accuracy of anomaly troubleshooting.

[0030] As Figure 1As shown in FIG. 1, which is a schematic diagram of an application scenario of an embodiment of the present specification, a system (hereinafter referred to as a service system) providing various services or functions may need to make various changes in the working process, such as code update, configuration modification, software and hardware upgrade, etc., which may cause the service system to be abnormal. At the same time, in order to ensure the normal operation of the service system, a monitoring system can be used to monitor the running state of each service system, and once an abnormality is found, a corresponding alarm event will be generated and an abnormality positioning system will be notified to position the alarm event. The abnormality positioning system will start the abnormality positioning process after receiving the abnormality positioning task to determine the positioning conclusion of the abnormality and recommend it to the operation and maintenance personnel, so that the operation and maintenance personnel can focus on the problem source based on the positioning conclusion, quickly determine the abnormality reason and repair the abnormality of the service system. For example, when the monitoring system detects that the service system 1 is abnormal, it can trigger a corresponding alarm event and send it to the abnormality positioning system. The abnormality positioning system can determine which applications may cause the abnormality based on the related information of the alarm event (such as alarm time, alarm type, alarm reason, etc.), select one application as the only root cause application, and then determine which object (such as a single machine, an OB, a message, etc.) in the only root cause application is abnormal and determine the only root cause object. Then, the abnormality positioning system can determine multiple change events based on the alarm event, obtain the running state data of the service system 1, determine the correlation index of the system running state data based on the running state data and the trigger time of each change event, and then determine the recommended order of the multiple change events based on the correlation index. Finally, the abnormality positioning system can generate a positioning conclusion based on the determined only root cause application, only root cause object and recommended order of the multiple change events, and display it to the operation and maintenance personnel. For example, the form of the positioning conclusion is as follows: application X; server Y; change event A, change event B, change event C, change event D.

[0031] The change event recommendation method provided by the embodiment of the present specification can be executed by an abnormality positioning system, which can be a positioning tool or platform for positioning and analyzing the abnormality of a system providing various services or functions. The abnormality positioning system can be deployed in an electronic device, which includes but is not limited to a physical server, a server cluster, a cloud server, etc.

[0032] The system in the embodiment of the present specification can be a system for providing various services or functions, such as a transaction system, a shopping system, a data storage system, a data processing system, a distributed computing system, etc.

[0033] As Figure 2 shown, the change event recommendation method can include the following steps:

[0034] S202, determine a plurality of change events based on the alarm event, the alarm event being triggered after detecting the system abnormality;

[0035] In step S202, after detecting the system abnormality, a corresponding alarm event can usually be generated. For example, the running state data of the system can be monitored, and when the running state data meets certain conditions, it is considered that the system is abnormal, and then an alarm event is generated. The abnormality positioning system can determine a plurality of change events based on the related information of the alarm event (such as alarm time, alarm type, alarm reason, etc.). The plurality of change events is a set of change events that may have caused the system abnormality.

[0036] S204, determine a relevance index of each change event in the plurality of change events, wherein the relevance index of each change event at least includes a system running state data relevance index, the system running state data relevance index being used to represent the closeness between the time node of the mutation of the system running state data and the triggering time node of the change event;

[0037] In step S204, after determining the plurality of change events, the relevance index of each change event in the plurality of change events can be determined. The relevance index can be various indexes used to evaluate the degree of relevance between the change event and the system abnormality. In order to accurately evaluate the degree of relevance between the change event and the system abnormality, the relevance index at least includes a system running state data relevance index. The system running state data relevance index is used to represent the closeness between the time node of the mutation of the system running state data and the triggering time node of the change event. The system running state data can be various data used to describe various performance indexes, resource usage, etc. of the system in the running process, which can reflect the health status, performance and running efficiency of the system. For example, the response time, throughput, interface call failure times, storage resource occupancy, etc. of the system. In some embodiments, the system running state data can be the system running state data related to the alarm event.

[0038] Since the mutation time point of the system running state data is usually the time point when the system appears abnormal, if the triggering time point of the change event is close to the mutation time point of the system running state data, it indicates that the change event may have a significant impact on the system in a short time. The closeness in time increases the possibility of the causal relationship between the change event and the system abnormality. Therefore, by representing the relevance between the mutation time point of the system running state data and the triggering time point of the change event, the change event related to the system abnormality can be more accurately determined.

[0039] Of course, in addition to the correlation index of the system's operating status data, other correlation indices that characterize the degree of correlation between change events and system anomalies can also be set, such as the proximity of the trigger time of a change event to the trigger time of an alarm event, etc. The specific indices can be flexibly set based on actual needs, and the embodiments in this specification do not impose any restrictions.

[0040] S206. Based on the correlation index, determine the recommended order of the multiple change events as change events that could trigger the system anomaly, and recommend the multiple change events based on the recommended order.

[0041] In step S206, after determining the relevance indicators of each change event, a recommended order of multiple change events as potential triggers for system anomalies can be determined based on these indicators, and the multiple change events can be recommended according to this recommended order. For example, the relevance of each change event to the system can be determined based on the relevance indicators, and then the multiple change events can be sorted in descending order of relevance to system anomalies as the recommended order, and the multiple change events can be recommended accordingly. For example, the change events ranked higher can be selected for recommendation to operations and maintenance personnel, or all change events can be recommended to operations and maintenance personnel according to the ranking, so that operations and maintenance personnel can prioritize the investigation of the change events ranked higher.

[0042] In some embodiments, to quantify the impact of change events on system operating status, when determining the correlation index of system operating status data, a target time node sequence can first be determined based on the trigger time node of the alarm event. This target time node sequence includes multiple time nodes near the trigger time node of the alarm event. For example, multiple time nodes before the trigger time node of the alarm event can be selected, or multiple time nodes before and after the trigger time node of the alarm event can be selected. The number of time nodes in the target time node sequence can be flexibly set based on actual needs, and this specification does not impose limitations on these embodiments. Then, system operating status data at each time node in the target time node sequence can be obtained to construct a system operating status data sequence, which describes the changing trend of the system operating status. Furthermore, the influence surface data of each change event at each time node in the target time node sequence can be determined to construct an influence surface data sequence. The influence surface data at each time node is used to characterize the degree of impact of the change event on the system at that time node, and the influence surface data sequence describes the changing trend of the influence surface of the change event on the system. After obtaining the system operation status data sequence and the influence surface data sequence, the correlation between the system operation status data sequence and the influence surface data sequence can be calculated, which serves as the correlation index of the system operation status data.

[0043] For example, assuming that the trigger time node of the alarm event is T 10 , the target time node sequence {T3, T4, T5, T6, T7, T8, T9, T 10 , T 11 , T 12} can be determined, and then the running state data of the system at each time node in the target time node sequence can be obtained. Assuming that the running state data is the number of XX interface call failures, the running state data sequence {N1, N2, N3, N4, N5, N6, N7, N8, N9, N 10 , N 11} can be obtained, and then the influence area data of the change event A at each time node in the target time node sequence can be determined, and the influence area data sequence {M1, M2, M3, M4, M5, M6, M7, M8, M9, M 10 , M 11} can be obtained, and then the correlation degree of the two sequences can be calculated as the system running state data correlation index.

[0044] By constructing the system running state data sequence and the influence area data sequence, and calculating their correlation degree, the influence of the change event on the system running state can be quantitatively evaluated. By analyzing the consistency of the change trend of the system running state and the change trend of the influence area of the change event, the correlation degree between the change event and the system abnormality can be more accurately evaluated, and the accuracy and efficiency of fault troubleshooting are improved.

[0045] In some embodiments, when determining the influence area data of each change event at each time node in the target time node sequence, if the change event only includes a single change operation, the influence area data of the time nodes before the trigger time node of the change event in the target time node sequence is a preset minimum influence degree value, the influence area data of the non-first time nodes after the trigger time node of the change event in the target time node sequence is a preset maximum influence degree value, and the influence area data of the first time node after the trigger time node of the change event in the target time node sequence is the product of the maximum influence degree value and a target proportion, wherein the target proportion is negatively correlated with the proximity of the trigger time node of the change event to the first time node. The minimum influence degree value and the maximum influence degree value can be a numerical value for measuring the influence degree of the change event on the system. Considering that normalization processing is performed when calculating the correlation degree of the two sequences, the specific value of the minimum influence degree value and the maximum influence degree value has little effect on the calculation result, as long as it can represent the relative relationship of different influence degrees. For example, the minimum influence degree value can be set to 0, the maximum influence degree value can be set to 1, and the intermediate influence degree value can be set to a value between 0 and 1.

[0046] For example, as shown in Figure 3 The association between the change event including only a single change operation and the system running state data is shown. For example, the alarm time node triggering the alarm event is time node a, a target time node sequence including 10 time nodes before time node a and 1 time node after time node a, that is, 11 time nodes in total can be constructed. The change event triggering time node is located between time nodes a-7 and a-6, therefore, the influence range data of time nodes a-9, a-8, a-7 before the triggering time node of the change event is the minimum influence degree value (assuming 0), the first time node a-6 after the triggering time node of the change event is between the minimum influence degree value and the maximum influence degree value (assuming 100), and the influence degree value can be determined based on the proximity of the triggering time node of the change event and the first time node, for example, if the triggering time node of the change event is closer to time node a-6, it means that the influence duration on the time node is shorter, and the influence degree value is smaller. Therefore, the target proportion can be determined based on the proximity of the triggering time node of the change event and the first time node, and then the influence range data of time node a-6 is obtained by multiplying the maximum influence degree value by the target proportion, assuming that the influence range data of time node a-6 is 80. The influence range data of time nodes a-5, a-4, a-3, a-2, a-1, a, a+1 after the triggering time node of the change event is the maximum influence degree value (100). Thus, the influence range data sequence of the change event {0, 0, 0, 80, 100, 100, 100, 100, 100, 100, 100} can be constructed.

[0047] Meanwhile, the system running state data of the system at the 11 time nodes can be obtained, assuming that it is the number of XX interface call failures, and the system running state data sequence {500, 520, 530, 650, 880, 980, 1000, 1000, 1100, 1011, 1010} can be constructed. Then the correlation of the two sequences can be calculated as the system running state data correlation index. For example, the Pearson correlation coefficient (0.921625073742963) of the two sequences can be calculated as the system running state data correlation index.

[0048] By accurately setting the influence range data of the change event at different time nodes, the dynamic influence of a single change operation on the system running state can be effectively quantified. In particular, by setting the influence range before the change as the minimum influence degree value, the influence range of the first time node after the change as the product of the maximum influence degree value and the target proportion, and the influence range of the subsequent time nodes as the maximum influence degree value, the change trend of the influence range of the change event, i.e., the starting, developing and stabilizing stages, can be accurately depicted. This detailed influence range data construction method ensures the accuracy in calculating the correlation degree of the two sequences, so that the influence trend of the change event can be more accurately reflected when evaluating the association between the change event and the system anomaly, thereby improving the efficiency and accuracy of fault troubleshooting.

[0049] In some embodiments, when determining the influence range data of each change event at each time node in the target time node sequence, if the change event includes multiple batches of change operations, the influence range data of the time nodes before the trigger time node of the change event in the target time node sequence is a preset minimum influence degree value, and the influence range data of each time node after the trigger time node of the change event in the target time node sequence is the cumulative value of the influence degree values corresponding to each batch of change operations that has occurred before the time node. For example, after obtaining the change operation related to the alarm event, the change operation can be processed first, and multiple change operations of the same type can be merged into one change event. For example, assuming that a distributed system includes multiple machines, if the hardware of the multiple machines is upgraded, and the hardware upgrade is performed in multiple batches, such as upgrading two machines each time, then the multiple batches of hardware upgrades (i.e., change operations) can be merged into one change event including multiple batches of change operations. Assuming that each change operation has a continuous influence on the system, for the change event including multiple batches of change operations obtained by merging, the influence of the change event on the system is continuously added.

[0050] For example, as shown in FIG. 4, the change event 401 includes three batches of change operations, i.e., the change operation 402, the change operation 403 and the change operation 404. The influence range of the change operation 402 is the influence range of the change operation 402, the influence range of the change operation 403 is the influence range of the change operation 403, and the influence range of the change operation 404 is the influence range of the change operation 404. The influence range of the change event 401 is the cumulative value of the influence ranges of the change operation 402, the change operation 403 and the change operation 404. Figure 4As shown, the association between the change event including multiple batches of change operations and the system running state data is demonstrated. It is assumed that the change event includes 4 batches of change operations, each of which changes 2 machines in the system. It is assumed that the alarm time node triggering the alarm event is time node a, and a target time node sequence can be constructed, which includes 10 time nodes before time node a and 1 time node after time node a, a total of 11 time nodes. Since the trigger time node of the first batch of change operations is located between time nodes a-9 and a-8, the influence range data of time node a-9 before the trigger time node of the change event can be set to the minimum influence degree value (assuming 0). The influence range data of each time node after the trigger time node of the change event is the cumulative value of the influence degree values corresponding to each batch of change operations that has occurred before the time node, for example, considering that each batch of change operations will gradually expand the influence range, because each batch changes two machines, so the influence range is also accumulated by 2, and thus the final influence range data sequence is: {0, 2, 2, 4, 4, 6, 6, 6, 8, 8, 8}.

[0051] Similarly, the system running state data of the system at the 11 time nodes can be obtained, which is assumed to be the number of XX interface call failures, and the system running state data sequence {500, 600, 630, 870, 880, 900, 930, 930, 1000, 1030, 1050} can be constructed. Then the correlation of the two sequences can be calculated as the system running state data correlation index. For example, the Pearson correlation coefficient (0.9655568584198055) of the two sequences can be calculated as the system running state data correlation index.

[0052] In some embodiments, the correlation of the system running state data sequence and the influence range data sequence can be represented by the Pearson correlation coefficient. The Pearson correlation coefficient, also known as the Pearson product-moment correlation coefficient, is a statistical index used to measure the degree of linear correlation between two variables. Its value ranges from -1 to 1, and is defined as follows: where 1 indicates complete positive linear correlation, meaning that when one variable increases, the other variable also increases completely in proportion, -1 indicates complete negative linear correlation, meaning that when one variable increases, the other variable completely decreases in proportion, and 0 indicates no linear correlation, meaning that there is no linear relationship between the variables.

[0053] By using the Pearson correlation coefficient to measure the degree of linear correlation between the system running state data sequence and the influence surface data sequence, the correlation between the change event and the system anomaly can be accurately evaluated in a quantitative manner. This quantitative method not only provides a clear correlation strength evaluation, but also allows objective comparison through standardized statistical indicators, enabling quick identification of which change events have a significant linear relationship with system anomalies.

[0054] In some embodiments, the correlation indicator of each change event further includes a positioning result correlation indicator for characterizing the degree of correlation between the abnormal positioning determination of the alarm event and the change event. Generally, after receiving an alarm event, abnormal positioning analysis can be performed on the alarm event to determine which applications and which objects in the applications have abnormalities. Therefore, when determining the correlation indicator of the change event, a positioning result correlation indicator for characterizing the degree of correlation between each change event and the determined abnormal applications can also be determined, and the recommended order of the change event is determined based on the indicator. For example, assuming that abnormal positioning is performed on the alarm event and it is determined that the abnormal applications include abnormal application X1, abnormal application X2, and abnormal application X3, the association between the change event and these abnormal applications can be analyzed, such as whether the change event will affect these abnormal applications. If the change event does not affect these abnormal applications, i.e., has no association with these abnormal applications, it means that the change event has little correlation with the system anomaly, and vice versa, which means that the change event is likely to be the change event that triggers the system anomaly.

[0055] By introducing the positioning result correlation indicator, the degree of association between the change event and the system anomaly can be refined. By combining the system running state data correlation indicator and the positioning result correlation indicator to evaluate the correlation between the change event and the system anomaly, the association between the change event and the system anomaly can be comprehensively evaluated in multiple dimensions, and the accuracy of change event recommendation can be improved.

[0056] In some embodiments, the positioning result correlation indicator includes an object correlation indicator and / or a type correlation indicator, wherein the object correlation indicator is used to characterize the consistency between the abnormal objects in the abnormal applications and the objects affected by the change event, and the type correlation indicator is used to characterize the degree of correlation between the abnormal types of the abnormal applications and the change types of the change event.

[0057] For example, for each abnormal application, its abnormal type can be determined, and then the correlation degree of the change type of the change event and the abnormal type can be analyzed. The abnormal type refers to the specific category of problems or failures that occur in the system during operation, which is usually classified according to the manifestation or impact range of the problem, such as performance degradation, high resource occupancy, configuration error, function failure, etc. The change type refers to the specific category of modifications or updates made to the system, which is usually classified according to the nature and impact range of the change, such as code modification, configuration modification, hardware upgrade, etc. By analyzing the correlation degree of the abnormal type and the change type, the association between the change event and the system abnormality can be reflected from a higher level.

[0058] The abnormal object can be a functional module, a single machine, a message, etc. in the abnormal application. The objects that will be affected by the change event, such as which functional modules, single machines, etc. can also be determined. Therefore, by analyzing the consistency between the abnormal objects in the abnormal application and the objects affected by the change event, the association between the change event and the system abnormality can be reflected from a more specific and detailed level. For example, suppose that through abnormal positioning, it is determined that the abnormal objects existing in the abnormal application X include objects Y1, Y2, and Y3, and through analysis of the influence of the change event, it is determined that the objects affected by the change event include objects Y1, Y2, Y3, and Y4. The consistency between the two is high, indicating that the change event is likely to be the change event that triggered the abnormality.

[0059] As shown in Figure 5 , assuming that the alarm event is analyzed by abnormal positioning, the abnormal object set determined includes N abnormal objects, and the influence object set affected by the change event A has N1 overlapping objects with the abnormal object set, then the object correlation index = N1 / N.

[0060] By further subdividing the positioning result correlation index into object correlation index and type correlation index, the correlation degree between the change event and the system abnormality can be reflected from different levels. This subdivision can quantify the relationship between the change event and the abnormality from multiple dimensions, thereby making more accurate change event recommendations.

[0061] In some embodiments, when determining the object correlation indicator, an impact object set composed of objects affected by the change event and an abnormal object set composed of abnormal objects in the abnormal application can be determined respectively, and for overlapping objects in the impact object set and the abnormal object set, a proportion of the number of the overlapping objects to the total number of abnormal objects in the abnormal object set can be determined as the object correlation indicator. Wherein, considering that different change events affect different objects, by considering the consistency of the located abnormal objects and the objects affected by the change event, the association between the change event and the system abnormality can be better measured. And by calculating the proportion of the number of overlapping objects between the impact object set and the abnormal object set, the object correlation between the change event and the system abnormality can be quantitatively evaluated. This method not only considers the consistency of the objects affected by the change event and the abnormal objects, but also reflects the overlapping degree in the form of proportion, so as to more accurately measure the association strength between the change event and the system abnormality.

[0062] In some embodiments, the abnormal application includes a plurality of applications, and the object correlation indicator includes a unique root cause object correlation indicator and a total object correlation indicator, wherein the unique root cause object correlation indicator is used to represent the consistency of abnormal objects in the unique root cause application and the objects affected by the change event, and the total object correlation indicator is used to represent the consistency of all abnormal objects in the plurality of abnormal applications and the objects affected by the change event.

[0063] For example, for alarm event abnormal positioning analysis, it is assumed that three applications X1, X2 and X3 are currently determined to possibly exist abnormality. The abnormal object set composed of objects possibly existing abnormality in the three abnormal applications can be determined first, such as objects 1, 2, 3, 4, 5 and 6 in the abnormal object set. Then, the impact object set composed of objects affected by the change event A is determined, such as objects 1, 4 and 5 in the impact object set. Among them, there are three overlapping objects (objects 1, 4 and 5), and the abnormal object correlation indicator of the change event A is 3 / 6. It is assumed that the unique root cause application determined from the three abnormal applications is application X1, and the objects possibly existing abnormality in application X1 are objects 1, 2 and 3. Among them, there is one overlapping object (object 1) with the objects affected by the change event A, and the unique root cause content correlation is 1 / 6.

[0064] Similarly, the type-related indicators include a unique root cause type-related indicator and an overall type-related indicator, the unique root cause type-related indicator being used to represent the degree of correlation between the unique root cause application and the change type of the change event, and the overall type-related indicator being used to represent the degree of correlation between each of the abnormal application and the change type of the change event.

[0065] By introducing the unique root cause object-related indicator and the overall object-related indicator, and the unique root cause type-related indicator and the overall type-related indicator, the degree of correlation between the change event and the system abnormality can be comprehensively evaluated from multiple levels. By introducing the unique root cause object-related indicator (unique root cause type-related indicator) and the overall object-related indicator (overall type-related indicator), the degree of correlation between the change event and the system abnormality can be comprehensively evaluated from both local and global levels. The unique root cause object-related indicator and the unique root cause type-related indicator focus on the core application (unique root cause application) that is most likely to cause the system abnormality, ensuring that the most critical problem source is prioritized; while the overall object-related indicator and the overall type-related indicator comprehensively consider the abnormal objects and types in all abnormal applications, providing a comprehensive perspective to prevent missing other possible correlations. This dual evaluation mechanism enables the operation and maintenance personnel to more accurately identify the change event most relevant to the system abnormality. Not only the correlation between the change event and the unique root cause application that is most likely to cause the system abnormality is considered, but also the overall correlation between the change event and the multiple potential abnormal applications determined by the abnormality positioning is considered. By evaluating the correlation between the change event and the unique root cause application and the overall correlation between the change event and the multiple potential abnormal applications respectively, the balance between local and global correlation analysis can be better achieved, and the accuracy of the correlation between the change event and the system abnormality determined finally can be improved.

[0066] In some embodiments, when determining the recommended order of the multiple change events as the change events that cause the system abnormality based on the correlation indicators, for each change event, a weight corresponding to each correlation indicator of the change event can be determined, where the weight is used to indicate the degree of influence of the correlation indicator on the recommended order of the change event. For example, the higher the risk level of the change event, i.e., the more likely the change event is to cause the abnormality, the more the change event needs to be recommended first, and the higher the reliability of the correlation indicator, the greater the weight of the correlation indicator should be. Then the correlation indicators of the change event can be weighted and averaged based on the weight to obtain an ordering value of the change event, and then the multiple change events can be sorted in descending order of the ordering value as the final recommended order.

[0067] By introducing a weight mechanism, the weights of different types of correlation indicators of different change events can be adaptively adjusted based on the characteristics of different change events and the importance of the correlation indicators, and then the correlation indicators of each change event are weighted and averaged to obtain a ranking value. The ranking value determined in this way can more accurately represent the degree of association between the change event and the system anomaly, and the recommendation order of the change event determined based on the ranking value is also more accurate, so that the operation and maintenance personnel can more efficiently focus on the change event that is most likely to cause the system anomaly, significantly improving the efficiency and accuracy of fault troubleshooting.

[0068] In some embodiments, the weight can be determined based on one or more of the following information: application type associated with the change event, execution object of the change event, running environment of the change event, risk level of the change event, determination method of the change event impact, reliability of each correlation indicator, etc.

[0069] In some embodiments, the influence of the application type associated with the change event on the weight is as follows: the weight of the change event associated with the unique root cause application > the weight of the change event associated with the alarm application triggering the alarm event > the weight of the change event associated with the specific dimension application related to the alarm event > the weight of the change event associated with the application with an anomaly determined in the abnormal positioning of the alarm event; wherein the unique root cause application is the application that is most likely to cause the system anomaly. Considering that the unique root cause application is the application that is most likely to cause the system anomaly, the change event associated with it should be given priority, and therefore the weight is the largest. The specific dimension application related to the alarm event is usually used to describe the impact range of the alarm event in a specific dimension or functional module. By setting a larger weight for the change event of such an application, the problem range can be more accurately located, helping the operation and maintenance personnel to quickly focus on the specific module or function affected.

[0070] In some embodiments, the weight of the correlation indicators of the change event can also be determined based on the position of the change event in the Trace link. The position in the Trace link refers to the specific position of the change event in the system call link. The Trace link is usually used to describe the calling relationship between different components in the system, helping to track the propagation path of the problem. By analyzing the position of the change event in the Trace link, the influence of the change event on the overall behavior of the system can be evaluated. For example, if the change event is located upstream of the critical call link, it may have a chain reaction on multiple downstream components, and therefore its weight may be higher.

[0071] In some embodiments, the influence of the execution object of the change event on the weight is as follows: the weight of the change event with the execution object being a person is greater than the weight of the change event with the execution object being a system. Since the change operation performed by a person usually involves more uncertainty and potential risks, the weight is higher.

[0072] In some embodiments, the influence of the running environment of the change event on the weight is as follows: the running environment consistent with the running environment corresponding to the alarm event > the online environment > the gray environment > the pre-release environment. If the change event occurs in the same running environment as the alarm event, it is more likely to be the cause of the anomaly. In addition, the change event of the online environment usually has a greater impact on the system, the change event of the gray environment has a relatively small impact range, and the change event of the pre-release environment usually does not directly affect the online environment, so the weight is the smallest.

[0073] In some embodiments, the higher the risk level of the change event, the greater the weight of the correlation indicator of the change event. High-risk change events (such as code updates involving critical functions) are more likely to trigger system anomalies than low-risk change events (such as document updates), so the weight is higher.

[0074] By comprehensively considering the above information to determine the weight, the degree of influence of each change event on the system anomaly can be more comprehensive and accurate. This method not only considers the properties of the change event (such as application type, execution object, running environment, risk level, determination method of impact surface, etc.), but also considers the importance and reliability of each correlation indicator. This multi-dimensional weight determination method can more accurately reflect the actual impact of the change event, making the recommended order more reasonable and reliable.

[0075] In some embodiments, the correlation indicators include system running state data correlation indicators, unique root cause object correlation indicators, all object correlation indicators, unique root cause type correlation indicators, and all type correlation indicators. When determining the ranking values of each change event, Borda counting can be used to determine the ranking values. Borda counting is a general voting scoring rule and a multiple-choice voting system used to determine one or more winners among multiple candidates. For each correlation indicator, its Borda ranking value can be determined in the following manner, taking the system running state data correlation indicator as an example:

[0076] Borda ranking value = system running state data correlation indicator * weight corresponding to application type * weight corresponding to execution object * weight corresponding to running environment * weight corresponding to risk level of change event

[0077] Assuming there are a total of 5 changes, the number of votes for the system running state data correlation indicator of the change event ranked first according to the Borda ranking value = 5, and the number of votes for the system running state data correlation indicator of the change event ranked second = 4. If the Borda ranking values of two change events are the same, the number of votes for their system running state data correlation indicators is equal.

[0078] The voting numbers of the unique root cause object correlation index, the total object correlation index, the unique root cause type correlation index, and the total type correlation index can be calculated in the above manner in turn. The final recommended ranking value of each change event is the total number of votes of each correlation index:

[0079] The total number of votes = system running state data correlation index + unique root cause object correlation index + total object correlation index + unique root cause type correlation index + total type correlation index

[0080] Then, the change events can be ranked and recommended based on the total number of votes.

[0081] In some embodiments, the plurality of change events includes a first type of change event related to an alarm application that triggered the alarm event, and a second type of change event related to an exception application that determined the abnormal positioning of the alarm event. Before determining the recommended order of the change events, the related change events are usually recalled. In related technologies, only high-risk change events related to the alarm application that triggered the alarm event are usually recalled. This recall method often results in a narrow range of recalled change events, which can easily miss the change events that actually caused the system exception. The embodiments of the present specification not only recall the first type of change event related to the alarm application that triggered the alarm event, but also recall the second type of change event related to the exception application that determined the abnormal positioning of the alarm event. The related change events are determined and recalled based on the abnormal positioning result obtained by analysis. For example, the range of recalled change events is expanded from the previous alarm application related to the positioning result related change events. The range of change events recalled is wider, not limited to high-risk changes, and low-risk changes are also within the recall range. The range of recalled change events can be expanded to comprehensively screen various change events that may cause system exceptions, and improve the recommendation accuracy.

[0082] In addition, the present application also provides a change event recommendation method, which comprises the following steps:

[0083] determining a plurality of change events based on an alarm event, the alarm event being triggered after detecting a system exception;

[0084] determining a correlation index of each change event in the plurality of change events, wherein the correlation index of each change event at least includes a positioning result correlation index, the positioning result correlation index being used to represent the correlation degree between an exception application that determines the abnormal positioning of the alarm event and the change event;

[0085] determine a recommended order of the plurality of change events as change events that cause the system anomaly based on the correlation indicators, and recommend the plurality of change events based on the recommended order.

[0086] The specific determination process of the positioning result correlation indicators can refer to the description in the above embodiments, which will not be described here.

[0087] Various technical features in the above embodiments can be combined in any manner as long as there is no conflict or contradiction between the features. However, due to the limited space, not all possible combinations are described, and any combination of the technical features in the above embodiments is within the scope of the present disclosure.

[0088] The above describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than the order described in the embodiments and still achieve the desired results. In addition, the processes depicted in the figures do not necessarily require the particular order shown or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing can be utilized or can be advantageous.

[0089] In some embodiments, the embodiments of the present disclosure also provide an electronic device, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor implements the method of any one of the above by running the executable instructions.

[0090] Figure 6 is a schematic structural diagram of an electronic device provided by an example embodiment. Please refer to Figure 6 At the hardware level, the electronic device includes a processor 602, an internal bus 604, a network interface 606, a memory 608, and a non-volatile memory 610, and of course can also include other hardware required for functionality. One or more embodiments of the present disclosure can be implemented in a software manner, such as by the processor 602 reading a corresponding computer program from the non-volatile memory 610 into the memory 608 and then running it. Of course, in addition to the software implementation, one or more embodiments of the present disclosure do not exclude other implementation manners, such as logic devices or a combination of software and hardware, etc. That is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or a logic device.

[0091] In some embodiments, the change event recommendation device can be applied to an electronic device as shown in Figure 6 The change event recommendation device can include:

[0092] a change event determination module configured to determine a plurality of change events based on an alarm event, the alarm event being triggered after detecting a system abnormality;

[0093] a correlation index determination module configured to determine a correlation index of each change event in the plurality of change events, wherein the correlation index of each change event comprises at least a system running state data correlation index, the system running state data correlation index being used to represent a closeness between a time node at which a system running state data mutates and a time node at which the change event is triggered;

[0094] a recommendation module configured to determine a recommended order of the plurality of change events as change events that cause the system abnormality based on the correlation index, and recommend the plurality of change events based on the recommended order.

[0095] In some embodiments, the system running state data correlation index is determined based on the following manner:

[0096] determining a target time node sequence based on a time node at which the alarm event is triggered, the target time node sequence comprising a plurality of time nodes around the time node at which the alarm event is triggered;

[0097] obtaining system running state data at each time node in the target time node sequence to construct a system running state data sequence;

[0098] determining influence area data of each change event at each time node in the target time node sequence to construct an influence area data sequence, wherein the influence area data is used to represent an influence degree of the change event on the system;

[0099] calculating a correlation degree between the system running state data sequence and the influence area data sequence as the system running state data correlation index.

[0100] In some embodiments, the determining of the influence area data of each change event at each time node in the target time node sequence comprises:

[0101] If the change event only includes a single change operation, the influence scope data of the time nodes in the target time node sequence before the trigger time node of the change event is a preset minimum influence degree value, the influence scope data of the time nodes in the target time node sequence after the first time node after the trigger time node of the change event is a preset maximum influence degree value, and the influence scope data of the first time node after the trigger time node of the change event is a product of the maximum influence degree value and a target proportion, wherein the target proportion is negatively correlated with the proximity between the trigger time node of the change event and the first time node.

[0102] and / or

[0103] If the change event includes multiple batches of change operations, the influence scope data of the time nodes in the target time node sequence before the trigger time node of the change event is a preset minimum influence degree value, and the influence scope data of each time node in the target time node sequence after the trigger time node of the change event is an accumulated value of the influence degree values corresponding to the change operations of each batch of change operations that have occurred before the time node.

[0104] In some embodiments, the relevance between the system running state data sequence and the influence scope data sequence is represented by a Pearson correlation coefficient.

[0105] In some embodiments, the relevance indicator of each change event further includes a positioning result relevance indicator, which is used to represent the relevance between the abnormal application of the abnormal positioning determination of the alarm event and the change event.

[0106] In some embodiments, the positioning result relevance indicator includes an object relevance indicator and / or a type relevance indicator, the object relevance indicator is used to represent the consistency between the abnormal object existing in the abnormal application and the object affected by the change event, and the type relevance indicator is used to represent the relevance between the abnormal type of the abnormal application and the change type of the change event.

[0107] In some embodiments, the object relevance indicator is determined based on the following manner:

[0108] respectively determining an influence object set composed of the objects affected by the change event and an abnormal object set composed of the abnormal objects existing in the abnormal application;

[0109] For the overlapping objects in the influence object set and the abnormal object set, determining the proportion of the number of the overlapping objects to the total number of the abnormal objects in the abnormal object set as the object relevance indicator.

[0110] In some embodiments, the multiple abnormal applications include a unique root cause application and multiple non-root cause applications, the object relevance indicators include a unique root cause object relevance indicator and a total object relevance indicator, the unique root cause object relevance indicator is used to represent consistency between an abnormal object in the unique root cause application and an object affected by the change event, and the total object relevance indicator is used to represent consistency between all abnormal objects in the multiple abnormal applications and the object affected by the change event.

[0111] The type relevance indicators include a unique root cause type relevance indicator and a total type relevance indicator, the unique root cause type relevance indicator is used to represent a degree of relevance between an abnormal type of the unique root cause application and a change type of the change event, and the total type relevance indicator is used to represent a degree of relevance between respective abnormal types of the multiple abnormal applications and the change type of the change event.

[0112] The unique root cause application is an application most likely to cause the system abnormality in the multiple abnormal applications.

[0113] In some embodiments, when the recommendation module is used to determine a recommended order of the multiple change events as change events causing the system abnormality based on the relevance indicators, the recommendation module is specifically configured to:

[0114] For each change event, determine a weight corresponding to each relevance indicator of the change event, and perform weighted average processing on each relevance indicator of the change event based on the weight to obtain an ordering value of the change event.

[0115] Sort the multiple change events in descending order of the ordering values as the recommended order.

[0116] The weight is determined based on one or more of the following information:

[0117] An application type associated with a change event, an execution object of a change event, a running environment of a change event, a risk level of a change event, a determination manner of a change event impact surface, and a reliability of the relevance indicators.

[0118] In some embodiments, the application type associated with the change event affects the weight as follows: a weight of a change event associated with a unique root cause application > a weight of a change event associated with an alarm application triggering the alarm event > a weight of a change event associated with an application of a specific dimension related to the alarm event > a weight of a change event associated with an application determined to have an abnormality based on the system positioning result; and the unique root cause application is an application most likely to cause the system abnormality.

[0119] The influence of the execution object of the change event on the weight is as follows: the weight of a change event whose execution object is a person is greater than the weight of a change event whose execution object is a system;

[0120] The influence of the running environment of the change event on the weight is as follows: the running environment corresponding to the alarm event is consistent with the running environment > online environment > gray environment > pre-release environment;

[0121] The higher the risk level of the change event, the greater the weight of the change event.

[0122] In some embodiments, the plurality of change events includes a first type of change event related to an alarm application that generates the alarm event, and a second type of change event related to an exception application that determines an exception positioning of the alarm event. The functions and roles of the various modules in the above device are implemented in the implementation process of the corresponding steps in the above method, which will not be described here.

[0123] Based on the same idea as the above method, the present specification also provides a computer readable storage medium, which stores computer instructions, and the instructions are executed by a processor to implement the steps of the method according to any of the above embodiments.

[0124] The computer readable medium includes permanent and non-permanent, removable and non-removable media, which can be implemented by any method or technology to store information. The information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, disk storage, quantum memory, graphene-based storage medium or other magnetic storage device, or any other non-transmission medium that can be used to store information that can be accessed by a computing device. According to the definition in this paper, computer readable medium does not include transitory computer readable medium, such as modulated data signals and carriers.

[0125] Based on the same idea as the above method, the present specification also provides a computer program product, which includes computer program / instructions, and the instructions are executed by a processor to implement the steps of the method according to any of the above embodiments.

[0126] The above description is only the preferred embodiment of one or more embodiments of the specification, and is not used to limit one or more embodiments of the specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of the specification should be included in the protection range of one or more embodiments of the specification.

Claims

1. A method for recommending change events, comprising: determining a plurality of change events based on an alarm event, the alarm event being triggered after detecting a system abnormality; determining a relevance indicator of each change event in the plurality of change events, wherein the relevance indicator of each change event comprises at least a system running state data relevance indicator, the system running state data relevance indicator being used to represent a closeness between a time node at which a system running state data mutates and a time node at which the change event is triggered; determining a recommended order of the plurality of change events as change events that cause the system abnormality based on the relevance indicators, and recommending the plurality of change events based on the recommended order.

2. The method of claim 1, wherein the system running state data relevance indicator is determined based on the following manner: determining a target time node sequence based on the time node at which the alarm event is triggered, the target time node sequence comprising a plurality of time nodes around the time node at which the alarm event is triggered; obtaining system running state data at each time node in the target time node sequence to construct a system running state data sequence; determining influence face data of each change event at each time node in the target time node sequence to construct an influence face data sequence, wherein the influence face data is used to represent an influence degree of the change event on the system; calculating a correlation degree between the system running state data sequence and the influence face data sequence as the system running state data relevance indicator.

3. The method of claim 2, wherein the determining of the influence face data of each change event at each time node in the target time node sequence comprises: if the change event comprises only a single change operation, the influence face data of a time node before the time node at which the change event is triggered in the target time node sequence is a preset minimum influence degree value, the influence face data of a time node after the first time node at which the change event is triggered in the target time node sequence is a preset maximum influence degree value, and the influence face data of the first time node after the time node at which the change event is triggered in the target time node sequence is a product of the maximum influence degree value and a target proportion, wherein the target proportion is negatively correlated with a closeness between the time node at which the change event is triggered and the first time node; and / or if the change event comprises a plurality of batches of change operations, the influence face data of a time node before the time node at which the change event is triggered in the target time node sequence is a preset minimum influence degree value, and the influence face data of each time node after the time node at which the change event is triggered in the target time node sequence is a cumulative value of influence degree values corresponding to the change operations of each batch that have occurred before the time node.

4. The method of claim 2 or 3, wherein the correlation degree between the system running state data sequence and the influence face data sequence is represented by a Pearson correlation coefficient. ​ 5. The method of claim 1, wherein the relevance indicator of each change event further comprises a positioning result relevance indicator, the positioning result relevance indicator being used to represent a degree of relevance of an abnormal application of an abnormal positioning determination of the alarm event to the change event.

6. The method of claim 5, wherein the positioning result relevance indicator comprises an object relevance indicator and / or a type relevance indicator, the object relevance indicator being used to represent a consistency of an abnormal object existing in the abnormal application and an object affected by the change event, and the type relevance indicator being used to represent a degree of relevance of an abnormal type of the abnormal application to a change type of the change event.

7. The method of claim 6, wherein the object relevance indicator is determined based on the following manner: determining, respectively, an impact object set composed of objects affected by the change event, and an abnormal object set composed of abnormal objects existing in the abnormal application; determining, for overlapping objects in the impact object set and the abnormal object set, a proportion of a number of the overlapping objects to a total number of abnormal objects in the abnormal object set as the object relevance indicator.

8. The method of claim 6, wherein the abnormal application comprises a plurality of abnormal applications, the object relevance indicator comprises a unique root cause object relevance indicator and a total object relevance indicator, the unique root cause object relevance indicator being used to represent a consistency of an abnormal object in a unique root cause application and an object affected by the change event, and the total object relevance indicator being used to represent a consistency of all abnormal objects in the plurality of abnormal applications and the object affected by the change event; the type relevance indicator comprises a unique root cause type relevance indicator and a total type relevance indicator, the unique root cause type relevance indicator being used to represent a degree of relevance of an abnormal type of a unique root cause application to a change type of the change event, and the total type relevance indicator being used to represent a degree of relevance of respective abnormal types of the plurality of abnormal applications to the change type of the change event; wherein, the unique root cause application is an application in the plurality of abnormal applications that is most likely to cause the system abnormality.

9. The method of claim 1, wherein the determining, based on the relevance indicators, a recommended order of the plurality of change events as change events causing the system abnormality comprises: for each change event, determining a weight corresponding to each relevance indicator of the change event, and performing a weighted average processing on each relevance indicator of the change event based on the weight to obtain an ordering value of the change event; ordering the plurality of change events in a descending order of the ordering values as the recommended order; and wherein the weight is determined based on one or more of the following information: an application type associated with the change event, an execution object of the change event, a running environment of the change event, a risk level of the change event, a determination manner of an impact surface of the change event, and a reliability of the relevance indicator.

10. The method of claim 9, wherein, The application type of the change event associated with the weight influences as follows: the weight of the change event associated with the unique root cause application > the weight of the change event associated with the alarm application triggering the alarm event > the weight of the change event associated with the specific dimension related to the alarm event > the weight of the change event associated with the application determined to have an anomaly based on the system positioning result; wherein the unique root cause application is the application most likely to cause the system anomaly; The execution object of the change event influences the weight as follows: the weight of the change event with the execution object being a human is greater than the weight of the change event with the execution object being a system; The running environment of the change event influences the weight as follows: the running environment consistent with the running environment corresponding to the alarm event > the online environment > the gray-scale environment > the pre-release environment; The higher the risk level of the change event, the greater the weight of the change event. 11.The method of claim 1, wherein the plurality of change events include a first type of change event related to an alarm application generating the alarm event, and a second type of change event related to an anomaly application determining an anomaly positioning of the alarm event.

12. An electronic device comprising: a processor; a memory for storing processor-executable instructions; wherein the processor implements the steps of the method of any one of claims 1-11 by running the executable instructions. 13.A computer readable storage medium having stored thereon computer instructions that, when executed by a processor, implement the steps of the method of any one of claims 1-11. 14.A computer program product comprising computer program / instructions that, when executed by a processor, implement the steps of the method of any one of claims 1-11.