Business abnormality processing method, device, computer equipment and storage medium
By obtaining and analyzing business data, preset hierarchical structure and business topology, the causes of business abnormalities in microservice design are accurately detected, and the problem of too many alerts in the system monitoring service is solved, and the effect of targeted processing of abnormal data is achieved to restore normal business operation.
Patent Information
- Application Number
- CN202010289094.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-04-14
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2040-04-14
AI Technical Summary
In the backend system designed by microservices, the system monitoring service faces the problem of a large number of alarms that cannot determine the root cause of the alarm being issued and cannot effectively handle exception information.
By acquiring business data, determining the exception data set, obtaining preset hierarchical structures and business topology, the target exception data that allows the alarm to be triggered and the reason for triggering the alarm is determined based on these structures.
Accurately detect the causes of business abnormalities, avoid unnecessary alarms, and process abnormal data in a targeted manner to restore normal business operation.
Smart Images

Figure CN113535443B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method, apparatus, computer equipment and storage medium for processing business exceptions. Background Art
[0002] With the widespread application of microservice design, the complexity of backend system design has increased dramatically. System monitoring services are facing increasing challenges, such as continuous alarm harassment of abnormalities within the system security range, frequent alarms for abnormalities that can be self-healed, and some sensitive abnormalities being ignored. However, when a large number of alarms appear, it is impossible to determine the root cause of the alarm, resulting in the inability to process the abnormal information that caused the alarm. Summary of the invention
[0003] Based on this, it is necessary to provide a method, device, computer equipment and storage medium for processing business anomalies in response to the above technical problems, which can accurately detect the causes of business anomalies.
[0004] A method for handling business exceptions, the method comprising:
[0005] Acquire business data, and determine abnormal data sets in the business data;
[0006] Acquire a preset hierarchical structure, and determine target abnormal data in the abnormal data set that is allowed to trigger an alarm according to the preset hierarchical structure;
[0007] A service topology structure is obtained, and a reason why the target abnormal data triggers an alarm is determined based on the service topology structure.
[0008] In one embodiment, determining the reason why the target abnormal data triggers an alarm according to the historical abnormal curve and the abnormal curve corresponding to the target abnormal data includes:
[0009] Determine whether the abnormal curve corresponding to the target abnormal data exists periodically in the historical abnormal curve;
[0010] When the abnormal curve corresponding to the target abnormal data exists periodically in the historical abnormal curve, the reason why the target abnormal data triggers an alarm is that the triggering is caused by the periodic abnormality.
[0011] A device for processing service exceptions, the device comprising:
[0012] An acquisition module, used to acquire business data and determine abnormal data sets in the business data;
[0013] A first determination module is used to obtain a preset hierarchical structure, and determine target abnormal data in the abnormal data set that is allowed to trigger an alarm according to the preset hierarchical structure;
[0014] The second determination module is used to obtain a service topology structure, and determine the reason why the target abnormal data triggers an alarm based on the service topology structure.
[0015] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0016] Acquire business data, and determine abnormal data sets in the business data;
[0017] Acquire a preset hierarchical structure, and determine target abnormal data in the abnormal data set that is allowed to trigger an alarm according to the preset hierarchical structure;
[0018] A service topology structure is obtained, and a reason why the target abnormal data triggers an alarm is determined based on the service topology structure.
[0019] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the following steps:
[0020] Acquire business data, and determine abnormal data sets in the business data;
[0021] Acquire a preset hierarchical structure, and determine target abnormal data in the abnormal data set that is allowed to trigger an alarm according to the preset hierarchical structure;
[0022] A service topology structure is obtained, and a reason why the target abnormal data triggers an alarm is determined based on the service topology structure.
[0023] The above-mentioned business anomaly processing method, device, computer equipment and storage medium obtain business data, determine the abnormal data set in the business data, obtain a preset hierarchical structure, and determine the target abnormal data in the abnormal data set that is allowed to trigger an alarm according to the preset hierarchical structure, so as to determine which abnormal data needs to trigger an alarm and which abnormal data does not need to trigger an alarm, so as to avoid any abnormal data triggering an alarm and causing the background to continuously issue an alarm. Obtain the business topology structure, and determine the reason why the target abnormal data triggers an alarm based on the business topology structure, so as to accurately detect the reason why the abnormal data triggers an alarm, so as to be able to perform targeted processing on the abnormal data that triggers the alarm to restore the normal operation of the business. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 An application environment diagram of a method for handling business exceptions in one embodiment;
[0025] Figure 2 The figure is a flowchart of a method for handling business exceptions in one embodiment;
[0026] Figure 3 A schematic diagram of a process for determining target abnormal data in an abnormal data set that is allowed to trigger an alarm according to a preset hierarchical structure in one embodiment;
[0027] Figure 4 A schematic diagram of a process for determining the cause of target abnormal data triggering an alarm based on a service topology structure in one embodiment;
[0028] Figure 5 is a schematic diagram of a method for handling business exceptions in an embodiment;
[0029] FIG6( a ) is a graph showing a small amount of accumulated abnormal data in one embodiment;
[0030] FIG6( b ) is a graph showing a sudden change in the number of abnormalities in one embodiment;
[0031] FIG6( c ) is a graph showing a sudden change in the number of abnormalities in another embodiment;
[0032] FIG6( d ) is a graph showing abnormal data oscillating in a short period of time in one embodiment;
[0033] FIG6( e ) is a graph showing abnormal data that cannot be self-healed in one embodiment;
[0034] FIG6( f ) is a graph corresponding to periodic abnormal data in one embodiment;
[0035] Figure 7 This is an interface diagram of historical work order processing results in an embodiment;
[0036] Figure 8 is a flowchart of a method for handling business exceptions in an embodiment;
[0037] Fig. 9 is a structural block diagram of a device for processing service exceptions in an embodiment;
[0038] Fig.10 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0039] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0040] The method for handling business exceptions provided in this application can be applied to Figure 1In the application environment shown, the terminal 102 communicates with the server 104 through a network. The terminal 102 can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, and portable wearable devices, and the server 104 can be implemented as an independent server or a server cluster consisting of multiple servers.
[0041] In this embodiment, the terminal 102 sends the business data corresponding to each business service to the server 104, and the server 104 receives the business data reported by each business service and determines the abnormal data set in the business data. Then, the server 104 obtains a preset hierarchical structure, and determines the target abnormal data in the abnormal data set that is allowed to trigger an alarm according to the preset hierarchical structure. Then, the server 104 obtains a business topology structure, and determines the reason why the target abnormal data triggers an alarm based on the business topology structure, so that the reason why the abnormal data triggers an alarm can be detected, so as to perform targeted processing on the abnormal data that triggers the alarm.
[0042] In one embodiment, Figure 2 As shown, a method for handling business exceptions is provided, and the method is applied to Figure 1 The server in the example is used to illustrate the following steps:
[0043] Step 202: Acquire business data and determine abnormal data sets in the business data.
[0044] The abnormal data set refers to a set of data with abnormalities in the business data. The abnormal data refers to data with abnormalities in the business data, for example, data that causes call timeout and data that causes access failure, but is not limited thereto.
[0045] Specifically, each business service reports its own business data to the server, and the server receives the reported business data. The business data includes normal data and abnormal data, and the server determines the abnormal data in the business data and forms an abnormal data set.
[0046] Step 204 , obtaining a preset hierarchical structure, and determining target abnormal data in the abnormal data set that is allowed to trigger an alarm according to the preset hierarchical structure.
[0047] The preset hierarchical structure refers to a pre-set logical structure for grading abnormal data. The preset hierarchical structure can be used to determine whether the abnormal data needs to trigger an alarm.
[0048] Specifically, the preset hierarchical structure divides the conditions corresponding to each layer. When the abnormal data meets the conditions corresponding to any layer in the preset hierarchical structure, it means that the abnormal data will trigger an alarm. The server obtains the preset hierarchical structure, compares the abnormal data set with the conditions corresponding to each layer in the preset hierarchical structure, and determines the abnormal data in the abnormal data set that meets the conditions corresponding to any layer in the preset hierarchical structure. The abnormal data in the abnormal data set that meets the conditions is used as the target abnormal data.
[0049] In this embodiment, when the server determines the target abnormal data in the abnormal data set, an alarm is triggered.
[0050] Step 206: Acquire the service topology structure, and determine the reason why the target abnormal data triggers an alarm based on the service topology structure.
[0051] The business topology structure refers to the connection or calling relationship between business objects. For example, A calls B and C, and B obtains the data reported by C.
[0052] Specifically, the server obtains the topology structure corresponding to the business service that reports the business data. Then, the server determines the cause of triggering the alarm based on the business topology structure and the target abnormal data.
[0053] In this embodiment, the server may obtain historical abnormal data, which may include: historical alarm results, historical work order results, and work order governance rule base, etc. The server may determine the specific reason why the target abnormal data triggers an alarm based on the business topology, historical abnormal data, and target abnormal data.
[0054] In this embodiment, the server can calculate the real-time calculation results of the business data in each dimension, and calculate the real-time calculation results of the abnormal data in each dimension, each dimension including the total amount of business service requests, time, business object and error code. Among them, the error code is a set of numbers (or a combination of letters and numbers), which will be associated with the error message and can be used to identify specific problems that occur in the program. For example, error code 400 indicates an error request, that is, the server does not understand the syntax of the request. It is understandable that the error message corresponding to the error code can be adjusted according to demand. The server can jointly determine the reason why the target abnormal data triggers an alarm based on the real-time calculation results of the business data in each dimension, the real-time calculation results of the abnormal data in each dimension, the business topology structure and the historical abnormal data.
[0055] In this embodiment, the server can use a stream computing framework to calculate the real-time computing results of business data in various dimensions. The stream computing framework can use Apache Storm (abbreviated as "Storm") or Apache Flink. The server stores the real-time computing results in a time series database for query verification when the server determines the cause of the abnormal data triggering the alarm.
[0056] In this embodiment, by acquiring business data, determining the abnormal data set in the business data, acquiring a preset hierarchical structure, and determining the target abnormal data in the abnormal data set that is allowed to trigger an alarm according to the preset hierarchical structure, it is possible to determine which abnormal data needs to trigger an alarm and which abnormal data does not need to trigger an alarm, thereby avoiding the situation where any abnormal data triggers an alarm and causes the background to continuously issue an alarm. The business topology structure is acquired, and the reason why the target abnormal data triggers an alarm is determined based on the business topology structure, so that the reason why the abnormal data triggers an alarm can be detected, so that the abnormal data that triggers the alarm can be targeted for processing to restore the normal operation of the business.
[0057] In one embodiment, Figure 3 As shown, the target abnormal data that is allowed to trigger an alarm in the abnormal data set is determined according to the preset hierarchical structure, including:
[0058] Step 302: Obtain hierarchical data corresponding to each layer in the preset hierarchical structure.
[0059] Step 304: input the abnormal data set into a preset hierarchical structure, and compare the abnormal data set with the hierarchical data corresponding to each layer.
[0060] Specifically, the server pre-sets hierarchical data corresponding to each layer in a preset hierarchical structure, and the server inputs the abnormal data set into each layer of the preset hierarchical structure. Then, the server obtains the hierarchical data corresponding to each layer, and compares each abnormal data in the abnormal data set with the hierarchical data corresponding to each layer to determine whether there is abnormal data in the abnormal data set that is the same as the hierarchical data corresponding to each layer.
[0061] Step 306, when there is abnormal data in the abnormal data set that is identical to the hierarchical data corresponding to any layer in the preset hierarchical structure, the abnormal data that is identical to the hierarchical data corresponding to any layer is used as the target abnormal data that is allowed to trigger an alarm; the types of target abnormal data include sensitive class, abrupt change class, oscillation class and long-lasting class.
[0062] Specifically, the server inputs the abnormal data set into the initial layer of the preset hierarchical structure to calculate the detection results of the abnormal data in each dimension in real time. Then, the server inputs the detection results of the abnormal data in each dimension output by the initial layer and the abnormal data set into the first layer of the preset hierarchical structure.
[0063] The server obtains the hierarchical data corresponding to the first layer, and compares each abnormal data in the abnormal data set and the detection result corresponding to the abnormal data with the hierarchical data corresponding to the first layer to determine whether there is abnormal data in the abnormal data set that is identical to the hierarchical data corresponding to the first layer. When there is abnormal data in the abnormal data set that is identical to the hierarchical data corresponding to the first layer, the abnormal data that is identical to the hierarchical data corresponding to the first layer is used as the target abnormal data, and an alarm is triggered.
[0064] It is understandable that the server inputs the abnormal data set and the detection results corresponding to the abnormal data into the preset hierarchical structure, and passes through each layer of the preset hierarchical structure in turn. When the current layer of the preset hierarchical structure cannot determine the target abnormal data in the abnormal data set, the abnormal data set and the detection results corresponding to the abnormal data are input to the next layer of the preset hierarchical structure. When the target abnormal data in the abnormal data set are all determined in a certain layer of the preset hierarchical structure, there is no need to enter the next layer, thereby completing the determination of the target abnormal data.
[0065] Furthermore, the server can filter out the target abnormal data in the abnormal data set through a preset hierarchical structure, and can determine the type of the target abnormal data. The types of the target abnormal data include sensitive, abrupt change, oscillation and long-term storage. The sensitive class refers to high-risk access, or long-term abnormalities, and the alarm is triggered when the number of abnormalities reaches a certain threshold. There are two types of abrupt change abnormal data. One is that the number of historical abnormalities is 0, and the number of abnormalities suddenly increases at a point in time or a period of time. The other is that the historical abnormal curve maintains a certain trend, but the trend of the abnormal curve changes greatly at a point in time or a period of time, such as the sudden abnormal curve showing an obvious upward or downward trend. Oscillation abnormal data refers to the repeated appearance of the abnormal curve corresponding to the abnormal data within a certain period of time, and the abnormal curve does not exist periodically in the historical abnormal curve. Long-term abnormal data refers to long-term abnormalities, and they do not exist periodically, and there is no obvious rule for the abnormal curve.
[0066] In this embodiment, hierarchical data corresponding to each layer in a preset hierarchical structure is obtained, an abnormal data set is input into the preset hierarchical structure, and the abnormal data set is compared with the hierarchical data corresponding to each layer. When there is abnormal data in the abnormal data set that is identical to the hierarchical data corresponding to any layer in the preset hierarchical structure, the abnormal data that is identical to the hierarchical data corresponding to any layer is used as target abnormal data that is allowed to trigger an alarm; the types of the target abnormal data include sensitive, abrupt change, oscillation, and long-lasting types, so that it is possible to determine which data needs to trigger an alarm and which data does not need to trigger an alarm by stratifying the abnormal data, and further determine the type of the target abnormal data that triggers the alarm, so as to perform targeted processing on each type of target abnormal data.
[0067] In one embodiment, the server can perform machine learning based on the alarm exception data and the alarm work order to obtain the alarm rules or alarm thresholds. Through the alarm rules or alarm thresholds, the target abnormal data that allows alarms to be issued can be screened out from the abnormal data set. In addition, the time for the business to return to normal can be determined through the historical abnormal data and the target abnormal data.
[0068] In one embodiment, determining the reason why target abnormal data triggers an alarm based on the service topology structure includes:
[0069] Based on the business topology and target abnormal data, the number of abnormalities of the target abnormal data within a preset unit time is determined; when the number of abnormalities of the target abnormal data within the preset unit time is greater than or equal to the threshold, the reason for the target abnormal data triggering an alarm is that the accumulation of abnormalities leads to the triggering.
[0070] The unit time refers to the time divided into units of seconds or minutes. The threshold refers to the number of abnormalities generated in a unit time that is preset. The unit time corresponding to the target abnormal data is the same as the unit time corresponding to the threshold.
[0071] Specifically, the server may obtain the business topology structure corresponding to the business service of the reported business data, and the time corresponding to the target abnormal data. The server determines the business object corresponding to the target abnormal data based on the business topology structure. Then, the server may calculate the number of abnormalities generated by the target abnormal data within a unit according to the business object and the time corresponding to the target abnormal data. The number of abnormalities generated within the unit may refer to the number of abnormalities generated by the target abnormal data within one minute, or the number of abnormalities generated within half a minute or 10 seconds. The unit time can be adjusted according to specific needs.
[0072] The server obtains the threshold of the number of abnormalities in a preset unit time, and calculates the number of abnormalities generated by the target abnormal data in the unit time according to the unit time corresponding to the threshold. Then, the server compares the number of abnormalities of the target abnormal data in the preset unit time with the threshold. When the number of abnormalities of the target abnormal data in the preset unit time is greater than or equal to the threshold, it is determined that the reason for the target abnormal data to trigger the alarm is that the abnormal number of times is obviously accumulated in a short period of time, resulting in the triggering.
[0073] In this embodiment, based on the business topology structure and the target abnormal data, the number of abnormalities of the target abnormal data within a preset unit time is determined, and the number of abnormalities of the target abnormal data within the preset unit time is compared with the threshold. When the number of abnormalities of the target abnormal data within the preset unit time is greater than or equal to the threshold, it is determined that the reason for the target abnormal data triggering the alarm is that the accumulation of abnormalities leads to the triggering. The reason why the target abnormal data triggers the alarm can be quickly and accurately determined, so that the abnormality can be processed in a targeted manner according to the cause.
[0074] In one embodiment, determining the reason why target abnormal data triggers an alarm based on the service topology structure includes:
[0075] Generate an abnormal curve corresponding to the target abnormal data based on the business topology structure; determine the change amount of the abnormal curve corresponding to the target abnormal data in different time periods. When the change amounts in different time periods are different, the reason why the target abnormal data triggers an alarm is that the number of abnormalities changes suddenly, causing the trigger.
[0076] Specifically, the server may obtain the business topology structure corresponding to the business service of the reported business data, and the time corresponding to the target abnormal data. The server determines the business object corresponding to the target abnormal data based on the business topology structure. Then, the server may generate an abnormal curve corresponding to the target abnormal data according to the business object and the time corresponding to the target abnormal data. The server may divide the abnormal curve into different time periods and calculate the change amount of the abnormal curve in different time periods.
[0077] Next, the server can convert the variation of the abnormal curve in different time periods into the variation in the same unit time, and determine whether the variation in the same unit time is the same. If it is not the same, it means that the reason why the target abnormal data triggers the alarm is that the number of abnormalities changes suddenly.
[0078] For example, the server may divide the abnormal curve into units of 2 minutes and calculate the change amount of the abnormal curve within 2 minutes. Then, the server determines whether the abnormal change amount of the abnormal curve within 2 minutes is the same. When the abnormal change amount of the abnormal curve within 2 minutes is not the same, it is determined that the reason for the target abnormal data to trigger the alarm is that the number of abnormalities suddenly increases significantly.
[0079] The server can divide the abnormal curve into units of 2 minutes, 3 minutes, and 5 minutes, and calculate the change of the abnormal curve within 2 minutes, the change of the abnormal curve within 3 minutes, and the change of the abnormal curve within 5 minutes. Then, the change of the abnormal curve within 2 minutes can be averaged to obtain the change of 1 minute; the change of the abnormal curve within 3 minutes can be averaged to obtain the change of 1 minute, and the change of the abnormal curve within 5 minutes can be averaged to obtain the change of 1 minute. Then, the change of 1 minute corresponding to the three units is calculated and compared. When the change of 1 minute corresponding to the three units is different, it is determined that the reason for the target abnormal data to trigger the alarm is that the number of abnormalities suddenly increased significantly.
[0080] In this embodiment, an abnormal curve corresponding to the target abnormal data is generated based on the business topology structure, and the change amount of the abnormal curve corresponding to the target abnormal data in different time periods is determined. When the change amount in different time periods is different, the reason why the target abnormal data triggers the alarm is that the number of abnormalities changes suddenly, causing the trigger. The reason why the target abnormal data triggers the alarm can be quickly and accurately determined, so that the abnormality can be processed in a targeted manner according to the cause.
[0081] In one embodiment, Figure 4 As shown in the figure, the reasons why the target abnormal data triggers an alarm are determined based on the business topology structure, including:
[0082] Step 402: Obtain a historical abnormality curve corresponding to the historical abnormal data.
[0083] Specifically, the server may obtain historical abnormal data, which includes historical alarm results, historical work order results, and work order governance rule bases, etc. The server may obtain a historical abnormal curve corresponding to the historical abnormal data. Alternatively, the server may generate a historical abnormal curve based on the historical alarm results in the historical abnormal data.
[0084] Step 404: Generate an abnormal curve corresponding to the target abnormal data based on the service topology structure.
[0085] Specifically, the server may obtain the business topology structure corresponding to the business service of the reported business data and the time corresponding to the target abnormal data. The server determines the business object corresponding to the target abnormal data based on the business topology structure. Then, the server may generate an abnormal curve corresponding to the target abnormal data according to the business object and the time corresponding to the target abnormal data.
[0086] Step 406: Determine the reason why the target abnormal data triggers an alarm based on the historical abnormal curve and the abnormal curve corresponding to the target abnormal data.
[0087] Specifically, the server may compare the historical abnormal curve with the abnormal curve corresponding to the target abnormal data, and determine the reason why the target abnormal data triggers an alarm according to the comparison result of the two.
[0088] Furthermore, the server may determine the reason why the target abnormal data triggers an alarm based on the amount of change of the historical abnormal curve and the abnormal curve corresponding to the target abnormal data within the same unit time, whether the historical abnormal curve and the abnormal curve corresponding to the target abnormal data are the same, etc.
[0089] In this embodiment, the historical abnormal curve corresponding to the historical abnormal data is obtained, and the abnormal curve corresponding to the target abnormal data is generated based on the business topology structure. The reason why the target abnormal data triggers an alarm is determined based on the historical abnormal curve and the abnormal curve corresponding to the target abnormal data. The reason for the current business abnormality can be analyzed in combination with the historical abnormal data to more accurately determine the reason for triggering the alarm.
[0090] In one embodiment, determining the reason why the target abnormal data triggers an alarm based on the historical abnormal curve and the abnormal curve corresponding to the target abnormal data includes:
[0091] Determine the amount of change of the abnormal curve corresponding to the target abnormal data within the first preset time period; when the amount of change of the abnormal curve corresponding to the target abnormal data within the first preset time period is higher than the amount of change of the historical abnormal curve within the first preset time period, the reason why the target abnormal data triggers the alarm is that the number of abnormalities changes suddenly, causing the trigger.
[0092] Specifically, the server can calculate the change amount of the abnormal curve corresponding to the target abnormal data within the first preset time period to obtain a first change amount. And calculate the change amount of the historical abnormal curve within the first preset time period to obtain a second change amount. Then, the server compares the first change amount with the second change amount. When the first change amount is greater than the second change amount, it is determined that the reason why the target abnormal data triggers an alarm is caused by a sudden and significant increase in the number of abnormalities.
[0093] For example, the time corresponding to the target abnormal data is 8:30-09:00, and the server calculates the change of the abnormal curve corresponding to the target abnormal data within 30 minutes, and determines the change of the historical abnormal curve within 8:00-8:30, that is, determines the change of the historical abnormal curve within 30 minutes closest to the time when the target abnormal data was generated. Determine whether the change of the abnormal curve corresponding to the target abnormal data within 8:30-09:00 and the change of the historical abnormal curve within 8:00-8:30 are the same. If the change within 8:30-09:00 is significantly greater than the change within 8:00-8:30, it means that the number of abnormalities has increased significantly within 30 minutes, and it can be determined that the reason for the target abnormal data to trigger the alarm is that the number of abnormalities has changed suddenly.
[0094] In this embodiment, by comparing the changes of the abnormal curve corresponding to the target abnormal data and the historical abnormal curve within the first preset time period, it can be determined whether the cause of the target abnormal data triggering the alarm is a sudden change in the number of abnormalities.
[0095] In one embodiment, determining the reason why the target abnormal data triggers an alarm based on the historical abnormal curve and the abnormal curve corresponding to the target abnormal data includes:
[0096] Determine whether the abnormal curve corresponding to the target abnormal data is repeatedly generated within the second preset time period; when the abnormal curve corresponding to the target abnormal data is repeatedly generated within the second preset time period, and there is no curve in the historical abnormal curve that is identical to the abnormal curve corresponding to the target abnormal data, the reason why the target abnormal data triggers the alarm is that the curve oscillates within the preset time.
[0097] Specifically, the server can determine whether the abnormal curve corresponding to the target abnormal data is repeatedly generated within the second preset time period. When the abnormal curve corresponding to the target abnormal data is repeatedly generated within the second preset time period, it means that the abnormal curve oscillates within a certain time period. The server then obtains the historical abnormal curve to determine whether the abnormal curve corresponding to the target abnormal data appears in the historical abnormal curve. When the abnormal curve corresponding to the target abnormal data is repeatedly generated within the second preset time period, and there is no curve in the historical abnormal curve that is the same as the abnormal curve corresponding to the target abnormal data, it means that the abnormal curve corresponding to the target abnormal data is not generated periodically, and it is determined that the reason for the target abnormal data to trigger the alarm is caused by the oscillation of the curve within the preset time.
[0098] In this embodiment, by determining whether the abnormal curve corresponding to the target abnormal data is repeatedly generated within the second preset time period, and whether there is a curve in the historical abnormal curve that is the same as the abnormal curve corresponding to the target abnormal data, it is possible to distinguish whether the curve corresponding to the target abnormal data is generated periodically or oscillates in a short period of time, so as to accurately distinguish the reason why the target abnormal data triggers an alarm.
[0099] In one embodiment, determining the reason why the target abnormal data triggers an alarm based on the historical abnormal curve and the abnormal curve corresponding to the target abnormal data includes:
[0100] Determine whether the abnormal curve corresponding to the target abnormal data exists periodically in the historical abnormal curve; when the abnormal curve corresponding to the target abnormal data exists periodically in the historical abnormal curve, the reason for triggering the alarm of the target abnormal data is that the periodic abnormality causes the trigger.
[0101] Specifically, the server may compare the abnormal curve corresponding to the target abnormal data with the historical abnormal curve to determine whether the abnormal curve corresponding to the target abnormal data appears in the historical abnormal curve. When the abnormal curve corresponding to the target abnormal data exists in the historical abnormal curve, it is determined whether the abnormal curve corresponding to the target abnormal data exists periodically in the historical abnormal curve. When the abnormal curve corresponding to the target abnormal data exists periodically in the historical abnormal curve, it indicates that the target abnormal data is a periodic abnormality, and the reason for the target abnormal data triggering the alarm is that the periodic abnormality causes the trigger.
[0102] In this embodiment, it is determined whether the abnormal curve corresponding to the target abnormal data exists periodically in the historical abnormal curve, so as to determine whether the cause of the target abnormal data triggering the alarm is caused by a periodic abnormality, so that the abnormal problem of the business can be handled accordingly according to the cause of the triggering alarm.
[0103] like Figure 5 As shown, it is a schematic diagram of a method for handling business anomalies in one embodiment. The preset hierarchical structure includes 0 to N+2 layers, and each layer corresponds to its own hierarchical data. The server can input the abnormal data set into the 0th layer of the preset hierarchical structure to calculate the detection results of the abnormal data in each dimension in real time. The abnormal data has an order of magnitude difference relative to the full data, and the calculation can be achieved by a single machine. Then, the server inputs the abnormal data set and the detection result into the first layer to determine whether there is a small amount of accumulated abnormal data in the abnormal data set. If there is, the abnormal data exceeding the threshold is used as the target abnormal data that is allowed to trigger an alarm.
[0104] When the number of abnormal data anomalies is significantly higher than the previous number of abnormal data, and significantly higher than the subsequent number of abnormal data, the abnormal data with a sudden change in the number of abnormal data anomalies is used as the target abnormal data that allows triggering an alarm. For example, the number of abnormal data in the black dashed box in Figure 6(a) is significantly higher than the number of abnormal data in the non-dashed box. Figure 6(a) shows a graph of a small amount of accumulated abnormal data in the first layer.
[0105] Next, the server may input the abnormal data other than the target abnormal data in the abnormal data set into the second layer to determine whether there is abnormal data with a sudden change in the number of abnormalities in the abnormal data set. When there is no small amount of accumulated abnormal data in the abnormal data set, the entire abnormal data set is input into the second layer. The server uses the abnormal data with a sudden change in the number of abnormalities as the target abnormal data that is allowed to trigger an alarm. The second layer contains two abnormal sudden changes, as shown in the curve graphs of Figure 6 (b) and Figure 6 (c).
[0106] Figure 6(b) shows a curve chart of abnormal data with a sudden and obvious increase in the number of historical abnormalities at the second layer, which is close to 0. That is, the number of historical abnormalities on the left side of the black dotted line in Figure 6(b) is all 0, while the number of abnormalities on the right side of the black dotted line has increased significantly. Figure 6(c) shows a schematic diagram of the historical abnormalities at the second layer that have always existed in a certain curve (such as downstream access timeout, incomplete user authentication parameters, etc.), but the curve trend at a certain point in time (the number of abnormalities in the dotted part of the figure has increased significantly) suddenly and significantly increased. That is, the abnormal curve on the left side of the black dotted line in Figure 6(c) has always maintained a relatively stable trend, while in the black solid line frame, the number of abnormalities at the time point corresponding to the black dotted line position has increased significantly relative to the number of historical abnormalities.
[0107] Next, the server may input the abnormal data other than the target data in the abnormal data set into the third layer to determine whether there is abnormal data that oscillates in a short period of time in the abnormal data set. When there is no abnormal data with a sudden change in the number of abnormalities in the abnormal data set, the entire abnormal data set is input into the third layer. The server uses the abnormal data that oscillates in a short period of time as the target abnormal data that is allowed to trigger an alarm. The server detects whether the abnormal curve has appeared repeatedly in the past certain period of time. If it appears repeatedly and there is no non-periodic abnormality, it is used as the target abnormal data that is allowed to trigger an alarm, as shown in the curve graph of Figure 6(d). As shown in Figure 6(d), the dotted line is the dividing line, and the historical abnormal curve corresponding to the historical abnormal data is on the left side of the dotted line (there is no abnormality on the left side of the dotted line in this figure, so there is no historical abnormal curve). The black dotted box on the right side of the dotted line is the abnormal curve corresponding to the abnormal data in the abnormal data set. It can be seen from the figure that the abnormal curve in the black dotted box has not appeared repeatedly in the past certain period of time.
[0108] Next, the server can input the abnormal data other than the target data in the abnormal data set into the fourth layer to determine whether each abnormal data in the abnormal data set can be self-healed in a short time. The server regards the abnormal data that cannot be self-healed in a short time as the target abnormal data that is allowed to trigger an alarm.
[0109] As shown in Figure 6(e), the curve of abnormal changes detected in the fourth layer can be self-healed within a certain period of time. If it cannot self-heal, an alarm is generated. If it can self-heal, it will detect whether it has recovered after a certain period of time (the abnormality becomes 0 or recovers to the amount before the change is recovery). If it has not recovered, an alarm will be issued. If it recovers, a report will be generated and pushed to the user regularly. Self-healing means automatically returning to normal or maintaining the same state as before. As shown in Figure 6(e), there is only one time point in the black dotted box where the abnormality occurs, and no abnormality occurs again before and after this time point, which means that the abnormality can be self-healed.
[0110] Next, the server can input the abnormal data other than the target data in the abnormal data set into the fifth layer to determine whether each abnormal data in the abnormal data set exists in the historical alarm results, and determine whether the abnormal data exists periodically. The server uses the periodic abnormal data in the abnormal data set that exists in the historical alarm results as the target abnormal data that is allowed to trigger the alarm. The fifth layer detects whether the abnormality is generated periodically, as shown in Figure 6(f), which shows the curve graph corresponding to the abnormal data with periodic abnormalities. That is, the abnormal curves in the two black dotted boxes in Figure 6(f) are the same, indicating that the abnormal data corresponding to the abnormal curve exists periodically.
[0111] Next, the server can input the abnormal data in the abnormal data set except the target data into the sixth layer to determine whether each abnormal data in the abnormal data set exists in the historical work order processing results. The server uses the abnormal data in the abnormal data set that appeared in the historical work order processing results as the target abnormal data that is allowed to trigger an alarm.
[0112] Understandably, Figure 6(a) to Figure 6(f) In the curve graph shown, the horizontal axis is time and the vertical axis is the number of abnormalities. That is, the horizontal axis corresponding to the point on the curve is the time point, and the vertical axis is the number of abnormalities corresponding to the time point.
[0113] like Figure 7 The figure shows an interface diagram of a historical work order processing result in an embodiment, and the historical work order processing result includes a method for processing the cause of the exception. By searching the historical work order processing result, it can be determined whether each abnormal data exists in the historical work order processing result.
[0114] Next, the server determines the reason why the target abnormal data triggers the alarm based on one or a combination of the business topology, historical alarm results, historical work order results, work order governance rule base, real-time calculation results of business data in various dimensions, real-time calculation results of abnormal data in various dimensions, and preset hierarchical structures. For example, when the abnormal ratio is small, it is determined whether it is caused by a single machine failure or a grayscale release. Grayscale release (also known as canary release) refers to a release method that can smoothly transition between black and white. A / B testing can be performed on it, that is, let some users continue to use product feature A, and some users start using product feature B. If the user has no objection to B, then gradually expand the scope and migrate all users to B. Grayscale release can ensure the stability of the overall system, and problems can be discovered and adjusted at the initial grayscale to ensure their impact. Alternatively, by detecting historical alarm results, historical work order results, work order governance rule base, or combining with business topology diagrams, determine whether there is a call exception.
[0115] Next, the server can combine historical abnormal data to determine the time required to handle the abnormality, thereby determining the time when the business returns to normal. The server can test at regular intervals to confirm whether the business abnormality has returned to normal.
[0116] In one embodiment, a method for handling business exceptions is provided, comprising:
[0117] The server obtains the business data and determines an abnormal data set in the business data.
[0118] Next, the server obtains a preset hierarchical structure, and obtains hierarchical data corresponding to each layer in the preset hierarchical structure.
[0119] Next, the server inputs the abnormal data set into the preset hierarchical structure, and compares the abnormal data set with the hierarchical data corresponding to each layer.
[0120] Furthermore, when there is abnormal data in the abnormal data set that is identical to the hierarchical data corresponding to any layer in the preset hierarchical structure, the server will use the abnormal data that is identical to the hierarchical data corresponding to any layer as the target abnormal data that is allowed to trigger an alarm; the types of the target abnormal data include sensitive class, abrupt change class, oscillation class and long-lasting class.
[0121] Next, the server obtains a service topology structure, and determines the number of abnormalities of the target abnormal data within a preset unit time based on the service topology structure and the target abnormal data.
[0122] When the number of abnormalities of the target abnormal data within the preset unit time is greater than or equal to the threshold, the reason for the target abnormal data triggering the alarm is that the abnormal number accumulates and causes the triggering.
[0123] Next, the server generates an abnormal curve corresponding to the target abnormal data based on the service topology structure.
[0124] Optionally, the server determines the amount of change of the abnormal curve corresponding to the target abnormal data in different time periods. When the amounts of change in different time periods are different, the reason why the target abnormal data triggers an alarm is that the number of abnormalities changes suddenly.
[0125] Furthermore, the server obtains a historical abnormality curve corresponding to the historical abnormal data.
[0126] Optionally, the server determines the amount of change of the abnormal curve corresponding to the target abnormal data within the first preset time period; when the amount of change of the abnormal curve corresponding to the target abnormal data within the first preset time period is higher than the amount of change of the historical abnormal curve within the first preset time period, the reason why the target abnormal data triggers the alarm is that the number of abnormalities changes suddenly, causing the trigger.
[0127] Optionally, the server determines whether the abnormal curve corresponding to the target abnormal data is repeatedly generated within a second preset time period; when the abnormal curve corresponding to the target abnormal data is repeatedly generated within the second preset time period, and there is no curve in the historical abnormal curve that is identical to the abnormal curve corresponding to the target abnormal data, the reason why the target abnormal data triggers the alarm is that the curve oscillates within the preset time period.
[0128] Optionally, the server determines whether the abnormal curve corresponding to the target abnormal data exists periodically in the historical abnormal curve; when the abnormal curve corresponding to the target abnormal data exists periodically in the historical abnormal curve, the reason why the target abnormal data triggers an alarm is that the trigger is caused by periodic abnormality.
[0129] In this embodiment, the target abnormal data in the abnormal data set that is allowed to trigger an alarm is determined according to the preset hierarchical structure, so that it is possible to determine which abnormal data needs to trigger an alarm and which abnormal data does not need to trigger an alarm. The reason why the target abnormal data triggers an alarm is determined based on the business topology structure, so that the reason why the abnormal data triggers an alarm can be detected, so that the abnormal data that triggers the alarm can be processed in a targeted manner.
[0130] In one embodiment, Figure 8 As shown, a method for handling business exceptions is provided, including:
[0131] Step 802: Acquire business data and determine abnormal data sets in the business data.
[0132] Specifically, each business service reports its own business data to the server, and the server receives the reported business data. The business data includes normal data and abnormal data, and the server determines the abnormal data in the business data and forms an abnormal data set.
[0133] Step 804 , obtaining a preset hierarchical structure, determining target abnormal data in the abnormal data set that is allowed to trigger an alarm according to the preset hierarchical structure, and triggering an alarm based on the target abnormal data.
[0134] Specifically, the preset hierarchical structure divides the conditions corresponding to each layer. When the abnormal data meets the conditions corresponding to any layer in the preset hierarchical structure, it means that the abnormal data will trigger an alarm. Then the server obtains the preset hierarchical structure, compares the abnormal data set with the conditions corresponding to each layer in the preset hierarchical structure, and determines the abnormal data in the abnormal data set that meets the conditions corresponding to any layer in the preset hierarchical structure. The abnormal data in the abnormal data set that meets the conditions is used as the target abnormal data. Then, when the server determines the target abnormal data in the abnormal data set, an alarm is triggered.
[0135] Step 806: Acquire the service topology structure, and determine the reason why the target abnormal data triggers an alarm based on the service topology structure.
[0136] Specifically, the server obtains the topology structure corresponding to the business service that reports the business data. Then, the server determines the cause of triggering the alarm based on the business topology structure and the target abnormal data.
[0137] In this embodiment, the server may obtain historical abnormal data, which may include: historical alarm structure, historical work order results, and work order governance rule base, etc. The server may determine the specific reason why the target abnormal data triggers an alarm based on the business topology, historical abnormal data, and target abnormal data.
[0138] In this embodiment, the server can calculate the real-time calculation results of the business data in each dimension, and calculate the real-time calculation structure of the abnormal data in each dimension, each dimension including time, business object and error code. The server can jointly determine the reason why the target abnormal data triggers an alarm based on the real-time calculation results of the business data in each dimension, the real-time calculation results of the abnormal data in each dimension, the business topology structure and the historical abnormal data.
[0139] Step 808: Output the target abnormal data and the reason why the target abnormal data triggers an alarm.
[0140] Specifically, the server will establish an association between each target abnormal data and the reason for triggering the alarm, and output the target abnormal data and the reason for the target abnormal data triggering the alarm to the monitoring service platform to instruct the monitoring manager to process each target abnormal data according to the reason for triggering the alarm.
[0141] In this embodiment, by acquiring business data, determining the abnormal data set in the business data, acquiring a preset hierarchical structure, and determining the target abnormal data in the abnormal data set that is allowed to trigger an alarm according to the preset hierarchical structure, it is possible to determine which abnormal data needs to trigger an alarm and which abnormal data does not need to trigger an alarm, thereby avoiding the situation where any abnormal data triggers an alarm and causes the background to continuously issue an alarm. The business topology structure is acquired, and the reason why the target abnormal data triggers an alarm is determined based on the business topology structure, so that the reason why the abnormal data triggers an alarm can be detected, and the target abnormal data and the reason why the target abnormal data triggers an alarm can be output, so that the abnormal data that triggers the alarm can be targeted for processing to restore the normal operation of the business.
[0142] It should be understood that although Figure 2-Figure 8The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 2-Figure 8 At least part of the steps may include multiple steps or multiple stages. These steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed in turn or alternately with other steps or at least part of the steps or stages in other steps.
[0143] In one embodiment, Fig. 9 As shown, a device for processing business exceptions is provided. The device can adopt a software module or a hardware module, or a combination of the two to become a part of a computer device. The device specifically includes: an acquisition module 902, a first determination module 904 and a second determination module 906, wherein:
[0144] The acquisition module 902 is used to acquire business data and determine abnormal data sets in the business data.
[0145] The first determination module 904 is configured to obtain a preset hierarchical structure, and determine target abnormal data in the abnormal data set that is allowed to trigger an alarm according to the preset hierarchical structure.
[0146] The second determination module 906 is used to obtain a service topology structure, and determine the reason why the target abnormal data triggers an alarm based on the service topology structure.
[0147] In this embodiment, by acquiring business data, determining the abnormal data set in the business data, acquiring a preset hierarchical structure, and determining the target abnormal data in the abnormal data set that is allowed to trigger an alarm according to the preset hierarchical structure, it is possible to determine which abnormal data needs to trigger an alarm and which abnormal data does not need to trigger an alarm, thereby avoiding the situation where any abnormal data triggers an alarm and causes the background to continuously issue an alarm. The business topology structure is acquired, and the reason why the target abnormal data triggers an alarm is determined based on the business topology structure, so that the reason why the abnormal data triggers an alarm can be accurately detected, so that the abnormal data that triggers the alarm can be targeted for processing to restore the normal operation of the business.
[0148] In one embodiment, the first determination module 904 is also used to: obtain hierarchical data corresponding to each layer in the preset hierarchical structure; input the abnormal data set into the preset hierarchical structure, and compare the abnormal data set with the hierarchical data corresponding to each layer; when there is abnormal data in the abnormal data set that is identical to the hierarchical data corresponding to any layer in the preset hierarchical structure, the abnormal data that is identical to the hierarchical data corresponding to any layer is used as target abnormal data that is allowed to trigger an alarm; the types of the target abnormal data include sensitive class, abrupt change class, oscillation class and long-lasting class.
[0149] In this embodiment, hierarchical data corresponding to each layer in a preset hierarchical structure is obtained, an abnormal data set is input into the preset hierarchical structure, and the abnormal data set is compared with the hierarchical data corresponding to each layer. When there is abnormal data in the abnormal data set that is identical to the hierarchical data corresponding to any layer in the preset hierarchical structure, the abnormal data that is identical to the hierarchical data corresponding to any layer is used as target abnormal data that is allowed to trigger an alarm; the types of the target abnormal data include sensitive, abrupt change, oscillation, and long-lasting types, so that it is possible to determine which data needs to trigger an alarm and which data does not need to trigger an alarm by stratifying the abnormal data, and further determine the type of the target abnormal data that triggers the alarm, so as to perform targeted processing on each type of target abnormal data.
[0150] In one embodiment, the second determination module 906 is also used to: determine the number of abnormalities of the target abnormal data within a preset unit time based on the business topology structure and the target abnormal data; when the number of abnormalities of the target abnormal data within the preset unit time is greater than or equal to a threshold, the reason for the target abnormal data triggering an alarm is that the accumulation of abnormalities leads to the triggering.
[0151] In this embodiment, based on the business topology structure and the target abnormal data, the number of abnormalities of the target abnormal data within a preset unit time is determined, and the number of abnormalities of the target abnormal data within the preset unit time is compared with the threshold. When the number of abnormalities of the target abnormal data within the preset unit time is greater than or equal to the threshold, it is determined that the reason for the target abnormal data triggering the alarm is that the accumulation of abnormalities leads to the triggering. The reason why the target abnormal data triggers the alarm can be quickly and accurately determined, so that the abnormality can be processed in a targeted manner according to the cause.
[0152] In one embodiment, the second determination module 906 is also used to: generate an abnormal curve corresponding to the target abnormal data based on the business topology structure; determine the change amount of the abnormal curve corresponding to the target abnormal data in different time periods. When the change amount in different time periods is different, the reason why the target abnormal data triggers an alarm is that the number of abnormalities changes suddenly, causing the trigger.
[0153] In this embodiment, an abnormal curve corresponding to the target abnormal data is generated based on the business topology structure, and the change amount of the abnormal curve corresponding to the target abnormal data in different time periods is determined. When the change amount in different time periods is different, the reason why the target abnormal data triggers the alarm is that the number of abnormalities changes suddenly, causing the trigger. The reason why the target abnormal data triggers the alarm can be quickly and accurately determined, so that the abnormality can be processed in a targeted manner according to the cause.
[0154] In one embodiment, the second determination module 906 is also used to: obtain a historical abnormal curve corresponding to the historical abnormal data; generate an abnormal curve corresponding to the target abnormal data based on the business topology structure; and determine the reason why the target abnormal data triggers an alarm based on the historical abnormal curve and the abnormal curve corresponding to the target abnormal data.
[0155] In this embodiment, the historical abnormal curve corresponding to the historical abnormal data is obtained, and the abnormal curve corresponding to the target abnormal data is generated based on the business topology structure. The reason why the target abnormal data triggers an alarm is determined based on the historical abnormal curve and the abnormal curve corresponding to the target abnormal data. The reason for the current business abnormality can be analyzed in combination with the historical abnormal data to more accurately determine the reason for triggering the alarm.
[0156] In one embodiment, the second determination module 906 is also used to: determine the amount of change of the abnormal curve corresponding to the target abnormal data within the first preset time period; when the amount of change of the abnormal curve corresponding to the target abnormal data within the first preset time period is higher than the amount of change of the historical abnormal curve within the first preset time period, the reason why the target abnormal data triggers the alarm is that the number of abnormalities changes suddenly, causing the trigger.
[0157] In this embodiment, by comparing the changes of the abnormal curve corresponding to the target abnormal data and the historical abnormal curve within the first preset time period, it can be determined whether the cause of the target abnormal data triggering the alarm is a sudden change in the number of abnormalities.
[0158] In one embodiment, the second determination module 906 is also used to: determine whether the abnormal curve corresponding to the target abnormal data is repeatedly generated within a second preset time period; when the abnormal curve corresponding to the target abnormal data is repeatedly generated within the second preset time period, and there is no curve in the historical abnormal curve that is identical to the abnormal curve corresponding to the target abnormal data, the reason why the target abnormal data triggers the alarm is that the curve oscillates within the preset time period.
[0159] In this embodiment, by determining whether the abnormal curve corresponding to the target abnormal data is repeatedly generated within the second preset time period, and whether there is a curve in the historical abnormal curve that is the same as the abnormal curve corresponding to the target abnormal data, it is possible to distinguish whether the curve corresponding to the target abnormal data is generated periodically or oscillates in a short period of time, so as to accurately distinguish the reason why the target abnormal data triggers an alarm.
[0160] In one embodiment, the second determination module 906 is also used to determine whether the abnormal curve corresponding to the target abnormal data exists periodically in the historical abnormal curve; when the abnormal curve corresponding to the target abnormal data exists periodically in the historical abnormal curve, the reason why the target abnormal data triggers the alarm is that the trigger is caused by periodic abnormality.
[0161] In this embodiment, it is determined whether the abnormal curve corresponding to the target abnormal data exists periodically in the historical abnormal curve, so as to determine whether the cause of the target abnormal data triggering the alarm is caused by a periodic abnormality, so that the abnormal problem of the business can be handled accordingly according to the cause of the triggering alarm.
[0162] In one embodiment, the service abnormality processing device further includes: an output module. The output module is used to: output the target abnormal data and the reason why the target abnormal data triggers an alarm.
[0163] In this embodiment, by acquiring business data, determining the abnormal data set in the business data, acquiring a preset hierarchical structure, and determining the target abnormal data in the abnormal data set that is allowed to trigger an alarm according to the preset hierarchical structure, it is possible to determine which abnormal data needs to trigger an alarm and which abnormal data does not need to trigger an alarm, thereby avoiding the situation where any abnormal data triggers an alarm and causes the background to continuously issue an alarm. The business topology structure is acquired, and the reason why the target abnormal data triggers an alarm is determined based on the business topology structure, so that the reason why the abnormal data triggers an alarm can be accurately detected, and the target abnormal data and the reason why the target abnormal data triggers an alarm can be output, so that the abnormal data that triggers the alarm can be targeted for processing to restore the normal operation of the business.
[0164] For the specific definition of the service exception processing device, please refer to the definition of the service exception processing method above, which will not be repeated here. Each module in the above service exception processing device can be implemented in whole or in part by software, hardware and a combination thereof. The above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0165] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Fig.10 As shown. The computer device includes a processor, a memory and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the determination of the cause of business anomalies and business anomaly monitoring data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for determining the cause of business anomalies and processing business anomalies is implemented.
[0166] Those skilled in the art will understand that Fig.10 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0167] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.
[0168] In one embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0169] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0170] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0171] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A method for handling business exceptions, characterized in that: The method comprises: Acquire business data, and determine abnormal data sets in the business data; Obtaining hierarchical data corresponding to each layer in a preset hierarchical structure, inputting the abnormal data set into the preset hierarchical structure, and comparing the abnormal data set with the hierarchical data corresponding to each layer, wherein the preset hierarchical structure is a pre-set logical structure for grading abnormal data; When there is abnormal data in the abnormal data set that is identical to hierarchical data corresponding to any layer in the preset hierarchical structure, the abnormal data that is identical to the hierarchical data corresponding to any layer is used as target abnormal data that is allowed to trigger an alarm; the types of the target abnormal data include sensitive, abrupt change, oscillating, and long-lasting types; Acquire a business topology structure, where the business topology structure is a connection or call relationship between business objects; Based on the service topology and the target abnormal data, determining the number of abnormalities of the target abnormal data within a preset unit time; When the number of abnormalities of the target abnormal data within the preset unit time is greater than or equal to a threshold, the reason for the target abnormal data to trigger an alarm is that the number of abnormalities accumulates and causes the triggering.
2. The method according to claim 1, characterized in that The method further comprises: Generate an abnormal curve corresponding to the target abnormal data based on the business topology structure; Determine the change amount of the abnormal curve corresponding to the target abnormal data in different time periods. When the change amounts in different time periods are different, the reason why the target abnormal data triggers the alarm is that the abnormal number of times changes suddenly.
3. The method according to claim 1, characterized in that The method further comprises: Obtain the historical abnormal curve corresponding to the historical abnormal data; Generate an abnormal curve corresponding to the target abnormal data based on the business topology structure; The reason why the target abnormal data triggers an alarm is determined according to the historical abnormal curve and the abnormal curve corresponding to the target abnormal data.
4. The method according to claim 3, characterized in that The determining, based on the historical abnormal curve and the abnormal curve corresponding to the target abnormal data, a reason why the target abnormal data triggers an alarm includes: Determine a change amount of the abnormal curve corresponding to the target abnormal data within a first preset time period; When the change amount of the abnormal curve corresponding to the target abnormal data within the first preset time period is higher than the change amount of the historical abnormal curve within the first preset time period, the cause of the target abnormal data triggering the alarm is that the number of abnormalities changes suddenly.
5. The method according to claim 3, characterized in that: The determining, based on the historical abnormal curve and the abnormal curve corresponding to the target abnormal data, a reason why the target abnormal data triggers an alarm includes: Determining whether an abnormal curve corresponding to the target abnormal data is repeatedly generated within a second preset time period; When the abnormal curve corresponding to the target abnormal data is repeatedly generated within the second preset time period, and there is no curve in the historical abnormal curves that is identical to the abnormal curve corresponding to the target abnormal data, the reason why the target abnormal data triggers an alarm is that the curve oscillates within the preset time.
6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: The target abnormal data and the reason why the target abnormal data triggers an alarm are output.
7. A device for processing business anomalies, characterized in that: The device comprises: An acquisition module, used to acquire business data and determine abnormal data sets in the business data; A first determination module is used to obtain hierarchical data corresponding to each layer in a preset hierarchical structure, input the abnormal data set into the preset hierarchical structure, and compare the abnormal data set with the hierarchical data corresponding to each layer, wherein the preset hierarchical structure is a pre-set logical structure for grading abnormal data; when there is abnormal data in the abnormal data set that is identical to the hierarchical data corresponding to any layer in the preset hierarchical structure, the abnormal data that is identical to the hierarchical data corresponding to any layer is used as target abnormal data that is allowed to trigger an alarm; the types of the target abnormal data include sensitive, abrupt change, oscillation, and long-lasting; The second determination module is used to obtain a business topology structure, which is a connection or call relationship between business objects; based on the business topology structure and the target abnormal data, determine the number of abnormalities of the target abnormal data within a preset unit time; when the number of abnormalities of the target abnormal data within the preset unit time is greater than or equal to a threshold, the reason why the target abnormal data triggers an alarm is that the accumulation of abnormalities leads to the triggering.
8. The device according to claim 7, characterized in that The second determination module is also used to generate an abnormal curve corresponding to the target abnormal data based on the business topology structure; determine the change amount of the abnormal curve corresponding to the target abnormal data in different time periods. When the change amounts in different time periods are different, the reason why the target abnormal data triggers an alarm is that the number of abnormalities changes suddenly, causing the trigger.
9. The device according to claim 7, characterized in that The second determination module is also used to obtain a historical abnormal curve corresponding to the historical abnormal data; generate an abnormal curve corresponding to the target abnormal data based on the business topology structure; and determine the reason why the target abnormal data triggers an alarm based on the historical abnormal curve and the abnormal curve corresponding to the target abnormal data.
10. The device according to claim 9, characterized in that The second determination module is also used to determine the amount of change of the abnormal curve corresponding to the target abnormal data within the first preset time period; when the amount of change of the abnormal curve corresponding to the target abnormal data within the first preset time period is higher than the amount of change of the historical abnormal curve within the first preset time period, the reason why the target abnormal data triggers the alarm is that the number of abnormalities changes suddenly, causing the trigger.
11. The device according to claim 9, characterized in that The second determination module is also used to determine whether the abnormal curve corresponding to the target abnormal data is repeatedly generated within a second preset time period; when the abnormal curve corresponding to the target abnormal data is repeatedly generated within the second preset time period, and there is no curve in the historical abnormal curve that is identical to the abnormal curve corresponding to the target abnormal data, the reason why the target abnormal data triggers the alarm is that the curve oscillates within the preset time period.
12. The device according to any one of claims 7 to 11, characterized in that The device also includes: The output module is used to output the target abnormal data and the reason why the target abnormal data triggers an alarm.
13. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
14. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Server exception handling method, device, and processor
CN109284200A
A service monitoring system and method
CN110287081A