Operation and maintenance index anomaly evaluation method, device and equipment

By establishing the probability distribution of abnormal deviation and abnormal duration of operation and maintenance KPI indicators, combining the characteristics of abnormal deviation and abnormal duration, abnormal scores are determined and alarms are generated, the problem of false alarms in the existing methods is solved, and the accuracy of alarms and the quality and efficiency of operation and maintenance work are improved.

CN119938477APending Publication Date: 2025-05-06BEIJING KEDONG ELECTRIC POWER CONTROL SYST CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510055355.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing abnormal evaluation methods in the operation and maintenance field are easily affected by random abnormal data, ignoring changes in overall trends, leading to high-level false alarms, and affecting the work efficiency of operation and maintenance personnel.

Method used

By establishing the probability distribution of the abnormal deviation sequence and abnormal duration sequence of the operation and maintenance KPI indicators, combining the two attribute characteristics of abnormal deviation and abnormal duration, the abnormal score is determined based on the preset abnormal scoring algorithm, and the alarm level is determined based on the abnormal score.

Benefits of technology

This method can more accurately reflect the abnormal information of real faults on the time series of the operation and maintenance KPI indicators, avoid or reduce high-level false alarms, and enable operation and maintenance personnel to understand the cause and level of the alarm more clearly, thereby making a more accurate response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938477A_ABST
    Figure CN119938477A_ABST
Patent Text Reader

Abstract

The invention provides an operation and maintenance index anomaly evaluation method, device and equipment. The method comprises the following steps: establishing probability distribution of an abnormal deviation sequence and an abnormal duration sequence of the operation and maintenance KPI according to a time sequence of the operation and maintenance KPI obtained in real time; respectively determining an abnormal deviation probability value and an abnormal duration probability value according to the probability distribution of the abnormal deviation sequence and the abnormal duration sequence of the operation and maintenance KPI; according to the abnormal deviation probability value and the abnormal duration probability value, determining an abnormal score based on a preset abnormal scoring algorithm; and determining an alarm level according to the abnormal score and a preset corresponding relationship between the abnormal score and the alarm level, and generating an alarm. According to the method, the overall trend on the time sequence of the operation and maintenance KPI can be reflected, the difference between the abnormal values on the time sequence of the operation and maintenance KPI can be reflected, the obtained abnormal score can better reflect the abnormal information of a real fault, and high-level false alarms are avoided or reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application provides a method, device and equipment for evaluating abnormal operation and maintenance indicators. Background Art

[0002] The abnormal evaluation of operation and maintenance indicators is mainly used in key performance indicators (KPI) monitoring scenarios. In the operation and maintenance scenario, multiple different KPI indicators are set for the operation and maintenance objects, and thresholds are set for these KPI indicators, including upper thresholds and lower thresholds. When the KPI indicator is higher than the upper threshold or lower than the lower threshold, an abnormality will occur. According to different alarm grading methods, the abnormality is converted into alarms of different levels, and alarm notifications are sent to the operation and maintenance engineers.

[0003] Alarm classification is widely recognized as an important and classic demand scenario in the field of operation and maintenance. The general levels include "normal", "serious" and "disaster". Operation and maintenance engineers give priority to paying attention to and solving high-level alarms.

[0004] In the current operation and maintenance field, static thresholds and dynamic thresholds are usually used to determine whether an indicator abnormality occurs, and then determine the alarm level corresponding to the abnormality and generate an alarm. Among them:

[0005] The static threshold method refers to setting a fixed threshold (including an upper threshold and a lower threshold) in advance, and then comparing the indicator value of the KPI indicator at the current time point with the fixed threshold to determine whether it is abnormal.

[0006] The dynamic threshold method dynamically calculates a dynamic threshold based on the historical data of the KPI indicator within a preset time period before the current time point, and then compares the indicator value of the KPI indicator at the current time point with the dynamic threshold to determine whether it is abnormal.

[0007] The above two abnormality evaluation methods are to identify the abnormality at a single time point. After the abnormality is identified, the alarm level of the abnormality can be determined based on the ratio of the indicator value (abnormal value) at that time point to the threshold (fixed threshold or dynamic threshold), and an alarm can be generated. Alternatively, after the abnormality is identified, the alarm level of the abnormality can be determined based on the ratio of the number of abnormalities in the most recent preset time period at that time point, and an alarm can be generated. Summary of the invention

[0008] In order to better implement operation and maintenance alarms, this application provides an operation and maintenance indicator abnormality evaluation method, device and equipment. The technical solution proposed in this application is as follows:

[0009] In a first aspect, the present application provides a method for evaluating abnormal operation and maintenance indicators, comprising:

[0010] According to the time series of operation and maintenance KPI indicators obtained in real time, the probability distribution of abnormal deviation series and abnormal duration series of operation and maintenance KPI indicators is established;

[0011] According to the probability distribution of the abnormal deviation sequence and the abnormal duration sequence of the operation and maintenance KPI indicators, the abnormal deviation probability value and the abnormal duration probability value are determined respectively;

[0012] Determine an abnormality score according to the abnormal deviation probability value and the abnormal duration probability value based on a preset abnormality scoring algorithm;

[0013] According to the abnormality score and the preset corresponding relationship between the abnormality score and the alarm level, the alarm level is determined and an alarm is generated.

[0014] In combination with the first aspect above, in a possible implementation, establishing the probability distribution of the abnormal deviation sequence and the abnormal duration sequence of the operation and maintenance KPI indicator according to the time series of the operation and maintenance KPI indicator obtained in real time includes:

[0015] According to the indicator values ​​of each observation point in the time series of the operation and maintenance KPI indicators obtained in real time and the preset baseline threshold sequence, the abnormal deviation of each observation point is determined, and whether each observation point is an abnormal value is determined;

[0016] Obtaining an abnormal deviation sequence of the operation and maintenance KPI indicator according to the deviation value of each observation point in the time series of the operation and maintenance KPI indicator, and determining a probability distribution of the abnormal deviation sequence of the operation and maintenance KPI indicator;

[0017] According to the abnormal values ​​in all observation points in the time series of the operation and maintenance KPI indicator, an abnormal duration sequence of the operation and maintenance KPI indicator is obtained, and a probability distribution of the abnormal duration sequence of the operation and maintenance KPI indicator is determined.

[0018] In combination with the first aspect above, in a possible implementation, the method further includes: for any observation point in the time series of the operation and maintenance KPI indicator, determining the abnormal deviation of the observation point by the following method:

[0019] deviation=value-normal;

[0020] Among them, deviation represents the abnormal deviation of the observation point, value represents the indicator data of an observation point in the time series of the operation and maintenance KPI indicator, and normal represents the preset baseline threshold of the corresponding observation point.

[0021] In combination with the first aspect above, in a possible implementation, determining the abnormality score according to the abnormal deviation probability value and the abnormality duration probability value based on a preset abnormality scoring algorithm includes:

[0022] The abnormality score is calculated based on the following formula 1 according to the abnormal deviation probability value and the abnormal duration probability value, as well as the weight of the abnormal deviation and the weight of the abnormal duration:

[0023]

[0024] Among them, significance(t) represents the abnormal score, w deviation Indicates the weight of the abnormal deviation of the pre-set operation and maintenance KPI indicator, 0≤w deviation ≤1,w duration Indicates the weight of the abnormal duration of the pre-set operation and maintenance KPI indicator, 0≤w duration ≤1, P(de<=deviation(t)) represents the probability value of abnormal deviation, and P(du<=duration(t)) represents the probability value of abnormal duration.

[0025] In combination with the first aspect above, in a possible implementation, before determining the alarm level according to the abnormality score and the preset abnormality score and alarm level correspondence rule, and generating the alarm, the method further includes:

[0026] The anomaly score is processed in the following manner to obtain a processed anomaly score:

[0027] S(t)=significance(t) / max(significance(t))*100;

[0028] Among them, S(t) represents the anomaly score after processing, significance(t) represents the anomaly score, and max(significance(t)) represents the maximum value of the anomaly score at all times in the time series of the KPI indicator to be evaluated.

[0029] In combination with the first aspect above, in a possible implementation, the preset correspondence between the abnormality score and the alarm level includes multiple alarm levels and an abnormality score value range corresponding to each alarm level;

[0030] The determining the alarm level and generating the alarm according to the abnormality score and the preset abnormality score and alarm level correspondence relationship includes:

[0031] According to the correspondence between the processed anomaly score and the preset anomaly score and the alarm level, the anomaly score value range to which the processed anomaly score belongs is determined, the corresponding alarm level is obtained, and an alarm corresponding to the alarm level is generated.

[0032] In a second aspect, the present application provides an operation and maintenance indicator abnormality evaluation device, comprising:

[0033] A probability distribution determination module is used to establish the probability distribution of the abnormal deviation sequence and the abnormal duration sequence of the operation and maintenance KPI indicators according to the time series of the operation and maintenance KPI indicators obtained in real time;

[0034] A probability value determination module is used to determine the abnormal deviation probability value and the abnormal duration probability value respectively according to the probability distribution of the abnormal deviation sequence and the abnormal duration sequence of the operation and maintenance KPI indicator;

[0035] an abnormality score determination module, configured to determine an abnormality score according to the abnormal deviation probability value and the abnormal duration probability value based on a preset abnormality score algorithm;

[0036] The alarm module is used to determine the alarm level and generate an alarm according to the abnormality score and the preset corresponding relationship between the abnormality score and the alarm level.

[0037] In a third aspect, the present application provides a computer-readable storage medium, in which instructions are stored. When the instructions are executed on a terminal, the terminal executes the operation and maintenance indicator abnormality evaluation method described in the first aspect.

[0038] In a fourth aspect, the present application provides a computer program product comprising instructions, which, when executed on a computer device, enables the computer device to execute the operation and maintenance indicator abnormality evaluation method as described in the first aspect.

[0039] In a fifth aspect, the present application provides a computer device, characterized in that it includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus;

[0040] Memory, used to store computer programs;

[0041] The processor is used to implement the operation and maintenance indicator abnormality evaluation method as described in the first aspect when executing the program stored in the memory.

[0042] The descriptions of the second to fifth aspects of the present application can refer to the detailed description of the first aspect; and the beneficial effects of the descriptions of the second to fifth aspects can refer to the beneficial effect analysis of the first aspect, which will not be repeated here.

[0043] Based on the above technical solution, the beneficial effects of this application compared with the prior art are as follows:

[0044] The operation and maintenance indicator anomaly evaluation method provided in the embodiment of the present application can not only reflect the overall trend of the time series of the operation and maintenance KPI indicator, but also reflect the difference between each abnormal value in the time series of the operation and maintenance KPI indicator by comprehensively considering the two attribute characteristics of abnormal deviation and abnormal duration, and considering the abnormal amplitude of all time points in the time series of the operation and maintenance KPI indicator and the duration of all abnormalities in the time series. The obtained abnormal score can better reflect the abnormal information of the real fault in the time series of the operation and maintenance KPI indicator, avoid or reduce high-level false alarms, and enable the operation and maintenance personnel to understand the cause and level of the alarm more clearly, so as to make a more accurate response, which is conducive to improving the quality and efficiency of the operation and maintenance work.

[0045] Compared with the traditional anomaly evaluation method, the operation and maintenance indicator anomaly evaluation method provided in the embodiment of the present application can more accurately capture and evaluate the changes of KPI indicators in the actual operation and maintenance process by comprehensively considering the two attribute characteristics of anomaly deviation and anomaly duration, thereby closely fitting the implementation logic of anomaly detection, and has better robustness and higher accuracy of alarm levels, thereby improving the accuracy and practicality of alarms. In addition, when calculating anomaly scores, no complex models or large amounts of computing resources are required, achieving lightweight and efficient anomaly evaluation, reducing operation and maintenance costs, and improving system performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 A flowchart of the method for evaluating abnormal operation and maintenance indicators provided in an embodiment of the present application;

[0047] Figure 2 A schematic diagram of a real-time update process of the probability distribution of abnormal deviation and abnormal duration in the abnormal scoring process provided in an embodiment of the present application;

[0048] Figure 3 A schematic diagram of an abnormal example provided in an embodiment of the present application;

[0049] Figures 4a to 4c A schematic diagram of abnormal evaluation of operation and maintenance indicators of a certain system's call count indicator provided in the application embodiment;

[0050] Figure 5 A schematic diagram of the structure of the operation and maintenance indicator abnormality evaluation device provided in an embodiment of the present application;

[0051] Figure 6 A structural block diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0052] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0053] The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.

[0054] The terms "first" and "second" and the like in the specification and drawings of this application are used to distinguish different objects, or to distinguish different processing of the same object, rather than to describe a specific order of objects.

[0055] In addition, the terms "including" and "having" and any variations thereof mentioned in the description of the present application are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device comprising a series of steps or units is not limited to the listed steps or units, but may optionally include other steps or units that are not listed, or may optionally include other steps or units that are inherent to these processes, methods, products or devices.

[0056] It should be noted that, in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.

[0057] In the description of the present application, unless otherwise specified, “plurality” means two or more.

[0058] The inventors found that in the prior art, when performing abnormal evaluation, since the abnormality is determined based on the KPI indicator data at a single time point, the alarm classification is performed based on the ratio of the abnormal value to the threshold, which is easily affected by random abnormal data and ignores the change of the overall trend. The alarm level is set according to the ratio of the number of abnormalities in the most recent preset time period, and each abnormality is treated equally (that is, only determining whether the indicator data at each time point is an abnormal value), but in fact the deviation scale of each abnormality may be different. When encountering KPI indicators with frequent noise, it is easy to generate many high-level false alarms caused by noise. The above two methods are both prone to high-level false alarms, while ignoring the abnormal information reflecting the real fault, affecting the work efficiency of the operation and maintenance personnel. Based on this, the embodiment of the present application provides an abnormal evaluation method for operation and maintenance indicators. The following will be described in detail with reference to the drawings of the specification.

[0059] First, several nouns used in the embodiments of the present application are explained as follows:

[0060] 1. Anomaly (including abnormal events, continuous anomalies or group anomalies): composed of multiple single-point anomalies.

[0061] 2. Anomaly score significance: used to describe the importance of the anomaly. The more important the anomaly is, the more attention it deserves. The higher the anomaly score is, the higher the alarm level corresponding to the anomaly is.

[0062] 3. Two attribute characteristics of anomalies:

[0063] (1) Deviation: The degree to which the anomaly deviates from the normal value / threshold.

[0064] (2) Abnormal duration: the duration of the abnormality.

[0065] Reference Figure 1 As shown, the operation and maintenance indicator abnormality evaluation method provided in this application includes the following steps:

[0066] S101: establishing a probability distribution of an abnormal deviation sequence and an abnormal duration sequence of the operation and maintenance KPI indicator according to the time series of the operation and maintenance KPI indicator obtained in real time;

[0067] S102: Determine an abnormal deviation probability value and an abnormal duration probability value respectively according to the probability distribution of the abnormal deviation sequence and the abnormal duration sequence of the operation and maintenance KPI indicator;

[0068] S103: determining an abnormality score according to the abnormal deviation probability value and the abnormal duration probability value based on a preset abnormality scoring algorithm;

[0069] S104: Determine an alarm level according to the abnormality score and a preset correspondence between the abnormality score and the alarm level, and generate an alarm.

[0070] Regarding the changes in KPI indicators during system operation, that is, during the normal operation of the system, it will be affected by the noise of many external factors, such as network jitter, memory release, etc., which will cause the KPI indicators to fluctuate violently and briefly, or produce subtle and long-term fluctuations. When a real failure occurs in the system, the KPI indicators often fluctuate violently and for a long time. The inventor considers the characteristics of KPI indicator anomalies in the above-mentioned operation and maintenance scenarios, and proposes an operation and maintenance indicator anomaly evaluation method of the embodiment of the present application. In the operation and maintenance anomaly evaluation process, the abnormal deviation is innovatively introduced to measure the degree of abnormality at all time points in the time series of the operation and maintenance KPI indicator. By combining the two attribute characteristics of the abnormal deviation and the abnormal duration of the KPI indicator anomaly, the anomalies generated by the KPI indicator monitoring are scored, and the alarm level is determined according to the abnormal score and the preset abnormal score and alarm level correspondence. The higher the abnormal score, the higher the corresponding alarm level, thereby generating an alarm that can better reflect the actual abnormality of the KPI indicator.

[0071] The operation and maintenance indicator anomaly evaluation method provided in the embodiment of the present application, when performing anomaly evaluation, comprehensively considers the two attribute characteristics of anomaly deviation and anomaly duration, and simultaneously considers the anomaly amplitude at all time points in the time series of the operation and maintenance KPI indicator and the duration of all anomalies in the time series. It can not only reflect the overall trend of the time series of the operation and maintenance KPI indicator, but also reflect the difference between each anomaly value in the time series of the operation and maintenance KPI indicator. The obtained anomaly score can better reflect the abnormal information of the real fault in the time series of the operation and maintenance KPI indicator, avoid or reduce high-level false alarms, and enable the operation and maintenance personnel to understand the cause and level of the alarm more clearly, so as to make a more accurate response, which is conducive to improving the quality and efficiency of the operation and maintenance work.

[0072] Compared with the traditional anomaly evaluation method, the operation and maintenance indicator anomaly evaluation method provided in the embodiment of the present application can more accurately capture and evaluate the changes of KPI indicators in the actual operation and maintenance process, and improve the accuracy and practicality of the alarm, by comprehensively considering the two attribute characteristics of anomaly deviation and anomaly duration, thereby closely fitting the implementation logic of anomaly detection. In addition, when calculating the anomaly score, no complex model or a large amount of computing resources are required, which realizes lightweight and efficient anomaly evaluation, reduces operation and maintenance costs, and improves system performance.

[0073] In an optional embodiment, in the above step S101, the probability distribution of the abnormal deviation sequence and the abnormal duration sequence of the operation and maintenance KPI indicator is established according to the time series of the operation and maintenance KPI indicator obtained in real time, specifically including:

[0074] According to the indicator values ​​of each observation point in the time series of the operation and maintenance KPI indicators obtained in real time and the preset baseline threshold sequence, the abnormal deviation of each observation point is determined, and whether each observation point is an abnormal value is determined;

[0075] Obtaining an abnormal deviation sequence of the operation and maintenance KPI indicator according to the deviation value of each observation point in the time series of the operation and maintenance KPI indicator, and determining a probability distribution of the abnormal deviation sequence of the operation and maintenance KPI indicator;

[0076] According to the abnormal values ​​in all observation points in the time series of the operation and maintenance KPI indicator, an abnormal duration sequence of the operation and maintenance KPI indicator is obtained, and a probability distribution of the abnormal duration sequence of the operation and maintenance KPI indicator is determined.

[0077] In the embodiment of the present application, the time length of the time series of the operation and maintenance KPI indicator can be set according to the actual situation, and no specific limitation is made here. For example, for a KPI indicator with a daily cycle feature, the time length of the time series can be set to 5 minutes.

[0078] In a specific embodiment, for any observation point in the time series of the operation and maintenance KPI indicator, the abnormal deviation of the observation point is determined in the following manner:

[0079] deviation=value-normal;

[0080] Among them, deviation represents the abnormal deviation of the observation point, value represents the indicator data of an observation point in the time series of the operation and maintenance KPI indicator, and normal represents the preset baseline threshold of the corresponding observation point.

[0081] In a specific embodiment, for any observation point in the time series of the operation and maintenance KPI indicator, if the indicator data of the observation point is within the baseline threshold range of the corresponding observation point, then the observation point is a normal value (recorded as 0), otherwise the observation point is an abnormal value (recorded as 1).

[0082] In the embodiment of the present application, by calculating the deviation value of each observation point in the time series of the operation and maintenance KPI indicator, the deviation values ​​of all observation points are arranged in chronological order to obtain the abnormal deviation sequence of the operation and maintenance KPI indicator. In addition, by calculating the abnormal values ​​of all observation points in the time series of the operation and maintenance KPI indicator, the normal value 0 and the abnormal value 1 of all observation points are arranged in chronological order to obtain the abnormal duration sequence of the operation and maintenance KPI indicator.

[0083] In the embodiment of the present application, a probability distribution is established according to the abnormal deviation sequence of the operation and maintenance KPI indicator and the abnormal duration sequence of the operation and maintenance KPI indicator. In the embodiment of the present application, the type of the above probability distribution can be uniform distribution, Gaussian distribution, mixed Gaussian distribution, etc. Of course, it can also be other distribution types recorded in the prior art, and no specific limitation is made in the embodiment of the present application.

[0084] Furthermore, based on step S102, according to the probability distribution of the abnormal deviation sequence of the operation and maintenance KPI indicator and the probability distribution of the abnormal duration sequence of the operation and maintenance KPI indicator, the abnormal deviation probability value P (de<=deviation(t)) and the abnormal duration probability value (du<=duration(t)) are calculated respectively.

[0085] For example, assuming that the probability distribution of the abnormal deviation sequence of the operation and maintenance KPI indicator is normally distributed, then the probability distribution function of the normal distribution is:

[0086]

[0087] Correspondingly, the abnormal deviation probability value P(de<=deviation(t)) is calculated based on the following formula:

[0088]

[0089] Similarly, assuming that the probability distribution of the abnormal duration of the operation and maintenance KPI indicator is also normally distributed, the abnormal deviation probability value P (de<=deviation(t)) can be calculated in the same way according to the probability distribution of the abnormal duration.

[0090] In an optional embodiment, in the above step S103, the abnormality score is determined based on the abnormal deviation probability value and the abnormality duration probability value based on a preset abnormality scoring algorithm, specifically including:

[0091] The abnormality score is calculated based on the following formula 1 according to the abnormal deviation probability value and the abnormal duration probability value, as well as the weight of the abnormal deviation and the weight of the abnormal duration:

[0092]

[0093] Among them, significance(t) represents the abnormal score, w deviation Indicates the weight of the abnormal deviation of the pre-set operation and maintenance KPI indicator, 0≤w deviation ≤1,w duration Indicates the weight of the abnormal duration of the pre-set operation and maintenance KPI indicator, 0≤w duration ≤1, P(de<=deviation(t)) represents the probability value of abnormal deviation, and P(du<=duration(t)) represents the probability value of abnormal duration.

[0094] In the embodiment of the present application, since it is necessary to monitor each KPI indicator in real time in the operation and maintenance KPI indicator monitoring scenario to determine whether a new data point generated every minute is abnormal, therefore, refer to Figure 2 As shown in the figure, during the real-time monitoring process, each KPI indicator detection task will record the probability distribution of the abnormal deviation sequence deviation of the KPI indicator and the distribution type and key parameters (including the size of the time window and the parameters corresponding to the distribution type) of the probability distribution of the abnormal duration sequence duration, that is, Figure 2 The deviation scoring model and duration scoring model in the above formula 1 are used to calculate the deviation score model and duration score model when a new anomaly occurs. Figure 2 The comprehensive scoring model in is used to calculate the anomaly score in real time. And after the score calculation is completed, the probability distribution is constructed with the new anomaly deviation sequence and anomaly duration sequence of the KPI indicator and the probability distribution parameters are updated, that is, Figure 2 The deviation scoring model and duration scoring model are updated as shown in the figure to obtain the updated deviation scoring model and duration scoring model, so that new abnormal deviation probability values ​​and abnormal duration probability values ​​can be re-acquired, and new abnormal scores are calculated according to the abnormal score significance formula to achieve automatic update and adaptation of the KPI indicator data distribution. In other words, the probability distribution of the abnormal deviation sequence and the abnormal duration sequence is dynamic. Therefore, the scoring formula can continuously adapt to the new model to obtain the corresponding abnormal score.

[0095] In the embodiment of the present application, the weight w of the abnormal deviation of the above-preset operation and maintenance KPI indicator can be adjusted according to the actual operation and maintenance scenario. deviation and the weight w of the abnormal duration of the pre-set operation and maintenance KPI indicator duration If there is annotation information for anomalies or alarms, the scoring formula can be automatically optimized based on these annotations by adjusting the weight w of the anomaly deviation. deviation and the weight of the abnormal duration wduration , can reduce the level of false alarms and improve the accuracy of alarms. This adaptive optimization capability enables this method to continuously learn and improve, and better adapt to different operation and maintenance scenarios.

[0096] Based on the abnormal evaluation method of operation and maintenance indicators provided in the embodiment of the present application, an abnormal / alarm annotation feedback mechanism can be introduced. After receiving the alarm notification, the operation and maintenance personnel can score the accuracy of the alarm level. For example, if the operation and maintenance personnel believe that many long duration low deviations are marked as false alarms, the weight of the abnormal deviation w can be adjusted in the weighted integration process of the abnormal score. deviation and the weight of the abnormal duration w duration , to reduce the impact of abnormal deviation duration on abnormal score.

[0097] In an embodiment of the present application, the above-mentioned baseline threshold may include an upper threshold and a lower threshold. In the process of abnormal evaluation (i.e., abnormal conversion alarm), according to the difference in KPI indicator types, those skilled in the art may give different degrees of attention to the upper limit abnormality (i.e., the KPI indicator data exceeds the upper threshold) and the lower limit abnormality (i.e., the KPI indicator data is lower than the lower threshold). The abnormalities of the same KPI indicator are divided into upper limit abnormalities and lower limit abnormalities. The upper limit abnormalities and the lower limit abnormalities respectively construct their own probability distributions. The specific implementation method may not be specifically limited here.

[0098] In an optional embodiment, before executing the above step S104, the method may further include:

[0099] The anomaly score is processed in the following manner to obtain a processed anomaly score:

[0100] S(t)=significance(t) / max(significance(t))*100;

[0101] Among them, S(t) represents the anomaly score after processing, significance(t) represents the anomaly score, and max(significance(t)) represents the maximum value of the anomaly score at all times in the time series of the KPI indicator to be evaluated.

[0102] In the embodiment of the present application, the above-mentioned processing method can be used to normalize the anomaly score of any scenario in the operation and maintenance scenario to within 0 to 100. The operation and maintenance personnel can convert the anomaly into different levels of alarms according to the size of the anomaly score through a unified conversion rule, without setting different alarm conversion rules according to the value ranges of different indicators. Compared with the traditional anomaly evaluation method, in which the alarm is graded according to the ratio of the anomaly value to the threshold, or the alarm level is set according to the proportion of the number of anomalies in the most recent preset time period, by normalizing the anomaly score of any scenario in the operation and maintenance scenario to between 0 and 100, the operation and maintenance personnel can determine the value range of the anomaly score corresponding to each alarm level according to the divided alarm levels, and obtain the preset correspondence between the anomaly score and the alarm level, that is, the preset correspondence between the anomaly score and the alarm level includes multiple alarm levels and the value range of the anomaly score corresponding to each alarm level. For example, the alarm levels include "normal", "serious" and "catastrophic", and the corresponding abnormality score ranges are 0-30, 31-70, and 70-100, respectively, so as to more reasonably determine the abnormal alarm level.

[0103] For example, refer to Figure 3 As shown, it is an example of an anomaly in the anomaly evaluation process in a KPI indicator, wherein the red dot represents a single-point anomaly corresponding to the observation point on the time series of the KPI indicator, the blue dot represents the baseline threshold of the observation point on the time series of the KPI indicator, and the blue box represents the identified anomaly. Exemplarily, for the observation point a1, the corresponding anomaly deviation value (deviation) is the indicator value corresponding to the red dot minus the upper threshold value in the baseline threshold corresponding to the blue dot, and it can be seen that the duration value of the abnormal duration (duration) of the anomaly is 1, and by determining the abnormal deviation value and abnormal duration of all observation points including a1 to a8, the corresponding abnormal deviation sequence and abnormal duration sequence are obtained, and the probability distribution of the abnormal deviation sequence and the abnormal duration sequence is established, and the abnormal deviation probability value and the abnormal duration probability value are calculated, and then based on the following formula 1, the abnormal score is calculated.

[0104] In an optional embodiment, in the above step S104, determining the alarm level and generating the alarm according to the abnormality score and the preset correspondence between the abnormality score and the alarm level specifically includes:

[0105] According to the correspondence between the processed anomaly score and the preset anomaly score and the alarm level, the anomaly score value range to which the processed anomaly score belongs is determined, the corresponding alarm level is obtained, and an alarm corresponding to the alarm level is generated.

[0106] Exemplarily, assuming that the value of the processed anomaly score is 40, the preset correspondence between the anomaly score and the alarm level includes three levels: "normal", "serious", and "disaster", and the corresponding anomaly score value ranges are 0-30, 31-70, and 70-100, respectively. Then the alarm level is determined to be "serious", thereby generating an alarm corresponding to the "serious" alarm level, and sending an alarm message to the operation and maintenance personnel.

[0107] In an embodiment of the present application, by recording and precipitating the distribution of anomalies during the anomaly evaluation process, business knowledge can be quickly applied to other sequences of the same KPI indicator. For example, the distribution of anomalies of the KPI indicator CPU utilization in an IT enterprise is precipitated as business knowledge and applied to the anomaly evaluation of the time series of the KPI indicator CPU utilization in a financial enterprise, thereby achieving the sharing and reuse of operation and maintenance knowledge and improving operation and maintenance efficiency.

[0108] In order to more clearly illustrate the specific implementation process of the embodiment of the present application, Figures 4a to 4c Taking the call count indicator monitoring process of a certain system as an example, the specific implementation process of abnormal evaluation of operation and maintenance indicators is described in detail as follows:

[0109] Figure 4a The blue line in the call count indicator of the system shown is the number of calls per minute, which has a daily cycle characteristic. The blue shaded part is the upper and lower threshold range of the indicator, and the red dot is the abnormal point where the indicator exceeds the upper threshold. Figure 4a The call number indicator sequence shown in FIG. 1 is based on the above step S101. First, the indicator data in the call number indicator is subtracted from the upper threshold to obtain a residual value, such as Figure 4b Then, according to Figure 4b The residual value results are used to determine the abnormal deviation sequence and abnormal duration sequence corresponding to the call number indicator, and the probability distribution of the abnormal deviation sequence and abnormal duration sequence is established.

[0110] Next, according to the above step S102, the abnormal deviation probability value and the abnormal duration probability value are calculated.

[0111] Next, according to the above step S103, the anomaly score (significance) of each anomaly is determined, and the anomaly score of each anomaly of the call number indicator is finally obtained as follows: Figure 4c The anomaly score of each anomaly is scored by combining the anomaly deviation and anomaly duration of the anomaly. The larger the anomaly deviation and anomaly duration, the larger the anomaly score.

[0112] Finally, based on the above step S104, the alarm level of the abnormality is determined, an alarm corresponding to the alarm level is generated, and the alarm information is sent to the operation and maintenance personnel.

[0113] Embodiment 2

[0114] Based on the same inventive concept, the present application also provides an abnormal operation and maintenance indicator evaluation device, referring to Figure 5 As shown, the device comprises:

[0115] The probability distribution determination module 101 is used to establish the probability distribution of the abnormal deviation sequence and the abnormal duration sequence of the operation and maintenance KPI indicator according to the time series of the operation and maintenance KPI indicator obtained in real time;

[0116] The probability value determination module 102 is used to determine the abnormal deviation probability value and the abnormal duration probability value respectively according to the probability distribution of the abnormal deviation sequence and the abnormal duration sequence of the operation and maintenance KPI indicator;

[0117] An anomaly score determination module 103, configured to determine an anomaly score according to the anomaly deviation probability value and the anomaly duration probability value based on a preset anomaly score algorithm;

[0118] The alarm module 104 is used to determine the alarm level and generate an alarm according to the abnormality score and the preset corresponding relationship between the abnormality score and the alarm level.

[0119] Embodiment 3

[0120] Based on the same inventive concept, an embodiment of the present application also provides a computer-readable storage medium, in which instructions are stored. When the instructions are executed on a terminal, the terminal executes the above-mentioned operation and maintenance indicator abnormality evaluation method.

[0121] Among them, the computer readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a register, a hard disk, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above, or any other form of computer readable storage medium known in the art. An exemplary storage medium is coupled to a processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an application-specific integrated circuit (ASIC). In the embodiments of the present application, a computer-readable storage medium may be any tangible medium that contains or stores a program, which may be used by or in conjunction with an instruction execution system, apparatus, or device.

[0122] Embodiment 4

[0123] Based on the same inventive concept, an embodiment of the present application also provides a computer program product comprising instructions. When the computer program product is run on a computer, the computer executes the above-mentioned operation and maintenance indicator abnormality evaluation method.

[0124] Embodiment 5

[0125] Based on the same inventive concept, an embodiment of the present application further provides a computer device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;

[0126] Memory, used to store computer programs;

[0127] The processor is used to implement the above-mentioned operation and maintenance indicator abnormality evaluation method when executing the program stored in the memory.

[0128] Figure 6A possible structural diagram of the computer device involved in the above embodiment is shown. The computer device includes: a processor 1002 and a communication interface 1003. The processor 1002 is used to control and manage the actions of the computer device, for example, to execute the above-mentioned operation and maintenance indicator abnormality evaluation method, and / or to execute other processes of the technology described in this article. The communication interface 1003 is used to support the communication between the computer device and other network entities, for example, to execute the steps performed by the above-mentioned communication unit 902. The computer device may also include a memory 1001 and a bus 1004, and the memory 1001 is used to store program code and data of the computer device.

[0129] Among them, the memory 1001 can be a memory in a computer device, etc. The memory may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a read-only memory, a flash memory, a hard disk or a solid-state drive; the memory may also include a combination of the above types of memory.

[0130] The processor 1002 may be a processor that implements or executes various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of the present application. The processor may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field programmable gate array, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of the present application. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.

[0131] The bus 1004 may be an Extended Industry Standard Architecture (EISA) bus, etc. The bus 1004 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0132] Through the description of the above implementation methods, technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device and unit described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0133] Since the computer-readable storage medium, computer program product, and computer device in the embodiments of the present application can be applied to the above-mentioned method, the technical effects that can be obtained can also refer to the above-mentioned method embodiments, and the embodiments of the present application will not be repeated here.

[0134] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0135] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0136] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0137] The above are only specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A method for evaluating abnormal operation and maintenance indicators, characterized in that: include: According to the time series of operation and maintenance KPI indicators obtained in real time, the probability distribution of abnormal deviation series and abnormal duration series of operation and maintenance KPI indicators is established; According to the probability distribution of the abnormal deviation sequence and the abnormal duration sequence of the operation and maintenance KPI indicators, the abnormal deviation probability value and the abnormal duration probability value are determined respectively; Determine an abnormality score according to the abnormal deviation probability value and the abnormal duration probability value based on a preset abnormality scoring algorithm; According to the abnormality score and the preset corresponding relationship between the abnormality score and the alarm level, the alarm level is determined and an alarm is generated.

2. The method according to claim 1, characterized in that The method of establishing the probability distribution of the abnormal deviation sequence and the abnormal duration sequence of the operation and maintenance KPI indicator according to the time series of the operation and maintenance KPI indicator obtained in real time includes: According to the indicator values ​​of each observation point in the time series of the operation and maintenance KPI indicators obtained in real time and the preset baseline threshold sequence, the abnormal deviation of each observation point is determined, and whether each observation point is an abnormal value is determined; Obtaining an abnormal deviation sequence of the operation and maintenance KPI indicator according to the deviation value of each observation point in the time series of the operation and maintenance KPI indicator, and determining a probability distribution of the abnormal deviation sequence of the operation and maintenance KPI indicator; According to the abnormal values ​​in all observation points in the time series of the operation and maintenance KPI indicator, an abnormal duration sequence of the operation and maintenance KPI indicator is obtained, and a probability distribution of the abnormal duration sequence of the operation and maintenance KPI indicator is determined.

3. The method according to claim 1, characterized in that Also includes: For any observation point in the time series of the operation and maintenance KPI indicator, the abnormal deviation of the observation point is determined in the following way: deviation=value-normal; Among them, deviation represents the abnormal deviation of the observation point, value represents the indicator data of an observation point in the time series of the operation and maintenance KPI indicator, and normal represents the preset baseline threshold of the corresponding observation point.

4. The method according to claim 1, characterized in that The step of determining the abnormality score according to the abnormal deviation probability value and the abnormal duration probability value based on a preset abnormality scoring algorithm includes: The abnormality score is calculated based on the following formula 1 according to the abnormal deviation probability value and the abnormal duration probability value, as well as the weight of the abnormal deviation and the weight of the abnormal duration: Among them, significance(t) represents the abnormal score, w deviation Indicates the weight of the abnormal deviation of the pre-set operation and maintenance KPI indicator, 0≤w deviation ≤1,w duration Indicates the weight of the abnormal duration of the pre-set operation and maintenance KPI indicator, 0≤w duration ≤1, P(de<=deviation(t)) represents the probability value of abnormal deviation, and P(du<=duration(t)) represents the probability value of abnormal duration.

5. The method according to claim 4, characterized in that Before determining the alarm level according to the abnormality score and the preset abnormality score and alarm level correspondence rule and generating an alarm, the method further includes: The anomaly score is processed in the following manner to obtain a processed anomaly score: S(t)=significance(t) / max(significance(t))*100; Among them, S(t) represents the anomaly score after processing, significance(t) represents the anomaly score, and max(significance(t)) represents the maximum value of the anomaly score at all times in the time series of the KPI indicator to be evaluated.

6. The method according to claim 5, characterized in that The preset correspondence between the abnormality score and the alarm level includes multiple alarm levels and the abnormality score value range corresponding to each alarm level; The determining the alarm level and generating the alarm according to the abnormality score and the preset abnormality score and alarm level correspondence relationship includes: According to the correspondence between the processed anomaly score and the preset anomaly score and the alarm level, the anomaly score value range to which the processed anomaly score belongs is determined, the corresponding alarm level is obtained, and an alarm corresponding to the alarm level is generated.

7. A device for evaluating abnormal operation and maintenance indicators, characterized in that: include: A probability distribution determination module is used to establish the probability distribution of the abnormal deviation sequence and the abnormal duration sequence of the operation and maintenance KPI indicators according to the time series of the operation and maintenance KPI indicators obtained in real time; A probability value determination module is used to determine the abnormal deviation probability value and the abnormal duration probability value respectively according to the probability distribution of the abnormal deviation sequence and the abnormal duration sequence of the operation and maintenance KPI indicator; an abnormality score determination module, configured to determine an abnormality score according to the abnormal deviation probability value and the abnormal duration probability value based on a preset abnormality score algorithm; The alarm module is used to determine the alarm level and generate an alarm according to the abnormality score and the preset corresponding relationship between the abnormality score and the alarm level.

8. A computer-readable storage medium, wherein instructions are stored in the computer-readable storage medium. When the instructions are executed on a terminal, the terminal executes the operation and maintenance indicator abnormality evaluation method according to any one of claims 1 to 6.

9. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the operation and maintenance indicator abnormality evaluation method described in any one of claims 1 to 6 is implemented.

10. A computer device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory, used to store computer programs; The processor is used to implement the operation and maintenance indicator abnormality evaluation method as described in any one of claims 1 to 6 when executing the program stored in the memory.

Citation Information

Cited By

  • Early warning method and system for operation and maintenance delivery abnormal event based on cloud platform

    CN120811863A

  • User tag anomaly detection method and device, equipment and medium

    CN121637325A

  • User tag anomaly detection method, device, equipment and medium

    CN121637325B