A fault detection method and system

By calculating anomaly scores and using multi-model learning, fault information is filtered and extracted, solving the problem of difficult fault location in large IT enterprises and achieving efficient and accurate fault detection.

CN115061838BActive Publication Date: 2025-11-21JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202210316935.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-28
Publication Date
2025-11-21
Estimated Expiration
2042-03-28

AI Technical Summary

Technical Problem

In large IT enterprises, when the system is large in scale and the applications are diverse, existing technologies are difficult to efficiently locate faults, resulting in a large number of noise alarms and difficulties in troubleshooting for operations and maintenance engineers.

Method used

By calculating the anomaly score of alarm data, an unsupervised learning model is used to filter suspected fault alarm data, and a supervised learning model is combined to further extract fault information, including the root cause of the fault and prediction information.

Benefits of technology

It effectively reduces troubleshooting time, improves the efficiency and accuracy of fault detection, and can quickly filter out key fault information from massive alarm data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115061838B_ABST
    Figure CN115061838B_ABST
Patent Text Reader

Abstract

The disclosure provides a fault detection method and system, the method comprising: acquiring a large amount of alarm data; calculating the abnormality score of each alarm data, and screening the alarm data with an abnormality score higher than a first threshold as abnormal alarm data; performing first processing on the abnormal alarm data based on a pre-stored first model to obtain suspected fault alarm data; performing second processing on the suspected fault alarm data based on a pre-stored second model to obtain alarm fault data, and extracting fault information from the alarm fault data. The method disclosed in the disclosure can efficiently obtain relatively accurate fault information from a large amount of alarm data, and reduce the troubleshooting time of the fault.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of computer application, and particularly relates to a fault detection method and system. BACKGROUND

[0002] The business of large IT enterprises is complex, and there are numerous monitoring indicators. In order to ensure the safe and stable operation of the service system, it is necessary to detect possible system faults in real time. Developers often design a large number of monitoring rules in the script execution process, thereby issuing an alarm for abnormal indicators.

[0003] The prior art usually uses a clustering algorithm to perform alarm clustering summary on a large amount of alarm, and then hands over the alarm to an operation and maintenance engineer for analysis. When the system is large in scale and the application is various in type, the alarm data obtained after clustering is still large in scale and is mixed with a large amount of noise alarm, thereby increasing the difficulty of troubleshooting for the operation and maintenance engineer and making it very difficult to locate the fault.

[0004] Therefore, how to reduce the troubleshooting time of the fault and realize efficient fault detection is an important topic that needs to be solved in the industry. SUMMARY

[0005] The fault detection method and system provided by the present disclosure solve the defects that the prior art has a large difficulty and low efficiency in locating the fault of the alarm data when the system is large in scale and the application is various in type, so that the troubleshooting time of the fault can be reduced and the efficiency of fault detection can be improved.

[0006] The present disclosure provides a fault detection method, comprising:

[0007] Obtaining a large amount of alarm data;

[0008] Calculating a first abnormality score of each of the alarm data, and screening the alarm data with a first abnormality score higher than a first threshold value as abnormal alarm data;

[0009] Performing first processing on the abnormal alarm data based on a pre-stored first model to obtain suspected fault alarm data;

[0010] Performing second processing on the suspected fault alarm data based on a pre-stored second model to obtain alarm fault data, and extracting fault information from the alarm fault data.

[0011] According to the fault detection method provided by the present disclosure, the first abnormality score of each alarm data is calculated, including: performing hierarchical analysis on the alarm data based on a pre-stored alarm template to obtain alarm template data; performing periodic analysis on the alarm template data to obtain first classification data, wherein the first classification data includes periodic alarm data and non-periodic alarm data; performing rarity analysis on the alarm template data to obtain second classification data; wherein the second classification data includes high-frequency alarm data and low-frequency alarm data; and calculating the first abnormality score based on the first classification data and the second classification data.

[0012] According to the fault detection method provided by the present disclosure, the periodic analysis on the alarm template data to obtain first classification data includes: extracting minute-level aggregation features of the alarm template data; and dividing the alarm template data into the periodic alarm data and the non-periodic alarm data based on the minute-level aggregation features.

[0013] According to the fault detection method provided by the present disclosure, the rarity analysis on the alarm template data to obtain second classification data includes: performing aggregation processing on repeatedly occurring data in the alarm template data to obtain high-frequency alarm data, and taking other data in the alarm template data except the high-frequency alarm data as low-frequency alarm data.

[0014] According to the fault detection method provided by the present disclosure, the calculation of the first abnormality score based on the first classification data and the second classification data includes: extracting a trend component and a residual of the periodic alarm data in the first classification data, and calculating a second abnormality score of the periodic alarm data according to the trend component and the residual; in the case that the alarm data is the non-periodic alarm data and the low-frequency alarm data, increasing the second abnormality score by a first proportion to obtain the first abnormality score; and in the case that the alarm data is the periodic alarm data and the high-frequency alarm data, reducing the second abnormality score by a second proportion to obtain the first abnormality score.

[0015] According to the fault detection method provided by the present disclosure, the first processing of the abnormal alarm data based on a pre-stored first model to obtain suspected fault alarm data includes: extracting alarm features in the abnormal alarm data; inputting the alarm features into the first model for recall processing of fault alarms, and obtaining the suspected fault alarm data from the abnormal alarm data according to a result of the recall processing.

[0016] According to the fault detection method provided in the present disclosure, the abnormal alarm data is one-hot encoded to obtain an encoding processing result, and alarm timing features and alarm state distribution features are extracted from the encoding processing result; the alarm timing features and the alarm state distribution features are taken as the alarm features; wherein the alarm timing features include minute-level granularity aggregation features, minute-level alarm application numbers, and minute-level maximum application numbers; and the alarm state distribution features include alarm timing number distribution features and alarm timing frequency distribution features.

[0017] According to the fault detection method provided in the present disclosure, the alarm features are input into the first model for recall processing of fault alarms, and the suspected fault alarm data is obtained from the abnormal alarm data according to a result of the recall processing, which includes: performing the recall processing on the alarm features by using the first model, and extracting abnormal time points containing suspected fault information from a result of the recall processing; performing fault scoring on each of the abnormal time points to obtain a fault score corresponding to each of the abnormal time points; and in a case where the fault score is higher than a second threshold, taking alarm data corresponding to the abnormal time points as the suspected fault alarm data.

[0018] According to the fault detection method provided in the present disclosure, the suspected fault alarm data is subjected to a second processing based on a pre-stored second model to obtain fault alarm information, which includes: inputting the suspected fault alarm data into the second model for screening to obtain fault alarm data, and obtaining the fault alarm information from the fault alarm data; wherein the fault alarm information includes fault root cause information and fault prediction information.

[0019] The present disclosure further provides a fault detection system, which includes:

[0020] An alarm data obtaining unit is configured to obtain a large amount of alarm data; an abnormal alarm data obtaining unit is configured to calculate abnormality scores of the alarm data, and screen the alarm data with abnormality scores higher than a first threshold as abnormal alarm data; a suspected fault alarm data obtaining unit is configured to perform a first processing on the abnormal alarm data based on a pre-stored first model to obtain suspected fault alarm data; and an alarm fault information obtaining unit is configured to perform a second processing on the suspected fault alarm data based on a pre-stored second model to obtain alarm fault data, and extract fault information from the alarm fault data.

[0021] The present disclosure further provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the fault detection method according to any one of the above when executing the program.

[0022] The disclosure also provides a non-transitory computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, implements any of the above-mentioned fault detection methods.

[0023] The disclosure also provides a computer program product comprising a computer program, which, when executed by a processor, implements any of the above-mentioned fault detection methods.

[0024] The fault detection method and system provided by the disclosure can quickly filter out abnormal alarm data from a large amount of alarm data through the abnormality scores of the alarm data, and perform unsupervised learning on the abnormal alarm data by using a first model to obtain suspected fault alarm data, thereby ensuring the recall rate of the suspected fault alarm data. In addition, supervised learning is performed on the suspected fault alarm data by using a second model, further filtering out alarm fault data from the suspected fault alarm data, and extracting the required fault information from the alarm fault data, so as to locate and predict the fault according to the fault information. The method disclosed by the disclosure can efficiently extract accurate fault information from a large amount of alarm data, and reduce the troubleshooting time of the fault. BRIEF DESCRIPTION OF DRAWINGS

[0025] In order to more clearly illustrate the technical solutions in the disclosure or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are some embodiments of the disclosure, and other drawings can also be obtained by those skilled in the art without creative labor.

[0026] Figure 1 is one of the flowcharts of the fault detection method provided by the embodiments of the disclosure;

[0027] Figure 2 is the second flowchart of the fault detection method provided by the embodiments of the disclosure;

[0028] Figure 3 is the third flowchart of the fault detection method provided by the embodiments of the disclosure;

[0029] Figure 4 is the fourth flowchart of the fault detection method provided by the embodiments of the disclosure;

[0030] Figure 5 is the interaction diagram of the fault detection method provided by the embodiments of the disclosure;

[0031] Figure 6 is the fifth flowchart of the fault detection method provided by the embodiments of the disclosure;

[0032] Figure 7is a structural schematic diagram of a fault detection system provided by an embodiment of the present disclosure.

[0033] Figure 8 is a structural schematic diagram of an electronic device provided by the present disclosure. DETAILED DESCRIPTION

[0034] For the purposes of the present disclosure, technical solutions and advantages, the technical solutions in the present disclosure will be described in detail below with reference to the drawings in the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present disclosure.

[0035] The fault detection method provided by an embodiment of the present disclosure will be described below, which comprises: Figure 1 The fault detection method provided by an embodiment of the present disclosure will be described below, which comprises:

[0036] Step 110, acquiring a large amount of alarm data.

[0037] Alarm refers to an event report composed of a notification issued by a monitored object when a specific event occurs, used to deliver alarm information. Alarm data is time-sensitive, indicating alarm data dynamically generated when a specific event occurs within a certain time period. These alarm data contain various types of information, including characteristic information closely related to the occurrence of the event, and noise information unrelated to the occurrence of the event; for example, for a certain furniture lamp that does not light up under the action of the alarm, there can be many possible reasons, such as burnt filament, circuit voltage loss, temporary power failure, arrears and other information, etc.

[0038] In this step, the alarm data is an alarm message issued by the monitoring center of the system when the server fails. The alarm message can include the corresponding server name, the time when the alarm message is issued, and the alarm topic, etc.

[0039] In this embodiment, one kind of alarm data acquired is "

P3

Warning

External data interface initial version monitoring---field monitoring alarm

[0040] Step 120, calculating the abnormality score of each alarm data, and screening the alarm data with an abnormality score higher than a first threshold value as abnormal alarm data.

[0041] In this step, all alarm data are scored for abnormal alarms, i.e. the abnormality score of each alarm data is calculated, which is used to distinguish the abnormality degree of each alarm data.

[0042] The abnormality alarm scoring of the alarm data can be scoring according to a default security level of the alarm data, scoring according to high and low frequencies of occurrence of each alarm data, scoring according to periodic characteristics of each alarm data, or scoring in other manners, which is not limited in the embodiment.

[0043] In this step, the first threshold value can be a critical value set by human being. When the abnormality score of the alarm data exceeds the critical value, it can be considered that the abnormality degree of the alarm data is large, and the alarm data can contain the required fault information. The alarm data should be extracted for the screening process in the subsequent step. The threshold value set by human experience has better reliability in abnormality scoring of the alarm data that has occurred in history. The first threshold value can also be a dynamic value obtained by intelligent analysis of different time and different types of alarm data. The first threshold value is adaptively adjusted according to whether the alarm data is new alarm data. The dynamic value is suitable for a wider range of data types in abnormality.

[0044] It should be noted that the first threshold value is in the range of 8.5-9.5.

[0045] In this step, Figure 2 In the embodiment shown in FIG. 8, the abnormality scores of each alarm data are calculated, and the alarm data with an abnormality score higher than 9.2 is regarded as abnormal alarm data, otherwise, the alarm data is removed.

[0046] Step 130, performing first processing on the abnormal alarm data based on a pre-stored first model to obtain suspected fault alarm data.

[0047] In this step, the first model is an unsupervised learning model, which is used to learn suspected fault information in the abnormal alarm data and extract corresponding suspected fault alarm data. The first processing refers to a process of inputting the abnormal alarm data into the first model for unsupervised learning, and extracting suspected fault alarm data according to the unsupervised learning result.

[0048] In this step, the first model can be a single unsupervised model, such as a local outlier factor detection model (LOF) and an isolation forest model (IForest), or a comprehensive detection model obtained by combining multiple unsupervised models.

[0049] In this step, Figure 3In the illustrated embodiment, the abnormal alarm data is input into the LOF model, the IForest model, the KNN algorithm and the ABOD algorithm for recall processing, and after multi-layer abnormality detection voting and multi-dimensional weighted fault scoring, the required suspected fault alarm data is screened out.

[0050] Optionally, the unsupervised anomaly detection model comprises a local outlier factor detection model, an isolated forest model, a nearest neighbor model and an angle-based abnormal time point detection model.

[0051] The local outlier factor (LOF) detection model reflects the abnormality degree of a sample by calculating the "local reachable density". The greater the local reachable density of a sample point, the more likely it is an abnormal time point. The isolated forest (IForest) extracts a plurality of samples to construct a plurality of binary trees (iTree), and then calculates the abnormal score of each data point by synthesizing the data points generated by the plurality of binary trees. The basic principle is that data quickly divided into leaf nodes is abnormal data. The KNN model is used to filter abnormal time points according to different K-neighbor distances. The basic idea of the ABOD model for detecting abnormal time points is to calculate the variance of the angle formed by each sample and all other samples. Abnormal time points are far away from normal points, so the variance changes little. The four detection models are used simultaneously in this embodiment to detect abnormal time points from different angles such as density, division hyperplane, distance and angle.

[0052] It can be understood that the recall model belongs to a kind of unsupervised model, which can quickly filter out valuable suspected fault information from a large amount of abnormal alarm data to solve the problem of data overload. It can also be used to fuse the data recalled by multiple channels to obtain a refined suspected fault data set, which can solve the problem of single recall feature, small information quantity and poor diversity.

[0053] It should be noted that in this embodiment, the abnormality score corresponding to the abnormal alarm data can be used as the initial weight of the recall model to obtain the abnormal time points of the abnormal alarm data, and the voting mechanism is used to score the fault of each abnormal time point, so as to obtain the recalled suspected fault alarm data.

[0054] Step 140, performing second processing on the suspected fault alarm data based on the pre-stored second model to obtain alarm fault data, and extracting fault information from the alarm fault data.

[0055] In this step, the second model is a supervised learning model for learning fault information of suspected fault alarm data and extracting corresponding fault alarm data; the second processing refers to the process of inputting suspected fault alarm data into the second model for supervised learning and extracting fault alarm data according to the supervised screening result.

[0056] In this step, the first model can be a single supervised model, such as a multi-classification support vector machine, an XGBoost model, and a deep model, wherein the deep model can be a convolutional neural network or a comprehensive detection model obtained by combining multiple unsupervised models.

[0057] It should be noted that the supervised model is a model obtained by taking historical fault alarm data as a training set and performing supervised training, which is used for fault classification of newly input alarm data.

[0058] In this step, for the above recalled suspected fault alarm data, the fault time information contained therein needs to be further screened and classified by fault type to extract fault alarm data containing determined fault information, wherein the fault information can be fault root cause information for fault localization, and can also be fault prediction information for predicting unknown faults.

[0059] In Figure 4 In the embodiment shown, suspected fault alarm data is input into the trained XGBoost model for classification screening, and fault alarm data is screened from the classification result, and finally fault root cause information and fault prediction information are extracted from the fault alarm data.

[0060] The fault detection method provided by the present disclosure can quickly screen out abnormal alarm data from a large amount of alarm data through the abnormality score of each alarm data, and use the first model to perform unsupervised learning on the abnormal alarm data to obtain suspected fault alarm data, thereby ensuring the recall rate of the suspected fault alarm data. The second model is used to perform supervised learning on the suspected fault alarm data, further screen out fault alarm data from the suspected fault alarm data, and extract the required fault information from the fault alarm data, so as to locate and predict the fault according to the fault information. The method described in the present disclosure can efficiently extract accurate fault information from a large amount of alarm data, and reduce the troubleshooting time of the fault.

[0061] Optionally, the alarm data is hierarchically parsed based on a pre-stored alarm template to obtain alarm template data; the alarm template data is periodically analyzed to obtain first classification data, wherein the first classification data includes periodic alarm data and non-periodic alarm data; the alarm template data is analyzed for rarity to obtain second classification data; wherein the second classification data includes high-frequency alarm data and low-frequency alarm data; and a first abnormality score is calculated based on the first classification data and the second classification data.

[0062] Specifically, the embodiment is based on a frequent template tree (FT-tree) model, and massive alarm data is first hierarchically parsed, and the division attribute of each layer is the inherent category attribute of the alarm data. Then, the alarm data after multi-layer parsing is obtained according to multiple layers to obtain a structured alarm template. The template is essentially an alarm data set, the number of samples in the data set is the alarm data, and the number of layers in the hierarchical parsing is the number of characteristics corresponding to each alarm data.

[0063] In this embodiment, the parsing types of each layer of the FT-tree model are set as extracting alarm time, application name, alarm application name and topic, and the alarm data is input into the FT-tree model according to the parsing types to constitute an alarm template.

[0064] It should be noted that there are a large number of noise alarms in the alarm data. These noise alarms often appear in the form of periodic alarms, and the density of the occurrence time is high. The occurrence of noise alarms is generally not related to faults, but the number of noise alarms is large and the interference is strong. Therefore, the embodiment can use periodic analysis of alarm data to eliminate the interference of these noise alarms.

[0065] In this embodiment, the alarm data in the alarm template is divided into periodic data and non-periodic data by performing Fourier series analysis and ACF autocorrelation function analysis on the alarm template data.

[0066] In addition, rare alarm data may carry important fault information, but rare alarms are easily submerged in a large number of noise alarms due to their small number. Therefore, the embodiment can also analyze the rarity of this type of rare alarm data according to the frequency of occurrence of the alarm data to give more attention to this type of alarm data.

[0067] In this embodiment, the alarm template data is aggregated according to the frequency of occurrence, and the alarm data is divided into high-frequency alarm data and low-frequency alarm data according to the number of alarm data of the same kind contained in the aggregation result.

[0068] In the embodiment, the alarm template data is classified into first classification data and second classification data according to periodicity analysis and rarity analysis, that is, the alarm data can be classified into four categories: periodic high-frequency alarm data, periodic low-frequency alarm data, non-periodic high-frequency alarm data and non-periodic low-frequency alarm data. According to experience, it is known that the non-periodic low-frequency alarm data is more likely to contain fault information, and the periodic high-frequency alarm data is usually regular alarm information. Therefore, a higher abnormal score needs to be given to the non-periodic low-frequency alarm data in the alarm data, and the abnormality score of the periodic high-frequency alarm data is reduced, so as to realize alarm noise reduction and at the same time retain important rare alarm data, and the retained alarm data after noise reduction is used as abnormal alarm data.

[0069] In the embodiment, the alarm data is structured to obtain alarm template data, and then the alarm data is classified into first classification data and second classification data according to the periodicity characteristics and rarity characteristics of each data in the template, and different abnormality scores are given according to the category of each data, so as to improve the attention of low-frequency alarm data and exclude the interference of noise alarm data on subsequent recall processing and data screening.

[0070] In some embodiments, the alarm template data includes alarm time, application name and alarm application name of the alarm data.

[0071] It can be understood that when a large amount of alarm data is hierarchically analyzed, the alarm time, application name and alarm application name of the alarm data can be used as the basis for dividing levels to form structured alarm template data.

[0072] The embodiment provides a specific composition method of alarm template data, so that the structured template composed of alarm data corresponding to the above three attributes has a unified type format, which provides convenience for the subsequent data screening process.

[0073] Optionally, minute-level aggregation features of the alarm template data are extracted; and the alarm template data is classified into periodic alarm data and non-periodic alarm data based on the minute-level aggregation features.

[0074] In Figure 5 In the embodiment shown in FIG. 8, one piece of alarm template data is “app_name=‘graph-backend’”. In this embodiment, the data is aggregated at a minute level to obtain aggregation features of different time periods, then a Fourier series analysis is used to generate a potential period length of the corresponding alarm application, and then a self-correlation function is calculated according to the period length to obtain a periodicity score of the corresponding period length. Then, all alarm data is divided into periodic alarm data and non-periodic alarm data according to a preset periodicity score threshold.

[0075] In Figure 5 In the embodiment shown, after obtaining the periodic alarm data, the seasonal decomposition method is used to extract information that is obviously different from noise alarms in the alarm data, and the abnormality degree of the alarm is scored according to the information.

[0076] The embodiment provides a specific method for dividing the alarm template data into periodic alarm data and non-periodic alarm data, which can exclude the interference of noise alarms on subsequent recall processing and data screening.

[0077] Optionally, the data that repeatedly appears in the alarm template data is aggregated to obtain high-frequency alarm data, and the data other than the high-frequency alarm data in the alarm template data is taken as low-frequency alarm data.

[0078] It can be understood that, since rare alarm data may carry important fault information, but rare alarms are easily submerged in a large number of noise alarms due to the small quantity, the embodiment can perform rarity analysis on the rare alarm data according to the frequency of the alarm data, so as to give higher attention to the alarm data.

[0079] In the embodiment, the alarm template data is aggregated according to the number of occurrences, and a plurality of aggregation results can be obtained, which all contain the same alarm template data; the threshold range for distinguishing high and low frequencies is set to 3-20 according to the number of data in the aggregation result, and the threshold is taken as 5 in the embodiment, that is, when there are at least 5 same alarm template data in the aggregation result, the alarm template data is divided into high-frequency alarm data, and similarly, when there are less than 5 same alarm template data in the aggregation result, the alarm template data is low-frequency alarm data.

[0080] The embodiment provides a specific method for dividing the alarm template data into high-frequency alarm data and low-frequency alarm data, which can improve the attention of the low-frequency alarm data, so as to improve the probability of screening fault information.

[0081] Optionally, the trend component and residual of the periodic alarm data in the first classification data are extracted, and the second abnormality score of the periodic alarm data is calculated according to the trend component and the residual; in the case that the alarm data is non-periodic alarm data and low-frequency alarm data, the second abnormality score is increased by a first proportion to obtain a first abnormality score; in the case that the alarm data is periodic alarm data and high-frequency alarm data, the second abnormality score is reduced by the first proportion to obtain the first abnormality score.

[0082] It can be understood that the noise alarm data generally appears periodically, and for the obtained periodic alarm data, a feature that is obviously different from the noise alarm data needs to be obtained to exclude the interference of the noise alarm, and the trend component and the residual in the periodic alarm calculated according to the time length (the length of a season) can be used to calculate the alarm abnormality degree corresponding to each alarm data, that is, the second abnormality score.

[0083] It should be noted that after the arithmetic abnormality score of each alarm data is calculated, the alarm data needs to be scored comprehensively in combination with the rarity characteristics of the alarm data to improve the importance of the low-frequency word alarm feature.

[0084] It should be noted that the value range of the first ratio and the second ratio is between 0.5 and 1.5.

[0085] In this embodiment, the difference score of the alarm data is the product of the arithmetic abnormality score and the first ratio, for example, the arithmetic abnormality score of a non-periodic alarm data is 8, and the alarm data also meets the low-frequency alarm data, and the first ratio is set to 1.2, so the difference score of the alarm data is 9.6; the arithmetic abnormality score of a periodic alarm data is 7, and the alarm data also meets the low-frequency alarm data, and the second ratio is set to 0.7, so the difference score of the alarm data is 4.9.

[0086] The embodiment provides a method for performing periodic analysis and rarity analysis on alarm data to determine an abnormality score, and through the abnormality score, noise alarm data contained in periodic alarm can be effectively excluded and the attention to low-frequency word alarm features can be improved.

[0087] Optionally, an alarm feature in the abnormal alarm data is extracted; the alarm feature is input into a first model for recall processing of fault alarm, and suspected fault alarm data is obtained from the abnormal alarm data according to a result of the recall processing.

[0088] It can be understood that the abnormal alarm data contains various data types, and an effective alarm feature needs to be constructed to enable the first model to recall alarm data containing suspected fault information.

[0089] In this embodiment, the alarm feature can be a time sequence feature of the alarm data, which is used to reflect real-time updated alarm information, or a state distribution feature, which is used to aggregate minute-level alarm data into minute-level alarm state distribution.

[0090] In this embodiment, the abnormal alarm data containing the above alarm features are respectively input into the LOF model, the IForest model, the KNN model and the ABOD model for recall processing, and after multi-layer abnormal detection voting and multi-dimensional weighted fault scoring, the required suspected fault alarm data are screened out.

[0091] The embodiment provides a method for extracting alarm data containing fault information by using a recall model, and can screen out alarm data corresponding to all possible fault information from abnormal alarm data.

[0092] Optionally, the abnormal alarm data are subjected to one-hot encoding processing to obtain an encoding processing result, and alarm timing features and alarm state distribution features are extracted from the encoding processing result; the alarm timing features and the alarm state distribution features are taken as alarm features; wherein the alarm timing features include minute-level granularity aggregation features, minute-level alarm application numbers and minute-level maximum application numbers; and the alarm state distribution features include alarm timing number distribution features and alarm timing frequency distribution features.

[0093] It should be noted that in the above alarm template data, a plurality of fields are template library fields, and therefore in the process of building the feature system, first, the structured alarm data are subjected to 0 / 1 one-hot encoding based on a plurality of non-continuous fields such as an application name and an alarm application name; the 0 / 1 one-hot encoding is to convert non-continuous segments into corresponding continuous data for use by a related machine learning algorithm, the alarm features extracted in the embodiment are taken as input features of an unsupervised detection model, the non-continuous fields in the template should be converted into continuous digital types by using the 0 / 1 one-hot encoding mode, and then the alarm features required for unsupervised learning are obtained from the encoding result; since the alarm batch update speed of the system is relatively fast, it is necessary to capture the state distribution features corresponding to the alarm information in real time as much as possible, and in order to improve the sensitivity of the system to the change of alarm timing, it is necessary to extract the timing distribution features in the abnormal alarm data, and the above state distribution features and timing distribution features are taken as input features of a subsequent detection model for recall processing of abnormal alarm data.

[0094] The embodiment provides a method for data preprocessing and extracting alarm features, and provides input features for recall processing of a subsequent detection model.

[0095] In some embodiments, the alarm timing features include encoded minute-level granularity aggregation features, minute-level alarm application numbers and minute-level maximum application numbers; and the alarm state distribution features include alarm timing number distribution features and alarm timing frequency distribution features.

[0096] It can be understood that, in order to capture the state distribution of the alarm information in real time, the embodiment encodes the minute-level granularity aggregation features of the alarm, the minute-level application number of the alarm and the minute-level maximum application number as the state distribution features of the abnormal alarm data. The above three state distribution features are feature quantities with a minute as a statistical interval, have good timeliness, and are applicable to alarm data. In order to improve the sensitivity of the system to the timing change of the alarm, the embodiment generates alarm number statistics under different lengths of sliding windows as alarm timing quantity distribution features, and generates the total number of seconds of the alarm in a minute as a timing frequency distribution feature. The above two features constitute the timing distribution features of the alarm. For example, for the timing quantity distribution features, the length of the sliding window can be: 1 min, 2 min, 5 min, 10 min, 20 min, 30 min, and then the number of alarms under different windows is counted to generate the timing quantity distribution features, and the timing frequency distribution features are the number of alarm seconds in a minute.

[0097] The embodiment provides an extraction method of alarm timing features and alarm state distribution features, so that the extracted alarm features can better reflect the timeliness of the alarm information and the sensitivity of the timing change of the alarm.

[0098] Optionally, the first model is used for recall processing on the alarm features, and an abnormal time point containing suspected fault information is extracted from the result of the recall processing. Fault scores are calculated for each abnormal time point to obtain fault scores corresponding to each abnormal time point. In a case where the fault score is higher than a second threshold, the alarm data corresponding to the abnormal time point is taken as the suspected fault alarm data.

[0099] It can be understood that the recall model is a kind of unsupervised detection model, and the recall model is used for recall processing on the abnormal alarm data. The abnormal time points corresponding to the suspected fault information of the fault are screened out according to the alarm features, and the voting mechanism of the recall model is used for fault scoring on the abnormal time points. The suspected fault alarm data corresponding to the abnormal time points with fault scores higher than a second threshold is recalled.

[0100] In Figure 6 In the embodiment shown in the figure, the abnormal alarm data is preprocessed (normalized, etc.), then the alarm features (alarm timing features and alarm state distribution features) are extracted, and the alarm data containing the alarm features is input into the recall model. After the abnormal detection voting of the recall model, the multidimensional weighted fault scores of each abnormal time point are calculated. Finally, in a case where the fault score is higher than a second threshold, the suspected fault alarm data corresponding to the abnormal time point is recalled, and the recalled suspected fault alarm data is sorted in descending order according to the above fault scores, for subsequent further screening of the suspected fault alarm data by the supervised model.

[0101] The embodiment provides a method for recall processing of abnormal alarm data by using alarm features, which can filter out corresponding suspected fault alarm data according to abnormal time points containing suspected fault information.

[0102] Optionally, the suspected fault alarm data is input into the second model for screening to obtain fault alarm data, and the fault alarm information is obtained from the fault alarm data; wherein the fault alarm information includes fault root cause information and fault prediction information.

[0103] It should be noted that the supervised model is a model obtained by taking historical fault alarm data as a training set and performing supervised training, and is used for fault classification of newly input alarm data.

[0104] In the embodiment, the suspected fault alarm data is input into the supervised model, and the fault types contained in the suspected fault alarm data are classified to obtain fault information corresponding to a fault occurrence time; in addition, based on each returned fault time point in the fault alarm data, time sequence correlation analysis is performed on the multi-application alarm data of the nearby time, the calling chain information is combined, and root cause positioning analysis is performed on the fault alarm data to determine the fault information of the fault alarm data, i.e. the root cause information of the fault alarm data.

[0105] It should be noted that the embodiment can also perform potential fault prediction analysis on the fault alarm data, i.e. combining the application calling relationship in the calling chain and the historical abnormal alarm data, the future possible alarm data triggered by the associated application and the corresponding fault are predicted in advance.

[0106] The embodiment provides a supervised detection method for obtaining a fault type corresponding to a fault occurrence time from suspected fault alarm information.

[0107] In some embodiments, the creation of the second model includes: obtaining a training set based on suspected fault alarm data; and performing supervised training on the XGBoost model by using the training set to obtain a fault alarm screening model.

[0108] It can be understood that the embodiment first combines artificial means to further label the fault types of historical fault data to generate a training set, and then uses the XGBoost model to perform supervised classification on suspected fault alarm data to obtain a fault type corresponding to each abnormal time point.

[0109] The embodiment provides a specific supervised classification model for performing second screening on suspected fault alarm data, and classifying the fault types contained in each suspected fault alarm data to determine the fault type of each suspected fault alarm data at an abnormal time.

[0110] In addition to the method described above, the disclosure also provides a fault detection system for implementing the method described above.

[0111] In combination Figure 7 A fault detection system provided by an embodiment of the disclosure is described below, which can be referred to in combination with the fault detection method described above.

[0112] The disclosure also provides a fault detection system, comprising:

[0113] The alarm data acquisition unit 710 is configured to acquire a large amount of alarm data; the abnormal alarm data acquisition unit 720 is configured to calculate the abnormality scores of the alarm data and filter alarm data with abnormality scores higher than a first threshold as abnormal alarm data; the suspected fault alarm data acquisition unit 730 is configured to perform first processing on the abnormal alarm data based on a pre-stored first model to acquire suspected fault alarm data; and the alarm fault information acquisition unit 740 is configured to perform second processing on the suspected fault alarm data based on a pre-stored second model to obtain alarm fault data and extract fault information from the alarm fault data.

[0114] The fault detection method and system provided by the disclosure can acquire a large amount of alarm data as detection data through the alarm data acquisition unit 710, quickly filter abnormal alarm data from the large amount of alarm data according to the abnormality scores of the alarm data through the abnormal alarm data acquisition unit 720, perform unsupervised learning on the abnormal alarm data through the suspected fault alarm data acquisition unit 730 using a first model to obtain suspected fault alarm data, ensure the recall rate of the suspected fault alarm data, perform supervised learning on the suspected fault alarm data through the alarm fault information acquisition unit 740 using a second model, further filter alarm fault data from the suspected fault alarm data, and extract the required fault information from the alarm fault data, so as to locate and predict the fault according to the fault information. The device described in this embodiment can efficiently extract accurate fault information from a large amount of alarm data and reduce the troubleshooting time of the fault.

[0115] Figure 8 An example of an entity structure diagram of an electronic device is shown in FIG. 1, which includes a processor 100, a memory 200, and a communication interface 300. Figure 8As shown, the electronic device can include a processor 810, a communications interface 820, a memory 830, and a communications bus 880, wherein the processor 810, the communications interface 820, and the memory 830 complete mutual communication through the communications bus 880. The processor 810 can invoke a logic instruction in the memory 830 to execute a fault detection method, which includes: obtaining a large amount of alarm data; calculating the abnormality score of each alarm data, and screening alarm data with an abnormality score higher than a first threshold value as abnormal alarm data; performing first processing on the abnormal alarm data based on a pre-stored first model to obtain suspected fault alarm data; performing second processing on the suspected fault alarm data based on a pre-stored second model to obtain alarm fault data, and extracting fault information from the alarm fault data.

[0116] In addition, the logic instruction in the memory 830 described above can be implemented in the form of a software functional unit and sold or used as an independent product, which can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present disclosure essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present disclosure. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0117] On the other hand, the present disclosure also provides a computer program product, which includes a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program is executed by a processor, so that the computer can execute a fault detection method provided by each method described above, which includes: obtaining a large amount of alarm data; calculating the abnormality score of each alarm data, and screening alarm data with an abnormality score higher than a first threshold value as abnormal alarm data; performing first processing on the abnormal alarm data based on a pre-stored first model to obtain suspected fault alarm data; performing second processing on the suspected fault alarm data based on a pre-stored second model to obtain alarm fault data, and extracting fault information from the alarm fault data.

[0118] In another aspect, this disclosure also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, is implemented to perform a fault detection method provided by the methods described above. The method includes: acquiring a massive amount of alarm data; calculating the anomaly score of each alarm data and filtering alarm data with an anomaly score higher than a first threshold as abnormal alarm data; performing a first processing on the abnormal alarm data based on a pre-stored first model to obtain suspected fault alarm data; performing a second processing on the suspected fault alarm data based on a pre-stored second model to obtain alarm fault data, and extracting fault information from the alarm fault data.

[0119] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0120] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments.

[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this disclosure, and are not intended to limit them. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure.

Claims

1. A fault detection method, characterized in that, include: Acquire massive amounts of alarm data; Calculate a first anomaly score for each alarm data, and filter alarm data with a first anomaly score higher than a first threshold as abnormal alarm data; wherein, alarm data includes server name, alarm message sending time, alarm topic, alarm severity level, and alarm content; abnormal alarm scoring includes scoring according to the default security level of alarm data, scoring according to the high and low frequency of occurrence of each alarm data, and scoring according to the periodic characteristics of each alarm data; The abnormal alarm data is processed based on a pre-stored first model to obtain suspected fault alarm data. The first model includes the LOF model, the IFOreest model, the nearest neighbor algorithm, and an angle-based abnormal time point detection algorithm. The first model detects abnormal time points by density, partitioning hyperplane, distance, and included angle, respectively. The suspected fault alarm data is processed based on the pre-stored second model to obtain alarm fault data, and fault information is extracted from the alarm fault data; the fault information includes fault root cause information and fault prediction information. The calculation of the first anomaly score for each of the alarm data includes: The alarm data is parsed hierarchically based on the pre-stored alarm templates to obtain alarm template data; The alarm template data is periodically analyzed to obtain first category data, wherein the first category data includes periodic alarm data and non-periodic alarm data; based on the first category data and the second category data, the first anomaly score is calculated. The alarm template data is periodically analyzed to obtain first-category data, including: Extract the minute-level aggregated features of the alarm template data; the minute-level aggregated features are the aggregated features of different time periods obtained by aggregating the alarm template data at minute intervals; Based on the minute-level aggregation feature, the alarm template data is divided into periodic alarm data and non-periodic alarm data; the division of the alarm template data into periodic alarm data and non-periodic alarm data based on the minute-level aggregation feature includes: Fourier series analysis is performed on the minute-level aggregated features to generate the potential cycle length of the corresponding alarm application. The autocorrelation function is used to calculate the potential cycle length to obtain the periodicity score of the corresponding cycle length. Then, according to the preset periodicity score threshold, all alarm data are divided into periodic alarm data and non-periodic alarm data according to the application category.

2. The fault detection method according to claim 1, characterized in that, The second category of data is obtained through the following steps: Rarity analysis is performed on the alarm template data to obtain second-category data; wherein, the second-category data includes high-frequency alarm data and low-frequency alarm data.

3. The fault detection method according to claim 2, characterized in that, Rarity analysis is performed on the alarm template data to obtain second-category data, including: The data that appears repeatedly in the alarm template data are aggregated to obtain high-frequency alarm data, and the other data in the alarm template data other than the high-frequency alarm data are taken as low-frequency alarm data.

4. The fault detection method according to claim 1, characterized in that, The second category of data includes high-frequency alarm data and low-frequency alarm data; The step of calculating the first anomaly score based on the first classification data and the second classification data includes: Extract the trend components and residuals of the periodic alarm data in the first classification data, and calculate the second anomaly score of the periodic alarm data based on the trend components and residuals; When the alarm data consists of the non-periodic alarm data and the low-frequency alarm data, the second anomaly score is increased by a first ratio to obtain the first anomaly score. When the alarm data consists of the periodic alarm data and the high-frequency alarm data, the second anomaly score is reduced by a second ratio to obtain the first anomaly score.

5. The fault detection method according to claim 1, characterized in that, The first processing of the abnormal alarm data based on the pre-stored first model to obtain suspected fault alarm data includes: Extract alarm features from the abnormal alarm data; The alarm features are input into the first model for fault alarm recall processing, and the suspected fault alarm data is obtained from the abnormal alarm data based on the result of the recall processing.

6. The fault detection method according to claim 5, characterized in that, The extraction of alarm features from the abnormal alarm data includes: The abnormal alarm data is subjected to one-hot encoding to obtain the encoding result, and alarm timing features and alarm status distribution features are extracted from the encoding result. The alarm timing characteristics and the alarm status distribution characteristics are used as the alarm characteristics; The alarm timing characteristics include minute-level granularity aggregation characteristics, minute-level alarm application number, and minute-level maximum application number; the alarm status distribution characteristics include alarm timing quantity distribution characteristics and alarm timing frequency distribution characteristics.

7. The fault detection method according to claim 5 or 6, characterized in that, The step of inputting the alarm features into the first model for fault alarm recall processing, and obtaining the suspected fault alarm data from the abnormal alarm data based on the result of the recall processing, includes: The first model is used to perform the recall process on the alarm features, and abnormal time points containing suspected fault information are extracted from the result of the recall process. A fault score is calculated for each of the aforementioned abnormal time points to obtain the fault score corresponding to each of the aforementioned abnormal time points. If the fault score is higher than the second threshold, the alarm data corresponding to the abnormal time point will be used as the suspected fault alarm data.

8. The fault detection method according to any one of claims 1-6, characterized in that, The suspected fault alarm data is processed a second time based on a pre-stored second model to obtain fault alarm information, including: The suspected fault alarm data is input into the second model for filtering to obtain fault alarm data, and the fault alarm information is obtained from the fault alarm data; The fault alarm information includes fault root cause information and fault prediction information.

9. A fault detection system, employing the fault detection method as described in claim 1, characterized in that, The system includes: Alarm data acquisition unit, used to acquire massive amounts of alarm data; An abnormal alarm data acquisition unit is used to calculate the abnormality score of each alarm data and filter the alarm data with the abnormality score higher than a first threshold as abnormal alarm data. The suspected fault alarm data acquisition unit is used to perform a first processing on the abnormal alarm data based on a pre-stored first model to obtain suspected fault alarm data. The first model includes the LOF model, the IFOreest model, the nearest neighbor algorithm, and the angle-based abnormal time point detection algorithm. The first model detects abnormal time points by density, partitioning hyperplane, distance, and included angle, respectively. The alarm and fault information acquisition unit is used to perform a second processing on the suspected fault alarm data based on a pre-stored second model to obtain alarm and fault data, and extract fault information from the alarm and fault data; the fault information includes fault root cause information and fault prediction information.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the fault detection method as described in any one of claims 1-8.

11. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the fault detection method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Method for assessing whether or not a measured value of a physical parameter of an aircraft engine is normal

    CN106233115A

  • False alarm information identification method and device, storage medium and electronic terminal

    CN109379228A

  • Dynamic alarm grading method and device, electronic equipment and storage medium

    CN111338915A

  • Abnormal data detection method and device and computer readable storage medium

    CN112464051A

  • Training method, fault prediction method, related device and equipment

    CN113010389A