Fault Determination Method and Device, Electronic Device, and Computer-Readable Storage Medium

By obtaining the target reference data and the index data to be detected, and using similarity comparison and clustering algorithms, the problem that the fault detection method in the prior art cannot establish the association between the fault type and the index type, achieving more accurate and widely applicable fault prediction.

CN114298221BActive Publication Date: 2025-07-22CHINA CONSTRUCTION BANK
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111625223.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-28
Publication Date
2025-07-22
Estimated Expiration
2041-12-28

AI Technical Summary

Technical Problem

In the prior art, fault discovery methods usually only compare within the same indicators, and fail to establish a potential relationship between the fault type and the indicator type, resulting in deviations in the discovery of complex fault situations and cannot effectively compare between different software and hardware devices.

Method used

By obtaining the target reference data and the indicator data to be detected, using similarity comparison and clustering algorithms, establish the correlation between the fault type and the indicator type, and perform the fusion of multiple indicators to predict whether the target fault category will occur in the equipment to be detected.

Benefits of technology

The fusion of multiple correlation indicators is achieved, which improves the accuracy and applicability of fault prediction, can predict the occurrence of complex faults in a wider range of scenarios, and reduces the deviation of fault detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114298221B_ABST
    Figure CN114298221B_ABST
Patent Text Reader

Abstract

The present disclosure provides a fault determination method, an apparatus, an electronic device, and a computer-readable storage medium, which can be applied to the technical field of data center operation and maintenance, and can also be used in the financial technology field. The fault determination method includes: obtaining target reference data, where the target reference data includes: index data of at least one associated index of a faulty device within a preset historical fault time period, and at least one associated index is associated with a target fault category; obtaining a plurality of to-be-detected index data, where each to-be-detected index data includes: index data of one of the to-be-detected indexes of the to-be-detected device within the to-be-detected time period; and determining a detection result for characterizing whether the to-be-detected device will have a fault of the target fault category within a preset future time period according to the plurality of to-be-detected index data and the target reference data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of data center operation and maintenance, and more particularly, to a method and device for fault determination, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] Currently, with the wide application of new technologies such as virtualization and cloud computing, the scale of IT infrastructure within enterprise data centers has increased exponentially, the scale of computer hardware and software has been continuously expanding, and corresponding computer failures have occurred frequently, increasing the difficulty of handling for operation and maintenance personnel.

[0003] In the process of implementing the concept of the present disclosure, the inventors found that there are at least the following problems in the related art. Currently, the fault discovery method usually only makes comparisons within the same metrics, without establishing a potential association between the fault type and the metric type, and cannot perform mixed calculations and comparisons on multiple metrics, resulting in deviations in the discovery of more complex fault situations. Summary of the Invention

[0004] In view of the above problems, the present disclosure provides a method, device, equipment, medium, and program product for fault determination.

[0005] According to a first aspect of the present disclosure, there is provided a method for fault determination, including:

[0006] Obtaining target reference data, where the target reference data includes: metric data of at least one associated metric of a faulty device within a preset historical fault time period, where at least one associated metric is associated with a target fault category;

[0007] Obtaining a plurality of to-be-detected metric data, where each to-be-detected metric data includes: metric data of one of the to-be-detected metrics of the to-be-detected device within the to-be-detected time period;

[0008] Determining a detection result for characterizing whether a fault of the target fault category will occur in the to-be-detected device within a future preset time period according to the plurality of to-be-detected metric data and the target reference data.

[0009] According to an embodiment of the present disclosure, determining a detection result for characterizing whether a fault of the target fault category will occur in the to-be-detected device within a future preset time period according to the plurality of to-be-detected metric data and the target reference data includes:

[0010] Comparing the similarity of each to-be-detected metric data with the target reference data respectively to obtain a first similarity comparison result;

[0011] Determining a detection result for characterizing whether a fault of the target fault category will occur in the to-be-detected device within a future preset time period according to the first similarity comparison result.

[0012] According to an embodiment of the present disclosure, wherein, comparing the similarity of each index data to be detected with the target reference data respectively to obtain the similarity comparison result includes:

[0013] Calculating the cluster center of the index data of at least one associated index to obtain the target cluster center;

[0014] Using a preset similarity algorithm, comparing the similarity of each index data to be detected with the target cluster center respectively to obtain the similarity comparison result.

[0015] According to an embodiment of the present disclosure, it further includes: determining at least one associated index associated with the target fault category;

[0016] Wherein determining at least one associated index associated with the target fault category includes:

[0017] Obtaining the index data of multiple initial indexes of the faulty device during a preset historical fault time period;

[0018] Performing clustering processing on the index data of multiple initial indexes of the faulty device to obtain multiple faulty index data sets;

[0019] Determining the faulty index data set that meets the preset discrimination condition as the target faulty index data set;

[0020] Extracting at least one associated index associated with the target fault category from the target faulty index data set.

[0021] According to an embodiment of the present disclosure, wherein performing clustering processing on the index data of multiple initial indexes of the faulty device to obtain multiple faulty index data sets includes:

[0022] In the case where, according to the current cluster center of at least one current data set, the current data in the index data of multiple initial indexes is completed for the current clustering to form at least one next data set, iteratively execute: updating the cluster center of the next data set, so that according to the updated cluster center of the next data set, the next data in the index data of multiple initial indexes is completed for the next clustering until a preset termination condition is reached, and finally determining multiple faulty index data sets.

[0023] According to an embodiment of the present disclosure, wherein, according to the current cluster center of at least one current data set, the current data in the index data of multiple initial indexes is completed for the current clustering to form at least one next data set includes:

[0024] Using a preset similarity algorithm, comparing the similarity of the current data with at least one current cluster center respectively to obtain at least one second similarity comparison result;

[0025] Add the current data to the target current data set, so as to form at least one next data set through the target current data set and the unupdated current data set, where the comparison result of the second similarity between the current cluster center of the target current data set and the current data is greater than the preset similarity threshold.

[0026] According to an embodiment of the present disclosure, wherein the preset similarity threshold is 0.9.

[0027] According to an embodiment of the present disclosure, wherein obtaining the target reference data includes:

[0028] Collect the original target data, where the original target data includes: the original index data of at least one associated index of the faulty device within a preset historical fault time period;

[0029] Normalize the original index data of at least one associated index respectively to obtain the target reference data.

[0030] According to an embodiment of the present disclosure, wherein the target reference data further includes: the index data of the preselected indexes under the predefined fault type;

[0031] Wherein, obtaining the target reference data includes:

[0032] Obtain a fault curve associated with the predefined fault type, where the fault curve is used to characterize the trend of the index value of the preselected index changing with time;

[0033] Extract a preset number of index values from the fault curve at preset time intervals and determine them as the target reference data.

[0034] A second aspect of the present disclosure provides a fault determination device, including a first acquisition module, a second acquisition module and a first determination module.

[0035] Wherein, the first acquisition module is used to acquire the target reference data, where the target reference data includes: the index data of at least one associated index of the faulty device within a preset historical fault time period, and at least one associated index is associated with the target fault category;

[0036] The second acquisition module is used to acquire a plurality of index data to be detected, where each index data to be detected includes: the index data of one of the indexes to be detected of the device to be detected within the time period to be detected;

[0037] The first determination module is used to determine a detection result for characterizing whether the device to be detected will have a fault of the target fault category within a preset time period in the future according to the plurality of index data to be detected and the target reference data.

[0038] According to an embodiment of the present disclosure, wherein the first determination module includes a comparison unit and a first determination unit.

[0039] Among them, a comparison unit is configured to separately compare the index data of each index to be detected with target reference data to obtain a first similarity comparison result;

[0040] A first determination unit is configured to determine a detection result for characterizing whether a target fault category fault will occur in a preset future time period for the device to be detected according to the first similarity comparison result.

[0041] According to an embodiment of the present disclosure, among them, the comparison unit includes a calculation subunit and a comparison subunit.

[0042] Among them, the calculation subunit is configured to calculate the cluster center of the index data of at least one associated index to obtain a target cluster center;

[0043] The comparison subunit is configured to use a preset similarity algorithm to separately compare the index data of each index to be detected with the target cluster center to obtain a similarity comparison result.

[0044] According to an embodiment of the present disclosure, it further includes: a second determination module configured to determine at least one associated index associated with the target fault category.

[0045] Among them, the second determination module includes a first acquisition unit, a clustering unit, a second determination unit, and a first extraction unit.

[0046] Among them, the first acquisition unit is configured to acquire the index data of multiple initial indexes of a faulty device within a preset historical fault time period;

[0047] The clustering unit is configured to perform clustering processing on the index data of multiple initial indexes of the faulty device to obtain multiple faulty index data sets;

[0048] The second determination unit determines the faulty index data set that meets the preset discrimination condition as the target faulty index data set;

[0049] The first extraction unit is configured to extract at least one associated index associated with the target fault category from the target faulty index data set.

[0050] According to an embodiment of the present disclosure, among them, the clustering unit includes a clustering subunit, which is configured to, in the case of completing the current clustering of the current data in the index data of multiple initial indexes according to the current cluster center of at least one current data set to form at least one next data set, iteratively execute: updating the cluster center of the next data set, so as to complete the next clustering of the next data in the index data of multiple initial indexes according to the updated cluster center of the next data set until a preset termination condition is reached, and obtaining multiple finally determined faulty index data sets.

[0051] According to an embodiment of the present disclosure, in the clustering subunit, according to the current cluster centers of at least one current data set, the current data in the index data of multiple initial metrics is used to complete the current clustering to form at least one next data set, including:

[0052] Using a preset similarity algorithm, compare the current data with at least one current cluster center respectively to obtain at least one second similarity comparison result;

[0053] Add the current data to the target current data set, so as to form at least one next data set through the target current data set and the unupdated current data set, where the second similarity comparison result between the current cluster center of the target current data set and the current data is greater than the preset similarity threshold.

[0054] According to an embodiment of the present disclosure, the preset similarity threshold is 0.9.

[0055] According to an embodiment of the present disclosure, the first acquisition module includes an acquisition unit and a processing unit.

[0056] Among them, the acquisition unit is used to acquire the original target data, where the original target data includes: the original index data of at least one associated metric of the faulty device within a preset historical fault time period;

[0057] The processing unit is used to perform normalization processing on the original index data of at least one associated metric respectively to obtain the target reference data.

[0058] According to an embodiment of the present disclosure, the target reference data further includes: the index data of the preselected metrics under the predefined fault type.

[0059] Among them, the first acquisition module includes a second acquisition unit and a second extraction unit.

[0060] Among them, the second acquisition unit is used to acquire the fault curve associated with the predefined fault type, where the fault curve is used to characterize the trend of the index value of the preselected metric changing with time;

[0061] The second extraction unit is used to extract a preset number of index values from the fault curve at a preset time interval and determine them as the target reference data.

[0062] A third aspect of the present disclosure provides an electronic device, including: one or more processors; a memory for storing one or more programs, where when the one or more programs are executed by the one or more processors, the one or more processors execute the above-mentioned fault determination method.

[0063] A fourth aspect of the present disclosure further provides a computer-readable storage medium having executable instructions stored thereon, and when the instructions are executed by a processor, the processor is caused to execute the above-mentioned fault determination method.

[0064] A fifth aspect of the present disclosure further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the above-mentioned fault determination method is implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, the above-mentioned content and other objects, features, and advantages of the present disclosure will become clearer. In the drawings:

[0066] Figure 1 Schematically shows an application scenario diagram of a fault determination method, apparatus, device, medium, and program product according to an embodiment of the present disclosure;

[0067] Figure 2 Schematically shows a flowchart of a fault determination method in the related art;

[0068] Figure 3 Schematically shows a flowchart of a fault determination method according to an embodiment of the present disclosure;

[0069] Figure 4 Schematically shows a schematic diagram of data cleaning for collected data according to an embodiment of the present disclosure;

[0070] Figure 5 Schematically shows a flowchart of clustering processing for index data of multiple initial indexes of a faulty device according to an embodiment of the present disclosure;

[0071] Figure 6 Schematically shows a structural block diagram of a fault determination apparatus according to an embodiment of the present disclosure; and

[0072] Figure 7 Schematically shows a block diagram of an electronic device suitable for implementing a fault determination method according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0073] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present disclosure.

[0074] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0075] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0076] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning commonly understood by those of ordinary skill in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).

[0077] It should be noted that in the technical solutions of the present disclosure, before obtaining or collecting the user's personal information, the user's authorization or consent has been obtained.

[0078] Embodiments of the present disclosure provide a fault determination method, including:

[0079] Obtain target reference data, where the target reference data includes: index data of at least one associated index of a faulty device within a preset historical fault time period, where at least one associated index is associated with a target fault category;

[0080] Obtain multiple to-be-detected index data, where each to-be-detected index data includes: index data of one of the to-be-detected indexes of the to-be-detected device within the to-be-detected time period;

[0081] According to the multiple to-be-detected index data and the target reference data, determine a detection result for characterizing whether a fault of the target fault category will occur in the to-be-detected device within a future preset time period.

[0082] Figure 1 Schematically shows an application scenario diagram of a fault determination method, device, equipment, medium, and program product according to an embodiment of the present disclosure.

[0083] As Figure 1As shown, the system architecture 100 according to this embodiment may include a system to be detected 101, a data processing system 102, and a fault handling system 103. The system to be detected 101, the data processing system 102, and the fault handling system 103 can communicate with each other through a network, and the network can include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0084] In the application scenario of the embodiments of the present disclosure, the system to be detected 101 may be various software and hardware devices to be detected in a data center, which are used to provide various types of business services in response to user requests. The system to be detected 101 may include one or more business servers, or may be one or more business service clusters.

[0085] The servers in the system to be detected 101 may be servers that provide various services. For example, a background management server (only for example) that supports the websites browsed by users using terminal devices. The background management server can analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device. Users can use the terminal device to interact with the server through the network to receive or send messages, etc.

[0086] The data processing system 102 is used to collect various index data (such as CPU usage rate, memory usage rate, disk I / O speed, etc.) of each device in the system to be detected 101, and combine with a preset algorithm to perform data processing on the various index data of each device. For example, it can be based on the time nodes when faults occurred in history and the index data generated before, combined with the index values sampled in real time, compare the two, and make a preliminary discovery of faults according to the comparison results; at the same time, the index values sampled in real time can be compared with a preset alarm threshold, and when the index value is greater than the preset threshold, an alarm notification is sent to the fault handling system 103.

[0087] The fault handling system 103 is used to receive the alarm notification sent by the data processing system 102, and perform fault emergency handling on the faulty devices, such as forwarding user requests from the faulty server to the normally operating business server, etc.

[0088] The following will be based on Figure 1 the described scenario, and through Figures 2 to 7 describe in detail the fault determination method of the disclosed embodiments.

[0089] As the scale of the IT infrastructure within an enterprise data center has grown exponentially, the scale of computer hardware and software has been continuously expanding, and correspondingly, computer failures have occurred frequently, increasing the difficulty of handling for operation and maintenance personnel. However, in the work of data center operation and maintenance, how to detect the signs and trends of failures before they occur and then take corresponding measures before the failures occur is an urgent problem to be solved. Among them, how to pre-detect failures through metric data is the main focus in related technologies.

[0090] In related technologies, in the daily operation and maintenance work of a data center, generally, a mechanism for detecting software and hardware failures in the data center is constructed through a basic monitoring system and an application monitoring system. In this process, various software and hardware will generate a large amount of periodic metric data. By setting metric thresholds, a direct judgment of the occurrence of a failure can be made. In the application scenario of failure pre-detection, how to quickly and accurately judge that a failure is about to occur. In related technologies, for example, based on the failure occurrence nodes in history and the metric values generated before, combined with the current metric values, a failure prediction can be made after comparison.

[0091] Figure 2 Schematically shows a flowchart of a failure determination method in related technologies. As Figure 2 shown, the method for predicting and determining data center failures in related technologies can adopt the way of mean regression.

[0092] As Figure 2 shown, the main idea of using the mean regression method for failure determination is: Assume that the business access volume received by the business system is constant in the same time period. For example, at 10 am on each working day, the access volume received by the business is similar. Therefore, it is speculated that in similar time periods, the system pressure (CPU usage rate, memory usage rate, network usage rate, etc.) of the business is also similar. Therefore, the system pressure values at the same time point in the previous several days can be statistically counted, and mean calculation is performed, and then weighted by a certain coefficient. The calculated result data is used as a parameter value for reference comparison with the system pressure value at the current time point. If the system pressure value at the current time point is much higher or much lower than the previous mean value (it may not have reached the alarm threshold yet), but it can still be inferred that there is an abnormal situation in the system pressure (or business anomaly). Based on this situation, a failure warning can be given. In addition, because the comparison is made for the same metric of the same software and hardware device, generally, data cleaning operations such as data normalization are not required.

[0093] In the process of implementing the concept of the present disclosure, the inventors found that there are at least the following problems in related technologies:

[0094] The current fault detection methods usually only perform comparisons within the same metrics, and the fault warning thresholds calculated based on the average values of metrics in the same historical time period are derived from normal data, generally unable to reflect relatively special extreme situations and the situations before a fault is about to occur; the calculated metrics can only rely on the data of the same software and hardware devices in the same service environment, so the applicable range is relatively narrow and cannot be compared between the same software and hardware devices; the metric data generally refers to the metric data of a fixed past time for comparison, and the range of samples obtained is very small, and the algorithm cannot cover more samples; in short, the fault detection methods in the related technologies do not establish a potential association between the fault types and the metric types, cannot perform mixed calculations and comparisons on multiple metrics, and there are deviations in the detection of relatively complex fault situations.

[0095] In view of this, embodiments of the present disclosure provide a fault determination method.

[0096] Figure 3 A flowchart of the fault determination method according to an embodiment of the present disclosure is schematically shown.

[0097] As Figure 3 shown, the fault determination method of this embodiment includes operation S310 to operation S330.

[0098] In operation S310, target reference data is obtained, where the target reference data includes: within a preset historical fault time period, the metric data of at least one associated metric of a faulty device, and at least one associated metric is associated with a target fault category.

[0099] In operation S320, multiple to-be-detected metric data are obtained, where each to-be-detected metric data includes: within a to-be-detected time period, the metric data of one of the to-be-detected metrics of the to-be-detected device.

[0100] In operation S330, according to the multiple to-be-detected metric data and the target reference data, a detection result is determined for characterizing whether a fault of the target fault category will occur in the to-be-detected device within a preset future time period.

[0101] According to an embodiment of the present disclosure, the target reference data is the historical metric data of faulty devices in a business system, including the data of multiple associated metrics associated with a certain target fault category, which can characterize the change rules of the relevant metric data under this fault category. The faults of faulty devices in the business system can be divided into multiple different fault categories, such as network faults, storage faults, middleware faults, application program faults, etc. The target fault category can be any one of the above fault categories and is the fault category currently being analyzed.

[0102] According to an embodiment of the present disclosure, the associated metrics can be any software or hardware metrics of the faulty device. For example, they can be CPU usage rate, memory usage rate, disk I / O speed, and so on.

[0103] According to an embodiment of the present disclosure, for each fault category, one or more metrics are respectively associated. For example, the metrics associated with disk - related faults (such as disk jitter) can include CPU usage rate and disk I / O speed; for example, the metrics associated with memory - related faults (such as overly frequent memory release) can include CPU usage rate, memory usage rate, and so on.

[0104] According to an embodiment of the present disclosure, obtaining the target reference data can be, when the types of the associated metrics associated with the target fault category are known, extracting the historical data of the corresponding associated metrics from the historical metric data of the faulty device.

[0105] According to an embodiment of the present disclosure, obtaining the target reference data can be, when the types of the associated metrics associated with the target fault category are unknown, first analyzing and processing all relevant metric data of the faulty device to determine the types of the associated metrics associated with the target fault category, and then extracting the historical data of the corresponding associated metrics from the historical metric data of the faulty device.

[0106] According to an embodiment of the present disclosure, obtaining the target reference data can also be, according to the user - defined fault category, first obtaining the fault curve associated with the user - defined fault type, and then extracting a preset number of metric values from the fault curve as the values of the associated metrics associated with the user - defined fault category.

[0107] According to an embodiment of the present disclosure, the preset historical fault time period is the time period before the faulty device fails, rather than the metric data at the time of or after the fault occurs. Because in the several time periods before the fault occurs, the relevant metrics are likely to have shown some abnormal conditions. At this time, the metric values may not have exceeded the warning threshold and may not have been discovered and paid attention to. However, the changes in the metric data at this stage have shown a certain degree of abnormality, indicating a certain degree of fault warning information. Therefore, the data at this stage plays a relatively important reference role in fault pre - detection.

[0108] According to an embodiment of the present disclosure, multiple metrics to be detected can include, for example, CPU usage rate data, memory usage rate data, disk I / O speed data, etc., of the device to be detected.

[0109] According to an embodiment of the present disclosure, based on multiple data of indexes to be detected and target reference data, a detection result for characterizing whether a target failure category of failure will occur in a future predicted time period for the device to be detected can be obtained by comparing each data of the indexes to be detected with the target reference data respectively to obtain a comparison result, and then determining whether the device will have a failure of the target failure category according to the comparison result. According to an embodiment of the present disclosure, comparing each data of the indexes to be detected with the target reference data respectively can be, for example, performing similarity comparison, difference comparison, ratio comparison, etc. Determining whether the device will have a failure of the target failure category according to the comparison result can be, for example, based on the similarity value obtained from the similarity comparison, determining that the device may have a failure of the target failure category when the similarity value is greater than a preset threshold; or based on the difference obtained from the difference comparison, determining that the device may have a failure of the target failure category when the difference is less than the preset threshold, etc.

[0110] According to an embodiment of the present disclosure, the target reference data includes data of multiple associated indexes associated with the target failure category, which characterizes the potential association relationship between the failure type and the index type. By determining whether the device to be detected will have a failure of the target failure category based on multiple data of the indexes to be detected and the target reference data, the fusion processing of multiple associated indexes is realized, rather than just comparing within the same indexes. The failure prediction result is more accurate. In addition, the data of the indexes to be detected includes multiple index data of the device to be detected, taking into account the influence of multiple index types that may be associated with the occurrence of the failure. The calculated indexes do not depend on the data of the same software and hardware device in the same service environment, and have a wider application range, and can be used in more complex failure scenarios, and the failure prediction result is more referenceable.

[0111] According to an embodiment of the present disclosure, obtaining the target reference data can be, when the types of the associated indexes associated with the target failure category are known, extracting the historical index data of the corresponding associated indexes from the historical index data of the failed device. Specifically, obtaining the target reference data includes:

[0112] Collecting original target data, where the original target data includes: the original index data of at least one associated index of the failed device within a preset historical failure time period;

[0113] Respectively performing data cleaning on the original index data of at least one associated index, that is, performing data screening, normalization processing, etc., to obtain the target reference data.

[0114] Due to the large number of devices in the data center and the relatively complex business functions, the data volume of the collected original target data is relatively large. It is necessary to use the operation and maintenance big data platform as the technical support to clean the data. At the same time, it is necessary to locate and classify the failures to facilitate sample classification.

[0115] Figure 4 Schematically shown is a schematic diagram of data cleaning for the collected data according to an embodiment of the present disclosure. The following will be combined with Figure 4 to describe the method of the embodiment of the present disclosure.

[0116] The operation and maintenance big data platform can use preset big data components (such as hdfs, kudu, hive, elasticsearch, etc.) to store a large amount of operation and maintenance related data (such as alarm data, log data, and metric data, etc.). Due to the large amount of data, cold and hot data separation can be set, with recent data as hot data and data from a relatively long time ago as cold data, and indexes are added.

[0117] As Figure 4 shown, according to an embodiment of the present disclosure, before data extraction, fault location needs to be performed, that is, through the alarm information on the operation and maintenance big data platform, determine the software and hardware and time when the fault occurs.

[0118] According to an embodiment of the present disclosure, after fault location, fault classification also needs to be performed. Determine the identified faults through the alarm information of the operation and maintenance big data platform, and classify them into different fault categories according to the content of the alarm, such as network faults, storage faults, middleware faults, application program faults, etc. For different types of faults, the change patterns of their metric values may be different, so different types of faults can be processed separately.

[0119] According to an embodiment of the present disclosure, fault metric extraction refers to extracting the historical metric data of corresponding associated metrics from the historical metric data of the faulty device, that is, collecting the original target data (the original metric data of at least one associated metric of the faulty device).

[0120] According to an embodiment of the present disclosure, during the process of sampling and extracting the historical metric data of the faulty software and hardware, because the sampling intervals of different metrics are different (commonly sampled once every 1 minute or 5 minutes), the required time advance amounts are also different (for some faults, the situation 1 hour before the fault needs to be considered, and for some faults, the situation 12 hours before the fault needs to be considered), and the required sampling points are also different (some need to sample 30 metric values, and some need to sample 60 metric values). Therefore, for the process of metric data extraction, the efficiency of obtaining metrics can be improved by setting a configuration file, so that in the case where the designed data extraction range changes, the data extraction method can be quickly changed.

[0121] According to an embodiment of the present disclosure, in the process of sampling and extracting historical index data of faulty software and hardware, for each associated index, the values of the n index sampling points before the occurrence of the fault can be set and stored in an n-dimensional matrix in ascending order of time. The following matrix shows the sampling data of one of the indexes X:

[0122] X = [x1, x2, x3,..., x n

[0123] For example, assuming that the index sampling frequency is once per minute and n = 30, in matrix X, x1 represents the index value 30 minutes ago, x2 represents the index value 29 minutes ago, and so on. For example, the CPU usage index of a certain Linux machine is sampled once per minute, and the values sampled in the last 30 minutes are X = [46.0, 51.4, 49.3,..., 48.3].

[0124] According to an embodiment of the present disclosure, after extracting the original index data of one or more associated indexes, it is also necessary to perform data cleaning on the original index data of these associated indexes respectively, that is, perform data screening, normalization processing, etc., to obtain target reference data. Normalization is a dimensionless processing method that makes the absolute value of the physical system numerical value become a certain relative value relationship. When the data needs to be put into a calculation space that intersects with other data, it is necessary to perform normalization processing on the data, which can make the calculated numerical values be compared on the same dimension, and at the same time, the sample data will not be too sparse and scattered during the calculation.

[0125] According to an embodiment of the present disclosure, the method for normalizing the original index data of the above-mentioned associated indexes is shown in the following formula (1):

[0126]

[0127] Wherein, X min represents the minimum value in matrix X, and X max represents the maximum value in matrix X. Among them, X, X′, X min , X max are all n-dimensional matrices.

[0128] According to an embodiment of the present disclosure, when it is necessary to perform a comprehensive calculation on multiple associated index columns associated with the target fault category and output the values of multiple associated indexes as a whole, the values of multiple associated indexes that have been normalized can be further normalized to obtain target reference data.

[0129] ​According to an embodiment of the present disclosure, the above operation may be, for example, assuming that there is original index data of three index items X, Y, and Z. After single-index normalization, X′, Y′, and Z′ are obtained. Using the average weighting algorithm (or a weighting method with specified weighting coefficients), the normalized values of the three indexes are obtained as shown in the following formula:

[0130] K = (X′ + Y′ + Z′) / 3 Formula (2)

[0131] According to an embodiment of the present disclosure, obtaining target reference data may be, when the type of the associated index associated with the target fault category is unknown, first analyzing and processing all relevant index data of the faulty device to determine the type of the associated index associated with the target fault category, and then extracting the historical data of the corresponding associated index from the historical index data of the faulty device to obtain the target reference data. Among them, the operation of extracting index data may refer to the methods of sampling original index data and normalizing the original index data in the above embodiments, which will not be elaborated here. The following focuses on the method of analyzing and processing all relevant index data of the faulty device to determine the type of the associated index associated with the target fault category when the type of the associated index associated with the target fault category is unknown.

[0132] According to an embodiment of the present disclosure, determining at least one associated index associated with the target fault category includes:

[0133] First, obtain the index data of multiple initial indexes of the faulty device within a preset historical fault time period; for example, it may be to extract all types of index data within the preset historical time period from the historical index data of the faulty device.

[0134] Perform clustering processing on the index data of multiple initial indexes of the faulty device to obtain multiple faulty index data sets; for example, various clustering algorithms may be used to perform clustering processing on all types of index data to obtain multiple faulty index data sets.

[0135] Determine the faulty index data set that meets the preset discrimination condition as the target faulty index data set; for example, it may be to use the faulty index data set with the largest number of index types after clustering as the target faulty index data set. Because the target faulty index data set aggregates the largest number of index types, it characterizes the aggregation of these indexes in a certain characteristic, and the aggregation degree is relatively tight. It can be considered that these indexes are relatively closely related to the target fault category.

[0136] Finally, extract at least one associated index associated with the target fault category from the target faulty index data set.

[0137] According to an embodiment of the present disclosure, for example, the above method may be that a fault of disk jitter occurs in the faulty device A, and the target fault category is a disk - type fault. All types of metric data within the last hour before the fault occurs can be extracted from the historical metric data of the faulty device A, including CPU usage rate, memory usage rate, disk I / O speed, and so on. Then, clustering processing is performed on these metric data to obtain multiple fault metric data sets. For example, the fault metric data set 1 includes metrics: CPU usage rate, memory usage rate; the fault metric data set 2 includes the metric: CPU usage rate; the fault metric data set 3 includes metrics: CPU usage rate, memory usage rate, disk I / O speed. Then, the fault metric data set 3 can be determined as the target fault metric data set, and at least one associated metric associated with the target fault category can be extracted from the target fault metric data set as: CPU usage rate, memory usage rate, disk I / O speed.

[0138] According to an embodiment of the present disclosure, when the type of the associated metric associated with the target fault category is unknown, the type of the associated metric associated with the target fault category can be determined by the above - mentioned method of the embodiment of the present disclosure. Compared with the method of determining the type of the associated metric associated with the target fault category according to operation and maintenance experience, it overcomes the defect that the empirical method has inaccurate fault location due to strong subjectivity, and is convenient for discovering potential associated fault metrics under the target fault type. In addition, by performing data processing on a large amount of fault metrics through density clustering, rapid automatic classification of fault metrics can be achieved, which is convenient for quickly establishing the association between fault types and metric types.

[0139] According to an embodiment of the present disclosure, in the process of determining the type of the associated metric associated with the target fault category, the method of performing clustering processing on the metric data of multiple initial metrics of the faulty device to obtain multiple fault metric data sets specifically includes:

[0140] When, according to the current cluster centers of at least one current data set, the current data in the metric data of multiple initial metrics is completed for the current clustering to form at least one next data set, iteratively execute: update the cluster centers of the next data set, so that according to the updated cluster centers of the next data set, the next data in the metric data of multiple initial metrics is completed for the next clustering until a preset termination condition is reached, and finally determine multiple fault metric data sets.

[0141] Among them, according to the current cluster centers of at least one current data set, completing the current clustering of the current data in the metric data of multiple initial metrics to form at least one next data set includes:

[0142] Using a preset similarity algorithm, compare the current data with at least one current cluster center respectively to obtain at least one second similarity comparison result;

[0143] Add the current data to the target current data set, so as to form at least one next data set through the target current data set and the unupdated current data set, where the second similarity comparison result between the current cluster center of the target current data set and the current data is greater than a preset similarity threshold.

[0144] Figure 5 Schematically shows a flowchart of clustering the index data of multiple initial indexes of a faulty device according to an embodiment of the present disclosure. The following combines Figure 5 to give an exemplary illustration of the above method.

[0145] As Figure 5 shown, the operations of the above clustering process include the following steps:

[0146] Step (1), sample the index data of multiple initial indexes of the faulty device from the historical index data of the known faulty device, and obtain the values of multiple normalized indexes X1, X2,..., X n , these indexes are all n-dimensional matrices, and are data of multiple indexes of the same data dimension that may be associated with the known faults of the faulty device (it can be a single index or multiple indexes);

[0147] Step (2), calculate the similarity between the data of these indexes and the current cluster centers of the existing current data sets respectively. For example, use the cosine similarity algorithm to calculate the similarity (the first input index forms a single cluster and no calculation is required), and obtain multiple similarity comparison results;

[0148] Step (3), if the cosine value of the angle calculated between the current index data and the cluster center of a certain data set is greater than a preset similarity threshold (this threshold can be adjusted), then add this index data X i to this data set and update the calculation of the cluster center of this data set; if the cosine value of the angle calculated between the current index data and the cluster center of a certain data set is less than or equal to the preset similarity threshold (this threshold can be adjusted), then use this index X i as a new cluster, and set X i as the cluster center of the new cluster.

[0149] According to an embodiment of the present disclosure, the preset similarity threshold can be 0.9. Setting the preset similarity threshold to 0.9 is an empirical value obtained through a large number of experiments and verified. During a large number of experiments, it is found that when performing a clustering algorithm on a large number of samples, when the similarity between two data is 0.9, the clustering effect is better, and vice versa, the clustering effect is worse. For example, when the similarity is less than 0.85, the clustering is too large; when the similarity is greater than 0.95, the clustering is too loose.

[0150] In step (iii) above, calculate the cluster center X of each data set core The method shown in the following formula (iii) can be adopted:

[0151]

[0152] where X i represents the data of each index in the data set

[0153] In step (iv), loop through all the metrics E1, E2,..., En, and finally obtain a series of clusters. For clusters with only one sample or a small number of samples, the cluster data set can be excluded, and the minimum threshold of the number of samples in the data set can be specified

[0154] According to the embodiments of the present disclosure, through the above clustering algorithm, it is possible to examine the continuity between samples from the perspective of sample density, and continuously expand the clustering clusters based on the connectable samples to obtain the final clustering result. Clustering is performed based on the density of the data set in the spatial distribution, and finally multiple fault metric data sets are obtained. The advantage of clustering through the above clustering algorithm is that when the density of the sample set is uneven, it will not affect the clustering effect, and several clusters will still be formed (there may be more clusters in dense areas and fewer clusters in sparse areas).

[0155] According to the embodiments of the present disclosure, after obtaining the target reference data and the multiple to-be-detected metric data of the device to be detected, the detection result for characterizing whether the device to be detected will have a fault of the target fault category within a preset future time period based on the multiple to-be-detected metric data and the target reference data includes:

[0156] Compare the similarity of each to-be-detected metric data with the target reference data respectively to obtain the first similarity comparison result

[0157] According to the first similarity comparison result, determine the detection result for characterizing whether the device to be detected will have a fault of the target fault category within a preset future time period. For example, when the similarity value is greater than the preset threshold, it can be predicted that the device may have a fault of the target fault category, and when the similarity value is less than or equal to the preset threshold, it can be predicted that the device will not have a fault of the target fault category

[0158] According to the embodiments of the present disclosure, in the above operation, comparing the similarity of each to-be-detected metric data with the target reference data respectively to obtain the similarity comparison result specifically includes:

[0159] Calculate the cluster center of the metric data of at least one associated metric to obtain the target cluster center

[0160] Using a preset similarity algorithm, compare the data of each index to be detected with the target cluster center respectively to obtain the similarity comparison result.

[0161] The following is an exemplary description of the operation method of the above similarity comparison:

[0162] According to an embodiment of the present disclosure, the target reference data includes, for example, a plurality of associated index data associated with the target fault category, and each associated index data may be data that has been normalized, for example:

[0163] Same type of fault 1: Index X1, Index X2, Index X3,...

[0164] Same type of fault 2: Index Y1, Index Y2, Index Y3,...

[0165] Same type of fault 3: Index Z1, Index Z2, Index Z3,...

[0166] Among them, the fault indicators X, fault indicators Y, fault indicators Z, etc. are all n-dimensional matrices, respectively representing the same normalized index (which can be a single index or multiple indices) of a certain type of fault. The target reference data can be, for example, any one of the above fault indicators X, fault indicators Y, fault indicators Z...

[0167] According to an embodiment of the present disclosure, the data of multiple indices to be detected of the device to be detected can be obtained from the operation and maintenance big data platform. The operation and maintenance big data platform performs data cleaning on the index data of the software and hardware of all the devices to be detected that are being monitored, including index dimension consistency processing (the same sampling times) and normalization processing (single-index normalization, multi-index normalization), to obtain a series of index sequences of faults to be determined, denoted as: A1, A2, A3,..., An (all n-dimensional matrices);

[0168] The following matrix shows the sampling data of one of the indices:

[0169] A i =[x1, x2, x3,..., x n

[0170] According to an embodiment of the present disclosure, first, the cluster centers of the index data of a plurality of associated indices (fault indicators X, fault indicators Y, fault indicators Z...) respectively associated with each fault category can be calculated to obtain a plurality of reference cluster centers: X core , Y core , Z core ..., where the method for calculating the target cluster center can adopt the method shown in formula (4), which will not be elaborated here. Among them, the target cluster center is the above X core , Y core , Z core ​……, which is related to the fault category that the user is currently dealing with.

[0171] Then, using a preset similarity algorithm, the data of each index to be detected is respectively compared with the target cluster center for similarity to obtain a similarity comparison result. For example, the cosine of the angle algorithm can be used to calculate the distances between A1, A2, A3, ..., An and the target cluster center (X core , Y core , Z core ……, one of them), and the cosine of the angle algorithm is as follows:

[0172]

[0173] According to an embodiment of the present disclosure, the similarity value calculated by the above cosine of the angle algorithm is within the range of [0, 1]. When the similarity value is greater than a preset threshold of 0.9 (this threshold can be adjusted), it is considered that the distance between a certain index Ai of the device to be detected and the cluster center of the target fault category is very close, indicating that this index can be classified into the target fault category, and it means that the probability of the device having the target fault in the future is very high. On the contrary, when the similarity value is less than or equal to the preset threshold of 0.9 (this threshold can be adjusted), it is predicted that the probability of the device having a fault of the target fault category is relatively small.

[0174] According to an embodiment of the present disclosure, the above fault determination method exemplifies a method for predicting whether a device to be detected will have a target fault, that is, a method for predicting whether a specific type of fault will occur. The above fault determination method can be extended to the situation of respectively predicting whether a device to be detected will have various types of faults. In this situation, for example, the cosine of the angle algorithm can be used to calculate the distances between A1, A2, A3, .., An and the index cluster centers of multiple fault categories (X core , Y core , Z core ……, one of them), and according to the result of the similarity calculation, it is respectively determined whether the device to be detected will have each type of fault.

[0175] According to an embodiment of the present disclosure, in the process of similarity calculation, it is not necessary to calculate the cosine of the angle between A1, A2, A3, ..., An and X core , Y core , Z core …… one by one in a loop. The cosine of the angle values that need to be calculated can be quickly calculated by changing A1, A2, A3, ..., An and X core , Y core , Z core …… into two two-dimensional matrices respectively and calculating the dot product of the two two-dimensional matrices, so as to speed up the calculation speed and save calculation resources.

[0176] According to an embodiment of the present disclosure, by using a preset similarity algorithm, the similarity between each data of the index to be detected and the target cluster center is compared respectively to obtain the similarity comparison result, realizing the quantitative calculation of the distance between the index to be measured and the fault index, and can more accurately realize the pre-discovery of faults.

[0177] According to an embodiment of the present disclosure, obtaining the target reference data may also be, according to the user-defined fault category, first obtaining the fault curve associated with the user-defined fault type, and then extracting a preset number of index values from the fault curve as the values of the associated index associated with the user-defined fault category. Specifically, in this case, the target reference data includes: the index data of the preselected index under the predefined fault type. Obtaining the target reference data includes:

[0178] Obtaining the fault curve associated with the predefined fault type, where the fault curve is used to characterize the trend of the index value of the preselected index changing with time; the fault curve associated with the predefined fault type may be a curve drawn by the operation and maintenance personnel according to the operation and maintenance experience, and it is preset that this curve represents a special fault type.

[0179] According to the preset time interval, extract a preset number of index values from the fault curve and determine them as the target reference data. For example, this curve can be placed in a two-dimensional coordinate system (horizontal X, vertical Y), evenly divided into n segments, extract the corresponding y values, and after normalization, obtain the index R, that is, obtain the target reference data R.

[0180] According to an embodiment of the present disclosure, after obtaining the target reference data R, the current monitoring indexes A1, A2,..., An of all software and hardware targets can be calculated with R respectively for the cosine value of the included angle; if the cosine value of the included angle exceeds the threshold, it is considered that this software and hardware has the risk of an R-type fault occurring.

[0181] According to an embodiment of the present disclosure, through the above method, obtaining the fault curve associated with the predefined fault type, determining the target reference data according to the fault curve, and using the target reference data as the reference object to perform fault prediction on the device to be measured, the purpose of discovering faults according to the user-defined fault type is realized, providing a mechanism for quickly constructing a fault pattern and finding similar faults, which is convenient for discovering potential fault types that have not been characterized yet.

[0182] Based on the above fault determination method, the present disclosure also provides a fault determination device. The following will be combined with Figure 6 Describe this device in detail.

[0183] Figure 6 Schematically shows the structural block diagram of the fault determination device according to an embodiment of the present disclosure. As Figure 6As shown, the fault determination device of this embodiment includes a first acquisition module 610, a second acquisition module 620, and a first determination module 630.

[0184] Among them, the first acquisition module 610 is configured to acquire target reference data, where the target reference data includes: index data of at least one associated index of a faulty device within a preset historical fault time period, and at least one associated index is associated with the target fault category;

[0185] The second acquisition module 620 is configured to acquire a plurality of index data to be detected, where each index data to be detected includes: index data of one index to be detected of a device to be detected within a time period to be detected;

[0186] The first determination module 630 is configured to determine a detection result for characterizing whether a fault of the target fault category will occur in the device to be detected within a preset future time period according to the plurality of index data to be detected and the target reference data.

[0187] According to an embodiment of the present disclosure, the target reference data acquired by the first acquisition module 610 includes data of a plurality of associated indexes associated with the target fault category, which characterizes the potential association relationship between the fault type and the index type. The first determination module 630 determines whether a fault of the target fault category will occur in the device to be detected according to the plurality of index data to be detected and the target reference data, realizing the fusion processing of multiple associated indexes, rather than just comparing within the same index. The fault prediction result is more accurate. In addition, the index data to be detected obtained by the second acquisition module 620 includes a plurality of index data of the device to be detected, considering the influence of multiple index types that may be associated with the occurrence of a fault. The calculated indexes do not depend on the data of the same software and hardware device in the same service environment, and have a wider application range. They can be used in more complex fault scenarios, and the fault prediction result is more referenceable.

[0188] According to an embodiment of the present disclosure, the first determination module includes a comparison unit and a first determination unit.

[0189] Among them, the comparison unit is configured to respectively compare the similarity of each index data to be detected with the target reference data to obtain a first similarity comparison result;

[0190] The first determination unit is configured to determine a detection result for characterizing whether a fault of the target fault category will occur in the device to be detected within a preset future time period according to the first similarity comparison result.

[0191] According to an embodiment of the present disclosure, the comparison unit includes a calculation subunit and a comparison subunit.

[0192] Among them, a calculation subunit is configured to calculate the cluster center of the index data of at least one associated index to obtain a target cluster center;

[0193] A comparison subunit is configured to use a preset similarity algorithm to separately compare the similarity of each index data to be detected with the target cluster center to obtain a similarity comparison result.

[0194] According to an embodiment of the present disclosure, it further includes: a second determination module configured to determine at least one associated index associated with the target fault category.

[0195] Among them, the second determination module includes a first acquisition unit, a clustering unit, a second determination unit, and a first extraction unit.

[0196] Among them, the first acquisition unit is configured to acquire the index data of multiple initial indexes of a faulty device within a preset historical fault time period;

[0197] The clustering unit is configured to perform clustering processing on the index data of multiple initial indexes of the faulty device to obtain multiple faulty index data sets;

[0198] The second determination unit determines the faulty index data set that meets the preset discrimination condition as the target faulty index data set;

[0199] The first extraction unit is configured to extract at least one associated index associated with the target fault category from the target faulty index data set.

[0200] According to an embodiment of the present disclosure, among them, the clustering unit includes a clustering subunit configured to, in the case of completing the current clustering of the current data in the index data of multiple initial indexes according to the current cluster center of at least one current data set to form at least one next data set, iteratively execute: updating the cluster center of the next data set, so as to complete the next clustering of the next data in the index data of multiple initial indexes according to the updated cluster center of the next data set until a preset termination condition is reached, and obtaining multiple finally determined faulty index data sets.

[0201] According to an embodiment of the present disclosure, among them, in the clustering subunit, completing the current clustering of the current data in the index data of multiple initial indexes according to the current cluster center of at least one current data set to form at least one next data set includes:

[0202] Using a preset similarity algorithm to separately compare the similarity of the current data with at least one current cluster center to obtain at least one second similarity comparison result;

[0203] Add the current data to the target current data set so as to form at least one next data set through the target current data set and the unupdated current data set, wherein the comparison result of the second similarity between the current cluster center of the target current data set and the current data is greater than 0.9.

[0204] According to an embodiment of the present disclosure, wherein the first acquisition module includes an acquisition unit and a processing unit.

[0205] Among them, the acquisition unit is used to acquire original target data, where the original target data includes: original index data of at least one associated index of a faulty device within a preset historical fault time period;

[0206] The processing unit is used to perform normalization processing on the original index data of at least one associated index respectively to obtain target reference data.

[0207] According to an embodiment of the present disclosure, wherein the target reference data further includes: index data of preselected indexes under a predefined fault type.

[0208] Among them, the first acquisition module includes a second acquisition unit and a second extraction unit.

[0209] Among them, the second acquisition unit is used to acquire a fault curve associated with a predefined fault type, where the fault curve is used to characterize the trend of the index value of the preselected index changing with time;

[0210] The second extraction unit is used to extract a preset number of index values from the fault curve at preset time intervals and determine them as target reference data.

[0211] According to embodiments of the present disclosure, any plurality of the first acquisition module 610, the second acquisition module 620, and the first determination module 630 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to embodiments of the present disclosure, at least one of the first acquisition module 610, the second acquisition module 620, and the first determination module 630 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system in package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in any suitable combination of several of them. Alternatively, at least one of the first acquisition module 610, the second acquisition module 620, and the first determination module 630 may be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.

[0212] Figure 7 A block diagram of an electronic device suitable for implementing a fault determination method according to an embodiment of the present disclosure is schematically shown.

[0213] As Figure 7 shown, the electronic device 700 according to an embodiment of the present disclosure includes a processor 701, which can perform various appropriate actions and processes according to a program stored in a read only memory (ROM) 702 or a program loaded from a storage section 708 into a random access memory (RAM) 703. The processor 701 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 701 may also include on-board memory for caching purposes. The processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0214] In the RAM 703, various programs and data required for the operation of the electronic device 700 are stored. The processor 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. The processor 701 performs various operations of the method flow according to the embodiments of the present disclosure by executing the programs in the ROM 702 and / or the RAM 703. It should be noted that the programs may also be stored in one or more memories other than the ROM 702 and the RAM 703. The processor 701 may also perform various operations of the method flow according to the embodiments of the present disclosure by executing the programs stored in the one or more memories.

[0215] According to an embodiment of the present disclosure, the electronic device 700 may further include an input / output (I / O) interface 705, and the input / output (I / O) interface 705 is also connected to the bus 704. The electronic device 700 may further include one or more of the following components connected to the I / O interface 705: an input portion 706 including a keyboard, a mouse, etc.; an output portion 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 708 including a hard disk, etc.; and a communication portion 709 including a network interface card such as a LAN card, a modem, etc. The communication portion 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed so that a computer program read from it can be installed into the storage portion 708 as needed.

[0216] The present disclosure also provides a computer-readable storage medium, which may be included in the device / device / system described in the above embodiments; or may exist separately without being assembled into the device / device / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present disclosure is implemented.

[0217] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include one or more memories other than the above-described ROM 702 and / or RAM 703 and / or ROM 702 and RAM 703.

[0218] An embodiment of the present disclosure also includes a computer program product, which includes a computer program, and the computer program contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the fault determination method provided by the embodiment of the present disclosure.

[0219] When the computer program is executed by the processor 701, it executes the above functions defined in the system / apparatus of the embodiment of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0220] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 709, and / or be installed from the removable medium 711. The program code included in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0221] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 709, and / or be installed from the removable medium 711. When the computer program is executed by the processor 701, it executes the above functions defined in the system of the embodiment of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0222] According to embodiments of the present disclosure, program code for executing the computer programs provided by the embodiments of the present disclosure may be written in any combination of one or more programming languages. Specifically, these computing programs may be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).

[0223] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a portion of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0224] Those skilled in the art will appreciate that the features recited in the various embodiments and / or claims of the present disclosure may be combined or combined in a variety of ways, even if such combinations or combinations are not explicitly recited in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features recited in the various embodiments and / or claims of the present disclosure may be combined and combined in a variety of ways. All such combinations and / or combinations fall within the scope of the present disclosure.

[0225] The embodiments of the present disclosure have been described above. However, these embodiments are merely for illustrative purposes and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and these substitutions and modifications should fall within the scope of the present disclosure.

Claims

1. A fault determination method, comprising: Obtaining target reference data, where the target reference data includes: index data of at least one associated index of a faulty device within a preset historical fault time period, and the at least one associated index is associated with a target fault category; Obtaining a plurality of index data to be detected, where each index data to be detected includes: index data of one index to be detected of a device to be detected within a time period to be detected; Determining a detection result for characterizing whether a fault of the target fault category will occur in the device to be detected within a preset future time period according to the plurality of index data to be detected and the target reference data; Further comprising, when the type of the associated index associated with the target fault category is unknown, performing the following operations: Obtaining index data of a plurality of initial indexes of the faulty device within the preset historical fault time period, and performing normalization to obtain values of a plurality of normalized indexes, where each index is an n-dimensional matrix and is data of a plurality of indexes in the same data dimension that may be associated with known faults of the faulty device; Performing clustering processing on the index data of the plurality of initial indexes of the faulty device to obtain a plurality of fault index data sets; Taking the fault index data set with the largest number of index types as the target fault index data set; Extracting at least one associated index associated with the target fault category from the target fault index data set.

2. The method according to claim 1, where the determining a detection result for characterizing whether a fault of the target fault category will occur in the device to be detected within a preset future time period according to the plurality of index data to be detected and the target reference data includes: Performing a similarity comparison between each index data to be detected and the target reference data respectively to obtain a first similarity comparison result; Determining a detection result for characterizing whether a fault of the target fault category will occur in the device to be detected within a preset future time period according to the first similarity comparison result.

3. The method according to claim 2, wherein The performing a similarity comparison between each index data to be detected and the target reference data respectively to obtain a similarity comparison result includes: Calculating the cluster center of the index data of the at least one associated index to obtain a target cluster center; Using a preset similarity algorithm to perform a similarity comparison between each index data to be detected and the target cluster center respectively to obtain the similarity comparison result.

4. The method according to claim 1, where the performing clustering processing on the index data of the plurality of initial indexes of the faulty device to obtain a plurality of fault index data sets includes: When, according to the current cluster center of at least one current data set, the current data in the index data of the plurality of initial indexes is completed for the current clustering to form at least one next data set, iteratively execute: updating the cluster center of the next data set, so that according to the updated cluster center of the next data set, the next data in the index data of the plurality of initial indexes is completed for the next clustering until a preset termination condition is reached, and finally determining a plurality of the fault index data sets.

5. The method according to claim 4, wherein, Completing the current clustering of the current data in the metric data of the multiple initial metrics according to the current cluster centers of at least one current data set to form at least one next data set includes: Using a preset similarity algorithm, comparing the similarity of the current data with at least one of the current cluster centers respectively to obtain at least one second similarity comparison result; Adding the current data to the target current data set, so as to form the at least one next data set through the target current data set and the unupdated current data set, where the second similarity comparison result between the current cluster center of the target current data set and the current data is greater than a preset similarity threshold.

6. The method according to claim 5, wherein: The preset similarity threshold is 0.

9.

7. The method according to claim 1, wherein, The obtaining of the target reference data includes: Collecting original target data, where the original target data includes: the original metric data of at least one of the associated metrics of the faulty device during the preset historical fault time period; Performing normalization processing on the original metric data of at least one of the associated metrics respectively to obtain the target reference data.

8. The method according to claim 1, wherein The target reference data further includes: the metric data of the preselected metrics under the predefined fault type; Wherein, the obtaining of the target reference data includes: Obtaining a fault curve associated with the predefined fault type, where the fault curve is used to characterize the trend of the metric value of the preselected metric changing with time; Extracting a preset number of metric values from the fault curve at preset time intervals and determining them as the target reference data.

9. A fault determination device, comprising: A first obtaining module, configured to obtain target reference data, where the target reference data includes: the metric data of at least one associated metric of a faulty device during a preset historical fault time period, where the at least one associated metric is associated with a target fault category; A second obtaining module, configured to obtain a plurality of metric data to be detected, where each of the metric data to be detected includes: the metric data of one of the metrics to be detected of the device to be detected during the time period to be detected; A first determining module, configured to determine a detection result for characterizing whether the device to be detected will have a fault of the target fault category within a preset future time period according to the plurality of metric data to be detected and the target reference data; A second determining module, configured to determine at least one associated metric associated with the target fault category when the type of the associated metric associated with the target fault category is unknown. The second determining module includes: A first obtaining unit, configured to obtain the metric data of a plurality of initial metrics of a faulty device during a preset historical fault time period, and perform normalization to obtain the values of a plurality of normalized metrics, and each metric is an n-dimensional matrix, which is the data of a plurality of metrics in the same data dimension that may be associated with the known faults of the faulty device; A clustering unit, configured to perform clustering processing on the metric data of a plurality of initial metrics of a faulty device to obtain a plurality of fault metric data sets; A second determining unit, configured to use one of the fault metric data sets with the largest number of metric types as the target fault metric data set; A first extraction unit, configured to extract at least one associated metric associated with a target fault category from a target fault metric dataset.

10. An electronic device, comprising: One or more processors; A storage device for storing one or more programs, wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the method according to any one of claims 1 to 8.

11. A computer-readable storage medium, having stored thereon executable instructions that, when executed by a processor, cause the processor to execute the method according to any one of claims 1 to 8.

12. A computer program product, comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Fault prediction method and device, storage medium and electronic equipment

    CN110349289A

  • Fault early warning method and device and computer storage medium

    CN110740061A

  • Fault prediction method and device, computing device and computer readable storage medium

    CN110851342A