Classification device, classification method, and program

The classification device effectively distinguishes between known and unknown failures by clustering observability data, allowing for autonomous and assisted recovery processes.

JP7710147B2Active Publication Date: 2025-07-18NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024519148
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-05-02
Publication Date
2025-07-18
Estimated Expiration
2042-05-02

AI Technical Summary

Technical Problem

Existing technologies fail to differentiate between known and unknown failures in maintenance operations, leading to difficulties in autonomous recovery, as they lack a method to classify failures and formulate appropriate countermeasures.

Method used

A classification device that extracts abnormal observability data, clusters it with past known events, calculates ratios, and classifies failures as known or unknown, enabling automatic recovery for known failures and assisted recovery for unknown failures.

Benefits of technology

Enables rapid classification of failures as known or unknown, facilitating autonomous recovery for known failures and assisted recovery for unknown failures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007710147000001
    Figure 0007710147000001
  • Figure 0007710147000002
    Figure 0007710147000002
  • Figure 0007710147000003
    Figure 0007710147000003
Patent Text Reader

Abstract

A classification device 10 for classifying failures that occurred in a service being monitored comprises a failure data extraction unit 11 and an event classification unit 12. The failure data extraction unit 11 extracts a failure data item from observability data items at a time when failures occurred in the service being monitored. The event classification unit 12, for respective known events that were addressed in the past, classifies observability data items and failure data items of the known events into clusters, calculates the proportions of the failure data items included in the clusters for the known events, and classifies, on the basis of the proportions calculated for the respective known events, the occurred failure into either a known event or an unknown event which was not addressed in the past.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a classification device, a classification method, and a program.

Background Art

[0002] As a known maintenance automation technology for assisting the judgment of maintenance personnel in maintenance operations, an autonomous control loop method has been proposed. The autonomous control loop method is a technology in which each operation component operates autonomously by componentizing and autonomizing the functions of maintenance operations. In the autonomous control loop method, the target service is monitored, and the aim is to achieve autonomous follow-up to the addition of new functions and specification changes to the service and automatic recovery when a failure occurs.

[0003] In Non-Patent Document 1, an information acquisition method for observability data called Logs / Metrics / Tracing has been proposed in the autonomous control loop method, and in Non-Patent Document 2, a method for processing observability data and exploring the cause at the time of a failure has been proposed. These technologies can support rapid recovery by maintenance personnel by acquiring observability data from the monitored service and analyzing the observability data.

Prior Art Documents

Non-Patent Documents

[0004]

Non-Patent Document 1

Non-Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0005] Non-Patent Documents 1 and 2 mention the data acquisition method and the exploration of factors at the time of failure, but do not mention the response method at the time of failure. Autonomous recovery for all failures is difficult to achieve because a device that can judge all operations such as the pursuit of factors and the formulation of countermeasures is required. Therefore, for known failures that have been handled by maintenance personnel in the past, the monitored service is autonomously recovered by performing recovery processing without the intervention of the maintenance personnel, and for unknown failures that the maintenance personnel have no experience in handling, the recovery of the failure is assisted by presenting factors related to the failure, aiming at quick recovery.

[0006] In off-the-shelf technologies such as random forest used for classification, it is possible to learn data with pre-assigned labels and assign the input data to one of the pre-learned labels. However, since the input data is always assigned to one of the learned labels, there is a problem that unknown events are classified into existing labels.

[0007] The present invention has been made in view of the above, and an object thereof is to classify whether the occurred failure is a failure that has been handled in the past or a failure that has not been handled.

Means for Solving the Problems

[0008] A classification device according to an aspect of the present invention is a classification device that classifies a failure that has occurred in a monitored service, and includes an extraction unit that extracts abnormal observability data from the observability data at the time of failure of the monitored service, and for each known event that has been handled in the past, classifies the observability data of the known event and the abnormal observability data into clusters, calculates the ratio of the abnormal observability data included in the cluster of the known event, and based on the ratio calculated for each known event, classifies the occurred failure into either an unknown event that has not been handled in the past or the known event.

Effects of the Invention

[0009] According to the present invention, it is possible to classify whether the occurred failure is a failure that has been dealt with in the past or a failure that has not been dealt with.

Brief Description of Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Embodiments for Carrying Out the Invention

[0011] Hereinafter, embodiments of the present invention will be described with reference to the drawings.

[0012] [Configuration] Referring to FIG. 1, the configuration of the classification device 10 according to the present embodiment will be described. The classification device 10 shown in FIG. 1 analyzes the observability data of the monitored service at the time of a failure and determines whether the occurred failure is a known event or an unknown event. In the system of FIG. 1, when the failure is a known event, a recovery process is automatically performed without the intervention of a maintenance person, and when the failure is an unknown event, a cause search process is performed to estimate the cause of the failure and the result is presented to the maintenance person. A known event is a failure that has been dealt with in the past. An unknown event is a failure that has not been dealt with in the past and is being dealt with for the first time.

[0013] A data storage unit 30 is connected to the classification device 10. The data storage unit 30 stores the observability data acquired from the management target that provides the monitored service. The management target is, for example, a device or container used to provide a service, software operating on the device or container, and the like. When the monitored service is a system using an autonomous control loop method, the observability data may be acquired from each operation component of the autonomous control loop method by the method of Non-Patent Document 1. The observability data is, for example, a log, metrics, and trace that can be acquired from the management target. The observability data for each management target may be combined into one in time units using the method of Non-Patent Document 2. Thereby, the time difference of the observability data is eliminated, and the possibility that each of the management targets affects other management targets can be considered.

[0014] The classification device 10 includes a failure data extraction unit 11, an event classification unit 12, and a known event data storage unit 13.

[0015] The failure data extraction unit 11 classifies the observable data at the time of failure into normal data and failure data, and extracts the failure data. Specifically, the failure data extraction unit 11 acquires the observable data near the time zone when the failure occurs from the data storage unit 30, classifies each of the observable data into normal data and failure data using a machine learning model, and outputs the failure data to the event classification unit 12. The normal data is normal observable data acquired when each of the management targets of the monitored service is operating normally. The failure data is abnormal observable data acquired when a failure occurs in the monitored service or when a failure is likely to occur. The machine learning model is a model that is learned to classify observable data into normal data and failure data using the observable data stored in the known event data storage unit 13 as teacher data. The failure data extraction unit 11 may learn the machine learning model when extracting failure data from the observable data.

[0016] The event classification unit 12 compares the failure data with the observable data for each known event, and determines whether the failure is a known event or an unknown event. If the failure is a known event, the event classification unit 12 determines the known event corresponding to the failure. Specifically, the event classification unit 12 acquires the observable data for each known event from the known event data storage unit 13, and for each known event, clusters the failure data and the observable data of the known event into two clusters, namely, a "known event cluster" and an "other cluster". If most of the failure data is classified into the other cluster in all trials for each known event, the event classification unit 12 determines the failure as an unknown event. In addition, the event classification unit 12 determines the known event with a proportion of the failure data classified higher than the threshold as the known event of the failure. The event classification unit 12 may further reclassify the failure data determined as a known event using the machine learning model to improve the determination accuracy.

[0017] The known event data storage unit 13 stores data with event labels attached to the observability data. The event label is information indicating the classification of known events. Hereinafter, the classification of known events is also referred to as the failure type. The known event data storage unit 13 stores not only failure data but also normal data. The observability data stored in the known event data storage unit 13 is used as teacher data for a machine learning model used when classifying the observability data into normal data and failure data. It is also used when the event classification unit 12 determines the failure type of the failure data.

[0018] The classification result of the classification device 10 can be used by the cause search processing unit 50 and the recovery processing operation component 60. For example, when classifying a failure as an unknown event, the classification device 10 transmits the failure data to the cause search processing unit 50. The cause search processing unit 50 analyzes the failure data, estimates the cause of the failure, and presents it to the maintainer. For the processing of the cause search processing unit 50, for example, the method of Non-Patent Document 2 can be used. When classifying a failure as one of the known events 1-N, the classification device 10 instructs the recovery processing operation component 60 corresponding to the known events 1-N to perform the recovery processing of the failure. The recovery processing operation component 60 executes the recovery processing according to the procedures determined for each of the known events 1-N, such as restarting the software or restarting the device. Note that even if the cause search processing unit 50 and the recovery processing operation component 60 are not provided, prompt recovery can be expected by presenting to the maintainer whether it is an unknown event or a known event.

[0019] [Operation] Next, with reference to the flowchart of FIG. 2, an example of the learning process of a machine learning model for classifying observability data into normal data and failure data will be described. The process shown in FIG. 2 may be executed when executing the process of classifying the observability data into normal data and failure data described later, or may be executed when the data in the known event data storage unit 13 is updated.

[0020] In step S11, the failure data extraction unit 11 acquires observability data from the known event data storage unit 13. Fig. 3 shows an example of the observability data stored in the known event data storage unit 13. In the example of Fig. 3, the observability data acquired from each of the management targets of the monitored service is combined into one for each time, and an event label is assigned to each of the observability data. For example, as the event label, 0 is assigned to normal data, and a numerical value corresponding to the failure type is assigned to failure data. When the failure data extraction unit 11 acquires the observability data, it converts the event label of the observability data into a binary value of normal or abnormal. Specifically, the event label of normal data remains 0, and the event labels of failure data other than normal data are all converted to 1.

[0021] In step S12, the failure data extraction unit 11 applies principal component analysis to the acquired observability data for dimensionality reduction and calculates the eigenvectors. There may be more than 100 types of data in the observability data. When the number of items is large, the number of dimensions of the data (the number of columns in the table of Fig. 3) is reduced by principal component analysis. Fig. 4 shows an example of the data obtained by applying principal component analysis to the observability data of Fig. 3 for dimensionality reduction. Fig. 4 shows the first principal component (PC1) and the second principal component (PC2) obtained by principal component analysis. The principal component scores from the third principal component onwards may also be used. The eigenvectors obtained by the principal component analysis in the learning process are used to calculate the principal component scores of the observability data to be classified in the process of classifying the observability data into normal data and failure data, which will be described later.

[0022] In step S13, the failure data extraction unit 11 uses the observability data dimensionally reduced in step S12 as training data to train a machine learning model for classifying the observability data into normal data and failure data. For example, the failure data extraction unit 11 uses random forest, which is one of the classification methods, to create a machine learning model for classifying the observability data into normal data and failure data.

[0023] Next, referring to the flowchart of FIG. 5, an example of the process of classifying the observability data at the time of failure into normal data and failure data will be described. The process shown in FIG. 5 is executed when a failure in the service to be monitored is detected.

[0024] In step S21, the failure data extraction unit 11 acquires the observability data near the time zone when the failure occurred from the data storage unit 30. FIG. 6 shows an example of the observability data at the time of failure. In the example of FIG. 6, the observability data combined into one at 10-second intervals is shown.

[0025] In step S22, the failure data extraction unit 11 compresses the dimension of the observability data at the time of failure acquired in step S21 using the eigenvector calculated in step S12 of FIG. 2. FIG. 7 shows an example of the dimension compression of the observability data in FIG. 6. The failure data extraction unit 11 calculates PC1 and PC2 shown in FIG. 7 based on the observability data and the eigenvector in FIG. 6. Note that in FIG. 7, the determination result of the next step S23 is also shown.

[0026] In step S23, the failure data extraction unit 11 inputs the dimension-compressed observability data into the machine learning model learned in the process of FIG. 2, and classifies the observability data into normal data and failure data. In the example of FIG. 7, the two rows of observability data indicated by the arrow are classified as failure data.

[0027] In step S24, the failure data extraction unit 11 extracts the observability data before dimension compression of the observability data classified as failure data, and outputs it to the event classification unit 12. Specifically, the failure data extraction unit 11 outputs the two rows of observability data determined as failure data among the observability data in FIG. 6 to the event classification unit 12.

[0028] Next, referring to the flowchart of FIG. 8, an example of the process of classifying the failure event will be described. The process shown in FIG. 8 is executed when the failure data is input from the failure data extraction unit 11.

[0029] In step S31, the event classification unit 12 acquires the observability data of one known event from the known event data storage unit 13. That is, the event classification unit 12 acquires the observability data with the same event label from the known event data storage unit 13. For example, in the first execution of the loop, the observability data with the event label 1 is acquired, and in the second execution, the observability data with the event label 2 is acquired. The processing from step S31 to S35 is repeated N times for each of the known events 1-N except for normal.

[0030] In step S32, the event classification unit 12 combines the observability data of the known event acquired in step S31 and the failure data.

[0031] In step S33, the event classification unit 12 applies principal component analysis to the combined data for dimensionality reduction. Fig. 9 shows an example of the data obtained by combining the observability data of the known event and the failure data in the row direction and performing dimensionality reduction. The data in the upper three rows of the example in Fig. 9 is the observability data of the same known event acquired from the known event data storage unit 13. The data in the lower two rows of the example in Fig. 9 is the failure data to be classified.

[0032] In step S34, the event classification unit 12 clusters the combined data into two clusters. For clustering, for example, Minibatch K-means, which is one of the unsupervised learning methods, can be used. Among the two clusters, one cluster is a known event cluster including the observability data of the existing event, and the other cluster is the other cluster. Fig. 10 shows the state of clustering the observability data of the known event and the failure data for each known event. As shown in Fig. 10, for each of the known events 1-N, the data obtained by combining the observability data of the known event and the failure data is clustered into two clusters. Note that since the event classification unit 12 clusters the observability data of the known event and the failure data so that two clusters are formed, when a failure of the known event occurs, the observability data or failure data of the known event may be classified into the other cluster.

[0033] In step S35, the event classification unit 12 calculates the ratio of the failure data classified into the known event cluster. If a large amount of failure data is included in the known event cluster, the failure can be estimated to be the known event.

[0034] The processes from step S31 to S35 are executed for each of the known events 1-N, and the ratio of the failure data included in each cluster of the known events 1-N is calculated.

[0035] In step S36, the event classification unit 12 classifies the occurred failure into either an unknown event or one of the known events 1-N based on the ratio of the failure data included in each calculated cluster for each of the known events 1-N. Specifically, if the ratio of the failure data classified into the known event cluster is lower than the threshold for all of the known events 1-N, the event classification unit 12 classifies the occurred failure as an unknown event. Also, the event classification unit 12 determines the known event of the occurred failure as the known event 1-N for which the ratio of the failure data classified into the known event cluster is higher than the threshold.

[0036] When the event classification unit 12 classifies the occurred failure into one of the known events 1-N, it may confirm that the failure data is classified into the known event 1-N by re-executing classification such as random forest on the failure data.

[0037] As described above, the present embodiment is a classification device 10 that classifies failures occurring in a service to be monitored, and includes a failure data extraction unit 11 and an event classification unit 12. The failure data extraction unit 11 extracts failure data from observable data at the time of occurrence of a failure in the service to be monitored. The event classification unit 12 classifies the observable data and the failure data of each known event that has been dealt with in the past into clusters, calculates the ratio of the failure data included in the cluster of the known event, and based on the ratio calculated for each known event, classifies the occurred failure into either an unknown event that has not been dealt with in the past or a known event. Thereby, it is possible to determine whether the failure that has occurred in the service to be monitored is an unknown event that has not been dealt with in the past or a known event that has been dealt with in the past, and rapid failure recovery becomes possible.

[0038] For the classification device 10 described above, for example, a general-purpose computer system including a central processing unit (CPU) 901, a memory 902, a storage 903, a communication device 904, an input device 905, and an output device 906 as shown in FIG. 11 can be used. In this computer system, the classification device 10 is realized by the CPU 901 executing a predetermined program loaded on the memory 902. This program can be recorded on a computer-readable recording medium such as a magnetic disk, an optical disk, or a semiconductor memory, or can be distributed via a network.

Explanation of Signs

[0039] 10 Classification device 11 Failure data extraction unit 12 Event classification unit 13 Known event data storage unit 30 Data storage unit 50 Cause exploration processing unit 60 Recovery processing operation component

Claims

1. A classification device for classifying faults occurring in a service to be monitored, comprising: an extraction unit that extracts abnormal observability data from the observability data at the time of occurrence of a fault in the service to be monitored; for each known event that has been dealt with in the past, classifying the observability data of the known event and the abnormal observability data into clusters, calculating the ratio of the abnormal observability data included in the cluster of the known event, and classifying the occurred fault into either an unknown event that has not been dealt with in the past or the known event based on the ratio calculated for each known event. Classification device.

2. The classification device according to claim 1, comprising: a generation unit that generates a machine learning model for classifying observability data into normal observability data and abnormal observability data of known events, using the normal observability data and the abnormal observability data of known events as training data; the extraction unit inputs the observability data at the time of occurrence of the fault into the machine learning model to extract the abnormal observability data. Classification device.

3. The classification device according to claim 1, comprising: a storage unit that stores observability data with labels classifying the known events; the classification unit acquires the observability data with the same label from the storage unit, and classifies the acquired observability data and the abnormal observability data into two clusters. Classification device.

4. The classification device according to claim 1, wherein: when the classification unit classifies the occurred fault into a known event, it uses another classification method to confirm that the abnormal observability data is classified into the known event. Classification device.

5. The classification device according to claim 1, wherein: the observability data is data obtained by combining data acquired from the service to be monitored at predetermined times into one. Classification device.

6. A classification method for classifying faults occurring in a service to be monitored, comprising: a computer: extracts abnormal observability data from the observability data acquired from the service to be monitored at the time of occurrence of a fault; for each known event that has been dealt with in the past, classifies the observability data of the known event and the abnormal observability data into clusters, and calculates the ratio of the abnormal observability data included in the cluster of the known event; classifies the occurred fault into either an unknown event that has not been dealt with in the past or the known event based on the ratio calculated for each known event. Classification method.

7. A program for operating a computer as each part of the classification device according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method for clustering data

    JP2002278762A

  • Classification device, classification method, and program

    JP2020046883A

  • System analysis device and system analysis method

    WO2014125796A1

  • Log analysis system, log analysis method, and program recording medium

    WO2016132717A1

  • Log analysis system and method, and recording medium

    WO2017081865A1