Fault information estimation device, fault information estimation method, and fault information estimation program

The failure information estimation device efficiently estimates failure information in monitoring target systems by extracting relevant data, converting timestamps, and using encoders to handle asynchronous data, thereby reducing downtime and improving detection accuracy.

JP7694702B2Active Publication Date: 2025-06-18NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023564300
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-11-30
Publication Date
2025-06-18
Estimated Expiration
2041-11-30

AI Technical Summary

Technical Problem

Existing failure information estimation methods in monitoring target systems are inefficient and take too long to estimate failure information after a failure occurs, which can adversely impact users.

Method used

A failure information estimation device and method that includes a data acquisition unit, a pruning unit, a time series data encoder, a metadata encoder, and a failure information estimation unit. This system acquires time-series data and metadata from monitoring targets, extracts relevant data related to failures, converts timestamps to relative time, and estimates failure information using encoded data.

Benefits of technology

The proposed solution efficiently estimates failure information in a short time, reducing the adverse impact on users and improving the accuracy of failure detection by handling asynchronous time-series data uniformly and coping with dynamic changes in monitored metrics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007694702000001
    Figure 0007694702000001
  • Figure 0007694702000002
    Figure 0007694702000002
  • Figure 0007694702000003
    Figure 0007694702000003
Patent Text Reader

Abstract

This failure information estimation device comprises: a data acquisition unit that acquires data including time-series data and metadata of a plurality of metrics of a plurality of monitored objects in a monitored system; a pruning unit that extracts failure-related metric data from the data of the plurality of metrics; and a failure information estimation unit that estimates failure information about a monitored object experiencing a failure on the basis of the data extracted by the pruning unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a failure information estimation device, a failure information estimation method, and a failure information estimation program.

Background Art

[0002] In service maintenance operations, when a failure occurs in a service, data is acquired from a large number of monitoring targets (devices, applications, etc.) in the monitoring target system and analyzed to estimate failure information such as the status and cause of the failure of the monitoring target where the failure has occurred.

Prior Art Documents

Non-Patent Documents

[0003]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In a monitoring target system, in order to minimize the adverse impact on users after a failure occurs, it is desired that the estimation of failure information be performed efficiently and in a short time.

[0005] The present invention has been made paying attention to the above circumstances, and an object thereof is to provide a failure information estimation device, a failure information estimation method, and a failure information estimation program that can efficiently estimate failure information of a monitoring target where a failure has occurred in a short time.

Means for Solving the Problems

[0006] One aspect of the present invention is a failure information estimation device. The failure information estimation device includes a data acquisition unit that acquires data having time-series data and metadata of a plurality of metrics of a plurality of monitoring targets in a monitoring target system, a pruning unit that extracts data of metrics related to a failure from among the data of the plurality of metrics, and a failure information estimation unit that estimates failure information of a monitoring target in which a failure has occurred based on the data extracted by the pruning unit. And a time series data encoder that converts a timestamp representing the absolute time of the time series data extracted by the pruning unit into a timestamp representing the relative time within the time window It has.

[0007] One aspect of the present invention is a failure information estimation method. The failure information estimation method includes acquiring data having time-series data and metadata of a plurality of metrics of a plurality of monitoring targets in a monitoring target system, extracting data of metrics related to a failure from among the data of the plurality of metrics, and estimating failure information of a monitoring target in which a failure has occurred based on the data of the metrics related to the failure. And converting a timestamp representing the absolute time of the time series data of the metrics related to the failure into a timestamp representing the relative time within the time window It has.

[0008] One aspect of the present invention is a failure information estimation program. The failure information estimation program causes a computer to execute the functions of each component of the above failure information estimation device.

Advantages of the Invention

[0009] According to the present invention, there are provided a failure information estimation device, a failure information estimation method, and a failure information estimation program that can efficiently estimate failure information of a monitoring target in which a failure has occurred in a short time.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Embodiments for Carrying Out the Invention

[0011] Hereinafter, embodiments according to the present invention will be described with reference to the drawings.

[0012] [Configuration Example] (Functional Configuration) First, the functional configuration of the failure information estimation device according to the embodiment will be described. FIG. 1 is a block diagram showing an example of the functional configuration of a failure information estimation device 30 according to the embodiment. In FIG. 1, in addition to the failure information estimation device 30, a node 10 in the monitoring target system and a monitoring system 20 are shown together. Although there are a large number of nodes 10 in the monitoring target system, only one representative node 10 is shown in FIG. 1 for convenience.

[0013] As shown in FIG. 1, each node 10 has an application 11, a monitoring agent 12, and a data recording unit 13. The monitoring agent 12 collects time-series data and metadata of monitoring items related to the application 11 arranged in the same node 10, and records this in the data recording unit 13. The monitoring agent 12 also transmits the time-series data and metadata recorded in the data recording unit 13 to the monitoring system 20 by polling / telemetry.

[0014] The monitoring system 20 collects data of metrics of each monitoring target from a plurality of nodes 10 within the monitoring target system. Hereinafter, the data of the metrics is also referred to as metrics data for convenience.

[0015] The fault information estimation device 30 is a device that acquires a plurality of metrics data of a plurality of monitoring targets from the monitoring system 20, estimates fault information, and outputs a fault report.

[0016] The fault information estimation device 30 has a data acquisition unit 31, a pruning unit 33, a time-series data encoder 34, a metadata encoder 35, a fault information estimation unit 36, and a fault report output unit 37.

[0017] The data acquisition unit 31 acquires data of a plurality of metrics of a plurality of monitoring targets from the monitoring system 20. Each metrics data has time-series data and metadata. Each time-series data is composed of a set of a timestamp and other data values at each time. Each metadata is composed of text information such as a name given to the metric, a variable name, and a container name.

[0018] The pruning unit 33 extracts (prunes) only the metric data related to the failure from among the plurality of metric data acquired by the data acquisition unit 31. For example, the pruning unit 33 extracts several tens of metric data from several thousands of metric data. Thereby, the metric data used for estimating the failure information is reduced. The metric data related to the failure is time-series data with abnormal fluctuations during the time window and the corresponding metadata. The extraction of the metric data is performed, for example, by calculating an anomaly score using a one-dimensional time-series anomaly detection model for the time-series data. For one-dimensional time-series anomaly detection, methods such as Spectral Residual (SR method) and Fourier transform-based anomaly detection methods can be used. The pruning unit 33 supplies the extracted metric data to the time-series data encoder 34 and the metadata encoder 35.

[0019] The time-series data encoder 34 encodes the time stamp and data value of the time-series data simultaneously. The encoding includes the conversion of the time stamp of the time-series data. The conversion of the time stamp converts the time stamp representing the absolute time into a time stamp representing the relative time within the time window. Further, for each metric, the time-series data encoder 34 calculates a vector representation from the time stamp representing the relative time and other data values and aggregates them. Thereby, asynchronous time-series data can be handled uniformly. The time-series data encoder 34 supplies the encoding result to the metadata encoder 35.

[0020] The metadata encoder 35 learns the time-series data supplied from the time-series data encoder 34 and the metadata supplied from the pruning unit 33 for each metric simultaneously. Thereby, the meaning of the time-series data can be grasped from the text information of the metadata. Also, the relationship between time-series data can be grasped. The metadata encoder 35 supplies the encoding result to the failure information estimation unit 36.

[0021] Based on the encoding results of the time series data encoder 34 and the metadata encoder 35, the failure information estimation unit 36 estimates failure information such as the status and cause of the failure of the monitoring target where a failure has occurred. The failure information estimation unit 36 also creates a failure report based on the estimation result and supplies this to the failure report output unit 37.

[0022] The failure report output unit 37 receives the failure report from the failure information estimation unit 36 and outputs this.

[0023] (Hardware Configuration) Next, the hardware configuration of the failure information estimation device 30 will be described. The failure information estimation device 30 is configured by a computer. For example, the failure information estimation device 30 is configured by a personal computer, a server computer, or the like.

[0024] FIG. 2 is a block diagram showing an example of the hardware configuration of the failure information estimation device 30 according to the embodiment. As shown in FIG. 2, the failure information estimation device 30 includes an input device 41, a CPU 42, a storage device 45, and an output device 48. The failure information estimation device 30 may further include other peripheral devices in addition to these.

[0025] The input device 41, the CPU 42, the storage device 45, and the output device 48 are electrically connected to each other via a bus 49 and exchange data and instructions via the bus 49.

[0026] The input device 41 is a device that receives data from the monitoring system 20. For example, the input device 41 is configured by a receiving device or the like. The input device 41 is not limited to this and may be configured by any other input device.

[0027] The output device 48 is a device that outputs a failure report. For example, the output device 48 is configured by a display, a transmission device, or the like. The output device 48 is not limited to this and may be configured by any other output device.

[0028] The storage device 45 stores programs and data necessary for the processes executed by the CPU 42. The CPU 42 performs various processes by reading and executing the necessary programs and data from the storage device 45.

[0029] The storage device 45 includes a main storage device 46 and an auxiliary storage device 47. The main storage device 46 and the auxiliary storage device 47 exchange programs and data with each other.

[0030] The main storage device 46 stores programs and data that are temporarily necessary for the processing of the CPU 42. For example, the main storage device 46 is composed of a volatile memory such as a RAM (Random Access Memory).

[0031] The auxiliary storage device 47 stores programs and data supplied via external devices or a network, and provides the programs and data that are temporarily necessary for the processing of the CPU 42 to the main storage device 46. For example, the auxiliary storage device 47 is composed of a non-volatile memory such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive).

[0032] The CPU 42 is a processor and is hardware that processes data and instructions. The CPU 42 includes a control device 43 and an arithmetic device 44.

[0033] The control device 43 controls the input device 41, the arithmetic device 44, the storage device 45, and the output device 48.

[0034] The arithmetic device 44 reads programs and data from the main storage device 46, executes the programs to process the data, and provides the processed data to the main storage device 46.

[0035] In such a hardware configuration, the input device 41 constitutes the data acquisition unit 31. The CPU 42 and the storage device 45 constitute the pruning unit 33, the time-series data encoder 34, the metadata encoder 35, and the failure information estimation unit 36. The output device 48 constitutes the failure report output unit 37.

[0036] For example, the CPU 42 reads a program that executes the functions of the pruning unit 33, the time-series data encoder 34, the metadata encoder 35, and the failure information estimation unit 36 from the auxiliary storage device 47 into the main storage device 46, and executes the read program, thereby performing the operations of the pruning unit 33, the time-series data encoder 34, the metadata encoder 35, and the failure information estimation unit 36.

[0037] [Operation Example] (Process of Estimating Failure Information) Next, with reference to FIG. 3, the flow of the process of estimating failure information executed by the failure information estimation device 30 will be described. FIG. 3 is a diagram schematically showing the flow of the process of estimating failure information executed by the failure information estimation device 30.

[0038] In the input layer, the data acquisition unit 31 acquires a plurality of metric data. Each metric data has time-series data and metadata.

[0039] In the pruning layer, the pruning unit 33 applies one-dimensional time-series anomaly detection to the time-series data (td1) by Spectral Residual (p1) to calculate an anomaly score (td2).

[0040] FIG. 4 is a diagram schematically showing an overview of an example in which an anomaly is detected by one-dimensional time-series anomaly detection. The left side of FIG. 4 shows the time-series data which is the input data. The time-series data is a set of timestamps and data values at each time. The center of FIG. 4 shows a graph of the time-series data obtained for this input data. This graph includes a point a1 having a particularly high value compared to other points due to the occurrence of a failure. The right side of FIG. 4 shows the anomaly score obtained by processing the graph in the center of FIG. 4. This anomaly score includes an anomaly point a2 having a particularly high value while other points have substantially zero values due to the occurrence of a failure.

[0041] Next, the pruning unit 33 extracts time-series data (td3), anomaly scores (td4), and metadata (md2) related to the failure based on the anomaly score (td2) by Pruning (p3). Pruning (p3) compares the anomaly score with a predetermined threshold value to determine the presence or absence of outliers, and extracts the anomaly score (td4) including outliers, the corresponding time-series data (td3), and metadata (md2).

[0042] Next, in the encoding layer shown in FIG. 3, the time-series data encoder 34 encodes the time stamps and data values of the time-series data (td3, td4) simultaneously using Transformer (p4) or other models. In this encoding, the time stamp representing the absolute time is converted into a time stamp representing the relative time within the time window. Thereby, asynchronous time-series data can be uniformly handled. Further, for each metric, a vector representation is calculated from the time stamp representing the relative time and other data values and these are aggregated.

[0043] FIG. 5 is a diagram schematically showing an example of the conversion of the time stamp. The left side of FIG. 5 shows the time-series data before the conversion of the time stamp, and the right side of FIG. 5 shows the time-series data after the conversion of the time stamp.

[0044] The time stamp of the time-series data after the conversion is obtained by subtracting a certain time stamp (1628143990) from the time stamp of the time-series data before the conversion. For example, the time stamp after the conversion of the first row is 1628142121 - 1628143990 = -1866.

[0045] Furthermore, in the encoding layer shown in FIG. 3, the metadata encoder 35 simultaneously learns the time-series data and metadata (md2) using Transformer (p3) or other models. As a result, the encoding results (d1) by the time-series data encoder 34 and the metadata encoder 35 are obtained.

[0046] The series of processes described so far are performed for each metric. This series of processes is shown enclosed by a dashed rectangle in FIG. 3. If the number of metrics is M, this series of processes is repeated M times.

[0047] FIG. 6 is a diagram schematically showing an example of pruning of metric data by the pruning unit 33. The left side of FIG. 6 shows the metric data before pruning, and the right side of FIG. 6 shows the metric data after pruning. In the metric data before pruning on the left side of FIG. 6, the time series graph and the anomaly score obtained by one-dimensional time series anomaly detection described with reference to FIG. 4 are drawn together.

[0048] As can be seen from FIG. 6, the metric data after pruning is composed of time series data corresponding to the anomaly scores having abnormal values and metadata corresponding to the time series data. Further, the time series data after pruning is composed of the time series data before pruning and the anomaly scores.

[0049] Generally, the number of metrics to be monitored in the monitoring target system is extremely large. Further, the data of those metrics includes a large number of time series data not related to failures. This is a factor that increases the time required for the analysis work of estimating failure information.

[0050] In the embodiment, in the pruning layer, metric data related to failures is extracted from among the plurality of metric data acquired in the input layer. Thereby, the metric data used for the analysis work of estimating failure information is reduced. This contributes to shortening the time required for the analysis work of estimating failure information.

[0051] Since the time series data of the metrics to be monitored distributed in the monitoring target system is collected asynchronously, the timestamps do not match. Therefore, missing values occur when aggregating the time series data into a matrix form. In that case, preprocessing of the missing values, for example, interpolation of the missing values or correction of the data is required. This is a factor that increases the labor and cost required for the analysis work of estimating failure information.

[0052] The time series data encoder 34 converts a timestamp representing an absolute time into a timestamp representing a relative time, and calculates and aggregates vector representations from the timestamp representing the relative time and other data values. As a result, asynchronous time series data can be uniformly handled without processing missing values. Therefore, the relationship between asynchronous metrics can be captured.

[0053] The number and types of metrics to be monitored in the monitoring target system may change dynamically. Causes of metric changes include abnormal termination of an application, container scaling out, etc. Fig. 7 schematically shows an example of the state of abnormal termination of an application. Also, Fig. 8 schematically shows an example of the state of container scaling out. When the metrics change, the meaning of the time series data cannot be grasped without the metadata of the metrics.

[0054] The metadata encoder 35 learns time series data and metadata simultaneously. As a result, the meaning of the time series data can be captured from the text information of the metadata. Also, the relationship between time series data can be captured. Thereby, it is possible to cope with dynamic changes in the number and types of metrics. That is, even if the number and types of metrics change, the correspondence before and after the change can be grasped.

[0055] Next, in the encoding layer shown in Fig. 3, the failure information estimation unit 36 estimates failure information (d2) such as the situation and cause of the failure based on the encoding result (d1) using Transformer (p5) or other models. Subsequently, the failure information estimation unit 36 creates a failure report (d3) based on the failure information (d2) using Fault Report Decorder (p6) or other models.

[0056] Next, in the output layer, the failure report output unit 37 outputs the failure report.

[0057] FIG. 9 is a diagram schematically showing an example of input and output in the failure information estimation device 30 according to the embodiment. An example of the input metric data, that is, time-series data and metadata in FIG. 9 is shown, and an example of the output failure report on the right side of FIG. 9 is shown.

[0058] (Flowchart) Next, with reference to FIG. 10, the processing procedure and processing content of estimating failure information executed by the failure information estimation device 30 will be described. FIG. 10 is a flowchart showing the processing procedure and processing content of estimating failure information executed by the failure information estimation device 30 according to the embodiment.

[0059] In step S1, the data acquisition unit 31 acquires a plurality of metric data, that is, time-series data and metadata, from the monitoring system 20.

[0060] In step S2, the pruning unit 33 extracts only the time-series data related to the failure from among the plurality of metric data. The pruning unit 33 supplies the extracted time-series data and the corresponding metadata to the time-series data encoder 34 and the metadata encoder 35. Thereby, the metric data used for estimating the failure information is reduced.

[0061] In step S3, the time-series data encoder 34 encodes the time-series data and the time stamp at the same time. In this encoding, the time stamp representing the absolute time is converted into a time stamp representing the relative time within the time window. Further, for each metric, a vector representation is calculated from the time stamp representing the relative time and other data values and these are aggregated. Thereby, asynchronous time-series data can be uniformly handled.

[0062] In step S4, the metadata encoder 35 encodes the metadata. In this encoding, the time-series data and the metadata are learned simultaneously. Thereby, the meaning of the time-series data can be grasped from the text information of the metadata. Also, the relationship between the time-series data can be grasped.

[0063] In step S5, the failure information estimation unit 36 estimates failure information such as the situation and cause of a failure occurring in the failure monitoring system based on the encoding results of the time-series data encoder 34 and the metadata encoder 35. The failure information estimation unit 36 also creates a failure report based on the estimation result.

[0064] In step S6, the failure report output unit 37 receives the failure report from the failure information estimation unit 36 and outputs the failure report.

[0065] [Effect] In the embodiment, the pruning unit 33 extracts the metrics data related to the failure from among the plurality of metrics data acquired by the data acquisition unit 31. As a result, the metrics data used for the analysis operation of estimating the failure information is reduced, and the time required for the analysis operation of estimating the failure information is shortened.

[0066] Also, the time-series data encoder 34 converts a timestamp representing an absolute time into a timestamp representing a relative time, calculates a vector representation from the timestamp representing the relative time and other data values, and aggregates them. As a result, asynchronous time-series data can be uniformly handled without processing missing values, and the relationship between asynchronous metrics can be grasped.

[0067] Furthermore, the metadata encoder 35 simultaneously learns the time-series data and the metadata. As a result, the meaning of the time-series data can be grasped from the text information of the metadata, and the relationship between the time-series data can be grasped. Thereby, it is possible to cope with dynamic changes in the number and type of metrics.

[0068] As a result, the applicable range of the monitoring target of the monitoring system is expanded, leading to a reduction in development costs. Furthermore, the accuracy of failure detection using metrics is improved.

[0069] Note that the present invention is not limited to the above-described embodiments, and various modifications can be made without departing from the gist thereof at the implementation stage. Also, the respective embodiments may be implemented in appropriate combination, and in that case, the combined effects can be obtained. Furthermore, the above-described embodiments include various inventions, and various inventions can be extracted by combinations selected from a plurality of disclosed constituent elements. For example, even if some constituent elements are deleted from all the constituent elements shown in the embodiments, if the problem can be solved and the effects can be obtained, the configuration from which these constituent elements are deleted can be extracted as an invention.

Explanation of Reference Numerals

[0070] 10…Node 11…Application 12…Monitoring Agent 13…Data Recording Unit 20…Monitoring System 30…Fault Information Estimation Device 31…Data Acquisition Unit 33…Pruning Unit 34…Time-Series Data Encoder 35…Metadata Encoder 36…Fault Information Estimation Unit 37…Fault Report Output Unit 41…Input Device 42…CPU 43…Control Device 44…Arithmetic Unit 45…Storage Device 46…Main Storage Device 47…Auxiliary Storage Device 48…Output Device 49…Bus

Claims

1. A data acquisition unit that acquires data having time-series data and metadata of a plurality of metrics of a plurality of monitoring targets in a monitoring target system; A pruning unit that extracts data of metrics related to a failure from among the data of the plurality of metrics; A failure information estimation unit that estimates failure information of a monitoring target in which a failure has occurred based on the data extracted by the pruning unit; It has a time-series data encoder that converts a timestamp representing an absolute time of the time-series data extracted by the pruning unit into a timestamp representing a relative time within a time window. A failure information estimation device.

2. The pruning unit extracts data by calculating an anomaly score using a one-dimensional time-series anomaly detection model for the time-series data. The failure information estimation device according to claim 1.

3. The time-series data encoder further calculates a vector representation from a timestamp representing a relative time and other data values for each metric and aggregates them. The failure information estimation device according to claim 1 or claim 2.

4. It further has a metadata encoder that simultaneously learns the time-series data supplied from the time-series data encoder and the metadata supplied from the pruning unit for each metric. The failure information estimation device according to any one of claims 1 to 3.

5. Acquiring data having time-series data and metadata of a plurality of metrics of a plurality of monitoring targets in a monitoring target system; Extracting data of metrics related to a failure from among the data of the plurality of metrics; Estimating failure information of a monitoring target in which a failure has occurred based on the data of the metrics related to the failure; Converting a timestamp representing an absolute time of the time-series data of the metrics related to a failure into a timestamp representing a relative time within a time window. Failure information estimation method.

6. A failure information estimation program that causes a computer to execute the functions of each component of the failure information estimation device according to any one of Claims 1 to 4.

Citation Information

Patent Citations

  • Anomaly aggregation method

    JP2009076056A

  • Behavior clustering analysis and alerting system for computer applications

    US20150205692A1