Estimation device, estimation method, and program

The estimation device addresses the challenge of manual root cause analysis by calculating anomaly scores from metrics and traces to automatically identify failed services and their root causes, enhancing operational efficiency in APM.

JP7695588B2Active Publication Date: 2025-06-19NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023573690
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-01-12
Publication Date
2025-06-19
Estimated Expiration
2042-01-12

AI Technical Summary

Technical Problem

Existing APM tools can detect services with failures but require manual analysis of metrics, traces, and logs to investigate the root cause, and they do not effectively utilize metrics for root cause estimation.

Method used

An estimation device that calculates an anomaly score from metrics and traces to identify failed services and determine the root cause of failures, using a multivariate time series model to analyze processing time and call order information.

Benefits of technology

Enables automated estimation of failed services and their root causes, reducing operator workload and shortening mean time to recovery by effectively combining and analyzing metrics and trace data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007695588000001
    Figure 0007695588000001
  • Figure 0007695588000002
    Figure 0007695588000002
  • Figure 0007695588000003
    Figure 0007695588000003
Patent Text Reader

Abstract

An estimation device 1 comprises: an abnormality score calculation unit 13 that calculates an abnormality score indicating a degree of deviation from normal on the basis of metrics that quantify the activity of each of a plurality of services and traces that record time information and the call sequence of processing for each of the plurality of services; a faulty service estimation unit 15 that estimates, on the basis of the abnormality score, a service in which a fault has occurred; and a root cause estimation unit 16 that estimates a root cause on the basis of the abnormality score of the metrics of the service in which a fault has occurred.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an estimation device, an estimation method, and a program.

Background Art

[0002] In recent years, a microservices architecture that configures applications with a combination of fine-grained services has attracted attention. In the microservices architecture, while an improvement in development speed and ease of scaling can be expected, operation management tends to become complicated. To support operation management, Application Performance Management (APM) tools for collectively managing monitoring data and methods for automatically detecting failures have been proposed.

[0003] APM tools aggregate three types of monitoring data: metrics, traces, and logs, and support operator monitoring. In some APM tools, failure detection based on metrics is possible. In the technology of Non-Patent Document 1, failure detection and estimation of the failed service are performed based on the service response time included in the trace. In the technology of Non-Patent Document 2, failure detection is performed based on the service response time and service call order included in the trace. Further, in the technology of Non-Patent Document 3, failure detection and estimation of the failed service are performed based on the service response time, service call information, and service response code included in the metrics and traces.

Prior Art Documents

Non-Patent Documents

[0004]

Non-Patent Document 1

Non-Patent Document 2

Non-Patent Document 3

Summary of the Invention

Problems to be Solved by the Invention

[0005] In the prior art, APM tools can be used to detect services where failures occur. However, in order to investigate the root cause, operators had to analyze metrics, traces, and logs themselves. Metrics are required for estimating the root cause, but the prior art has not been able to utilize metrics. In Non-Patent Documents 1 and 2, since metrics are not used, it is impossible to estimate the root cause. Non-Patent Document 3 uses metrics and traces in combination, but it only stops at the utilization for estimating services where failures occur and has not considered the utilization for estimating the root cause.

[0006] The present invention has been made in view of the above, and an object thereof is to estimate a service where a failure occurs and the root cause of the failure.

Means for Solving the Problems

[0007] An estimation device according to an aspect of the present invention is an estimation device that estimates a service in which a failure has occurred in a monitoring target service configured by combining a plurality of services and estimates the root cause of the failure, wherein a metric that quantifies the activities of each of the plurality of services, a processing time information of each of the plurality of services, and an anomaly score calculation unit that calculates an anomaly score indicating the degree of deviation from the normal state from a trace recording the call order, a failure occurrence service estimation unit that estimates a service in which a failure has occurred based on the anomaly score, and a root cause estimation unit that estimates the root cause based on the anomaly score of the metric of the service in which the failure has occurred.

[0008] An estimation method according to an aspect of the present invention is an estimation method that estimates a service in which a failure has occurred in a monitoring target service configured by combining a plurality of services and estimates the root cause of the failure occurrence, wherein a computer calculates an anomaly score indicating the deviation from the normal state from a metric that quantifies the activities of each of the plurality of services, processing time information of each of the plurality of services, and a trace recording the call order, estimates a service in which a failure has occurred based on the anomaly score, and estimates the root cause based on the anomaly score of the metric of the service in which the failure has occurred.

Advantages of the Invention

[0009] According to the present invention, it is possible to estimate the service in which a failure has occurred and the root cause of the failure.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

[0011] Hereinafter, embodiments of the present invention will be described with reference to the drawings.

[0012] FIG. 1 is a functional block diagram showing an example of the configuration of the estimation device according to the present embodiment. The estimation device 1 shown in the figure is a device that estimates a failed service from the monitoring data collected from the service to be monitored 5 and estimates the root cause of the failure. The service to be monitored 5 is, for example, a service using a microservice architecture configured by combining a plurality of fine-grained services. The monitoring data is metrics and traces collected from the service to be monitored 5. Metrics are data that quantify the activities of each service. For example, metrics include CPU usage rate, memory usage amount, or communication volume. A trace is data that records the time information and call order of the processing of each service. The metrics collection device 31 collects metrics, and the trace collection device 32 collects traces. Commercial technologies of open source software can be used for the metrics collection device 31 and the trace collection device 32.

[0013] The estimation device 1 shown in FIG. 1 includes a processing unit 11, a data storage unit 12, an anomaly score calculation unit 13, an anomaly score storage unit 14, a failure occurrence service estimation unit 15, a root cause estimation unit 16, an aggregation unit 17, and a display unit 18.

[0014] The processing unit 11 stores the metrics in the data storage unit 12 for each time point, and converts the trace into data of the response time of each service for each time point and stores it in the data storage unit 12. An example of a metric is shown in FIG. 2, and an example of a trace is shown in FIG. 3. The trace shown in FIG. 3 is data in JSON format. The processing unit 11 converts the JSON format trace into the response time of each service.

[0015] Here, referring to FIG. 4, an example of the trace processing by the processing unit 11 will be described. A trace is data that records the processing in each service in the form of a span for a series of processes from a request to a response for the monitored service 5. A span is data that records the time information of the processing of each service and the call order. In the upper part of FIG. 4, the span is represented by a rectangle. The horizontal length of the rectangle indicates the response time. The vertical alignment of the rectangle indicates the call order. The frame including a plurality of spans in the upper part of FIG. 4 is one trace, which is a series of processes from a request to a response for the monitored service 5.

[0016] The processing unit 11 removes unnecessary spans with low importance in failure detection from the traces received from the trace collection device 32. Unnecessary spans are, for example, spans that record only the processing related to the request / response between services without recording the processing of the service itself. By removing unnecessary spans, the number of dimensions (number of columns in the table) can be reduced, and the curse of dimensionality in the learning of the multivariate time series model of the anomaly score calculation unit 13 described later can be avoided.

[0017] After removing the unnecessary spans, the processing unit 11 extracts the response time of each service from the trace and generates data in a table format showing the response time of each service for each time point of the trace. Each row of the table corresponds to one trace.

[0018] Since the multivariate time series model of the anomaly score calculation unit 13 does not allow missing values, the processing unit 11 performs interpolation processing such as linear interpolation on the missing parts in the table, and stores the processed trace in the data storage unit 12. The thick frame in the left table at the lower part of FIG. 4 indicates the part where the missing values are interpolated.

[0019] Note that the processing unit 11 may combine the metrics and the processed trace for each time and store the combined data in the data storage unit 12. For example, the processing unit 11 may combine the metrics according to the time of the trace, or may combine the trace and the metrics at a predetermined time interval.

[0020] The anomaly score calculation unit 13 calculates an anomaly score indicating the degree of deviation from the normal state for each metric of each service and the response time of each service from the metrics and traces stored in the data storage unit 12 using a multivariate time series model. As shown in FIG. 5, the anomaly score calculation unit 13 learns the normal behavior by inputting the normal metrics and traces into the multivariate time series model after preprocessing. Thereby, it is possible to grasp the correlation relationship across the data types (column direction) and time (row direction), and accurate learning becomes possible.

[0021] At the time of estimation, the anomaly score calculation unit 13 operates at the timing when the monitoring data is generated, and outputs an anomaly score for one time period for the input for a plurality of time periods. For example, if the timing when the monitoring data is generated is time t, the anomaly score calculation unit 13 inputs the metrics and traces for M time periods from time t - M to time t into the multivariate time series model, and outputs the anomaly score at time t. M is the window size in the multivariate time series model. The anomaly scores up to time t are accumulated in the anomaly score storage unit 14. FIG. 6 shows an example of the anomaly scores. Each row is the anomaly score for one time period. The larger the numerical value, the greater the deviation from the normal state. Note that the thick frame in the anomaly score is the part that the failure-occurring service estimation unit 15 and the root cause estimation unit 16 described later focus on.

[0022] The failure-occurrence service estimation unit 15 estimates the failure-occurrence service using the anomaly scores accumulated in the anomaly score storage unit 14. Specifically, the failure-occurrence service estimation unit 15 focuses on the response time, which is an indicator susceptible to the impact of failures, and searches for portions in the anomaly scores where the response time of each service exceeds the threshold to estimate the failure-occurrence service. In the example of Fig. 6, since the anomaly score at the thick-frame portion of the response time of Service A exceeds a predetermined threshold, the failure-occurrence service estimation unit 15 estimates that a failure has occurred in Service A during the time period when the anomaly score exceeds the threshold.

[0023] The root cause estimation unit 16 calculates the average anomaly score of each metric for the service and time period determined by the failure-occurrence service estimation unit 15 to have experienced a failure, and estimates the root cause based on the average anomaly score. For example, the root cause estimation unit 16 estimates as the root cause either the one with the average anomaly score exceeding the threshold or the one with the largest average anomaly score. In the example of Fig. 6, the average anomaly score of the metrics of Service A within the thick frame is calculated. Fig. 7 shows an example of the calculated average anomaly score. In the example of Fig. 7, since the anomaly score of the CPU usage rate is large, the root cause estimation unit 16 estimates that the large load on the CPU of the server or virtual server of Service A is the root cause.

[0024] The aggregation unit 17 aggregates the failure information obtained by the failure-occurrence service estimation unit 15 and the root cause estimation unit 16. The aggregation unit 17 may aggregate the metrics and traces related to the failure, or may aggregate the logs obtained from the monitored service 5.

[0025] The display unit 18 presents the failure information in a format that is easy for the operator to understand. Fig. 8 shows an example of the display. On the failure list screen, the failure occurrence time and the failure information are displayed so that the situation can be immediately confirmed. The failure information shows the service where the failure occurred and the root cause estimated by the estimation device 1. When the operator selects a failure for which they want to check the details, the failure details are displayed. In the failure details, the abnormality degree of the root cause, the transition of the measured values, the services whose anomaly scores increased during the same time period, and their metrics can be confirmed as related information. Regarding the related information as well, by checking "display", the transition can also be confirmed together.

[0026] Next, an example of the operation of the estimation device 1 of the present embodiment will be described.

[0027] Fig. 9 is a sequence diagram showing an example of the process flow from collecting and storing metrics and traces from the monitored service 5.

[0028] In steps S11 and S12, the metrics collection device 31 collects metrics from the monitored service 5 and transfers them to the processing unit 11.

[0029] In steps S13 and S14, the trace collection device 32 collects traces from the monitored service 5 and transfers them to the processing unit 11.

[0030] In step S15, the processing unit 11 processes the traces into a table format. The processing unit 11 may combine the metrics and the processed traces.

[0031] In steps S16 and S17, the processing unit 11 transfers the metrics and the processed traces to the data storage unit 12 and stores the data in the data storage unit 12.

[0032] Through the above processing, the monitoring data that can be used for the learning of the anomaly score calculation unit 13 or the calculation of the anomaly score is stored in the data storage unit 12. During learning, the anomaly score calculation unit 13 collectively captures the normal data and causes it to be learned in the multivariate time series model. During estimation, when the monitoring data is stored in the data storage unit 12, the monitoring data is transmitted to the anomaly score calculation unit 13, and the anomaly score is calculated.

[0033] FIG. 10 is a sequence diagram showing an example of the flow of processing for estimating a failure-occurring service and estimating a root cause.

[0034] When data is stored by the processing of FIG. 9, at step S21, the monitoring data necessary for the calculation of the anomaly score is transmitted from the data storage unit 12 to the anomaly score calculation unit 13.

[0035] At step S22, the anomaly score calculation unit 13 calculates the anomaly score.

[0036] At steps S23 and S24, the anomaly score calculation unit 13 transmits the calculated anomaly score to the anomaly score storage unit 14 and stores the anomaly score in the anomaly score storage unit 14.

[0037] At step S25, when the anomaly score is transmitted from the anomaly score storage unit 14 to the failure-occurring service estimation unit 15, at step S26, the failure-occurring service estimation unit 15 estimates the service in which a failure has occurred based on the anomaly score.

[0038] When estimating the service in which a failure has occurred, at step S27, the failure-occurring service information indicating the service in which a failure has occurred is transmitted from the failure-occurring service estimation unit 15 to the root cause estimation unit 16, and at the same time, the anomaly score is transmitted from the anomaly score storage unit 14 to the root cause estimation unit 16.

[0039] At step S28, the root cause estimation unit 16 estimates the root cause of the failure.

[0040] In step S29, the root cause is transmitted from the root cause estimation unit 16 to the aggregation unit 17, the service failure information is transmitted from the service failure estimation unit 15 to the aggregation unit 17, and the anomaly score is transmitted from the anomaly score storage unit 14 to the aggregation unit 17.

[0041] In step S30, the aggregation unit 17 aggregates the received information.

[0042] In step S31, the aggregated failure information is transmitted to the display unit 18, and in step S32, the display unit 18 displays the failure information.

[0043] Through the above processing, the service where the failure occurs and the root cause of the failure are estimated and presented to the operator.

[0044] As described above, the estimation device 1 of the present embodiment is an estimation device 1 that estimates the service where a failure occurs in the monitored service 5 configured by combining a plurality of services and estimates the root cause of the failure. The estimation device 1 includes an anomaly score calculation unit 13 that calculates an anomaly score indicating the degree of deviation from the normal state from the metrics obtained by quantifying the activities of each of the plurality of services, the time information of the processing of each of the plurality of services, and the trace recording the call order; a service failure estimation unit 15 that estimates the service in which a failure has occurred based on the anomaly score; and a root cause estimation unit 16 that estimates the root cause based on the anomaly score of the service metrics. The estimation device 1 can estimate the service where the failure occurs and its root cause by combining and analyzing the metrics and the trace, and present them to the operator. Thereby, the load on the operator can be reduced and the mean time to recovery can be shortened.

[0045] For the estimation device 1 described above, for example, a general-purpose computer system including a central processing unit (CPU) 901, a memory 902, a storage 903, a communication device 904, an input device 905, and an output device 906 as shown in FIG. 11 can be used. In this computer system, the estimation device 1 is realized by the CPU 901 executing a predetermined program loaded onto the memory 902. This program can be recorded on a computer-readable recording medium such as a magnetic disk, an optical disk, or a semiconductor memory, or can be distributed via a network.

Explanation of Signs

[0046] 1 Estimation device 11 Processing unit 12 Data storage unit 13 Abnormality score calculation unit 14 Abnormality score storage unit 15 Failure occurrence service estimation unit 16 Root cause estimation unit 17 Aggregation unit 18 Display unit

Claims

1. An estimation device that estimates a service in which a failure has occurred in a monitored service configured by combining a plurality of services and estimates the root cause of the failure, an anomaly score calculation unit that calculates an anomaly score indicating the degree of deviation from normal based on metrics quantifying the activities of each of the plurality of services, time information of the processing of each of the plurality of services, and a trace recording the call order; a failure-occurring service estimation unit that estimates a service in which a failure has occurred based on the anomaly score; and a root cause estimation unit that estimates the root cause based on the anomaly score of the metrics of the service in which the failure has occurred. Estimation device.

2. The estimation device according to claim 1, comprising a processing unit that converts the trace into the response time of each of the plurality of services for each time, wherein the anomaly score calculation unit calculates an anomaly score from the metrics and the trace for each time. Estimation device.

3. The estimation device according to claim 2, wherein the processing unit removes processes with low importance in failure detection from the trace, extracts the response time of each of the plurality of services, and interpolates the response time for services for which the response time cannot be extracted. Estimation device.

4. The estimation device according to any one of claims 1 to 3, wherein the anomaly score calculation unit inputs the normal-time metrics and trace into a multivariate time-series model to learn the normal-time behavior, and at the time of estimation, inputs the metrics and trace into the multivariate time-series model to calculate the anomaly score. Estimation device.

5. The estimation device according to any one of claims 1 to 4, The failure-occurrence service estimation unit estimates, as the service in which a failure has occurred, the service in which the abnormality score has exceeded a predetermined threshold value. The root cause estimation unit obtains the average of the abnormality scores of the metrics in the time period in which the failure has occurred in the service in which the failure has occurred, and estimates the root cause based on the obtained average value. Estimation device. Claim 6 An estimation method for estimating the service in which a failure has occurred in a monitored service configured by combining a plurality of services and estimating the root cause of the occurrence of the failure, wherein a computer calculates an abnormality score indicating the deviation from the normal state from the metrics quantifying the activities of each of the plurality of services, the time information of the processing of each of the plurality of services, and the trace recording the call order, estimates the service in which a failure has occurred based on the abnormality score, and estimates the root cause based on the abnormality score of the metrics of the service in which the failure has occurred. Estimation method. Claim 7 A program for operating a computer as each part of the estimation device according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Anomaly detection for microservices

    US20210058424A1

  • Automated root-cause analysis for distributed systems using tracing-data

    WO2020177854A1