Outlier detection device and method

The outlier detection system addresses the challenge of noise outliers in IT system data by using sliding windows and multiple sub-detectors to compare actual and predicted time series data, achieving effective outlier detection without supervised machine learning or user feedback.

JP7672926B2Active Publication Date: 2025-05-08HITACHI VANTARA LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2021143534
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-09-02
Publication Date
2025-05-08
Estimated Expiration
2041-09-02

AI Technical Summary

Technical Problem

Existing techniques for detecting outliers in IT system data require supervised machine learning and user feedback, which can lead to noise outliers being misidentified.

Method used

The proposed solution involves an outlier detection system that uses a window generator to create sliding windows for comparing actual and predicted time series data, employing multiple outlier sub-detectors to identify outliers without relying on supervised machine learning or user feedback.

Benefits of technology

This approach effectively reduces noise outliers by comparing actual and predicted data sets within specified windows, allowing for accurate outlier detection without the need for user feedback or supervised learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007672926000001
    Figure 0007672926000001
  • Figure 0007672926000002
    Figure 0007672926000002
  • Figure 0007672926000003
    Figure 0007672926000003
Patent Text Reader

Abstract

To realize an outlier detection in which noise has been reduced, without supervised machine learning which requires feedback data from a user.SOLUTION: An outlier detection apparatus creates first and second processing windows each having a designated window length and performs sliding alignment for sliding the second processing window relative to the first processing window by a designated sliding alignment length. The apparatus performs one or more types of outlier sub-detections. The outlier sub-detection includes comparing, by a method corresponding to the type of outlier sub-detection, an actual time-series dataset which is a data portion corresponding to the first processing window from among actual time-series data which is a time series of actual values, with a forecasted time-series dataset which is a data portion corresponding to the second processing window after the sliding alignment from among forecasted time-series data which is a time series of forecasted values. The apparatus determines whether an outlier candidate based on a result of the one or more types of outlier sub-detections is an outlier or not.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention generally relates to techniques for detecting outliers. [Background technology]

[0002] One method for automatically detecting outliers from data of an IT (Information Technology) system is to model the performance load of the IT system, predict the performance load from the model, and compare the predicted performance load with the actual performance load. If the actual performance load significantly deviates from the predicted performance load, an outlier that may be related to an abnormality in the IT system can be detected.

[0003] The outliers detected may be so-called noise outliers, that is, outliers that are not related to anomalies in the actual IT system.

[0004] Patent Document 1 discloses a technique for training an outlier classifier based on features extracted from a context-dependent time series pattern detector and implicit or explicit feedback data from a user. The trained outlier classifier can reduce noise outliers from the initially identified abnormal event candidates. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] US10,261,851 Summary of the Invention [Problem to be solved by the invention]

[0006] The technology disclosed in Patent Document 1 requires implicit or explicit feedback data from a user to train an outlier classifier using supervised machine learning. [Means for solving the problem]

[0007] The outlier detection device includes an outlier detector and an outlier determiner. The outlier detector includes a window generator and one or more types of outlier sub-detectors. The window generator generates a first processing window and a second processing window having a specified window length, and performs a sliding adjustment to slide the second processing window by a specified sliding adjustment length relative to the first processing window. Each of one or more types of outlier sub-detectors performs outlier sub-detection including comparing an actual time series data set, which is a data portion corresponding to the first processing window among actual time series data, which is a time series of actual values, with a predicted time series data set, which is a data portion corresponding to the second processing window after the sliding adjustment among predicted time series data, which is a time series of predicted values, in a manner corresponding to the type of the outlier sub-detector. The outlier determiner determines whether an outlier candidate based on a result of outlier sub-detection by the one or more types of outlier sub-detectors is an outlier. Effect of the Invention

[0008] According to the present invention, it is possible to realize noise-reduced outlier detection without supervised machine learning that requires feedback data from users. [Brief description of the drawings]

[0009] [Figure 1] 1 is a diagram illustrating an example of a functional configuration of a noise reduction and outlier detection device according to an embodiment of the present invention. [Figure 2A] 13 is a diagram showing an example of the configuration of an actual time-series data table in a time-series DB; FIG. [Figure 2B] FIG. 13 is a diagram illustrating an example of a configuration of a predicted time-series data table in a time-series DB. [Figure 3A] 13 is a diagram illustrating an example of the configuration of a parameter table in a parameter / threshold DB. FIG. [Figure 3B] FIG. 13 is a diagram illustrating an example of the configuration of a threshold table in a parameter / threshold DB. [Figure 4] 13 is a flowchart showing an example of the flow of a spiking load threshold calculation process. [Diagram 5] 13 is a flowchart showing an example of the flow of an outlier detection process. [Figure 6] 6 is a flowchart showing an example of the process of step S11002 in FIG. 5. [Figure 7] 6 is a flowchart showing an example of the process of step S11003 in FIG. 5. [Figure 8] 6 is a flowchart showing an example of S11004 in FIG. 5. [Figure 9] 6 is a flowchart showing an example of the process of step S11005 in FIG. 5. [Figure 10A] FIG. 13 is a diagram illustrating an example of a configuration of a window outlier table in a log DB. [Figure 10B] FIG. 13 is a diagram illustrating an example of the configuration of an outlier determination table in a log DB. [Figure 10C] FIG. 13 is a diagram illustrating an example of a configuration of a threshold table in a log DB. [Figure 11A] 11 is a part of a flowchart showing an example of the flow of an outlier determination process. [Figure 11B] 11 is the remainder of the flowchart illustrating an example of the flow of the outlier determination process. [Figure 12] FIG. 13 is a diagram showing an example of an outlier detection result screen. [Figure 13] FIG. 2 is a diagram illustrating an example of a hardware configuration of a noise reduction outlier detection device. [Figure 14] FIG. 11 is an explanatory diagram of an example of the significance of sliding adjustment. [Figure 15] FIG. 13 is an explanatory diagram of an example of the significance of point-based expected spike detection. [Figure 16] FIG. 1 illustrates an example of the significance of distribution-based expected spike detection. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] In the following description, an "interface unit" may refer to one or more interface devices. The one or more interface devices may be at least one of the following: One or more I / O (Input / Output) interface devices. The I / O (Input / Output) interface devices are interface devices to at least one of the I / O devices and a remote display computer. The I / O interface device to the display computer may be a communications interface device. The at least one I / O device may be a user interface device, e.g., either an input device such as a keyboard and a pointing device, or an output device such as a display device. One or more communication interface devices. The one or more communication interface devices may be one or more homogeneous communication interface devices (e.g., one or more NICs (Network Interface Cards)) or two or more heterogeneous communication interface devices (e.g., a NIC and an HBA (Host Bus Adapter)).

[0011] In the following description, a "memory" refers to one or more memory devices, which are an example of one or more storage devices, and may typically be a primary storage device. At least one memory device in the memory may be a volatile memory device or a non-volatile memory device.

[0012] In the following description, an "auxiliary storage device" may be one or more auxiliary storage devices, which are an example of one or more storage devices. The auxiliary storage device may typically be a non-volatile storage device (e.g., an auxiliary storage device), and specifically may be, for example, a hard disk drive (HDD), a solid state drive (SSD), a non-volatile memory express (NVMe) drive, or a storage class memory (SCM).

[0013] In the following description, the "storage device" may be at least one of memory and auxiliary storage device.

[0014] Furthermore, in the following description, a "processor" may be one or more processor devices. The at least one processor device may typically be a microprocessor device such as a CPU (Central Processing Unit), but may also be other types of processor devices such as a GPU (Graphics Processing Unit). The at least one processor device may be a single-core or multi-core. The at least one processor device may be a processor core. The at least one processor device may also be a broader processor device such as a circuit that is a collection of gate arrays written in a hardware description language that performs part or all of the processing (e.g., an FPGA (Field-Programmable Gate Array), a CPLD (Complex Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit)).

[0015] In the following description, information that produces an output for an input may be described using expressions such as "xxxDB" or "xxx table" ("DB" is an abbreviation for database), but the information may be data of any structure (for example, structured data or unstructured data) or a learning model such as a neural network that generates an output for an input. Therefore, "xxxDB" or "xxx table" may be referred to as "xxx information." In the following description, the configuration of each DB or table is an example, and one DB or one table may be divided into two or more DBs or two or more tables, or all or a part of two or more DBs or two or more tables may be one DB or one table.

[0016] In the following description, functions may be described using the expression "yyy device", but the functions may be realized by one or more computer programs being executed by a processor, or by one or more hardware circuits (e.g., FPGA or ASIC), or by a combination thereof. When a function is realized by a program being executed by a processor, the function may be at least a part of the processor, since the specified processing is performed using a storage device and / or an interface device, etc., as appropriate. Processing described with a function as the subject may be processing performed by a processor or a device having the processor. A program may be installed from a program source. The program source may be, for example, a program distribution computer or a computer-readable recording medium (e.g., a non-transitory recording medium). The description of each function is an example, and multiple functions may be combined into one function, or one function may be divided into multiple functions.

[0017] Hereinafter, embodiments will be described with reference to the drawings. Note that the embodiments described below do not limit the invention described in the claims. Furthermore, the various components and their combinations described in the embodiments are not necessarily essential to the present invention.

[0018] In the description of the embodiment, an "outlier" may refer to a sufficient difference between two types of data compared to each other, where one type of data (predicted time series data, described below) may represent an expected state (e.g., a normal state) and the other type of data (actual time series data, described below) may represent a current state.

[0019] A "noise outlier" may be a sufficient difference between two types of data being compared to each other, where one type of data represents an expected normal state, while the other type of data represents a current state that arises due to expected variations in the normal state that cannot be accurately represented in the data representing the normal state, and should not be taken as a problem.

[0020] "Real time series data" refers to data that is generated from an IT system (e.g., physical or logical calculator The actual time-series data may be a type of measurement data that indicates the current state of a monitored object such as a monitoring system (such as a monitoring target). In this embodiment, the actual time-series data is a time-series of actual measured values ​​(one example of actual values) of a performance load, but the actual measured values ​​in the time-series may be actual measured values ​​of a data item other than the performance load (for example, temperature or humidity).

[0021] The "predicted time series data" may be a type of measurement data that represents a predicted state (e.g., a normal state). In this embodiment, the predicted time series data is a time series of predicted values ​​of the performance load. The predicted values ​​in the time series may be predicted values ​​of a data item other than the performance load, similar to the actual measured values.

[0022] An "expected spike" may be a period of time in the forecasted time series data where the performance load values ​​are particularly high.

[0023] "Distance" may refer to a quantifiable measure of the difference between actual and predicted time series data.

[0024] "Direction" may refer to a measure for assessing whether the actual time series data is greater or less in value than the predicted time series data.

[0025] The term "processing window" refers to any period of time within the time series data for comparing actual time series data with predicted time series data and outputting outlier results. The length of the processing window may be, for example, a length of time.

[0026] The "time series data set" may be data within a range of time series data that corresponds to the processing window.

[0027] FIG. 1 illustrates an example of the functional configuration of a noise reduction and outlier detection device according to an embodiment.

[0028] The noise-reduced outlier detection device 100 is a device that performs outlier detection with reduced noise. The noise-reduced outlier detection device 100 may be a physical computer system (one or more physical computers) having the hardware configuration illustrated in Fig. 13, or may be a logical computer system (e.g., a cloud computing service system) based on a physical computer system (e.g., a cloud platform).

[0029] The noise-reducing outlier detection device 100 acquires actual time series data and predicted time series data stored in the time series DB 200, and parameters and thresholds stored in the parameter / threshold DB 300, compares the actual time series data with the predicted time series data to detect outliers, and visualizes the output results including the outliers on the display 400.

[0030] Actual time series data and predicted time series data are stored in the time series DB 200. Details will be described later with reference to Figures 2A and 2B.

[0031] The parameter / threshold DB 300 stores a parameter table and a threshold table that are externally defined by a user of the noise reduction outlier detection device 100. Details will be described later with reference to Figures 3A and 3B.

[0032] The display 400 is an output device for visualizing the results obtained by the noise reduction outlier detection apparatus 100 .

[0033] The noise reduction outlier detection device 100 includes an outlier detector 110, a spiking load threshold calculator 120, a log DB 130, and an outlier determiner 140. The outlier detector 110 includes a window generator 111, a predicted spike detector 112, a direction calculator 113, and a distance calculator 114.

[0034] In the noise-reducing outlier detection device 100, first, the outlier detector 110 processes the acquired actual time series data and predicted time series data. Specifically, for example, the outlier detector 110 divides the actual time series data and the predicted time series data into a plurality of processing windows (a plurality of time series data sets) by the window generator 111, and calculates the possibility of an outlier in each actual time series data set by the three types of outlier sub-detectors 112 to 114. The results obtained by this process are stored in the log DB 130. In further detail, the outlier detector 110 will be described later with reference to Figs. 5 to 9, and the log DB 130 will be described later with reference to Fig. 10.

[0035] The output obtained from the outlier detector 110 and stored in the log DB 130 is processed by the outlier determiner 140. That is, the outlier determiner 140 determines a final outlier based on the results of the outlier sub-detectors 112 to 114. A log message is generated by the outlier determiner 140 as necessary. The final outlier and the log message are stored in the log DB 130 and then visualized on the display 400. The outlier determiner 140 will be described in further detail later with reference to FIG. 11, and an example of the configuration of a screen displayed on the display 400 will be described later with reference to FIG. 12.

[0036] The predicted time series data is further processed by a spiking load threshold calculator 120 which calculates a threshold for expected spikes, the results of which are stored in a log DB 130. Further details are provided below with reference to FIG.

[0037] The noise-reduced outlier detection device 100 can realize noise-reduced outlier detection without supervised machine learning that requires feedback data from a user.

[0038] The time series DB 200 stores an actual time series data table 201, an example of which is shown in FIG. 2A, and a predicted time series data table 202, an example of which is shown in FIG. 2B.

[0039] As illustrated in Fig. 2A, the actual time series data table 201 stores a time series of actual performance loads (actual measured values ​​of performance loads), that is, actual time series data. The actual time series data table 201 includes columns such as date / time D20101 and performance load D20102. Date / time D20101 stores the actual date / time (e.g., a timestamp indicating the date / time) when the performance load was measured. The unit of "date / time" is year / month / date / hour / minute / second in this embodiment, but it may be a coarser or finer unit, or a different unit. Performance load D20102 stores the actual measured value of the performance load (e.g., a numerical value obtained from data indicating the performance metrics of the monitored IT system).

[0040] As illustrated in FIG. 2B, the predicted time series data table 202 stores a time series of predicted performance loads (predicted values ​​of performance loads), that is, predicted time series data. The predicted time series data table 202 includes columns such as date and time D20201 and predicted load D20202. The date and time D20201 stores a predicted date and time (e.g., a timestamp indicating the date and time) that is the date and time when the predicted performance load is predicted to be measured. The predicted load D20202 stores a value predicted as the performance load. The predicted time series data may be obtained by any method. For example, the predicted time series data may be data output from a machine learning model (e.g., a neural network) by inputting at least a part of time series data of actual time series data and past time series data (e.g., past actual time series data, or predicted time series data obtained in the past (predicted time series data whose predicted date and time is a past date and time)) into the machine learning model (or data after processing of the data). Alternatively, the predicted time series data may be data manually prepared based on past time series data or other data.

[0041] The parameter / threshold DB 300 stores a parameter table 301, an example of which is shown in FIG. 3A, and a threshold table 302, an example of which is shown in FIG. 3B.

[0042] 3A, the parameter table 301 is a table that stores defined parameters. The parameter table 301 includes columns such as an entry ID D30101, an actual window length D30102, a predicted window length D30103, a sliding adjustment length D30104, and a point / distribution-based classifier D30105. In one entry (row), the values ​​stored in the columns D30102 to D30105 are each a parameter.

[0043] The entry ID D30101 stores the ID of the entry.

[0044] The actual window length D30102 stores (a numerical value representing) the actual window length, which is the length of the actual window (the processing window for actual time series data). The actual window length may be expressed, for example, in time (for example, in units of minutes or seconds).

[0045] The prediction window length D30103 stores (a numerical value representing) the prediction window length, which is the length of the prediction window (the processing window of the predicted time series data). In one entry, the prediction window length may be the same as or different from the actual window length in the entry. When the actual window length and the prediction window length are different, a predetermined method may be used (for example, a method called Dynamic Time Warping may be used in distance calculation).

[0046] The sliding adjustment length D30104 stores a (numerical value representing) the sliding adjustment length, which is the length of the adjustment time difference (deviation) between the actual window and the predicted window. The sliding adjustment length may be expressed, for example, in time (for example, in units of minutes or seconds). Details of the sliding adjustment length are as follows. A sliding adjustment length of "0" means that there is no gap between the actual window and the predicted window. In other words, the start date and time of the actual window (for example, the window Date and Time The identifier) ​​and the start date and time of the prediction window are the same. A negative value for the sliding adjustment length means that the actual window prediction It means that the window slides relatively backward. For example, a sliding adjustment length of "-30" means that the start date and time of the predicted window is 30 time steps (e.g., 30 seconds) earlier than the start date and time of the actual window. A positive value for the sliding adjustment length means that the prediction window slides relatively far into the future with respect to the actual window. For example, a sliding adjustment length of “30” means that the start date and time of the prediction window is 30 time steps (e.g., 30 seconds) later than the start date and time of the actual window.

[0047] The point / distribution-based classifier D30105 stores a classifier (for example, a value such as "point" or "distribution") indicating whether point-based processing or distribution-based processing is used for outlier detection.

[0048] 3B, the threshold table 302 is a table that stores defined thresholds. The threshold table 302 includes columns such as an entry ID D30201, a distance threshold D30202, a direction threshold D30203, a spike threshold D30204, and an incidence threshold D30205.

[0049] Entry ID D30201 stores the ID of an entry. Entries (rows) of the threshold table 302 correspond one-to-one to entries of the parameter table 301. Therefore, for example, the parameter table entry storing the entry ID "1" and the threshold table entry storing the entry ID "1" are specified using the entry ID "1" as a key. For processing using various parameters corresponding to the entry ID "1", various thresholds corresponding to the entry ID "1" are used.

[0050] The distance threshold D30202 stores the distance threshold that is the threshold of the distance between the actual time series data set and the predicted time series data set. If distance calculation is not required to evaluate the outlier candidate, the distance threshold may be unnecessary (for example, undefined).

[0051] The direction threshold D30203 stores a direction threshold that is a threshold of the direction between the actual time series dataset and the predicted time series dataset. The "direction" depends on, for example, whether there are relatively more actual performance loads that are greater than the predicted performance loads between the actual time series dataset and the predicted time series dataset. The direction threshold may be any threshold value according to the direction calculation method used. If the direction is already obtained in the distance calculation or if the direction calculation is not necessary to evaluate the outlier candidate, the direction threshold may be unnecessary (for example, undefined (for example, a value of "0")).

[0052] The spike threshold D30204 stores a spike threshold, which is a threshold for expected spikes. Expected spikes are identified from the forecast time series dataset and are used to evaluate outlier candidates. If expected spikes are not required to evaluate outlier candidates, the spike threshold may be unnecessary (e.g., undefined (e.g., a value of "0")).

[0053] The incidence threshold D30205 stores an incidence threshold that is a threshold for the incidence of true values ​​(proportion of true values ​​among all Boolean values) obtained in point-based processing. If the processing corresponding to an entry is a distribution-based processing, the incidence threshold is not required for the entry (for example, it may be undefined (for example, a value of "None")).

[0054] An example of the processing performed in this embodiment will now be described.

[0055] 4 is a flowchart showing an example of the flow of the spiking load threshold calculation process. The spiking load threshold calculation process is a process performed by the spiking load threshold calculator 120.

[0056] In S12001, the spiking load threshold calculator 120 acquires predicted time series data from the time series DB 200.

[0057] In S12002, the spiking load threshold calculator 120 calculates the average value and the standard deviation for the entire predicted time series data acquired in S12001.

[0058] In S12003, the spiking load threshold calculator 120 calculates a spiking load threshold from the average value and standard deviation obtained in step S12002. An example of the spiking load threshold is a value obtained by adding k times the standard deviation to the average value.

[0059] S1 2 In step S12004, the spiking load threshold calculator 120 transmits the spiking load threshold calculated in step S12003 to the predicted spike detector 112 and stores the spiking load threshold in the log DB 130.

[0060] The spiking load threshold may be determined based on the predicted time series data in this manner. The predicted time series data is data based on past time series data and corresponds to expected actual time series data (expected value for actual time series data), so the spiking load threshold calculator 120 automatically calculates at what timing a spike is expected based on such predicted time series data. The spiking load threshold may be set manually.

[0061] 5 is a flowchart showing an example of the flow of the outlier detection process. The outlier detection process is a process performed by the outlier detector 110. Note that the actual time series data and the predicted time series data in this process may be acquired by the outlier detector 110 from, for example, the time series DB 200 at any timing. Furthermore, the actual time series data and the predicted time series data include data for the same period.

[0062] In S11001, the outlier detector 110 acquires all entry IDs defined in the parameter / threshold DB 300. Then, the following steps S11002 to S11005 are executed for each entry ID acquired in S11001. S11002 to S11005 will be explained using one entry ID as an example.

[0063] In S11002, the window generator 111 generates an actual window (an example of a first processing window) and a predicted window (an example of a second processing window).

[0064] In S11003, expected spike detector 112 detects expected spikes in the load.

[0065] In S11004, the direction calculator 113 calculates the direction.

[0066] In S11005, the distance calculator 114 calculates the distance.

[0067] FIG. 6 is a flowchart showing an example of the process of S11002 in FIG.

[0068] In S11101, the window generator 111 obtains parameters (actual window length, predicted window length, sliding adjustment length) corresponding to the entry ID from the parameter / threshold DB 300.

[0069] In S11102, the window generator 111 generates an actual window (for example, a rolling window). The length of the actual window is the actual window length obtained in S11101.

[0070] In S11103, the window generator 111 generates a prediction window (for example, a rolling window). The length of the prediction window is the prediction window length obtained in S11101.

[0071] In S11104, the window generator 111 slides the prediction window relatively to the actual window by the same length as the sliding adjustment length represented by the entry ID. In this way, the window generator 111 performs sliding adjustment, which is sliding the prediction window relatively to the actual window.

[0072] The multiple periods corresponding to the multiple actual time series data sets obtained using the actual window may be consecutive periods that are not overlapped with each other, or may overlap with each other partially. For example, when the actual window length is "30", data corresponding to 30 from the beginning of the actual time series data may be the first actual time series data set (first actual window), and data corresponding to the next 30 may be the next actual time series data set (next actual window). Of the actual time series data, data in a range corresponding to the actual window is the actual time series data set. Since multiple actual time series data sets are obtained using the actual window, it can be said that an actual window exists for each actual time series data set. The start date and time of each actual window is the start date and time of the actual time series data set corresponding to the actual window.

[0073] The multiple time periods corresponding to the multiple predicted time series data sets obtained using the prediction window may be consecutive non-overlapping periods, or may partially overlap each other. The data in the range corresponding to the prediction window among the predicted time series data is the predicted time series data set. Since the multiple predicted time series data sets are obtained using the prediction window, it can be said that a prediction window exists for each predicted time series data set. The start date and time of each prediction window is the start date and time of the predicted time series data set corresponding to the prediction window.

[0074] The actual window generated in S11102 and the prediction window generated in S11103 form a window set (a window pair). Therefore, the actual time series data set corresponding to the actual window and the prediction time series data set corresponding to the prediction window also form a pair, and a comparison is made between the data sets forming the pair.

[0075] An example of the significance of the sliding adjustment is as shown in FIG. 14. If a predetermined process (e.g., batch processing) in an IT system starts on time, a spike should occur at the date and time shown in the predicted time series data indicated by the dashed line. However, due to a cause such as the start of the predetermined process being earlier than the scheduled time, a spike occurs at a date and time earlier than the predicted date and time of the spike, as shown in the actual time series data indicated by the solid line. In one comparative example, a spike occurring at a date and time different from the predicted date and time of the spike can be detected as an outlier. This is because the difference between the actual performance load and the predicted performance load is large at that date and time. However, this outlier is a noise outlier. This is because the occurrence of the predicted spike is not abnormal, even though the occurrence date and time are different. In this embodiment, the above-mentioned sliding adjustment is performed, so that the predicted date and time of the spike and the actual date and time of the spike can be relatively overlapped, thereby avoiding detection of such a spike (noise outlier) as an outlier, that is, reducing noise.

[0076] FIG. 7 is a flowchart showing an example of the process of S11003 in FIG.

[0077] In S11201, the expected spike detector 112 obtains, from the parameter / threshold DB 300, a point / distribution-based classifier and a spike threshold corresponding to the entry ID.

[0078] In S11202, the expected spike detector 112 determines whether the spike threshold acquired in S11201 is a defined value. If the determination result is Yes, the process proceeds to S11203. If the determination result is No (for example, if the spike threshold value is an undefined value), the process ends.

[0079] S11203 to S11211 are executed for each window set (pair) of an actual window and a predicted window. In the explanation of S11203 to S11211, one window set is taken as an example. Note that for the window set, the sliding adjustment length of the actual window and the predicted window may be zero, or may be smaller (negative value) or larger (positive value) than zero. Therefore, for one window set, the date and time in the actual window (actual time series data set) and the date and time in the predicted window (predicted time series data set) "correspond" to each other when the dates and times are the same (for example, both are "2019-12-01 10:00:00") or are relatively shifted by the sliding adjustment length (for example, one date and time is "2019-12-01 10:00:00" and the other date and time is "2019-12-01 10:00:30"). Therefore, the correspondence between the actual performance load and the predicted performance load (in other words, the difference between the actual date and time of the actual performance load and the predicted date and time of the predicted performance load) also follows such a sliding adjustment length (time difference).

[0080] In S11203, expected spike detector 112 judges whether the point / distribution-based classifier acquired in S11201 is "point." If the result of this judgment is Yes, S11204 to S11206 are executed. If the result of this judgment is No (i.e., if the point / distribution-based classifier acquired in S11201 is "distribution"), S11207 to S11211 are executed.

[0081] In S11204, the expected spike detector 112 generates a Boolean sequence composed of Boolean true values ​​(i.e., a Boolean sequence in which all Boolean values ​​are the true value "1"). The Boolean sequence has a length equal to the actual window length, and is composed of a plurality of Boolean values ​​corresponding to a plurality of dates and times that constitute a period equal to the actual window length.

[0082] In S11205, for each of the multiple dates and times corresponding to the Boolean sequence generated in S11204, if the predicted performance load (predicted performance load in the predicted time series data set) of the predicted date and time corresponding to the date and time is a value greater than the spiking load threshold, the predicted spike detector 112 assigns a Boolean false value to the date and time. In other words, the Boolean value in the Boolean sequence that corresponds to the predicted performance load greater than the spiking load threshold is changed to a Boolean false value.

[0083] In S11206, the expected spike detector 112 adds the Boolean sequence after processing in S11205 to the log DB 130 (the point-based spike result list of the window outlier table 131).

[0084] In S11207 of FIG. 7, the expected spike detector 112 counts the number of actual performance loads in the actual time series data set that exceed the spiking load threshold.

[0085] In S11208, the predicted spike detector 112 counts the number of predicted performance loads in the predicted time series data set that exceed the spiking load threshold.

[0086] In S11209, the predicted spike detector 112 calculates a percentage by dividing the number of actual performance loads counted in S11207 by the number of predicted performance loads counted in S11208.

[0087] In S11210, the expected spike detector 112 returns a Boolean true value if the percentage calculated in S11209 is greater than the spike threshold obtained in S11201. On the other hand, if the percentage calculated in S11209 is less than or equal to the spike threshold obtained in S11201, the expected spike detector 112 returns a Boolean true value. false Return a value.

[0088] In S11211, the expected spike detector 112 adds the Boolean value (the value returned in S11210) to the log DB 130 (the distribution-based spike result list of the window outlier table 131).

[0089] In this manner, the expected spike detector 112 performs outlier sub-detection in terms of expected spike detection on a point-based or distribution-based basis. In the distribution-based approach, a data set (a data portion corresponding to a window) of the time series data is regarded as one group (a cluster). Specifically, when comparing the actual time series data set with the predicted time series data set, the performance load is not compared for each time point, but the number of corresponding actual performance loads (actual performance loads exceeding the spiking load threshold) is compared with the number of corresponding predicted performance loads (predicted performance loads exceeding the spiking load threshold). The spiking load threshold is a threshold calculated from the predicted time series data, and the predicted time series data is data representing a normal state to be compared with the actual time series data. For this reason, appropriate distribution-based expected spike detection is expected.

[0090] An example of the significance of point-based predicted spike detection is shown in FIG. 15. In general, the predicted performance load tends to be smaller than the spikes of the actual performance load because it is based on the average of the past measured performance loads. Therefore, even if the difference between the actual performance load and the predicted performance load is large enough to be detected as a spike, the corresponding prediction If the performance load is greater than the spiking load threshold, then the spike is a planned spike and is therefore a noise outlier. Point-based expected spike detection according to S11204-S11206 can reduce the likelihood of detecting such a noise outlier as an outlier.

[0091] An example of the significance of distribution-based expected spike detection is shown in FIG. 16. There may be many dates and times where the difference between the actual measured performance load and the corresponding predicted performance load is large enough to be determined as a spike. However, if such a large difference occurs due to a known reason such as low accuracy of the predicted time series data set, the actual performance load belonging to such a difference is likely to be a noise outlier. According to distribution-based expected spike detection according to S11207 to S11211, it is possible to reduce the possibility of detecting many noise outliers related to such many differences as outliers.

[0092] FIG. 8 is a flowchart showing an example of the process of S11004 in FIG.

[0093] In S11301, the direction calculator 113 obtains, from the parameter / threshold DB 300, a point / distribution-based classifier and a direction threshold corresponding to the entry ID.

[0094] In S11302, the direction calculator 113 judges whether or not the direction threshold is a defined value. If the result of this judgment is Yes, the process proceeds to S11303. If the result of this judgment is No, the process ends.

[0095] S11303 to S11308 are executed for each window set of an actual window and a predicted window. In the explanation of S11303 to S11308, one window set is taken as an example.

[0096] In S11303, the direction calculator 113 judges whether the point / distribution-based classifier is "point." If the result of this judgment is Yes, S11304 to S11305 are executed. If the result of this judgment is No, S11306 to S11308 are executed.

[0097] In S11304, the direction calculator 113 generates a Boolean sequence composed of Boolean values. The Boolean sequence has a length of the actual window length and is composed of a plurality of Boolean values ​​corresponding to a plurality of dates and times constituting a period of the actual window length. For each of the plurality of dates and times, if the actual performance load is greater than the corresponding predicted performance load, the Boolean value corresponding to the date and time is a true value, and if the actual performance load is equal to or less than the corresponding predicted performance load, the Boolean value corresponding to the date and time is a false value.

[0098] In S11305, the direction calculator 113 adds the Boolean sequence generated in S11304 to the log DB 130 (the point-based direction result list of the window outlier table 131).

[0099] In S11306, the direction calculator 113 calculates the percentage of the number of dates and times when the actual performance load is greater than the predicted performance load, relative to the number of dates and times that make up the period of the processing window length.

[0100] In S11307, if the percentage calculated in S11306 is greater than the direction threshold value obtained in S11301, the direction calculator 113 returns a Boolean true value. On the other hand, if the percentage calculated in S11306 is equal to or less than the direction threshold value obtained in S11301, the direction calculator 113 returns a Boolean false value.

[0101] In S11308, the direction calculator 113 adds the Boolean value returned in S11307 to the log DB 130 (the distribution-based direction result list of the window outlier table 131).

[0102] In this manner, the direction calculator 113 performs outlier sub-detection in terms of the direction of the difference between the actual time series data set and the predicted time series data set (whether there is a general tendency for the actual time series data set to be larger than the predicted time series data set) on a point-based or distribution-based basis.

[0103] FIG. 9 is a flowchart showing an example of the process of S11005 in FIG.

[0104] In S11401, the distance calculator 114 obtains, from the parameter / threshold DB 300, a point / distribution-based classifier and a distance threshold corresponding to the entry ID.

[0105] In S11402, the distance calculator 114 determines whether the distance threshold acquired in S11401 is a defined value. If the result of this determination is Yes, the process proceeds to S11403. If the result of this determination is No, the process ends.

[0106] S11403 to S11410 are executed for each window set of an actual window and a predicted window. In the explanation of S11403 to S11410, one window set is taken as an example.

[0107] In S11403, the distance calculator 114 judges whether the point / distribution-based classifier is "point." If the result of this judgment is Yes, S11404 to S11406 are executed. If the result of this judgment is No, S11407 to S11410 are executed.

[0108] In S11404, the distance calculator 114 calculates the distance (for example, the difference between the feature amounts) between the actual performance load and the predicted performance load for each date and time.

[0109] In S11405, if the distance calculated in S11404 for each date and time exceeds the distance threshold acquired in S11401, the distance calculator 114 determines a Boolean true value for that date and time. On the other hand, if the distance calculated in S11404 is equal to or less than the distance threshold acquired in S11401, the distance calculator 114 determines a Boolean false value for that date and time. In this manner, a Boolean sequence composed of multiple Boolean values ​​corresponding to multiple dates and times is generated.

[0110] In S11406, the distance calculator 114 adds the generated Boolean sequence to the log DB 130 (the point-based distance result list of the window outlier table 131).

[0111] In S11407, the distance calculator 114 converts the actual window (actual time series data set) and the predicted window (predicted time series data set) into summarized distributions using the same processing function. The distribution corresponding to the actual window is called the "actual distribution", and the distribution corresponding to the predicted window is called the "predicted distribution". Each of these distributions may be, for example, a histogram of the same bin size. The bin size (width of the bin) may be a range of performance loads, and the length of the bin may be the number of performance loads that belong to the range. Specifically, for example, the bin size is a fixed width (e.g., 10), and multiple bins are prepared to correspond to the range of performance loads (e.g., CPU utilization is between 0 and 100%, so 10 bins are required).

[0112] In S11408, the distance calculator 114 calculates the distance between the actual distribution and the predicted distribution.

[0113] In S11409, if the distance calculated in S11408 exceeds the distance threshold acquired in S11401, the distance calculator 114 returns a Boolean true value. On the other hand, if the distance calculated in S11408 is equal to or less than the distance threshold acquired in S11401, the distance calculator 114 returns a Boolean false value.

[0114] In S11410, the distance calculator 114 adds the Boolean value returned in S11409 to the log DB 130 (the distribution-based distance result list of the window outlier table 131).

[0115] In this manner, the distance calculator 114 performs outlier sub-detection in terms of the distance between the actual time series data set and the predicted time series data set on a point-by-point or distribution-by-distribution basis.

[0116] The various outlier sub-detections mentioned above vesselcan perform either point-based or distribution-based outlier detection, but need not be adapted to perform either one of them.

[0117] For a parameter set that includes the point / distribution-based classifier "point", the point-based outlier sub-detection is to detect whether each actual measurement in the actual time series data set is a candidate outlier based on each actual measurement in the actual time series data set and each predicted value in the predicted time series data set. If so, a Boolean truth value is output for the actual measurement as a candidate outlier.

[0118] The point-based outlier sub-detection makes it possible to determine whether each actual performance load is an outlier candidate. The point-based predicted spike detection (S11204 to S11206 in FIG. 7) has been described with reference to FIG. 15. The point-based direction calculation (S11304 to S11305 in FIG. 8) makes it possible to exclude from outlier candidates an actual performance load that is equal to or smaller than the predicted performance load. The point-based distance calculation (S11404 to S11406 in FIG. 9) makes it possible to exclude from outlier candidates an actual performance load whose distance from the predicted performance load is equal to or smaller than a distance threshold.

[0119] According to the distribution-based outlier sub-detection, it is possible to know whether there is an outlier candidate for the entire actual time series data set. The distribution-based predicted spike detection (S11207 to S11211 in FIG. 7) has been described with reference to FIG. 16. According to the distribution-based direction calculation (S11306 to S11308 in FIG. 8), it is possible to determine that there is no outlier candidate if the ratio of the actual performance load exceeding the predicted performance load is equal to or less than the direction threshold. According to the distribution-based distance calculation (S11407 to S11410 in FIG. 9), it is possible to determine that there is no outlier candidate for the actual time series data set corresponding to the actual distribution whose distance from the predicted distribution is equal to or less than the distance threshold.

[0120] The log DB 130 stores a window outlier table 131 illustrated in FIG. 10A, an outlier determination table 132 illustrated in FIG. 10B, and a threshold table 133 illustrated in FIG. 10C.

[0121] As shown in FIG. 10A, the window outlier table 131 has columns such as a window date and time identifier D13101, a point-based distance result list D13102, a point-based direction result list D13103, a point-based spike result list D13104, a distribution-based distance result list D13105, a distribution-based direction result list D13106, and a distribution-based spike result list D13107.

[0122] The window date and time identifier D13101 stores a window date and time identifier assigned to the actual window (for example, a value indicating the start date and time of a period equivalent to the actual window length).

[0123] The point-based distance result list D13102 stores a list of Boolean sequences output in a point-based distance calculation. The point-based direction result list D13103 stores a list of Boolean sequences output in a point-based direction calculation. The point-based spike result list D13104 stores a list of Boolean sequences output in a point-based expected spike detection. For each of these lists D13102-D13104, there is a Boolean sequence for each window date and time identifier (for each window set including an actual window identified from that window date and time identifier). For each window date and time identifier, the point-based Boolean sequence is made up of a plurality of Boolean values ​​corresponding to a plurality of date and time that make up a period of the length of the processing window corresponding to that window date and time identifier.

[0124] The distribution-based distance result list D13105 stores the Boolean values ​​output in the distribution-based distance calculation. The distribution-based direction result list D13106 stores the Boolean values ​​output in the distribution-based direction calculation. The distribution-based spike result list D13107 stores the Boolean values ​​output in the distribution-based expected spike detection. For each of these lists D13105-D13107, there is a Boolean sequence for each window date and time identifier (for each window set including the actual window identified from that window date and time identifier). For each window date and time identifier, the distribution-based Boolean sequence consists of one Boolean value output for the processing window corresponding to that window date and time identifier.

[0125] As shown in FIG. 10B, the outlier determination table 132 includes columns such as a window date and time identifier D13201, an outlier Boolean value D13202, a noise Boolean value D13203, a predicted spike Boolean value D13204, an adjustment Boolean value D13205, and a log message D13206.

[0126] The window date and time identifier D13201 stores the date and time identifier actually assigned to the window.

[0127] The outlier Boolean value D13202 stores a Boolean true value as a result value when the actual window is identified as an outlier (a Boolean false value if not).

[0128] The noise Boolean value D13203 stores a Boolean true value as a result value when the actual window is identified as a noise outlier (a Boolean false value otherwise).

[0129] The expected spike Boolean value D13204 stores a Boolean true value as a result value when the actual window is identified as a noise outlier based on the expected spike represented by the predicted time series data (a Boolean false value if not).

[0130] The adjustment Boolean value D13205 stores a Boolean true value as a result value when the actual window is evaluated based on parameters including a non-zero sliding adjustment length (a Boolean false value otherwise). In addition to or instead of a Boolean value, the adjustment Boolean value D13205 can also store information indicating the sliding adjustment length used and the direction of the adjustment (i.e., information including information on whether the actual window is early or late relative to the predicted window and information indicating the time difference between those windows).

[0131] The log message D13206 stores a text message describing some information discovered during the outlier detection process from data about the state of the IT system, for example whether a value is an outlier, a noise outlier, or not an outlier, and additional detailed information if necessary.

[0132] The threshold table 133 includes columns such as threshold information D13301 and value D13302, as shown in FIG. 10C, for example.

[0133] The threshold information D13301 stores a description of each type of additional threshold information (e.g., information for convenience or future reference) calculated in the noise-reduction outlier detection device 100. Examples of threshold information include a spiking load threshold, a point-based adjustment list, and a distribution-based adjustment list.

[0134] The value D13302 is the threshold information D 13301 The data values ​​assigned to the corresponding descriptions in the table are stored.

[0135] 11A and 11B are flowcharts showing an example of the flow of the outlier determination process. The outlier determination process is performed by the outlier determiner 140. The outlier determination process is performed by all of the outlier sub-detectors 112 of the outlier detector 110. ~114 and finally determining the outliers using the results of the processing in . The outlier determination process may include generating necessary log messages that can be output to the display 400.

[0136] In S14001, the outlier determiner 140 refers to the parameter / threshold DB 300 and evaluates all point-based entries (all entries including the point / distribution-based classifier “point”). If there is a point-based entry including a sliding adjustment length other than “0”, the outlier determiner 140 adds a Boolean true value (otherwise, a Boolean false value) to the point-based adjustment list of the threshold table 133 of the log DB 130. As an example, for a point-based entry including a sliding adjustment length “0”, a Boolean false value ([0]) is recorded in the point-based adjustment list as illustrated in FIG. 10C. Furthermore, if there is a point-based entry including a sliding adjustment length other than “0” in addition to the point-based entry including the sliding adjustment length “0”, the Boolean true value is added to the point-based adjustment list of the threshold table 133 (as a result, the list becomes [0,1]).

[0137] In S14002, the outlier determiner 140 refers to the parameter / threshold DB 300 and of The distribution base entries (all entries including the point / distribution-based classifier "distribution") are evaluated. If there is a distribution base entry including a sliding adjustment length other than "0", the outlier determiner 140 adds a Boolean true value (otherwise a Boolean false value) to the distribution base adjustment list of the threshold table 133 of the log DB 130. As an example, for a distribution base entry including a sliding adjustment length other than "0", therefore, as illustrated in FIG. 10C, a Boolean true value ([1]) is recorded in the distribution base adjustment list. Furthermore, if there is a distribution base entry including a sliding adjustment length "0" as a distribution base entry in addition to the distribution base entry including a sliding adjustment length other than "0", a Boolean false value is added to the distribution base adjustment list of the threshold table 133 (as a result, the list becomes [1,0]).

[0138] In S14003, the outlier determiner 140 acquires the window outlier table 131 from the log DB 130. S14004 to S14016 are executed for each window date and time identifier in the window outlier table 131. S14004 to S14006 and S14007 may be executed in parallel. Furthermore, S14004 to S14006 are executed for each point-based entry for which the Boolean value of the corresponding point-based adjustment list is set to "0" (false) (i.e., for each point-based entry including a sliding adjustment length of "0"). S14004 to S14006 will be explained using one window date and time identifier and one point-based entry (a point-based entry including a sliding adjustment length of "0") as an example. S14007 will be explained using one window date and time identifier as an example.

[0139] In S14004, the outlier determiner 140 outputs a single point-based Boolean sequence by calculating the AND relation of all point-based Boolean sequences (i.e., the result list of point-based distances, directions, and spikes) in the window outlier table 131. For example, if the Boolean values ​​in all point-based Boolean sequences for a given date and time are "1", the Boolean value in the single point-based Boolean sequence for that date and time will also be "1". On the other hand, if the Boolean values ​​in all point-based Boolean sequences for a given date and time are "0", or if the point-based Boolean sequences contain a mixture of Boolean values ​​of "1" and "0", the Boolean value in the single point-based Boolean sequence for that date and time will be "0".

[0140] In S14005, outlier determiner 140 calculates the incidence rate of Boolean true values ​​for the single Boolean sequence obtained in step S14004 (the ratio of Boolean true values ​​in the single Boolean sequence to the number of Boolean values ​​constituting the single Boolean sequence). For example, when the window length is "5" (when the number of dates and times (time points) belonging to one processing window is "5"), the Boolean sequence output in S14004 is composed of five Boolean values. When the Boolean sequence is [1, 0, 1, 0, 1], the incidence rate of Boolean true values ​​calculated in S14005 is 60%.

[0141] In S14006, the outlier determiner 140 returns a Boolean true value (otherwise a Boolean false value) if the incidence rate obtained in step S14005 is greater than the incidence rate threshold (incidence rate threshold corresponding to the entry ID of the point-based entry) of the parameter / threshold DB 300. For example, if the incidence rate of the Boolean true value calculated in S14005 is 60% and the incidence rate threshold is 70%, the incidence rate is smaller and therefore a Boolean false value is output.

[0142] In S14007, the outlier determiner 140 outputs a single distribution-based Boolean sequence by calculating the AND relationship of all distribution-based Boolean sequences (i.e., distribution-based distance, direction, and spike result lists) in the window outlier table 131 for distribution-based entries whose Boolean values ​​in the corresponding distribution-based adjustment list are false (distribution-based entries that include a sliding adjustment length of "0"). In distribution-based processing, the outliers for one processing window are value Since the result of the sub-detection is one Boolean value, the Boolean sequence output in S14007 consists of a single Boolean value.

[0143] In S14008, the outlier determiner 140 calculates an AND relationship between the point-based output, which is the output of the loop of S14004 to S14006, and the distribution-based output, which is the output of S14007, and finally returns a Boolean value of the outlier as a result. That is, in S14008, an AND relationship is calculated between a single Boolean value as the point-based output and a single Boolean value as the distribution-based output.

[0144] In S14009, the outlier determiner 140 determines whether the Boolean value of the final outlier is true. If the result of this determination is Yes, the process proceeds to S14010. If the result of this determination is No, the process proceeds to S14014. Also, if there is no false-value point-based or distribution-based target in the adjustment list of the threshold table 133, the process may proceed to S14010.

[0145] In S14010, the outlier determiner 140 determines whether or not any of the point-based or distribution-based adjustment lists in the threshold table 133 of the log DB 130 is a true value. If the result of this determination is Yes, the process proceeds to S14011. If the result of this determination is No, the process proceeds to S14013.

[0146] In S14011, the outlier determiner 140 calculates an AND relationship between all point-based incidence evaluation results corresponding to the point-based or distribution-based adjustment list having true values ​​(the list in the threshold table 133) and the distribution-based Boolean result (the Boolean sequence as the output of S14007), and returns the result as an output of the Boolean value of the outlier. Although the detailed AND relationship calculation is not described here, it may be the same as the calculation of S14004 to S14008 described so far. For example, the Boolean sequence as all point-based incidence evaluation results corresponding to the point-based or distribution-based adjustment list having true values ​​may be calculated in the same way as S14004 to S14006. The differences between S14008 and S14011 are as follows. That is, S14008 is processing for an entry with a sliding adjustment length of "0" (processing for the case where sliding adjustment is not performed), while S14011 is processing for a sliding adjustment length other than "0" (processing for the case where sliding adjustment is performed).

[0147] In S14012, the outlier determiner 140 determines whether the Boolean value of the outlier obtained in S14011 is true. If the result of this determination is Yes, the process proceeds to S14013. If the result of this determination is No, the process proceeds to S14015.

[0148] In S14013, the outlier determiner 140 determines whether OutsideThe outlier determiner 140 calculates the severity of the outlier and stores a log message and an outlier Boolean value in the log DB 130 (outlier determination table 132). For example, the outlier determiner 140 can quantify the difference between the actual performance load and the predicted performance load using a window date and time identifier of the currently considered processing window (e.g., a rolling window) and the actual time series data set and the predicted time series data set of the corresponding processing window. The outlier determiner 140 can then generate a log message based on this quantified information. Furthermore, the outlier determiner 140 may observe expected spikes present in a time period corresponding to the processing window and identify an actual time series data set classified as an outlier because the actually observed spiking load is sufficiently longer than the predicted expected spike.

[0149] In S14014, the processing window (currently considered time frame) that the outlier determiner 140 identified as a non-outlier without sliding adjustment is tested for noise outliers. For example, a noise outlier may be observed when a distance or direction-based outlier caused by a predicted spike is identified as a non-outlier. Then, the outlier determination vessel 140 may generate log messages providing information on how much larger / smaller the actual time series is compared to the predicted time series, the difference in length observed for expected spikes in the predicted and actual time series, and warning about noise outliers, where outlier determiner 140 may determine false as the adjustment Boolean value and true or false as the expected spike Boolean value depending on the results of this test.

[0150] In S14015, the outlier determiner 140 identifies a noise outlier (a non-outlier with sliding adjustment) for the processing window (the time frame currently being considered). In this case, the date and time identifier of such a processing window is identified as an outlier in S14009, and is subsequently identified as a non-outlier in S14012 taking the sliding adjustment into account. Therefore, it is known that the outlier identified in S14009 is a noise outlier. Furthermore, the outlier determiner 140 may test whether the non-outlier for this processing window has become a noise outlier due to an expected spike. Then, the outlier determiner 140 may generate a log message to warn, for example, about an actual spike that is earlier or later than the expected spike. Here, the outlier determiner 140 may determine a true value as the adjusted Boolean value and a true or false value as the expected spike Boolean value depending on the test result.

[0151] In S14016, the outlier determiner 140 stores the outlier Boolean value, the noise Boolean value, the expected spike Boolean value, the adjusted Boolean value, and the generated log message in the outlier determination table 132 of the log DB 130. The outlier Boolean value and the noise Boolean value are values ​​according to at least one of the results of S14009 and S14012. The expected spike Boolean value, the adjusted Boolean value, and the generated log message are values ​​resulting from S14014 or S14015.

[0152] In S14017, the outlier determiner 140 analyzes the actual outliers and the noise outliers (e.g., in the larger context of a period corresponding to several consecutive processing windows). This analysis is performed, for example, based on the outlier Boolean value, the noise Boolean value, the expected spike Boolean value, and the adjustment Boolean value in the log DB 130 (outlier determination table 132). For example, for an actual outlier (the performance load in the processing window corresponding to the outlier Boolean value "1" and the noise Boolean value "0" or "None"), the outlier determiner 140 may identify additional information such as the duration of the actual outlier. Also, for example, for a noise outlier (the performance load in the processing window corresponding to the noise Boolean value "1"), the outlier determiner 140 may identify the occurrence pattern of the expected spikes and how large the actual spikes are compared to the expected spikes, for example, based on the expected spike Boolean value and the adjustment Boolean value. The magnitude of the actual spike may be identified from the actual time series data based on the date and time identifier (and the magnitude of the sliding adjustment) corresponding to the noise outlier. The magnitude of the expected spike may be determined from the forecasted time series data based on the date and time identifier (and the magnitude of the sliding adjustment) corresponding to the noise outlier. In S14017, the outlier determiner 140 may generate a log message based on the analysis result and store the log message in the log DB 130.

[0153] FIG. 12 shows an example of the outlier detection result screen.

[0154] The outlier detection result screen 1200 is a GUI (Graphical User Interface) displayed on the display 400 by the noise reduction outlier detection device 100. The display contents of the outlier detection result screen 1200 may be updated periodically (e.g., frequently) by acquiring all log messages, outliers, and time-series information from the log DB 130 and the time-series DB 200, for example.

[0155] The outlier detection result screen 1200 has a graphical visualization area 401 and a log message output area 402 .

[0156] The graphical visualization area 401 displays, for example, a graph of the time series of the actual performance load and the predicted performance load based on the actual time series data and the predicted time series data of the time series DB 200. The graphical visualization area 401 may also display an outlier occurrence time zone (for example, a continuous range of date and time identifiers corresponding to the outlier Boolean value "1" and the noise Boolean value "0" or "None") identified based on the log DB 130 (for example, the outlier determination table 132).

[0157] Log message output Power E rear 402 In the display area 401, log text messages stored in the log DB 130 are displayed as explanatory alternative outputs for the display in the graphical visualization area 401.

[0158] The outlier detection result screen 1200 may be a UI other than a GUI. Furthermore, the display areas of the outlier detection result screen 1200 may not be limited to the graphical visualization area 401 and the log message output area 402, and these display areas may be separated into two or more areas or may be one display area, and each display area may be located at an arbitrary position.

[0159] 11A and 11B, a log message may be created both when an outlier is detected and when a non-outlier (e.g., a noise outlier) is detected. As a result, when a log message is displayed as illustrated in FIG. 12, an operator can distinguish, for example, whether a normal actual performance load at a certain date and time is normal because it was detected as a noise outlier, or whether it was not detected as a noise outlier but was originally normal. The log message may include a message indicating what steps were taken (which steps in the above-mentioned flowchart were taken) and what kind of outlier detection result was obtained.

[0160] FIG. 13 shows an example of the hardware configuration of the noise reduction outlier detection device 100.

[0161] Noise reduction outlier detection device 100 is, for example, a general computer, and includes memory 502, auxiliary storage device 503, communication interface 504, media interface 505, input / output device 506, and CPU 501 connected thereto. Each of interfaces 504 to 506 is an example of an interface device. CPU 501 is an example of a processor.

[0162] The communication interface 504 is an interface device for communicating with other devices (eg, an external database that stores data to be analyzed) via a network 508 .

[0163] The memory 502 is, for example, a RAM (Random Access Memory), and 5 The auxiliary storage device 503 is, for example, an HDD or SSD, and stores the programs executed by the CPU 501 and the data used by the CPU 501. The external storage medium 507 is detachably attached to the media interface 505, and the media interface 505 mediates the input and output of data to and from the external storage medium 507.

[0164] The console 500 is connected to an input / output device 506, which inputs and outputs information to and from the console 500. The console 500 includes the display 400, for example.

[0165] The CPU 501 executes programs stored in the memory 502 or the auxiliary storage device 503 , and executes various processes using data stored in the memory 502 or the auxiliary storage device 503 .

[0166] Each function implemented in the noise-reduction outlier detection device 100 may be realized by the CPU 501 executing a program stored in the auxiliary storage device 503 or the memory 502. Information such as the DB or table described above is stored in at least one of the memory 502, the auxiliary storage device 503, the external storage medium 507, and an external storage device accessible via the network 508.

[0167] Although one embodiment has been described above, this is merely an example for explaining the present invention, and the scope of the present invention is not limited to this embodiment. The present invention can be embodied in various other forms.

[0168] For example, the noise reduction outlier detection device 100 may be applied to a use case of operation management of an IT system, but may also be applied to other use cases in which similar data analysis by comparing actual time series data with predicted time series data is possible. Also, for example, the loop processing for each window set may be performed in parallel.

[0169] Also, for example, for at least one of the point-based processing and distribution-based processing, some of the outlier sub-detection among the expected spike detection, direction calculation, and distance calculation may be omitted, or other types of outlier sub-detection may be adopted instead of or in addition to at least some of the outlier sub-detection among the expected spike detection, direction calculation, and distance calculation.

[0170] Also, for example, the outlier detector 110 (expected spike detector 112) may automatically determine whether to perform expected spike detection on a point basis or a distribution basis. Specifically, for example, when data representing an event with a small difference in spike occurrence timing (for example, data representing that the difference between a predetermined start date and time of a predetermined process and an actual start date and time is equal to or less than a tolerance) is input to the outlier detector 110, the outlier detector 110 (expected spike detector 112) may determine to perform expected spike detection on a point basis. When data representing an event with a large difference in spike occurrence timing (for example, data representing that the difference between a predetermined start date and time of a predetermined process and an actual start date and time exceeds a tolerance) is input to the outlier detector 110, the outlier detector 110 (expected spike detector 112) may determine to perform expected spike detection on a distribution basis.

[0171] For example, the sliding adjustment length is planned but The expected spike detector 112 may automatically determine the expected spike time based on data representing the difference between a predicted start date and time and the actual start date and time (e.g., data representing the difference between a predetermined start date and time of a predetermined process and the actual start date and time).

[0172] Also, for example, only one type of outlier sub-detector may be prepared. Also, for example, only one entry ID (see Figures 3A and 3B) may be prepared for the actual time series data and the predicted time series data. In other words, only one of point-based processing and distribution-based processing may be performed for those time series data. For example, when there is only one type of outlier sub-detector and only one entry ID, the output of the outlier sub-detector may be the output of the outlier determiner 140. Also, multiple entry IDs may be provided for at least one of the point-based processing and the distribution-based processing. value Other types of information may be employed as the output of the sub-detectors instead of or in addition to Boolean values. [Explanation of symbols]

[0173] 100: Noise reduction outlier detection device

Claims

1. an outlier detector; Outlier judgement and having the outlier detector comprises a window generator and one or more types of outlier sub-detectors; the window generator generates a first processing window and a second processing window having a specified window length, and performs a sliding adjustment to slide the second processing window by a specified sliding adjustment length relative to the first processing window; each of the one or more types of outlier sub-detectors among the one or more types of outlier sub-detectors performs distribution-based outlier sub-detection, which is to detect whether or not there is an outlier candidate in the actual time series data set using information based on the entirety of an actual time series data set, which is a data portion corresponding to the first processing window among actual time series data, which is a time series of actual values, and information based on the entirety of a predicted time series data set, which is a data portion corresponding to the second processing window after the sliding adjustment among predicted time series data, which is a time series of predicted values; The outlier determiner determines whether an outlier candidate based on a result of the outlier sub-detection by the one or more types of outlier sub-detectors is an outlier. Outlier detector.

2. a plurality of parameter threshold sets for the actual time series data and the forecasted time series data; Each of the plurality of parameter threshold sets includes a parameter set including a window length and a sliding adjustment length, and a threshold set including one or more thresholds used in outlier sub-detection; For each of the plurality of parameter threshold sets, the window generator generating first and second processing windows having window lengths in the set, and performing sliding adjustments on the generated first and second processing windows according to the sliding adjustment lengths in the parameter set; each of the one or more types of outlier sub-detectors performs distribution-based outlier sub-detection using the set of thresholds; the outlier determiner determines whether an outlier candidate is an outlier based on the outlier sub-detection results of one or more types of outlier sub-detectors obtained for each of the plurality of parameter sets; The outlier detection device according to claim 1 .

3. a plurality of parameter sets in the plurality of parameter threshold sets include a point / distribution-based classifier that indicates whether point-based or distribution-based processing is to be performed; For each of the plurality of parameter threshold sets, the outlier sub-detector: If the point / distribution-based classifiers in the set represent a distribution-based process, perform distribution-based outlier sub-detection; performing a point-based outlier sub-detection, where if the point / distribution-based classifier in the set represents a point-based process, detecting whether each actual value in the actual time series data set is a candidate outlier based on a threshold in the set, each actual value in the actual time series data set corresponding to a first processing window having a window length in the set, and each predicted value in the predicted time series data set corresponding to a second processing window having the window length; The outlier detection device according to claim 2 .

4. the one or more types of outlier sub-detectors include a first type of outlier sub-detector; The first type of outlier sub-detector comprises: identifying a first number in the actual time series data set that is greater than a value threshold determined based on the predicted time series data; identifying a second number, the number of predicted values ​​greater than the value threshold; calculating a ratio of the first number to the second number as the distribution-based comparison; Detecting whether there is a potential outlier in the actual time series data set depending on the magnitude of the calculated proportion. The outlier detection device according to claim 1 .

5. the one or more types of outlier sub-detectors include a second type of outlier sub-detector; The second type of outlier sub-detector comprises: comparing the actual time series data set with the forecasted time series data set to identify a number of actual values ​​that are greater than the forecasted values; calculating a ratio of the determined number to a number of actual values ​​in the actual time series data set; Detecting whether there is a potential outlier in the actual time series data set depending on the magnitude of the calculated proportion. The outlier detection device according to claim 1 .

6. the one or more types of outlier sub-detectors include a third type of outlier sub-detector; The third type outlier sub-detector comprises: identifying a first distribution, the first distribution being a distribution of the actual time series data set; identifying a second distribution, the second distribution being a distribution of the forecast time series data set; Calculating a distance between the first distribution and the second distribution; Detecting whether there is a potential outlier in the actual time series data set depending on the magnitude of the calculated distance. The outlier detection device according to claim 1 .

7. An outlier detector; Outlier judgement and having the outlier detector comprises a window generator and one or more types of outlier sub-detectors; the window generator generates a first processing window and a second processing window having a specified window length, and performs a sliding adjustment to slide the second processing window by a specified sliding adjustment length relative to the first processing window; each of one or more types of outlier sub-detectors among the one or more types of outlier sub-detectors performs point-based outlier sub-detection, in which each actual value in an actual time-series data set, which is a data portion corresponding to the first processing window in actual time-series data, which is a time series of actual values, and each predicted value in a predicted time-series data set, which is a data portion corresponding to the second processing window after the sliding adjustment in predicted time-series data, which is a time series of predicted values, is detected as an outlier candidate; The outlier determiner determines whether an outlier candidate based on a result of the outlier sub-detection by the one or more types of outlier sub-detectors is an outlier. Outlier detector.

8. the one or more types of outlier sub-detectors include a first type of outlier sub-detector; The first type of outlier sub-detector comprises: identifying a forecast value in the forecast time series data set that is greater than a value threshold determined based on the forecast time series data; removing an actual value corresponding to the identified predicted value from the actual time series data set as an outlier candidate, and setting the remaining actual values ​​as actual value candidates; The outlier detection device according to claim 7 .

9. the one or more types of outlier sub-detectors include a second type of outlier sub-detector; The second type of outlier sub-detector comprises: Among the actual time series data set, actual values ​​that are greater than the predicted value are set as outlier candidates, and actual values ​​that are equal to or less than the predicted value are excluded from the outlier candidates. The outlier detection device according to claim 7 .

10. the one or more types of outlier sub-detectors include a third type of outlier sub-detector; The third type outlier sub-detector comprises: Calculating, for each date and time, a distance between an actual value in the actual time series data set and a predicted value in the predicted time series data set; For each date and time, depending on the magnitude of the calculated distance, it is detected whether an actual value corresponding to the date and time in the actual time series data set is an outlier candidate. The outlier detection device according to claim 7 .

11. for each of the plurality of parameter threshold sets, the outlier sub-detector performs, based on the set, either a distribution-based outlier sub-detection or a point-based outlier sub-detection, which is to detect whether each actual value in the actual time series data set is a candidate outlier; The outlier sub-detector is When there are one or more point-based outlier sub-detection results, a single outlier sub-detection result is calculated as an AND of the one or more outlier sub-detection results, and a single point-based result value is calculated based on an occurrence rate, which is a ratio of values ​​that are considered to be outlier candidates among the single outlier sub-detection result; If there are one or more distribution-based outlier sub-detection results, calculating a distribution-based result value that is the AND of the one or more outlier sub-detection results; determining whether the outlier candidate is an outlier based on the one point-based result value and the one distribution-based result value; The outlier detection device according to claim 2 .

12. Each of the one or more types of outlier sub-detectors outputs information representing a result of the outlier sub-detection to log information; the outlier determiner outputs information indicating a determination result as to whether the outlier candidate has been determined to be an outlier to the log information; The information output to the log information includes a log message regarding a result of the detection or determination, displaying result information including an outlier determination result and a log message based on the log information; An outlier detection device according to any one of claims 1 to 11.

13. A computer generates a first processing window and a second processing window having a specified window length; a computer performs a sliding adjustment to slide the second processing window relatively to the first processing window by a specified sliding adjustment length; The computer performs one or more types of outlier sub-detection among the one or more types of outlier sub-detection; Each of the one or more types of outlier sub-detection includes performing distribution-based outlier sub-detection, which is to detect whether there is an outlier candidate in the actual time series data set using information based on the entirety of an actual time series data set, which is a data portion corresponding to the first processing window among actual time series data, which is a time series of actual values, and information based on the entirety of a predicted time series data set, which is a data portion corresponding to the second processing window after the sliding adjustment among predicted time series data, which is a time series of predicted values; The computer determines whether the outlier candidate based on the result of the one or more types of outlier sub-detection is an outlier. Outlier detection methods.

14. A method for generating a first processing window and a second processing window having a specified window length, a computer performs a sliding adjustment to slide the second processing window relatively to the first processing window by a specified sliding adjustment length; The computer performs one or more types of outlier sub-detection among the one or more types of outlier sub-detection; Each of the one or more types of outlier sub-detection performs point-based outlier sub-detection to detect whether each actual value in the actual time-series data set is an outlier candidate based on each actual value in the actual time-series data set, which is a data portion corresponding to the first processing window in the actual time-series data, which is a time series of actual values, and each predicted value in the predicted time-series data set, which is a data portion corresponding to the second processing window after the sliding adjustment in the predicted time-series data, which is a time series of predicted values; The computer determines whether the outlier candidate based on the result of the one or more types of outlier sub-detection is an outlier. Outlier detection methods.

Citation Information

Patent Citations

  • Abnormality diagnosis method and abnormality diagnosis system using pattern library

    JP2011247695A

  • Abnormality detection device and program

    JP2015026252A

  • Traffic management device, traffic management method and program

    JP2018195929A

  • Resource allocation optimization system and method

    JP2019082801A

  • Malfunction detection method and malfunction detection program

    JP2021182287A