Anomaly detection system, information processing system, anomaly detection method, and anomaly detection program
The abnormality detection system addresses the challenge of detecting storage device issues in redundant configurations without speed loss by using performance data analysis, enhancing detection accuracy and broad applicability.
Patent Information
- Application Number
- JP2023578304
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-02-04
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-02-04
AI Technical Summary
Existing failure prediction devices struggle to detect abnormalities in redundant storage devices without causing a decrease in processing speed, and specialized measurement means are required for general-purpose storage devices.
An abnormality detection system that acquires performance data from redundant storage devices, calculates a score indicating differences using a calculation model, and detects abnormalities based on these scores without requiring additional commands, thus maintaining processing speed.
The system effectively detects abnormalities in redundant storage devices without slowing down operations, improving detection accuracy and applicability to general-purpose devices.
Smart Images

Figure 0007708224000001 
Figure 0007708224000002 
Figure 0007708224000003
Abstract
Description
Technical Field
[0001] The present invention relates to an abnormality detection system and the like.
Background Art
[0002] In improving the availability of an information processing system, the stable operation of a storage device is an important factor. The storage device is operated, for example, in a redundant configuration by RAID (Redundant Arrays of Inexpensive Disks). Further, in order to improve the availability of the information processing system, it is desirable to be able to detect an abnormality in a redundant storage device at an early stage. By detecting an abnormality at an early stage, for example, an administrator of the information processing system can replace or adjust the storage device in which the abnormality has been detected before the influence of the abnormality becomes large. Therefore, the development of a technique for detecting an abnormality in an operating storage device has been carried out.
[0003] The failure prediction device of Patent Document 1 measures the access time of a hard disk drive using an access pattern for confirming an abnormality. The failure prediction device of Patent Document 1 records an access target location as an abnormal location when the measured access time exceeds a threshold value.
[0004] The failure omen detection device of Patent Document 2 detects an abnormality based on the floating amount of the head of a hard disk drive.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0006] The failure prediction device of Patent Document 1 may have difficulty in detecting an abnormality in an operating storage device without causing a decrease in processing speed. The failure omen detection device of Patent Document 2 may have difficulty in detecting an abnormality in a general-purpose storage device in circulation because special measurement means needs to be implemented.
[0007] In order to solve the above problems, an object is to provide an abnormality detection system or the like that can detect an abnormality in a redundant-configured storage device without causing a decrease in processing speed.
Means for Solving the Problems
[0008] In order to solve the above problems, the abnormality detection system of the present invention includes an acquisition means for acquiring performance data of each of a plurality of storage devices having a redundant configuration, a calculation means for calculating a score indicating a difference between the performance data acquired by the acquisition means and the performance data in a normal state using a calculation model, a detection means for detecting an abnormality in the storage device based on the difference in scores between the storage devices using the scores of each of the plurality of storage devices calculated by the calculation means, and an output means for outputting the result of the detection.
[0009] The abnormality detection method of the present invention acquires performance data of each of a plurality of storage devices having a redundant configuration, calculates scores of each of the plurality of storage devices from the acquired performance data using a calculation model that calculates a score indicating a difference between the acquired performance data and the performance data in a normal state, detects an abnormality in the storage device based on the difference in scores between the storage devices using the calculated scores of each of the plurality of storage devices, and outputs the result of the detection.
[0010] The abnormal detection program that non-temporarily records on a computer causes the computer to execute a process of acquiring performance data of each of a plurality of storage devices having a redundant configuration, a process of calculating scores of each of the plurality of storage devices from the acquired performance data using a calculation model that calculates a score indicating a difference between the acquired performance data and the performance data during normal operation, a process of detecting an abnormality of a storage device based on a difference between the scores of the storage devices using the calculated scores of each of the plurality of storage devices, and a process of outputting the result of the detection.
Effect of the Invention
[0011] According to the present invention, it is possible to detect an abnormality of a storage device having a redundant configuration without causing a decrease in processing speed.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Embodiments for Carrying Out the Invention
[0013] The first embodiment of the present invention will be described in detail with reference to the drawings. FIG. 1 is a diagram showing an overview of the configuration of the information processing system of this embodiment. The information processing system includes an abnormality detection system 10 and a storage device 20. The information processing system includes a plurality of storage devices 20 having a redundant configuration. The storage device 20, for example, stores data and outputs the stored data based on a request from a server. The plurality of storage devices 20 are operated, for example, as a RAID (Redundant Arrays of Inexpensive Disks).
[0014] The abnormality detection system 10 detects an abnormality in the storage device 20. An abnormality in the storage device 20 is, for example, a state in which it is difficult to continue the operation of the storage device 20. The abnormality may include a sign of an abnormal state. A sign of an abnormal state means, for example, a state in which there is a high possibility of an abnormality occurring if the operation of the storage device 20 is continued. A sign of an abnormal state is, for example, a state in which it is necessary to replace the storage device 20, change the settings, or review the operation method in order to avoid the occurrence of an abnormality.
[0015] The abnormality detection system 10 calculates the degree of abnormality of the storage device 20 based on the difference in the performance data of the storage device 20 having a redundant configuration. The abnormality detection system 10 calculates a score using a calculation model that calculates a score indicating the difference between the performance data of the storage device 20 during normal operation and the performance data acquired when detecting an abnormality. Then, the abnormality detection system 10 calculates the degree of abnormality indicating the degree of abnormality of the storage device 20 based on the difference in scores between the storage devices 20 having a redundant configuration. The abnormality detection system 10 detects an abnormality of the storage device 20 when the degree of abnormality exceeds a preset criterion. The method for calculating the score and the method for calculating the degree of abnormality will be described later.
[0016] The abnormality detection system 10 may be applied to each of the configurations obtained by subdividing the actual redundant configuration applied to the plurality of storage devices 20. For example, when the storage device 20 is configured with RAID10, the abnormality detection system 10 may be applied to a pair of storage devices 20 to which mirroring is performed. The configuration of the abnormality detection system 10 will be described. FIG. 2 is a diagram showing an example of the configuration of the abnormality detection system 10. The abnormality detection system 10 includes an acquisition unit 11, a calculation unit 12, a detection unit 13, an output unit 14, a model generation unit 15, and a storage unit 16.
[0017] The acquisition unit 11 acquires the performance data of each of a plurality of storage devices 20 having a redundant configuration. The items of performance data acquired by the acquisition unit 11 are set when generating the calculation model. The acquisition unit 11 acquires, for example, as the performance data of the storage device to be monitored, at least one or more items of data among the average write response time, the maximum write response time, the average read response time, and the maximum read response time. The unit of the data of each item is, for example, microseconds. Further, the acquisition unit 11 may acquire the data of one or more items among the write count, the write transfer rate, the read count, the read transfer rate, the busy ratio, and the busy time as the performance data. The units of the data of the write transfer rate and the read transfer rate are, for example, kilobytes per second. The unit of the data of the busy ratio is, for example, percent. The unit of the data of the busy time is, for example, milliseconds. The examples of the performance data acquired by the acquisition unit 11 are not limited to the above.
[0018] When generating the calculation model, the acquisition unit 11 may acquire the performance data when the storage device 20 is operating normally. The acquisition unit 11 acquires, for example, as the performance data used for generating the calculation model, the time-series data of the performance data of the storage device 20 to be detected for abnormality. The acquisition unit 11 may acquire the performance data of the same model as the storage device 20 to be detected for abnormality as the performance data used for generating the calculation model. The same model includes models that can be regarded as the same. The acquisition unit 11 stores the acquired performance data in the storage unit 16, for example, when the performance data is acquired.
[0019] The calculation unit 12 calculates the score for each of the plurality of storage devices 20 from the acquired performance data using a calculation model that calculates a score indicating the difference between the performance data acquired by the acquisition unit 11 and the performance data during normal operation.
[0020] The calculation unit 12 inputs the data acquired by the acquisition unit 11 into the calculation model. The calculation model calculates a score indicating the difference from the performance data during normal operation for each storage device 20. The performance data during normal operation is stored, for example, in the storage unit 16. The calculation model calculates the score, for example, by calculating the distance between the performance data during normal operation and the performance data acquired by the acquisition unit 11.
[0021] The calculation model calculates, for example, the distance between the performance data of the storage device 20 during normal operation and the performance data at each time acquired by the acquisition unit 11 using the k-NN method (k-nearest neighbor algorithm). The calculation model calculates, for example, the distance in the feature space between the feature quantities extracted from the performance data of the storage device 20 during normal operation and the feature quantities extracted from the performance data at each time acquired by the acquisition unit 11. The calculation model calculates, for example, the Euclidean distance between the performance data of the storage device 20 during normal operation and the performance data at each time acquired by the acquisition unit 11. The distance between the performance data of the storage device 20 during normal operation and the performance data at each time calculated by the calculation model is not limited to the Euclidean distance. The calculation model is set, for example, to calculate the average distance for 10 neighbors as the score. The number of neighbors for calculating the average distance is not limited to 10. The calculation model calculates the distance in a multi-dimensional space according to the number of items when there are a plurality of performance data items.
[0022] The calculation unit 12 calculates the score, for example, using the calculation model generated by the model generation unit 15. The calculation unit 12 may calculate the score using a calculation model that has already been generated on an external server.
[0023] The detection unit 13 detects an abnormality in the storage device 20 based on the difference in scores between the storage devices using the scores of each of the plurality of storage devices calculated by the calculation unit 12. For example, the detection unit 13 calculates an abnormality degree indicating the degree of abnormality based on the scores calculated by the calculation unit 12. Then, the calculation unit 12 detects an abnormality in the storage device 20 when the abnormality degree exceeds a predetermined standard.
[0024] The predetermined standard for the abnormality degree for detecting an abnormality is set in advance, for example, based on the abnormality degree calculated based on the performance data during normal times in the past. The predetermined standard for the abnormality degree may be set in advance based on the abnormality degree calculated based on the performance data when an abnormality occurred in the past. The detection unit 13 may detect an abnormality in the storage device 20 using the normalized abnormality degree using the maximum value of the abnormality degree calculated based on the normal-time data. For example, the detection unit 13 detects an abnormality in the storage device 20 using the normalized abnormality degree for each storage device 20.
[0025] The predetermined standard for the abnormality degree may be set in multiple stages. For example, the predetermined standard for the abnormality degree may be set in two stages: a first standard for detecting a sign and a second standard for detecting an abnormality. Also, the predetermined standard may be set in three or more stages.
[0026] The detection unit 13 may calculate an abnormality after performing a predetermined process on the scores calculated by the calculation unit 12. For example, as the predetermined process, the detection unit 13 normalizes the time-series data of the scores based on the normal-time scores. For example, the detection unit 13 normalizes the score at each time included in the time-series data using the average value or the maximum value of the normal-time scores of the storage device 20.
[0027] The detection unit 13 may smooth the time-series data of the scores by calculating the statistic for the time-series data of the scores in each section where a predetermined time range is sequentially moved over the time-series data of the scores. For example, the detection unit 13 smooths the time-series data of the scores by moving the section of the predetermined time range every minute. The unit for moving the predetermined time range is not limited to one minute. Also, the predetermined time range is set in advance, for example, according to the time interval of the time-series data of the scores. The predetermined time range is set to, for example, 10 minutes. The predetermined time range is not limited to 10 minutes.
[0028] For example, the detection unit 13 calculates a predetermined statistic of the time-series data of the scores within a predetermined time range. The predetermined statistic is, for example, the maximum value. The predetermined statistic may be other than the maximum value. In each section where the predetermined time range is moved, the detection unit 13 replaces the score of each section with the calculated predetermined statistic. For example, the detection unit 13 smooths the time-series data by replacing the score at the first time among the time-series data included in each section of the predetermined time range with the predetermined statistic. The detection unit 13 may also smooth the time-series data by replacing the score at the central time of the section among the time-series data of the scores in each section of the predetermined time range with the predetermined statistic. Which score at which time among the time-series data of the scores in each section of the predetermined time range is replaced with the predetermined statistic is not limited to the above example. Also, the method of the normalization and smoothing processes is not limited to the above example.
[0029] For example, when two storage devices 20 are redundant, the detection unit 13 calculates the abnormality degree based on the difference in scores between the storage devices 20. For example, when RAID1 redundancy is implemented, the detection unit 13 calculates the abnormality degree based on the difference in scores between the storage devices 20. The configuration in which two storage devices 20 are redundant is not limited to RAID1.
[0030] The detection unit 13 calculates the abnormality degree based on, for example, the difference between the score of each storage device 20 and the average value of the scores of the redundant storage devices 20 at each time when three or more storage devices 20 are redundant. The detection unit 13 calculates the abnormality degree based on, for example, the difference between the score of each device and the average value in the case of RAID5. The configuration in which three or more storage devices 20 are redundant is not limited to RAID5.
[0031] FIG. 3 is a diagram schematically showing an example of a flow when the abnormality detection system 10 detects an abnormality of the storage device 20. In FIG. 3(a), the calculation unit 12 calculates a score using the performance data of each storage device 20 and the performance data in a normal state by a calculation model. In FIG. 3(b), the detection unit 13 normalizes the score calculated by the calculation unit 12. Further, in FIG. 3(c), the detection unit 13 smoothes the time-series data of the score. Then, in FIG. 3(d), the detection unit 13 calculates the abnormality degree of the storage device 20 based on the difference in scores between the storage devices 20. The detection unit 13 detects an abnormality of the storage device 20 based on the calculated abnormality degree.
[0032] FIG. 4 shows an example of time-series data of the scores of each storage device 20 when two storage devices 20 are redundant. The vertical axis of the graph in FIG. 4 indicates the score. The horizontal axis of the graph in FIG. 4 indicates the time. FIGS. 4(a) and (b) show the time-series data of the scores of the two storage devices 20 respectively.
[0033] In the example of FIG. 4, around time T1, the scores of both (a) and (b) are increasing. When redundancy is provided by two storage devices 20, since access is performed on the two storage devices 20 at the same timing, during normal operation, changes in the scores occur around the same time. On the other hand, around time T2, the score of (a) is increasing, but the score of (b) is not increasing. Since the tendencies of the score changes of the two storage devices 20 are different, around time T2, it is highly likely that the two storage devices 20 are operating differently. Since normal ones show the same tendency of change, when the tendencies of the score changes are different, it is highly likely that an abnormality has occurred in one of the storage devices 20. Therefore, the detection unit 13 can detect the storage device 20 in which an abnormality has occurred, for example, based on the difference in the scores of the two storage devices 20.
[0034] FIG. 5 shows an example of time-series data of the scores of the respective storage devices 20 when redundancy is provided by three storage devices 20. The vertical axis of the graph in FIG. 5 indicates the score. The horizontal axis of the graph in FIG. 5 indicates the time. FIG. 5(a) is a diagram showing an example of the time-series data of the scores of each of the three storage devices 20 when the change in the score is small. FIG. 5(b) is a diagram showing an example of the time-series data of the scores of each of the three storage devices 20 when the change in the score is large. FIG. 5 is an example of the time-series data of the scores during normal operation in both (a) and (b).
[0035] In a redundant configuration such as RAID5, since data is stored separately in each storage device 20, as shown in the examples of FIGS. 5(a) and 5(b), the change timings of the scores of the three redundant storage devices 20 differ by the time difference for dividing and sequentially writing the data. However, as shown in the examples of FIGS. 5(a) and 5(b), when viewed on a wide time scale, the changes in the scores of the three redundant storage devices 20 show the same tendency. On the other hand, the score of the storage device 20 in which an abnormality has occurred shows a tendency different from that of the other storage devices 20. Therefore, the detection unit 13 can detect the storage device 20 in which an abnormality has occurred, for example, based on the difference between the average value of the scores of the three storage devices 20 and the score of each storage device 20.
[0036] The output unit 14 outputs the detection result of the abnormality of the storage device 20. For example, the output unit 14 outputs the detection result of the abnormality to a server that manages the storage device 20. The output unit 14 may output the detection result of the abnormality to a terminal device held by the administrator of the storage device 20. The output unit 14 may output the detection result of the abnormality to the control unit of the storage device 20. The output unit 14 may output the detection result of the abnormality to a display device (not shown) connected to the abnormality detection system 10. The output destination of the detection result of the abnormality is not limited to the above example.
[0037] For example, when the output unit 14 detects an abnormality in the storage device 20, the output unit 14 outputs information indicating that the abnormality has been detected as the detection result of the abnormality. For example, the output unit 14 may output the degree of abnormality of the storage device 20 as the detection result. Further, the output unit 14 may output the score of the storage device 20 together with the degree of abnormality.
[0038] When the criteria for the degree of abnormality are set in multiple stages, the output unit 14 may output the detection result of the abnormality when the criteria for each stage are exceeded. For example, when two-stage criteria of detection of a sign of an abnormal state and detection of the occurrence of an abnormality are set, the output unit 14 may output, as the abnormal result, information indicating that the sign of the abnormal state or the occurrence of the abnormality has been detected when the degree of abnormality exceeds each criterion.
[0039] When generating a calculation model in the abnormality detection system 10, the model generation unit 15 learns the normal performance data and generates a calculation model for calculating a score. For example, the model generation unit 15 learns the normal performance data and generates a calculation model for calculating the distance between the performance data. For example, the model generation unit 15 calculates the distance by the k-NN method and generates a calculation model that outputs the distance as a score. The model generation unit 15 stores the generated calculation model and the normal performance data in the storage unit 16.
[0040] The calculation model may be generated by a method other than the above as long as it can calculate a score indicating the degree of deviation from the normal state. The model generation unit 15 may, for example, learn an invariant relationship for time-series data of performance data and generate a calculation model that calculates a score indicating a deviation from the normal state. The model generation unit 15 may, for example, generate a relational expression indicating the relationship between the values of the performance data. Then, the model generation unit 15 calculates the deviation from the normal state in the relational expression indicating the relationship between the values of the performance data as a score. The algorithm for generating the calculation model for calculating the score is not limited to the above example.
[0041] The model generation unit 15 may, for example, generate a calculation model for each storage device 20. The model generation unit 15 may generate a calculation model for each group of redundant configurations. The group of redundant configurations means, for example, when three storage devices 20 are redundant, the three storage devices 20 are regarded as one group. The model generation unit 15 may generate a calculation model for each type of RAID. Also, the model generation unit 15 may generate a calculation model for each storage device 20 of the same model. The same model includes models that can be regarded as the same. The classification of the target when generating the calculation model is not limited to the above example.
[0042] The model generation unit 15 may generate a calculation model according to the elapsed time since the start of use of the storage device 20. For example, the model generation unit 15 may generate a calculation model for the period immediately after the start of use and the period when aging changes occur.
[0043] The storage unit 16 stores, for example, the calculation model used by the calculation unit 12 for calculating the score and the performance data in the normal state. Also, the storage unit 16 may store the performance data of the storage device 20 to be detected for abnormality output by the acquisition unit 11. When generating the calculation model, the storage unit 16 may store the performance data in the normal state used by the model generation unit 15 for generating the calculation model.
[0044] The storage device 20 stores, for example, data sent from a server in a redundant state among a plurality of storage devices 20. The plurality of storage devices 20 are operated, for example, in a redundant configuration by RAID. Also, the storage device 20 reads out data requested from the server and outputs it to the server. The storage device 20 is, for example, a hard disk drive. The storage device 20 is not limited to a hard disk drive as long as it is used in a redundant manner.
[0045] The operation for detecting an abnormality of the storage device 20 in the abnormality detection system 10 of the information processing system will be described. FIG. 6 is a diagram showing an example of an operation flow when the abnormality detection system 10 detects an abnormality of the storage device 20.
[0046] The acquisition unit 11 acquires performance data of each of the storage devices 20 having a redundant configuration that is a detection target for an abnormality (step S11). When the performance data is acquired, the calculation unit 12 calculates a score indicating the difference between the acquired performance data and the performance data at normal times for each of the storage devices 20 (step S12). The calculation unit 12 calculates scores for each of the plurality of storage devices 20 from the acquired performance data using a calculation model that calculates a score indicating the difference between the performance data acquired by the acquisition unit 11 and the performance data at normal times.
[0047] When the score is calculated, the detection unit 13 normalizes the score for each of the storage devices 20 (step S13). When the score is normalized, the detection unit 13 smoothes the time-series data of the score (step S14).
[0048] When the time-series data of the score is smoothed, the detection unit 13 calculates an abnormality degree indicating the degree of abnormality of the storage device based on the difference between the storage devices of the smoothed score. When two storage devices 20 are redundant (No in step S15), the detection unit 13 calculates the abnormality degree based on the difference between the scores between the storage devices 20 using the scores of each of the plurality of storage devices 20 (step S16).
[0049] When the degree of abnormality is equal to or higher than the reference, the detection unit 13 detects that an abnormality has occurred in the storage device 20. When an abnormality in the storage device 20 is detected, the output unit 14 outputs the detection result of the abnormality (step S17).
[0050] Also, when the storage devices 20 are redundant with three or more units in step S15 (Yes in step S15), the detection unit 13 calculates the degree of abnormality based on the difference between a predetermined statistic of the scores of the plurality of storage devices 20 and the score of the storage device 20 to be detected for abnormality (step S18).
[0051] When the degree of abnormality is equal to or higher than the reference, the detection unit 13 detects that an abnormality has occurred in the storage device 20. When an abnormality in the storage device 20 is detected, the output unit 14 outputs the detection result of the abnormality (step S17).
[0052] In the abnormality detection system 10, the operation when generating the calculation model will be described. FIG. 7 is a diagram showing an example of the operation flow when the abnormality detection system 10 generates the calculation model.
[0053] The acquisition unit 11 acquires the normal performance data of the storage device 20 (step S21). When the normal performance data is acquired, the model generation unit 15 learns the normal performance data and generates a calculation model that calculates a score indicating the difference between the newly acquired performance data and the normal performance data (step S22). When the calculation model is generated, the model generation unit 15 stores the generated calculation model and the normal data in the storage unit 16 (step S23).
[0054] The anomaly detection system 10 of the information processing system according to this embodiment acquires performance data from the storage device 20, and calculates a score indicating the difference between the acquired performance data and the normal performance data using a calculation model. Further, the anomaly detection system 10 calculates an anomaly indicating the degree of anomaly of the storage device 20 based on the difference in scores between the storage devices 20. Then, the anomaly detection system 10 detects that an anomaly has occurred in the storage device 20 when the degree of anomaly exceeds a reference. In this way, the anomaly detection system 10 can detect an anomaly without causing the storage device 20 to execute a special command for anomaly detection during the operation of the storage device 20. Therefore, the anomaly detection system 10 can detect an anomaly in the storage device 20 without causing a decrease in processing speed by detecting an anomaly based on the performance data during the operation of the storage device 20 having a redundant configuration.
[0055] FIG. 8 is a diagram showing an example of a graph for calculating the AUC (Area Under the Curve) when an anomaly is detected based on the difference in scores between the storage devices 20. Further, FIG. 9 is a diagram showing an example of a graph for calculating the AUC when an anomaly is detected based on the scores of each storage device.
[0056] The vertical axis of the graphs in FIGS. 8 and 9 is the TPR (True Positive Rate). The TPR indicates the ratio of data that should be determined as Positive and has been correctly determined as Positive. That is, the TPR indicates the ratio of data that should be determined as abnormal and has been determined as abnormal. Also, the horizontal axis of the graphs in FIGS. 8 and 9 is the FPR (False Positive Rate). The FPR indicates the ratio of data that should be determined as Negative but has been determined as Positive. That is, the FPR indicates the ratio of data that should not be determined as abnormal but has been determined as abnormal. The AUC indicates the ratio of the area below the curve formed by the plotted points in the graphs of FIGS. 8 and 9. Therefore, the larger the value of the AUC, the higher the accuracy of anomaly detection. In the example of FIG. 9, the AUC is 0.78, whereas in the example of FIG. 8, the AUC is 0.85. Therefore, the anomaly detection system 10 can improve the accuracy of detecting anomalies in the redundant memory devices 20 by detecting anomalies based on the score differences between the redundant memory devices 20 which are in a redundant configuration.
[0057] Also, by calculating the degree of abnormality based on the statistical quantity of the scores of multiple memory devices and the difference between the scores of each memory device, the anomaly detection system 10 can detect anomalies in the memory devices 20 even in the case where data is divided and stored in multiple memory devices 20.
[0058] Moreover, since the anomaly detection system 10 can detect anomalies based on the performance data of the memory devices 20, there is no need to add a function for detecting anomalies to the memory devices 20. Therefore, anomalies in the memory devices 20 can be detected without causing complexity in the configuration of the memory devices 20. Also, it can be applied to general-purpose memory devices in circulation.
[0059] (Second Embodiment) The second embodiment of the present invention will be described in detail with reference to the drawings. FIG. 10 is a diagram showing an example of the configuration of the anomaly detection system 100 of this embodiment. The anomaly detection system 100 includes an acquisition unit 101, a calculation unit 102, a detection unit 103, and an output unit 104.
[0060] The acquisition unit 101 acquires the performance data of each of a plurality of storage devices having a redundant configuration. The calculation unit 102 calculates the score of each of the plurality of storage devices from the acquired performance data using a calculation model that calculates a score indicating the difference between the performance data acquired by the acquisition unit 101 and the performance data during normal operation. The detection unit 103 detects an abnormality of the storage device based on the difference in scores between the storage devices using the scores of each of the plurality of storage devices calculated by the calculation unit 102. The output unit 104 outputs the result of the detection.
[0061] Here, the acquisition unit 11 of the first embodiment is an example of the acquisition unit 101. Further, the acquisition unit 101 is an aspect of the acquisition means. The calculation unit 12 of the first embodiment is an example of the calculation unit 102. Further, the calculation unit 102 is an aspect of the calculation means. The detection unit 13 of the first embodiment is an example of the detection unit 103. Further, the detection unit 103 is an aspect of the detection means. The output unit 14 of the first embodiment is an example of the output unit 104. Further, the output unit 104 is an aspect of the output means.
[0062] The operation of the abnormality detection system 100 will be described. FIG. 11 is a diagram showing an example of the operation flow of the abnormality detection system 100.
[0063] The acquisition unit 101 acquires the performance data of each of a plurality of storage devices having a redundant configuration (step S101). When the performance data is acquired, the calculation unit 102 calculates the score of each of the plurality of storage devices from the acquired performance data using a calculation model that calculates a score indicating the difference between the performance data acquired by the acquisition unit 101 and the performance data during normal operation (step S102). When the score is calculated, the detection unit 103 detects an abnormality of the storage device based on the difference in scores between the storage devices using the scores of each of the plurality of storage devices calculated by the calculation unit 102 (step S103). When an abnormality of the storage device is detected, the output unit 104 outputs the result of the detection (step S104).
[0064] The abnormality detection system 100 of this embodiment calculates a score from the acquired performance data by using a calculation model that calculates a score indicating the difference between the performance data acquired by the acquisition unit 101 from a storage device with a redundant configuration and the performance data during normal operation. Then, in the detection unit 103, the abnormality detection system 100 detects an abnormality in the storage device 20 based on the difference in scores between the storage devices 20 by using the scores of each of the plurality of storage devices 20 calculated by the calculation unit 102. In this way, the abnormality detection system 100 can detect an abnormality in the storage device 20 with a redundant configuration by determining the abnormality in the storage device 20 based on the degree of abnormality calculated from the difference in scores between the storage devices 20 with a redundant configuration.
[0065] Each process in the abnormality detection system 10 of the first embodiment and the abnormality detection system 100 of the second embodiment can be realized by executing a computer program on a computer. FIG. 12 shows an example of the configuration of a computer 200 that executes a computer program for performing each process in the abnormality detection system 10 of the first embodiment and the abnormality detection system 100 of the second embodiment. The computer 200 includes a CPU (Central Processing Unit) 201, a memory 202, a storage device 203, an input / output I / F (Interface) 204, and a communication I / F 205.
[0066] The CPU 201 reads out and executes a computer program for each process from the storage device 203. The CPU 201 may be configured by a combination of a plurality of CPUs. Also, the CPU 201 may be configured by a combination of a CPU and other types of processors. For example, the CPU 201 may be configured by a combination of a CPU and a GPU (Graphics Processing Unit). The memory 202 is composed of a DRAM (Dynamic Random Access Memory) or the like, and temporarily stores the computer program executed by the CPU 201 and the data being processed. The storage device 203 stores the computer program executed by the CPU 201. The storage device 203 is composed of, for example, a non-volatile semiconductor storage device. Other storage devices such as a hard disk drive may be used for the storage device 203. The input / output I / F 204 is an interface that receives input from an operator and outputs display data and the like. The communication I / F 205 is an interface that transmits and receives data between the storage device 20 and other information processing devices.
[0067] The computer program used for the execution of each process can also be stored on a recording medium that non-temporarily records data and distributed. As the recording medium, for example, a magnetic tape for data recording or a magnetic disk such as a hard disk can be used. Also, as the recording medium, an optical disk such as a CD-ROM (Compact Disc Read Only Memory) can be used. A non-volatile semiconductor storage device may be used as the recording medium.
[0068] Some or all of the above embodiments can also be described as in the following supplementary notes, but are not limited thereto.
[0069] [Supplementary Note 1] acquisition means for acquiring performance data of each of a plurality of storage devices having a redundant configuration; Calculation means for calculating a score indicating the difference between the performance data acquired by the acquisition means and the performance data in normal operation, from the acquired performance data for each of the plurality of storage devices, using a calculation model; Detection means for detecting an abnormality of the storage device based on a difference in scores between the storage devices, using the scores for each of the plurality of storage devices calculated by the calculation means; Output means for outputting the result of the detection Comprising Anomaly detection system.
[0070] [Appendix 2] The detection means calculates an abnormality degree indicating the degree of abnormality of the storage device based on the difference in scores between the storage devices, and detects an abnormality of the storage device when the calculated abnormality degree exceeds a predetermined standard. The anomaly detection system according to Appendix 1.
[0071] [Appendix 3] The detection means calculates an abnormality degree based on the difference between a predetermined statistic of the scores of the plurality of storage devices and the score of each storage device, and detects an abnormality of the storage device when the calculated abnormality degree exceeds a predetermined standard. The anomaly detection system according to Appendix 1.
[0072] [Appendix 4] The detection means uses the abnormality degree calculated based on the data in normal operation, and detects an abnormality of the storage device when the normalized abnormality degree exceeds a standard. The anomaly detection system according to Appendix 2 or 3.
[0073] [Appendix 5] The detection means normalizes the time-series data of the scores based on the scores in normal operation. The anomaly detection system according to any one of Appendices 1 to 4.
[0074] [Appendix 6] The detection means smooths the time-series data of the score by sequentially moving a predetermined time range and calculating a predetermined statistic for the time-series data of the score in each interval where the predetermined time range is moved. The anomaly detection system according to any one of Appendices 1 to 5.
[0075] [Appendix 7] The calculation model calculates the score based on the distance between the performance data acquired by the acquisition means and the performance data in the normal state. The anomaly detection system according to any one of Appendices 1 to 6.
[0076] [Appendix 8] The performance data includes at least one of an average write response time, a maximum write response time, an average read response time, and a maximum read response time. The anomaly detection system according to any one of Appendices 1 to 7.
[0077] [Appendix 9] Model generation means for learning the performance data in the normal state and generating the calculation model for calculating the score The anomaly detection system according to any one of Appendices 1 to 8, further comprising.
[0078] [Appendix 10] A plurality of storage devices with a redundant configuration, The anomaly detection system according to any one of Appendices 1 to 9 and Comprising The acquisition means of the anomaly detection system acquires the performance data of each of the storage devices and detects an anomaly in the storage device. An information processing system.
[0079] [Appendix 11] Acquire the performance data of each of the plurality of storage devices with a redundant configuration, Using a calculation model that calculates a score indicating the difference between the acquired performance data and the performance data in the normal state, calculate the score of each of the plurality of storage devices from the acquired performance data, Using the scores of each of the calculated plurality of storage devices, detecting an abnormality in the storage device based on the difference in scores between the storage devices, outputting the result of the detection, Anomaly detection method.
[0080] [Appendix 12] A process of acquiring performance data of each of a plurality of storage devices having a redundant configuration, Using a calculation model for calculating a score indicating the difference between the acquired performance data and the performance data in a normal state, calculating the score of each of the plurality of storage devices from the acquired performance data, Using the scores of each of the calculated plurality of storage devices, detecting an abnormality in the storage device based on the difference in scores between the storage devices, A process of outputting the result of the detection A recording medium that non-temporarily records an anomaly detection program that causes a computer to execute the above.
[0081] As described above, the present invention has been described by taking the above-described embodiments as examples. However, the present invention is not limited to the above-described embodiments. That is, the present invention can apply various aspects that can be understood by those skilled in the art within the scope of the present invention.
Explanation of reference numerals
[0082] 10 Anomaly detection system 11 Acquisition unit 12 Calculation unit 13 Detection unit 14 Output unit 15 Model generation unit 16 Storage unit 20 Storage device 100 Anomaly detection system 101 Acquisition unit 102 Calculation unit 103 Detection unit 104 Output unit 200 Computer 201 CPU 202 Memory 203 Storage device 204 Input / Output I / F 205 Communication I / F
Claims
1. An acquisition means for acquiring performance data of each of a plurality of storage devices having a redundant configuration; A calculation means for calculating a score for each of the plurality of storage devices from the acquired performance data using a calculation model that calculates a score indicating a difference between the performance data acquired by the acquisition means and the performance data in a normal state; A detection means for detecting an abnormality of the storage device based on a difference in scores between the storage devices using the scores of each of the plurality of storage devices calculated by the calculation means; An output means for outputting the result of the detection Comprising An abnormality detection system.
2. The detection means calculates an abnormality degree indicating the degree of abnormality of the storage device based on a difference in scores between the storage devices, and detects an abnormality of the storage device when the calculated abnormality degree exceeds a predetermined criterion. The abnormality detection system according to claim 1.
3. The detection means calculates an abnormality degree based on a difference between a predetermined statistic of the scores of the plurality of storage devices and the score of each storage device, and detects an abnormality of the storage device when the calculated abnormality degree exceeds a predetermined criterion. The abnormality detection system according to claim 1.
4. The detection means uses the abnormality degree calculated based on data in a normal state, and detects an abnormality of the storage device when the normalized abnormality degree exceeds a criterion. The abnormality detection system according to claim 2 or 3.
5. The detection means normalizes the time series data of the scores based on the scores in a normal state. The abnormality detection system according to any one of claims 1 to 4.
6. The detection means sequentially moves a predetermined time range for the time series data of the scores, and smoothes the time series data of the scores by calculating a predetermined statistic for the time series data of the scores in each section where the predetermined time range is moved. The abnormality detection system according to any one of claims 1 to 5.
7. Further comprising a model generation means for learning performance data in a normal state and generating the calculation model for calculating the score. The abnormality detection system according to any one of claims 1 to 6.
8. A plurality of storage devices having a redundant configuration; The abnormality detection system according to any one of claims 1 to 7 Comprising The acquisition means of the abnormality detection system acquires performance data of each of the storage devices and detects an abnormality of the storage device. An information processing system.
9. A computer Acquires performance data of each of a plurality of storage devices having a redundant configuration Using a calculation model that calculates a score indicating the difference between the obtained performance data and the performance data in a normal state, calculate the score for each of the plurality of storage devices from the obtained performance data. Using the scores for each of the plurality of storage devices calculated above, detect an abnormality in the storage device based on the difference in scores between the storage devices. Output the result of the detection. An abnormality detection method.
10. A process of obtaining performance data for each of a plurality of storage devices having a redundant configuration, a process of calculating the score for each of the plurality of storage devices from the obtained performance data using a calculation model that calculates a score indicating the difference between the obtained performance data and the performance data in a normal state, a process of detecting an abnormality in the storage device based on the difference in scores between the storage devices using the scores for each of the plurality of storage devices calculated above, and a process of outputting the result of the detection An abnormality detection program for causing a computer to execute.
Citation Information
Patent Citations
Failure detection device, information processing method, and program
JP2012108708A
Failure symptom detection device, method and storage system
JP2017037694A
Failure prediction device, failure prediction method, and failure prediction program
JP2019164817A
Storage system and control method thereof
JP2021043891A