An abnormal analysis method, device and computer-readable storage medium
By using the target exception classification model and feature indicators for inference processing in cloud computing scenarios, the challenge of root cause analysis of storage volume performance abnormalities is solved, and efficient and accurate analysis results are achieved.
Patent Information
- Application Number
- CN202111666129.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-31
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-12-31
AI Technical Summary
In cloud computing scenarios, the root cause analysis of abnormal storage volume performance is very challenging, and it is difficult for the existing technology to effectively analyze abnormal storage volume performance caused by many indicators, large amount of data and changes in business scenarios.
An abnormality analysis method is provided, by obtaining the target abnormality classification model and target characteristic indicators, using these models and indicators for inference processing, and determining the abnormality analysis results, thereby achieving effective analysis of the root cause of storage volume performance abnormalities.
This method can effectively analyze the abnormal root causes of storage volume performance, improve analysis quality, reduce inference cost, speed up computing speed, and provide more reference analysis results.
Smart Images

Figure CN114416410B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data analysis technology, and in particular, to an anomaly analysis method, device, and computer-readable storage medium. Background Art
[0002] In a cloud computing scenario, users can directly experience the performance of storage volumes in virtual machines. Since there are a large number of storage volumes in the cloud computing platform and many factors can trigger abnormal storage volume performance, it is very challenging to analyze the root cause of abnormal storage volume performance.
[0003] In related technologies, methods such as fault tree analysis, rule engines, and anomaly detection are usually used to analyze the root cause of abnormal storage volume performance. However, there are many metrics related to storage volume performance, a large amount of data, and the business scenario is constantly changing. Therefore, it is difficult for such methods in related technologies to effectively analyze the root cause of abnormal storage volume performance. Summary of the Invention
[0004] To solve the above technical problems, embodiments of this application are expected to provide an anomaly analysis method, device, and computer-readable storage medium that can effectively analyze the root cause of abnormal storage volume performance.
[0005] The technical solution of this application is implemented as follows:
[0006] Embodiments of this application provide an anomaly analysis method, including:
[0007] Obtain a target anomaly classification model and acquire at least one target feature metric;
[0008] Based on the relationship between the number of the at least one target feature metric and a first preset threshold, determine an anomaly analysis processing procedure;
[0009] When it is determined that the number of the at least one target feature metric reaches the first preset threshold, determine that the anomaly analysis processing procedure is to perform an inference processing procedure using the target anomaly classification model;
[0010] Obtain data to be analyzed, and acquire target metric values corresponding to the target feature metric from the data to be analyzed. The data to be analyzed includes multiple feature metrics and metric values corresponding to the multiple feature metrics;
[0011] Input the target feature metric and the target metric values corresponding to the target feature metric into the target anomaly classification model for inference processing, and obtain at least one decision path and inference results corresponding to the at least one decision path;
[0012] Based on the decision path and the inference results corresponding to the decision path, determine an anomaly analysis result.
[0013] An embodiment of the present application provides an abnormal analysis device, including:
[0014] A memory for storing executable abnormal analysis instructions;
[0015] A processor for implementing the abnormal analysis method provided by the embodiment of the present application when executing the executable abnormal analysis instructions stored in the memory.
[0016] An embodiment of the present application provides a computer-readable storage medium storing executable abnormal analysis instructions, which are used to cause a processor to implement the abnormal analysis method provided by the embodiment of the present application when executed.
[0017] An embodiment of the present application provides an abnormal analysis method, device and computer-readable storage medium. With the technical solution, first, a target abnormal classification model is obtained, and at least one target feature index is acquired. Then, based on the relationship between the number of at least one target feature index and a first preset threshold, an abnormal analysis processing procedure is determined. When it is determined that the number of at least one target feature index reaches the first preset threshold, the abnormal analysis processing procedure is determined to be an inference processing procedure using the target abnormal classification model. Next, data to be analyzed is acquired, target index values corresponding to the target feature index are obtained from the data to be analyzed, and the target feature index and the target index values corresponding to the target feature index are input into the target abnormal classification model for inference processing to obtain at least one decision path and at least one inference result corresponding to the decision path. Finally, based on the decision path and the inference result corresponding to the decision path, an abnormal analysis result is determined. In this way, by inputting the target feature index and the target index values corresponding to the target feature index into the target abnormal classification model for analysis processing, and based on the output result of the target abnormal classification model, an analysis result of the storage volume performance abnormality is obtained, thereby effectively analyzing the root cause of the storage volume performance abnormality. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a schematic flowchart of an abnormal analysis method provided by an embodiment of the present application;
[0019] Figure 2 It is a schematic flowchart of a method for determining an abnormal analysis result provided by an embodiment of the present application;
[0020] Figure 3 It is a schematic flowchart of another method for determining an abnormal analysis result provided by an embodiment of the present application;
[0021] Figure 4 It is a schematic flowchart of still another method for determining an abnormal analysis result provided by an embodiment of the present application;
[0022] Figure 5Schematic flowchart of a method for determining abnormal root cause indicators provided by an embodiment of the present application;
[0023] Figure 6 Schematic flowchart of a method for obtaining a target abnormal classification model provided by an embodiment of the present application;
[0024] Figure 7 Schematic diagram of an importance evaluation value of training feature indicators provided by an embodiment of the present application;
[0025] Figure 8 Flowchart of a method for obtaining an initial sorting result provided by an embodiment of the present application;
[0026] Figure 9 Schematic flowchart of a method for analyzing the root cause of storage volume performance anomalies provided by an embodiment of the present application;
[0027] Figure 10 Schematic diagram of the principle of a method for analyzing the root cause of storage volume performance anomalies provided by an embodiment of the present application;
[0028] Figure 11 Schematic diagram of the structure of an abnormal analysis device provided by an embodiment of the present application. Detailed implementation manners
[0029] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application.
[0030] In order to make the purpose, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. The described embodiments should not be construed as limitations on the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.
[0031] In the following descriptions, "some embodiments / other embodiments" are mentioned, which describe subsets of all possible embodiments. However, it can be understood that "some embodiments / other embodiments" can be the same subsets or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0032] In the following descriptions, the terms "first / second / third / fourth" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third / fourth" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0033] In a cloud computing platform, a storage volume can represent a partition of a dynamic disk of a cloud server, which is used to store various data information during the cloud computing process and provide support for the efficient and orderly progress of cloud computing. Therefore, it is particularly important to ensure good performance of the storage volume.
[0034] However, due to various influencing factors, such as the overall performance of the machine decreasing due to computing resource competition, compressed or uncompressed volumes, performance factors of the host machine, performance factors of the fiber optic switch, performance factors of the Storage Area Network (SAN), performance factors of the SAN Volume Controller (SVC), performance factors of the Network Attached Storage (NAS) or Ceph (a distributed file system), etc., the performance of the storage volume may become abnormal. When the performance of the storage volume becomes abnormal, it is very difficult to locate the abnormality. At the same time, there are a large number of storage volumes in the cloud platform, making the root causes of the abnormal performance of the storage volume numerous and complex. Therefore, for the abnormal analysis of such high-dimensional data, an efficient enough analysis method is required to obtain a high analysis quality.
[0035] In the related art, the fault tree analysis method, rule engine method, anomaly detection method, etc. are usually used to analyze the root cause of the abnormal performance of the storage volume. However, for a large amount of abnormal analysis data, these methods are difficult to ensure a high analysis quality and obtain effective analysis results.
[0036] The embodiment of the present application provides an abnormal analysis method for analyzing the root cause of the abnormal performance of the storage volume, which can effectively analyze the root cause of the abnormal performance of the storage volume. Next, the abnormal analysis method provided by the embodiment of the present application will be described as Figure 1 shown in the schematic flowchart of an abnormal analysis method provided by an embodiment of the present application. The method includes the following steps:
[0037] S101. Obtain a target abnormal classification model and acquire at least one target feature index.
[0038] It should be noted that the target abnormal classification model can be obtained after training the abnormal classification model. The abnormal classification model can be a classification model constructed based on the random forest algorithm. In such a model, there is one or more decision trees, which can classify each index affecting the performance of the storage volume. The classification result can include abnormal or normal, indicating whether it is an index that causes the abnormal performance of the storage volume. Based on the classification results of one or more decision trees, a final judgment result can be obtained, so as to determine one or more indexes affecting the performance of the storage volume.
[0039] It is understandable that the anomaly classification model built based on the random forest or decision tree algorithm has the characteristics of strong interpretability, which can effectively analyze the decision path, thereby realizing effective analysis of storage volume performance.
[0040] It should be noted that the characteristic index may be an index used to analyze the index that affects the performance of the storage volume, such as the input / output operations per second (IOPS), bandwidth, latency, etc. The target characteristic index may be a plurality of indicators selected from the characteristic indicators, and these selected indicators have a greater impact on the performance of the storage volume. In the sorting results of the indicators that affect the performance of the storage volume, the target characteristic index is ranked first. In the process of training the anomaly classification model, the indicator that has a significant impact on the performance of the storage volume can be determined, thereby obtaining the target characteristic index among the plurality of indicators.
[0041] S102: Determine an abnormality analysis process based on a relationship between the quantity of at least one target characteristic indicator and a first preset threshold.
[0042] It should be noted that the first preset threshold value can be any pre-set positive integer, for example, the first preset threshold value can be set to 5, and the analysis and processing method to be performed can be determined according to the number of target feature indicators. The abnormal analysis and processing process can include performing an inference processing process using a target abnormal classification model and directly performing anomaly detection on target feature data.
[0043] S103: Determine whether the number of at least one target characteristic indicator reaches a first preset threshold, and determine that the abnormality analysis processing process is an inference processing process performed using a target abnormality classification model.
[0044] In some embodiments, the number of target characteristic indicators reaching the first preset threshold can be understood as satisfying the first preset threshold condition, or can be understood as the number of target characteristic indicators being greater than or equal to the first preset threshold. When the number of target characteristic indicators is greater than or equal to the first preset threshold, it means that the number of target characteristic indicators is large, and the indicators or indicator combinations that affect the storage volume performance abnormality cannot be directly determined. The target characteristic indicators can be inferred based on the target abnormality classification model.
[0045] S104: Acquire the data to be analyzed, and obtain the target indicator value corresponding to the target characteristic indicator from the data to be analyzed.
[0046] The data to be analyzed includes multiple characteristic indicators and the corresponding indicator values of the multiple characteristic indicators. The data to be analyzed can be test data for performing abnormal analysis of storage volume performance. The corresponding indicator value of a characteristic indicator can be the value of each indicator. For example, the corresponding indicator value of the indicator read latency is 0.3 milliseconds. Correspondingly, the target indicator value can be the value corresponding to the target characteristic indicator. Since the target characteristic indicator has been obtained during the training process of the abnormal classification model, it is only necessary to select the value corresponding to the target characteristic indicator from the data to be analyzed.
[0047] It can be understood that, based on the target characteristic indicator and the target indicator value corresponding to the target characteristic indicator obtained from the data to be analyzed, as the analysis data to be input into the target abnormal classification model, rather than all the data to be analyzed, the reduction of the data volume is achieved, so as to reduce the inference cost and speed up the operation speed when performing inference based on the target abnormal classification model subsequently.
[0048] S105. Input the target characteristic indicator and the target indicator value corresponding to the target characteristic indicator into the target abnormal classification model for inference processing, and obtain at least one decision path and the inference results corresponding to at least one decision path.
[0049] In some embodiments, the target characteristic indicator and the target indicator value corresponding to the target characteristic indicator can be used as inference data. The inference processing can be to analyze and process the inference data using the target abnormal classification model, including performing classification decisions on each target characteristic indicator and the target indicator value corresponding to each target characteristic indicator, and performing path decision analysis on the impact of storage volume performance determined based on multiple target characteristic indicators and the target indicator values corresponding to multiple target characteristic indicators.
[0050] It should be noted that the decision path can be one or more paths in a decision tree or a random forest. One decision path includes one or more nodes, and each node includes a target characteristic indicator and the target indicator value corresponding to the target characteristic indicator. The inference result corresponding to the decision path indicates a normal state or an abnormal state. When the inference result of a certain decision path is abnormal, it means that the target characteristic indicator and the target indicator value corresponding to the target characteristic indicator in this path cause the abnormal performance of the storage volume.
[0051] S106. Determine the abnormal analysis result based on the decision path and the inference result corresponding to the decision path.
[0052] The abnormal analysis result indicates the analysis result of the abnormal performance of the storage volume obtained after further analyzing the decision path and the inference result corresponding to the decision path.
[0053] In the embodiments of the present application, a target anomaly classification model is obtained, and at least one target feature index is acquired. Based on the relationship between the number of at least one target feature index and a first preset threshold, an anomaly analysis and processing process is determined. When it is determined that the number of at least one target feature index reaches the first preset threshold, the anomaly analysis and processing process is determined to be an inference processing process using the target anomaly classification model. Then, the data to be analyzed is obtained, and the target index value corresponding to the target feature index is acquired from the data to be analyzed. The target feature index and the target index value corresponding to the target feature index are input into the target anomaly classification model for inference processing, and at least one decision path and the inference result corresponding to the decision path are obtained. Finally, based on the decision path and the inference result corresponding to the decision path, an anomaly analysis result is determined. In this way, by inputting the target feature index and the target index value corresponding to the target feature index into the target anomaly classification model for analysis and processing, and based on the output result of the target anomaly classification model, an analysis result of the storage volume performance anomaly is obtained, thereby realizing an effective analysis of the root cause of the storage volume performance anomaly.
[0054] As Figure 2 shown, it is a schematic flowchart of a method for determining an anomaly analysis result provided by an embodiment of the present application. In some embodiments of the present application, based on the decision path and the inference result corresponding to the decision path, an anomaly analysis result is determined, that is, S106 can be implemented through S201 to S204 described below. Each step is described below.
[0055] S201. Based on the inference results corresponding to each decision path, determine the target decision path in at least one decision path.
[0056] It should be noted that each decision path has its own corresponding inference result. The inference result may include a normal state and an abnormal state. The target decision path is the decision path with an abnormal state as the inference result. In some embodiments, the target decision path may be one, and all the target indicators and the target index values corresponding to the target feature indicators in this target decision path cause the storage volume performance anomaly.
[0057] S202. Acquire at least one target index value corresponding to each target feature indicator in the target decision path.
[0058] After determining the target decision path, it is necessary to determine that the target feature indicator and the target index value corresponding to the target feature indicator in the target decision path are the influencing factors of the storage volume performance. First, the target feature indicator and the target index value corresponding to the target feature indicator in the target decision path can be acquired.
[0059] S203. Based on at least one target index value, detect one or more target feature indicators in the target decision path to obtain a detection result.
[0060] In some embodiments, after obtaining the target feature indicators in the target decision path and the target indicator values corresponding to the target feature indicators, one or more of these target feature indicators and the target indicator values corresponding to these target feature indicators can be detected in real time. For example, set the target feature indicators regarding the storage volume performance as three target feature indicators in the target decision path, and the target indicator values corresponding to the three target feature indicators respectively. In the observed time slice, determine the real-time state of the storage volume performance, so as to obtain the detection results of these three target feature indicators.
[0061] S204. Determine the abnormal analysis result based on the detection result.
[0062] It should be noted that the detection result can be that in the observed time slice, the storage volume performance is in an abnormal state or a normal state, and the abnormal analysis result can be the target feature indicator corresponding to the case where the detection result indicates that the storage volume is in an abnormal state.
[0063] In some embodiments of the present application, one or more target feature indicators in the target decision path are detected based on at least one target indicator value to obtain a detection result, and an abnormal analysis result is determined based on the detection result. That is, S203 and S204 can be further implemented by the following method.
[0064] Determine the current state of the storage volume performance based on one or more target feature indicators in the target decision path and the target indicator values corresponding to the one or more target feature indicators.
[0065] It should be noted that the current state of the storage volume performance can be the state of the storage volume performance obtained in real time in the observed time slice, including the normal state and the abnormal state. One or more target feature indicators in the target decision path can be any selected single target feature indicator or a combination of multiple target feature indicators. Exemplarily, the target feature indicators in the target decision path include A, B, E, F, and H. Each of these target feature indicators can be selected for detection respectively to determine the current state of the storage volume performance under this target feature indicator and the corresponding target indicator value. Or multiple combinations of target feature indicators, such as A, B, E; A, B, F; A, E, F, H, etc. can be detected respectively to obtain the current state of the storage volume performance under the combination of multiple target features and the target indicator values corresponding to the multiple target feature indicators.
[0066] Determine the current state of the storage volume performance as the detection result. If the current state of the storage volume performance is an abnormal state, mark one or more target feature indicators, and determine the marked one or more target feature indicators as the abnormal analysis result. The marked one or more target feature indicators indicate the indicators that affect the storage volume performance determined through real-time detection.
[0067] As shown Figure 3 in FIG. is a schematic flow chart of another method for determining an abnormal analysis result provided by an embodiment of the present application. In some embodiments of the present application, the abnormal analysis result further includes a target sorting result representing the importance of a target feature index. At this time, the method for determining the abnormal analysis result is further implemented through the following S301 to S303.
[0068] S301. Obtain an initial sorting result of the target abnormal classification model for the target feature index.
[0069] It should be noted that the initial sorting result may be a result of sorting according to the importance of the target feature index to the storage volume performance, and may be a result of sorting the target feature indexes from high to low according to the importance to obtain the initial sorting result.
[0070] In some embodiments, the initial sorting result may be obtained during the training completion process of the abnormal classification model. The initial sorting result may be obtained based on a feature importance measurement method of a random forest or a decision tree. For example, through the method of out-of-bag data detection, obtain the importance measurement value of each target feature index, and determine the initial sorting result representing the importance of the target feature index based on the magnitudes of the importance measurement values corresponding to all target feature indexes.
[0071] S302. Based on the target decision path and the initial sorting result, determine the target sorting result of the target feature index.
[0072] It should be noted that the target sorting result may be a sorting result of all target feature indexes in the target decision path. In some embodiments, the target sorting result may be obtained by deleting the target feature indexes not in the target decision path based on the initial sorting result.
[0073] In some embodiments, the initial sorting result is the same as the target sorting result. In this case, the target decision path includes all target feature indexes, and no target feature indexes in the initial sorting result need to be deleted.
[0074] In some other embodiments, the initial sorting result is different from the target sorting result. In this case, the number of types of target feature indexes in the target decision path is less than the number of types of target feature indexes in the initial sorting result. For example, the initial sorting result of the target feature indexes is B, F, A, C, D, and the target feature indexes in the target decision path are A, B, D, F, missing the feature index C. Then the target sorting result is B, F, A, D.
[0075] S303. Determine the target sorting result as the abnormal analysis result.
[0076] The target sorting result represents the importance sorting of target feature indicators in the decision path, which can be used to characterize the importance of target feature indicators on the performance of the storage volume. Therefore, it can also be used as the abnormal analysis result of the storage volume performance.
[0077] As Figure 4 shown, it is a schematic flowchart of another method for determining the abnormal analysis result provided by the embodiment of the present application. In some embodiments of the present application, the method for determining the abnormal analysis result can also be implemented through the following S401 to S403.
[0078] S401. Obtain the number of decision trees in the random forest determined based on the target abnormal classification model.
[0079] It should be noted that the abnormal classification model is constructed based on the random forest model. When the training of the abnormal classification model is completed, the target abnormal classification model is obtained. Therefore, at this time, the corresponding random forest structure in the target abnormal classification model has been determined. Correspondingly, the number of decision trees corresponding to the random forest can be obtained accordingly. In some embodiments, the number of decision trees corresponding to the random forest can be one or more.
[0080] S402. Determine that the number of at least one target feature indicator does not reach the first preset threshold, and the number of decision trees is greater than the second preset threshold, and determine that the abnormal analysis processing process is the target feature abnormal detection processing process.
[0081] In some embodiments, the number of target feature indicators not reaching the first preset threshold can be understood as not meeting the first preset threshold condition, or it can be understood that the number of target feature indicators is less than the first preset threshold. When the number of target feature indicators is less than the first preset threshold, it means that the number of target feature indicators is small.
[0082] Furthermore, the second preset threshold can be any positive integer set in advance, which is used to represent the number of decision trees. When the number of decision trees determined in S401 is less than the second preset threshold, it means that the random forest structure determined by the target feature indicator and the target indicator value corresponding to the target feature indicator is simple, that is, the state of the storage volume can be determined by a small number of target feature indicators and the target indicator values corresponding to the target feature indicators. At this time, there is no need to use the target abnormal classification model to perform the inference processing process on the target feature indicators, and directly perform the abnormal detection processing on the target features.
[0083] S403. Perform abnormal detection on one or more of the target feature indicators to obtain a detection result, and determine the abnormal analysis result based on the detection result.
[0084] In some embodiments, the execution process of S403 is similar to that of S203, except that S403 directly performs anomaly detection on the target feature indicators, rather than on the target feature indicators in the target decision path. Exemplarily, five target feature indicators regarding the storage volume performance are set, as well as the target indicator values respectively corresponding to the five target feature indicators. In the observed time slice, the real-time state of the storage volume performance is determined, so as to obtain the detection results of the five target feature indicators.
[0085] As Figure 5 shown, it is a schematic flowchart of a method for determining the anomaly root cause indicator provided by an embodiment of the present application. In some embodiments of the present application, after performing anomaly detection on the target feature indicators and marking the target features corresponding to the detection results indicating that the storage volume performance is in an abnormal state, the method for further determining the anomaly root cause indicator can be implemented through the following S501 to S504.
[0086] S501. Obtain the normal indicator ranges corresponding to at least one target feature indicator, and determine the target normal values respectively corresponding to at least one target feature indicator based on the normal indicator ranges corresponding to at least one target feature indicator.
[0087] It should be noted that the normal indicator ranges corresponding to the target feature indicators can be determined during the training process of the anomaly classification model. For example, through the analysis of the training process, the target feature indicator read I / O count (read_ios) is obtained, and the normal indicator range corresponding to this target feature indicator is 100 to 200 times per second. The target normal value can be any value within the normal indicator range, such as the upper limit, lower limit, median, or mean of the target normal indicator range. Exemplarily, when the target normal indicator range of the target feature indicator read_ios is [100, 200], the target normal value can be the upper limit 200 of this target normal indicator range, the lower limit 100 of this target normal indicator range, the median 150 of this target normal indicator range, or the mean 150 of this target normal indicator range.
[0088] S502. Update the target indicator values corresponding to the marked target feature indicators to the corresponding target normal values.
[0089] In some embodiments, after detecting the target feature indicators, one or more marked target feature indicators can be obtained based on the detection results, and all the marked target feature indicators are respectively updated to their corresponding target normal values. For example, the target indicator value corresponding to the marked target feature indicator read latency is 50 milliseconds, the normal indicator range of this indicator is [20, 42], and the target normal value determined based on this normal indicator range is 31 milliseconds. Then, after the update, the target indicator value corresponding to the target feature indicator read latency is 31 milliseconds.
[0090] S503. Inferentially process the labeled target feature indicators and the corresponding target normal values of the labeled target feature indicators based on the target anomaly classification model to obtain the corrected inference results corresponding to the labeled target feature indicators.
[0091] In some embodiments, the corrected inference result may represent the inference result obtained after updating the target index value corresponding to the target feature indicator and then inputting it into the target anomaly classification model for inferential processing. Inferentially processing the labeled target feature indicators and the corresponding target normal values of the labeled target feature indicators may be to perform inferential processing on each labeled target feature indicator and the corresponding target normal value of the labeled target feature indicator separately. Exemplarily, each labeled target feature indicator, the corresponding target normal value of the labeled target feature indicator, other target feature indicators except the labeled target feature indicator, and the corresponding target index values of the other target feature indicators are sequentially input into the target anomaly classification model for inferential processing to obtain the corrected inference results corresponding to each labeled target feature indicator.
[0092] S504. Determine the anomaly root cause indicators from the labeled target feature indicators based on the corrected inference results corresponding to the labeled target feature indicators.
[0093] It should be noted that the corrected inference result may be an abnormal state or a normal state. The anomaly root cause indicator may be an indicator that causes the detection result in the observed time slice to indicate an abnormal state during the detection of the labeled target feature indicators in the target decision path.
[0094] In some embodiments, when the corrected inference result corresponding to a certain target feature indicator is in a normal state, and the absolute value of the difference between the target index value corresponding to the target feature indicator and the index value obtained after updating the target index value is less than a preset threshold, it can be determined that the labeled target feature indicator causes the analysis indicator to be abnormal in this time slice. Exemplarily, for example, the storage volume storage IO latency is high because the CPU load of the SVC storage device is too high.
[0095] It can be understood that by updating the target index value corresponding to the labeled target feature indicator to ensure that the updated index value is within the normal index range corresponding to the target feature indicator, based on the corrected inference result of the labeled target feature indicator and the difference relationship between the target index value corresponding to the labeled target feature indicator and the updated target normal value, the root cause indicator that causes the detection result to be in an abnormal state in a certain observed time slice is determined.
[0096] Such as Figure 6As shown in the figure, it is a schematic flowchart of a method for obtaining a target anomaly classification model provided by an embodiment of the present application. In some embodiments of the present application, the anomaly analysis method may further include obtaining a target anomaly classification model, and the method for obtaining the target anomaly classification model can be implemented through the following S601 to S605. The following explains each step.
[0097] S601. Obtain training data and a preset anomaly classification model.
[0098] It should be noted that the preset anomaly classification model can be a decision tree model or a random forest model, and the training data can be data sources obtained from a cloud computing database, a monitoring database, or other databases, including storage volumes and attributes related to the performance of storage volumes, such as resource dependency relationships, key attributes, performance metrics, and anomaly event definitions.
[0099] In some embodiments, the training data includes training feature indicators, training feature values corresponding to the training feature indicators, and training feature labels. The training feature indicators can be indicators related to the performance of storage volumes, such as the number of read I / Os, read latency, and write latency, etc. The training feature values corresponding to the training feature indicators can be the specific values corresponding to the training feature indicators. For example, the training feature value of the training feature indicator read latency is 10 milliseconds. The training feature labels can be the states of the storage volumes, such as normal state and abnormal state.
[0100] S602. Train the preset anomaly classification model based on the training data to obtain an initial anomaly classification model.
[0101] In some embodiments, before training the preset anomaly classification model, it is necessary to establish a positive and negative sample data set for anomaly root cause analysis according to a certain ratio relationship. The positive sample data set can be a data set composed of data with normal training feature indicators or normal indicator values corresponding to the training feature indicators. Correspondingly, the negative sample data set can be a data set composed of data with abnormal training feature indicators or data with normal training feature indicators but abnormal indicator values corresponding to the training feature indicators.
[0102] It should be noted that a certain ratio relationship needs to be followed when establishing the positive and negative sample data sets for anomaly root cause analysis. The ratio relationship can be the ratio of the positive and negative sample data sets. For example, the ratio of the positive sample data set to the negative sample data set is 1:1.
[0103] It can be understood that maintaining a proper ratio of positive and negative sample datasets in the training data enables a relatively high-quality initial anomaly classification model to be obtained when using positive and negative samples to train a preset anomaly classification model. As a result, when obtaining a target anomaly classification model based on the initial anomaly classification model and using the target anomaly classification model to infer the data to be analyzed, a more accurate inference result can be obtained. Further, when using the inference result for anomaly analysis, the correctness of the storage volume performance anomaly analysis can be improved.
[0104] It should be noted that in the training data, the state of the storage volume determined by the training feature index and the training feature value corresponding to the training feature index has been determined, that is, the training label is known. During the training of the anomaly classification model, the model parameters in the preset anomaly classification model are continuously adjusted through the training label, so that the correct training label can be obtained based on the training feature index and the training feature value corresponding to the training feature index. After adjusting the model parameters, when the states of all the storage volumes determined by the training feature index and the training feature value corresponding to the training feature index are consistent with the training label, the training of the preset anomaly classification model is completed, and an initial anomaly classification model is obtained.
[0105] S603. Determine the importance evaluation value of each training feature index based on the initial anomaly classification model.
[0106] In some embodiments, the importance evaluation value can represent the degree of importance of a certain training feature index affecting the storage volume performance. After obtaining the initial anomaly classification model, the importance evaluation value of each training feature index can be calculated in turn.
[0107] Exemplarily, the importance evaluation value of the training feature index can be calculated based on the out-of-bag data detection method. For example, when calculating the importance evaluation value of a certain training feature index, first calculate the first inference error e1 obtained after inputting other indexes except this training feature index into the initial anomaly classification model. Then, modify the training feature value corresponding to this training feature index to other index values outside the normal index range, and input this training feature index and the modified index value corresponding to this training feature index into the initial anomaly classification model to obtain the second inference error e2. Assuming that there are N trees in the random forest corresponding to the initial anomaly classification model, the importance evaluation value α = (e2 - e1) / N.
[0108] S604. Delete the training feature indexes whose importance evaluation values are lower than the fourth preset threshold to obtain multiple target feature indexes.
[0109] It should be noted that the fourth preset threshold can be any real number set in advance. For example, -3.2, 0, 5, etc. If the importance evaluation value corresponding to a certain training feature index calculated is less than the fourth preset threshold, it is determined that the training feature index is a training feature index with relatively low importance and is deleted. On the contrary, if the importance evaluation value corresponding to a certain training feature index calculated is greater than or equal to the fourth preset threshold, it is determined that the training feature index is a training feature index with relatively high importance and is retained. The importance evaluation values corresponding to all training feature indexes are respectively compared with the fourth preset threshold to obtain the target feature indexes corresponding to the importance evaluation values greater than or equal to the fourth preset threshold.
[0110] Exemplarily, as Figure 7 shown, it is a schematic diagram of the importance evaluation value of a training feature index provided by an embodiment of the present application. Figure 7 In it, the abscissa represents each training feature index, and the ordinate represents the importance evaluation value. By calculating the importance evaluation values of 10 training feature indexes A, B, C, D, E, F, G, H, I, J twice, it can be seen that the importance evaluation values of the training feature indexes G, H, I, J are all close to 0. Therefore, in practice, if the fourth preset threshold is set to 0.3, the training feature indexes G, H, I, J can be deleted to obtain the target feature indexes A, B, C, D, E, F.
[0111] S605. Input the target feature indexes in the training data and the index values corresponding to the target feature indexes in the training data into the initial anomaly classification model, and continue to train the initial anomaly classification model until the target anomaly classification model is obtained.
[0112] After obtaining the target feature indexes in the training feature indexes, input the target feature indexes in the training data and the index values corresponding to the target feature indexes into the initial anomaly classification model for training. According to the output result of the initial anomaly classification model, continuously adjust the model parameters in the model, and finally obtain the trained anomaly classification model, that is, the target anomaly classification model.
[0113] It can be understood that after obtaining the initial anomaly classification model through training based on the preset anomaly classification model, deleting the training feature indexes with relatively low importance to obtain the target feature indexes makes the size of the model obtained when training the initial anomaly classification model based on the training target feature indexes and the index values corresponding to the target feature indexes smaller, reducing the training cost of the target anomaly classification model.
[0114] Such as Figure 8As shown in the figure, it is a flowchart of a method for obtaining an initial sorting result provided by an embodiment of the present application. In some embodiments of the present application, after deleting training feature indicators with importance evaluation values lower than the fourth preset threshold and obtaining multiple target feature indicators, that is, after S604, an initial sorting result can also be obtained. The method for obtaining the initial sorting result can be implemented through the following S701 to S702.
[0115] S701. Obtain the importance evaluation values of each target feature indicator.
[0116] It should be noted that the importance evaluation values of each training feature indicator have been determined in S603, and the target feature indicators in the training feature indicators have been obtained in S604. Therefore, by only selecting the importance evaluation values corresponding to the target feature indicators from the importance evaluation values of the training feature indicators, the importance evaluation values of each target feature indicator can be obtained.
[0117] S702. Sort the multiple target feature indicators based on the importance evaluation values of each target feature indicator to obtain the initial sorting result of the target feature indicators.
[0118] In some embodiments, sorting the multiple target feature indicators based on the importance evaluation values of each target feature indicator can be according to the magnitudes of the importance evaluation values of each target feature indicator. For example, sorting from largest to smallest according to the importance evaluation values, so as to obtain the initial sorting result corresponding to the target feature indicator corresponding to the importance evaluation value.
[0119] Exemplarily, assume that the importance evaluation values of target feature indicators A, B, C, D, and E are 0.2, 0.1, 1.6, 3.5, and 0.8 respectively. Arrange the importance evaluation values corresponding to each target feature indicator from largest to smallest as: 3.5, 1.6, 0.8, 0.2, 0.1. Then, according to the magnitudes of the importance evaluation values of each target feature indicator, the initial sorting result of the target feature indicators is: D, C, E, A, B.
[0120] It can be understood that after sorting the target feature indicators based on the importance evaluation values of the target feature indicators to obtain the initial sorting result corresponding to the target feature indicators, the feature indicators with a greater impact on the storage volume performance can be directly obtained. When inferring the data to be analyzed later, only the target feature indicators and the target indicator values corresponding to the target feature indicators in the initial sorting are selected, reducing the inference cost of the target anomaly classification model. In addition, based on the initial sorting result and the inference result of the target anomaly classification model, the importance of the target feature indicators can be further sorted to obtain the analysis result affecting the storage volume performance.
[0121] Next, the implementation process of the embodiment of the present application in an actual application scenario will be introduced.
[0122] In some embodiments, as Figure 9 shown, it is a schematic flowchart of a method for analyzing the root cause of storage volume performance anomalies provided by an embodiment of the present application. The method for analyzing the root cause of storage volume performance anomalies provided by an embodiment of the present application can be implemented through the following S801 to S808, including a classification model training process S801 to S803 and a classification model inference process S804 to S808. The following explains each step.
[0123] S801. Obtain data sources related to the storage volume.
[0124] In some embodiments, as Figure 10 shown, it is a schematic diagram of a method for analyzing the root cause of storage volume performance anomalies provided by an embodiment of the present application. The method for analyzing the root cause of storage volume performance anomalies provided by an embodiment of the present application can be that after the device performance has an anomaly, the operation and maintenance engineer conducts the work of analyzing the root cause of the storage volume performance anomaly, or the scheduled job starts to execute the analysis of the root cause of the storage volume performance anomaly.
[0125] When obtaining data sources related to the storage volume, different resource dependency relationships can be docked with the cloud computing database to obtain volume configuration information, such as compressed volumes and business attributes, etc., and data in the monitoring database can be obtained. The data in the monitoring database includes anomaly event information, such as monitoring metric ranges, metric calculation methods, alarm thresholds, and alarm levels, etc.
[0126] S802. Construct training data based on positive and negative samples in the data source.
[0127] In some embodiments, as Figure 10 shown, a positive and negative sample data set for root cause analysis of anomalies can be established according to a certain positive and negative sample ratio. In practice, training data for a period of time is sampled according to a certain positive and negative sample ratio. For example, the positive and negative sample ratio or ratio range is configured according to business experience or system built-in, so as to ensure that the positive and negative sample ratio is not too disparate.
[0128] S803. Train a classifier model (preset anomaly classification model) based on the training data to generate a target classifier (target anomaly classification model), and obtain the feature importance ranking result (initial ranking result).
[0129] In some embodiments, before performing model training, as Figure 10 shown, a classifier model can be established based on a random forest model, the classifier model is trained using the training data, the quality of the classifier is measured based on the training result of the classifier model, and the parameters of the classification model are continuously adjusted to obtain the target classifier.
[0130] During the training process, if the quality of the classifier model is good enough, analyze the importance indicators of the model, remove the indicators with lower importance to the model (delete the training feature indicators whose importance evaluation values are lower than the fourth preset threshold), and retrain the model (continue to train the initial anomaly classification model) to reduce the model size and the training and inference costs.
[0131] S804. Use the target classifier to perform inference on the inference data (perform inference processing on the target anomaly classification model) to obtain the inference result (the inference result corresponding to at least one decision path) and the decision path (at least one decision path).
[0132] In some embodiments, the inference data (the target feature indicator and the target indicator value corresponding to the target feature indicator) can be test data used to perform inference analysis on the storage volume performance to determine whether the inference data causes the storage volume anomaly. The inference result can be the result obtained after the target classifier performs inference on the inference data, including the abnormal state and the normal state.
[0133] The decision path can be the multiple decision paths corresponding to the respective decision trees of the random forest in the target classification model. In implementation, the random forest decision path of the inference data can be obtained based on the combination strategy of the random forest. For example, in the random forest, if the inference data in a decision path a in decision tree A is the same as the inference data in a decision path b in another decision tree B, only one decision path is retained.
[0134] S805. Obtain the effective decision path (the target decision path) based on the inference result and the decision path.
[0135] In some embodiments, the random forest may include multiple decision trees, and each decision tree has multiple decision paths. At this time, the effective decision trees in the random forest can be obtained based on the inference result of the inference data, and the effective decision paths in the effective decision trees can be obtained (determine the target decision path in the at least decision path based on the inference results corresponding to each decision path). The effective decision tree can be the decision tree corresponding to the abnormal state of the inference result, and the effective decision path can be the decision path corresponding to the abnormal state in the effective decision tree.
[0136] In some other embodiments, there may be only one decision tree in the random forest. At this time, the decision path whose inference result of the inference data is the abnormal state can be directly found in this decision tree, and the decision path corresponding to the abnormal state of the inference result is used as the effective decision path.
[0137] S806. Obtain the target feature importance ranking result (the target ranking result of the target feature indicators) based on the feature importance ranking result of the model and the index range in the effective decision path (one or more target feature indicators in the target decision path).
[0138] It should be noted that the index range can be a set of all feature indicators (target feature indicators) in the effective decision path. Through the effective decision path, a feature attribute set can be obtained, and the feature attribute set includes feature indicators and the corresponding index values of the feature indicators (the target index values corresponding to one or more target feature indicators in the target decision path). In practice, based on the feature importance ranking result and all feature indicators in the effective decision path, the feature indicators that do not exist in the effective decision path in the feature importance ranking result can be removed, so as to obtain the sorted target feature importance ranking result.
[0139] S807. Perform anomaly detection on one or more feature indicators in the index range (detect one or more target feature indicators in the target decision path) to obtain a detection result.
[0140] In some embodiments, perform anomaly detection on each feature indicator in the effective decision path, or perform anomaly detection after combining multiple feature indicators in the effective decision path. If the data of the observed time slice belongs to the abnormal state, label the feature indicator or the combination of feature indicators (if the current state of the storage volume performance is the abnormal state, mark one or more target feature indicators), and use the labeled feature indicator or the combination of feature indicators as the detection result (the one or more target feature indicators marked are determined as the anomaly analysis result).
[0141] S808. Output the analysis result (anomaly analysis result).
[0142] It should be noted that the analysis result can include the index range in the effective decision path (one or more target feature indicators in the target decision path), the labeled feature indicator or the combination of feature indicators (the one or more target feature indicators marked), and the target feature importance ranking (target ranking result), etc. Therefore, when actually outputting the analysis result, the index range, the labeled feature indicator or the combination of feature indicators in the anomaly detection, and the target feature importance ranking, etc. can be output in sequence.
[0143] It can be understood that the root cause analysis method for abnormal storage volume performance provided by the embodiments of the present application supports the detection of abnormal storage volume performance under high-dimensional data. Through model inference using a random forest, relatively high analysis quality can be obtained. At the same time, by leveraging the strong interpretability of the random forest to analyze the decision path and further perform abnormal detection on the decision path, the main reasons that may cause failures can be identified, the device performance problems can be more accurately located, and more reference information can be provided for operation and maintenance engineers.
[0144] The embodiments of the present application also provide an abnormal analysis device. Figure 11 As shown in the structural schematic diagram of an abnormal analysis device provided by the embodiments of the present application, Figure 11 as shown, the abnormal analysis device 1 includes: a memory 11 for storing executable abnormal analysis instructions; a processor 12 for implementing the method provided by the embodiments of the present application when executing the executable abnormal analysis instructions stored in the memory, for example, implementing the abnormal analysis method provided by the embodiments of the present application.
[0145] The embodiments of the present application provide a computer-readable storage medium storing executable abnormal analysis instructions, which are used to cause the processor 12 to implement the method provided by the embodiments of the present application when executed, for example, the abnormal analysis method provided by the embodiments of the present application.
[0146] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a device, or a computer program product. Therefore, the present application can take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories and optical memories, etc.) containing computer-usable program codes.
[0147] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices, or computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0148] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in one or more of the blocks or blocks.
[0149] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in one or more of the blocks or blocks.
[0150] As described above, the foregoing are only embodiments of the present application and are not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and scope of the present application are included within the scope of protection of the present application.
Claims
1. An anomaly analysis method, comprising: Obtaining a target anomaly classification model and acquiring at least one target feature index; The target feature index is a plurality of indexes selected from the feature indexes, and the feature indexes are indexes used to analyze the performance of the storage volume; Determining an anomaly analysis processing procedure based on the relationship between the number of the at least one target feature index and a first preset threshold; Determining that the number of the at least one target feature index reaches the first preset threshold, and determining that the anomaly analysis processing procedure is an inference processing procedure using the target anomaly classification model; Obtaining data to be analyzed, and acquiring target index values corresponding to the target feature indexes from the data to be analyzed, where the data to be analyzed includes a plurality of feature indexes and index values corresponding to the plurality of feature indexes; Inputting the target feature index and the target index value corresponding to the target feature index into the target anomaly classification model for inference processing, to obtain at least one decision path and inference results corresponding to the at least one decision path; Determining an anomaly analysis result based on the decision path and the inference result corresponding to the decision path; The determining an anomaly analysis result based on the decision path and the inference result corresponding to the decision path includes: determining a target decision path in the at least one decision path based on the inference results corresponding to each decision path, where the target decision path is a decision path with an abnormal state as the inference result; acquiring at least one target index value corresponding to each target feature index in the target decision path; detecting the at least one target feature index in the target decision path based on the at least one target index value to obtain a detection result; and determining an anomaly analysis result based on the detection result.
2. The method according to claim 1, where the detecting the at least one target feature index in the target decision path based on the at least one target index value to obtain a detection result, and determining an anomaly analysis result based on the detection result includes: Determining a current state of the storage volume performance based on the at least one target feature index in the target decision path and the target index value corresponding to the at least one target feature index; Determining the current state of the storage volume performance as the detection result; If the current state of the storage volume performance is an abnormal state, marking the at least one target feature index; Determining the marked at least one target feature index as the anomaly analysis result.
3. The method according to claim 1, where the anomaly analysis result further includes a target sorting result representing the importance of the target feature index, and the method further includes: Obtaining an initial sorting result of the target anomaly classification model for the target feature index; Determining a target sorting result of the target feature index based on the target decision path and the initial sorting result; Determining the target sorting result as the anomaly analysis result.
4. The method according to claim 1, the method further includes: Obtaining the number of decision trees in the random forest determined based on the target anomaly classification model; It is determined that the number of the at least one target feature index does not reach the first preset threshold and the number of decision trees is less than the second preset threshold, and it is determined that the abnormal analysis and processing process is a target feature abnormal detection and processing process; Perform abnormal detection on one or more of the target feature indexes to obtain a detection result, and determine an abnormal analysis result based on the detection result.
5. The method according to claim 2, the method further includes: Obtain the normal index range corresponding to the at least one target feature index, and determine the target normal value corresponding to each of the at least one target feature index based on the normal index range corresponding to the at least one target feature index; Update the target index value corresponding to each marked target feature index to the corresponding target normal value; Perform an inference process on the marked target feature index and the target normal value corresponding to the marked target feature index based on the target abnormal classification model to obtain a corrected inference result corresponding to the marked target feature index; Determine an abnormal root cause index from the marked target feature indexes based on the corrected inference result corresponding to the marked target feature index.
6. The method according to claim 1, the method further includes: Obtain training data and a preset abnormal classification model, where the training data includes training feature indexes, training feature values corresponding to the training feature indexes, and training feature labels; Train the preset abnormal classification model based on the training data to obtain an initial abnormal classification model; Determine the importance evaluation value of each training feature index based on the initial abnormal classification model; Delete the training feature indexes with importance evaluation values lower than the fourth preset threshold to obtain a plurality of target feature indexes; Input the target feature indexes in the training data and the index values corresponding to the target feature indexes in the training data into the initial abnormal classification model, and continue to train the initial abnormal classification model until a target abnormal classification model is obtained.
7. The method according to claim 6, the method further includes: Obtain the importance evaluation value of each target feature index; Sort the plurality of target feature indexes based on the importance evaluation values of the respective target feature indexes to obtain an initial sorting result of the target feature indexes.
8. An abnormal analysis device, including: A memory for storing executable abnormal analysis instructions; A processor, configured to implement the method according to any one of claims 1 to 7 when executing the executable abnormal analysis instructions stored in the memory.
9. A computer-readable storage medium, storing executable abnormal analysis instructions, configured to cause a processor to implement the method according to any one of claims 1 to 7 when executed.
Citation Information
Patent Citations
Alarm dimension mining method, device and equipment
CN110046179A
Vehicle door anomaly diagnosis method and device
CN111797944A