Anomaly detection method, electronic device, and storage medium
By acquiring cluster attribute parameters at different time periods and determining storage volumes using correlation and prediction algorithms, the problem of relying on historical sample data in the prior art is solved, and efficient and accurate storage response delay abnormality analysis is achieved.
Patent Information
- Application Number
- CN202210474144.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-29
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-04-29
AI Technical Summary
When analyzing the causes of abnormally increased storage response delay, the prior art needs to rely on a large amount of historical sample data, resulting in large fluctuations in analysis quality in newly launched systems or systems with large business changes, and low manual inspection efficiency.
By obtaining the attribute parameters of the target cluster at different time periods, using correlation algorithms and prediction algorithms to determine the storage volumes to be analyzed and abnormal storage volumes to be analyzed, narrowing the analysis scope, reducing computing resource consumption, and improving analysis efficiency.
It realizes the accurate determination of the cause of the abnormal increase in storage response delay exceptions without relying on a large amount of historical sample data, which improves the accuracy and efficiency of the analysis.
Smart Images

Figure CN115080289B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to an anomaly detection method, an electronic device, and a storage medium. Background Technique
[0002] With the rapid development of science and technology, computer application technology has become more and more widespread. To ensure the user experience effect, distributed storage has been widely applied in various industries. When evaluating storage performance, it is usually evaluated by storage response latency. Once the storage response latency increases abnormally, the performance of the services using this set of storage will be affected. Currently, excluding network factors, the most direct reason for the abnormal increase in storage response latency may be the storage capacity or a certain storage resource reaching its bottleneck, which is insufficient to support the surging service traffic, i.e., abnormal service traffic. Currently, the methods for analyzing the reasons for the abnormal increase in storage response latency caused by storage capacity mainly include: manual one-by-one troubleshooting analysis and the random forest analysis method that relies on historical data input and performs model training on it for anomaly detection.
[0003] However, the method of manual one-by-one troubleshooting analysis has a large troubleshooting scope, resulting in low troubleshooting efficiency. And the random forest analysis method requires a sufficient amount of historical sample data to perform model training for subsequent analysis. As a result, when applying the random forest method to analyze in some newly launched systems or systems with large fluctuations in business changes, the analysis quality fluctuates greatly. Summary of the Invention
[0004] To solve the above technical problems, the embodiments of this application are expected to provide an anomaly detection method, an electronic device, and a storage medium, which solve the problem that the current method for analyzing the reasons for the abnormal increase in storage response latency caused by storage capacity needs to rely on a large amount of historical sample data, and propose a method for determining abnormal storage capacity that causes the abnormal increase in storage response latency. By relying on historical sample data for analysis, the range of abnormal storage capacity is determined, ensuring the accuracy of anomaly analysis.
[0005] The technical solution of this application is implemented as follows:
[0006] In a first aspect, an anomaly detection method, the method includes:
[0007] Obtain the first cluster attribute parameters corresponding to the target cluster within a first preset time period;
[0008] Based on the first cluster attribute parameters and the reference storage volumes included in the target cluster, determine the storage volumes to be analyzed;
[0009] Obtain the second cluster attribute parameters corresponding to the target cluster within a second preset time period; wherein, the first preset time period is different from the second preset time period;
[0010] Based on the second cluster attribute parameters and the storage volume to be analyzed, determine the abnormal storage volumes; wherein, the number of the abnormal storage volumes is less than or equal to the number of the storage volumes to be analyzed.
[0011] Optionally, before obtaining the first cluster attribute parameters corresponding to the target cluster within a first preset time period, the method further includes:
[0012] Detect the input / output (IO) response duration of the target cluster;
[0013] Correspondingly, obtaining the first cluster attribute parameters corresponding to the target cluster within a first preset time period includes:
[0014] If it is detected that the IO response duration exceeds a preset response duration, obtain the first cluster attribute parameters within the first preset time period before the current moment; wherein, the first cluster attribute parameters at least include m storage resource performance parameters of the target cluster, the IO response duration parameter of the target cluster, and the first IO traffic parameter of the reference storage volume.
[0015] Optionally, determining the storage volume to be analyzed based on the first cluster attribute parameters and the reference storage volume included in the target cluster includes:
[0016] Calculate the correlation coefficients between each of the storage resource performance parameters and the IO response duration parameter by using a first preset correlation algorithm to obtain m first correlation coefficients;
[0017] Based on the m first correlation coefficients, determine n target resource performance parameters from the m storage resource performance parameters;
[0018] Based on the n target resource performance parameters and the first IO traffic parameter, determine the storage volume to be analyzed.
[0019] Optionally, the reference storage volume includes p first sub-storage volumes, the first IO traffic parameter includes p sub-traffic parameters, and the sub-traffic parameters correspond to the first sub-storage volumes one by one. Determining the storage volume to be analyzed based on the n target resource performance parameters and the first IO traffic parameter includes:
[0020] Calculate the correlation coefficients between each of the target resource performance parameters and each of the sub-traffic parameters by using a second preset correlation algorithm to obtain n groups of p second correlation coefficients;
[0021] Determine the storage volume to be analyzed from p of the first sub-storage volumes based on n sets of p of the second correlation coefficients.
[0022] Optionally, the second preset time period is before the first preset time period, and the second cluster attribute parameter includes the second IO traffic parameter of the reference storage volume. Determining the abnormal storage volume based on the second cluster attribute parameter and the storage volume to be analyzed includes:
[0023] Use a preset prediction algorithm to perform predictive calculations on the second IO traffic parameter to determine an upper threshold for IO traffic prediction;
[0024] Determine the abnormal storage volume based on the upper threshold for IO traffic prediction and the third IO traffic parameter corresponding to the storage volume to be analyzed in the first IO traffic parameter.
[0025] Optionally, determining the abnormal storage volume based on the upper threshold for IO traffic prediction and the third IO traffic parameter corresponding to the storage volume to be analyzed in the first IO traffic parameter includes:
[0026] Determine the residual distribution between the upper threshold for IO traffic prediction and each IO traffic parameter included in the third IO traffic parameter to obtain at least one residual distribution; wherein, the IO traffic parameters included in the third IO traffic parameter correspond one-to-one with the storage volumes included in the storage volume to be analyzed;
[0027] Based on at least one of the residual distributions, determine the abnormal storage volume from the storage volume to be analyzed.
[0028] Optionally, after determining the abnormal storage volume based on the second cluster attribute parameter and the storage volume to be analyzed, the method further includes:
[0029] Sort the second sub-storage volumes included in the abnormal storage volume based on the second correlation coefficients of the second sub-storage volumes included in the abnormal storage volume to obtain a sorting result;
[0030] Display the sorting result.
[0031] Optionally, sorting the second sub-storage volumes included in the abnormal storage volume based on the second correlation coefficients of the second sub-storage volumes included in the abnormal storage volume to obtain a sorting result includes:
[0032] Obtain the IO traffic parameters of each of the second sub-storage volumes within a third preset time period to obtain fourth IO traffic parameters; wherein, the third preset time period includes at least the first preset time period;
[0033] Calculate the eigenvalue of each I / O traffic parameter in the fourth I / O traffic parameter to obtain at least one eigenvalue;
[0034] Based on at least one of the eigenvalues and the second correlation coefficient of each of the second sub-storage volumes, sort the second sub-storage volumes included in the abnormal storage volume to obtain the sorting result.
[0035] In a second aspect, an electronic device includes: a memory, a processor, and a communication bus; wherein:
[0036] The memory is used to store executable instructions;
[0037] The communication bus is used to implement the communication connection between the processor and the memory;
[0038] The processor is used to execute the abnormal detection program stored in the memory to implement the steps of the abnormal detection method described in any one of the above.
[0039] In a third aspect, a storage medium stores an abnormal detection program, and when the abnormal detection program is executed by a processor, the steps of the abnormal detection method described in any one of the above are implemented.
[0040] The embodiments of the present application provide an abnormal detection method, an electronic device, and a storage medium. After obtaining the first cluster attribute parameters corresponding to the target cluster in the first preset time period, based on the first cluster attribute parameters and the reference storage volumes included in the target cluster, the storage volumes to be analyzed are determined. Then, the second cluster attribute parameters of the target cluster in the second preset time period are obtained. Finally, based on the second cluster attribute parameters and the storage volumes to be analyzed, the abnormal storage volumes are determined. In this way, by analyzing the first cluster attribute parameters of the target cluster in the first preset time period and the reference storage volumes, after determining the storage volumes to be analyzed, the second cluster attribute parameters in the second preset time period are used to analyze the storage volumes to be analyzed to determine the abnormal storage volumes, solving the problem that the current method for analyzing the reason for the abnormal increase in storage response latency caused by storage capacity requires a large amount of historical sample data. An abnormal storage capacity determination method for the abnormal increase in storage response latency is proposed. By relying on historical sample data for analysis, the range of abnormal storage capacity is determined, ensuring the accuracy of abnormal analysis. Description of the Drawings
[0041] Figure 1 It is a schematic flowchart of an abnormal detection method provided by an embodiment of the present application;
[0042] Figure 2 It is a schematic flowchart of another abnormal detection method provided by an embodiment of the present application;
[0043] Figure 3 Schematic flowchart of another anomaly detection method provided by an embodiment of the present application;
[0044] Figure 4 Schematic diagram of an application scenario provided by an embodiment of the present application;
[0045] Figure 5 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0046] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application.
[0047] An embodiment of the present application provides an anomaly detection method. Referring to Figure 1 as shown, the method is applied to an electronic device, and the method includes the following steps:
[0048] Step 101: Obtain first cluster attribute parameters corresponding to the target cluster within a first preset time period.
[0049] In the embodiment of the present application, the target cluster generally refers to a single service cluster. The first cluster attribute parameters include a series of parameter values of at least one attribute parameter of the cluster, and can be specifically represented by a set. The duration corresponding to the first preset time period can be an empirical duration determined according to the actual application scenario. In some application scenarios, the first preset time period can be set according to actual needs. The electronic device is a device used to monitor and analyze the target cluster, such as a computer device or a server device, which can be specifically determined according to the actual application scenario.
[0050] During the operation of the target cluster, the first cluster attribute parameters of the target cluster are collected and recorded according to the preset sampling time.
[0051] Step 102: Determine the storage volume to be analyzed based on the first cluster attribute parameters and the reference storage volumes included in the target cluster.
[0052] In the embodiment of the present application, the reference storage volumes included in the target cluster include at least one storage volume. By analyzing the first cluster attribute parameters, the storage volume to be analyzed is determined from the reference storage volumes included in the target cluster. The storage volume to be analyzed includes at least one storage volume.
[0053] Step 103: Obtain second cluster attribute parameters corresponding to the target cluster within a second preset time period.
[0054] Among them, the first preset time period is different from the second preset time period.
[0055] In the embodiments of the present application, the first cluster attribute parameter and the second cluster attribute parameter may be the same or different, and the second cluster attribute parameter may also be some of the attribute parameters in the first cluster attribute parameter. The duration corresponding to the second preset time period may be an empirical duration determined according to the actual application scenario. In some application scenarios, the second preset time period may be set according to actual requirements.
[0056] Step 104: Determine the abnormal storage volume based on the second cluster attribute parameter and the storage volume to be analyzed.
[0057] Among them, the number of abnormal storage volumes is less than or equal to the number of storage volumes to be analyzed.
[0058] In the embodiments of the present application, the second cluster attribute parameter is used to analyze the storage volume to be analyzed to determine the abnormal storage volume from the storage volume to be analyzed. In this way, by gradually narrowing the analysis scope of the storage volume, the consumption of computing resources is effectively reduced, the analysis efficiency is improved, and the abnormal storage volume is quickly determined.
[0059] The abnormal detection method provided by the embodiments of the present application includes, after obtaining the first cluster attribute parameter corresponding to the target cluster within the first preset time period, determining the storage volume to be analyzed based on the first cluster attribute parameter and the reference storage volume included in the target cluster, then obtaining the second cluster attribute parameter of the target cluster within the second preset time period, and finally determining the abnormal storage volume based on the second cluster attribute parameter and the storage volume to be analyzed. In this way, after analyzing the first cluster attribute parameter and the reference storage volume of the target cluster within the first preset time period to determine the storage volume to be analyzed, the second cluster attribute parameter within the second preset time period is used to analyze the storage volume to be analyzed to determine the abnormal storage volume, which solves the problem that the current method for analyzing the reason for the abnormal increase in the storage response delay caused by the storage capacity needs to rely on a large amount of historical sample data, and proposes a method for determining the abnormal storage capacity that causes the abnormal increase in the storage response delay. By relying on historical sample data for analysis, the range of abnormal storage capacity is determined, ensuring the accuracy of abnormal analysis.
[0060] Based on the foregoing embodiments, an embodiment of the present application provides an abnormal detection method, which is applied to an electronic device. The method includes the following steps:
[0061] Step 201: Detect the input / output (IO) response duration of the target cluster.
[0062] In the embodiments of the present application, the electronic device performs real-time detection on the input / output (IO) response duration of the target cluster.
[0063] Step 202: If it is detected that the IO response duration exceeds the preset response duration, obtain the first cluster attribute parameters within the first preset time period before the current moment.
[0064] Among them, the first cluster attribute parameters at least include m storage resource performance parameters of the target cluster, the IO response duration parameter of the target cluster, and the first IO traffic parameter of the reference storage volume. The storage resource performance parameter can specifically be a parameter representing the usage of the cluster storage resources. The cluster storage resources include a central processing unit, a cache, a disk, a network disk, etc. The usage parameter of the cluster storage resources can be, for example, the usage rate of the cluster central processing unit (CPU), the usage rate of the cluster cache, the usage rate of the cluster disk, the usage rate of the cluster network disk, etc.
[0065] In the embodiments of the present application, the preset response duration is an empirical value of the IO response duration obtained based on a large number of experiments. The storage resource performance parameter is used to identify the performance attributes of the target cluster, such as storage performance, computing performance, etc.; the IO response duration parameter is used to identify the response duration of the input and output of the target cluster; the first IO traffic parameter is used to identify the input and output traffic parameters of each storage volume in the reference storage volume. The input and output traffic parameters can include read traffic and write traffic parameters.
[0066] When it is detected that the IO response duration exceeds the preset response duration, it indicates that there is an abnormal situation in the target cluster and abnormal analysis needs to be carried out. Therefore, within the first preset time period before the current moment, obtain the m storage resource performance parameters, the IO response duration parameter of the target cluster, and the first IO traffic parameter of the reference storage volume included in the target cluster, and use the m storage resource performance parameters, the IO response duration parameter, and the first IO traffic parameter as the first cluster attribute parameters. Among them, each storage resource performance parameter includes the parameter value collected within the first preset time period corresponding to a cluster storage resource of the target cluster.
[0067] Step 203: Based on the first cluster attribute parameters and the reference storage volume included in the target cluster, determine the storage volume to be analyzed.
[0068] In the embodiments of the present application, analyze the m storage resource performance parameters, the IO response duration parameter, and the first IO traffic parameter, and determine the storage volume to be analyzed from the reference storage volume.
[0069] Step 204: Obtain the second cluster attribute parameters corresponding to the target cluster within the second preset time period.
[0070] Among them, the first preset time period is different from the second preset time period.
[0071] In an embodiment of the present application, obtain second cluster attribute parameters corresponding to a target cluster within a second preset time period before a first preset time period.
[0072] Step 205: Determine abnormal storage volumes based on the second cluster attribute parameters and the storage volumes to be analyzed.
[0073] Among them, the number of abnormal storage volumes is less than or equal to the number of storage volumes to be analyzed.
[0074] In an embodiment of the present application, analyze the second cluster attribute parameters, and determine abnormal storage volumes from the storage volumes to be analyzed.
[0075] Based on the foregoing embodiments, in other embodiments of the present application, step 203 may be implemented by steps 203a to 203c:
[0076] Step 203a: Calculate the correlation coefficients between each storage resource performance parameter and the IO response duration parameter by using a first preset correlation algorithm, and obtain m first correlation coefficients.
[0077] In an embodiment of the present application, the first preset correlation algorithm may be a pre-determined algorithm for performing correlation calculation, such as a Pearson correlation coefficient calculation algorithm, a regression analysis method, etc. Exemplarily, when m is 3, the corresponding m storage resource performance parameters are the parameter value set 1 corresponding to storage resource 1, the parameter value set 2 corresponding to storage resource 2, and the parameter value set 3 corresponding to storage resource 3. In this way, use the first preset correlation algorithm to calculate the first correlation coefficient between the parameter value set 1 and the IO response duration parameter, denoted as coefficient 1; use the first preset correlation algorithm to calculate the first correlation coefficient between the parameter value set 2 and the IO response duration parameter, denoted as coefficient 2; use the first preset correlation calculation algorithm to calculate the first correlation coefficient between the parameter value set 3 and the IO response duration parameter, denoted as coefficient 3. In this way, 3 first correlation coefficients can be obtained.
[0078] Step 203b: Determine n target resource performance parameters from the m storage resource performance parameters based on the m first correlation coefficients.
[0079] In an embodiment of the present application, analyze the m first correlation coefficients to determine n target resource performance parameters from the m storage resource performance parameters. Among them, n is an integer less than or equal to m and greater than or equal to 1. For example, it may be to determine, from the m first correlation coefficients, the first correlation coefficients whose first correlation coefficients are greater than the first preset coefficient threshold, obtain n first correlation coefficients exceeding the threshold, and determine the storage resource performance parameters corresponding to the n first correlation coefficients exceeding the threshold, to obtain n target resource performance parameters.
[0080] Exemplarily, a first correlation coefficient with a coefficient greater than a first preset coefficient threshold is determined from coefficient 1, coefficient 2, and coefficient 3. Assuming that coefficient 1 and coefficient 3 are greater than the first preset coefficient threshold, it can be determined that the performance parameter 1 corresponding to coefficient 1 is the target resource performance parameter, and the performance parameter 3 corresponding to coefficient 3 is the target resource performance parameter.
[0081] Step 203c: Determine the storage volume to be analyzed based on the n target resource performance parameters and the first IO traffic parameter.
[0082] In the embodiments of the present application, the n target resource performance parameters and the first IO traffic parameter corresponding to the reference storage volume are analyzed to determine the storage volume to be analyzed from the reference storage volume.
[0083] Exemplarily, the performance parameter 1, the performance parameter 3, and the first IO traffic parameter are analyzed to determine the storage volume to be analyzed from the reference storage volume.
[0084] Based on the foregoing embodiments, in other embodiments of the present application, the reference storage volume includes p first sub-storage volumes, the first IO traffic parameter includes p sub-traffic parameters, and the sub-traffic parameters correspond to the first sub-storage volumes one by one. Step 203c can be implemented by steps a11 to a12:
[0085] Step a11: Calculate the correlation coefficient between each target resource performance parameter and each sub-traffic parameter using a second preset correlation algorithm to obtain n groups of p second correlation coefficients.
[0086] In the embodiments of the present application, p is an integer greater than or equal to 1. That is to say, the reference storage volume includes at least one first sub-storage volume, the first IO traffic parameter includes p sub-traffic parameters, and thus, one first sub-storage volume corresponds to one sub-traffic parameter. The second preset correlation algorithm can be the same as or different from the first preset correlation algorithm, and can be specifically determined according to the actual situation. A sub-traffic parameter is the IO traffic collected for one first sub-storage volume at a preset sampling interval within a second preset time period.
[0087] Exemplarily, when p is 4, the first IO traffic parameter includes 4 sub-traffic parameters, which are denoted as sub-traffic parameter 1, sub-traffic parameter 2, sub-traffic parameter 3, and sub-traffic parameter 4 for example. Thus, the correlation coefficient between performance parameter 1 and sub-traffic parameter 1 is calculated using the second preset correlation algorithm, the correlation coefficient between performance parameter 1 and sub-traffic parameter 2 is calculated using the second preset correlation algorithm, the correlation coefficient between performance parameter 1 and sub-traffic parameter 3 is calculated using the second preset correlation algorithm, and the correlation coefficient between performance parameter 1 and sub-traffic parameter 4 is calculated using the second preset correlation algorithm, obtaining a set of second correlation coefficients. Similarly, the correlation coefficients between performance parameter 3 and sub-traffic parameter 1, sub-traffic parameter 2, sub-traffic parameter 3, and sub-traffic parameter 4 are calculated respectively using the second preset correlation algorithm, obtaining another set of second correlation coefficients.
[0088] Step a12: Based on n groups of p second correlation coefficients, determine the storage volume to be analyzed from p first sub-storage volumes.
[0089] In the embodiment of the present application, analyze n groups of p second correlation coefficients, determine the second correlation coefficients greater than the second preset coefficient threshold, and determine the storage volumes corresponding to the second correlation coefficients greater than the second preset coefficient threshold from p first sub-storage volumes, obtaining the storage volume to be analyzed.
[0090] Based on the foregoing embodiments, in other embodiments of the present application, the second preset time period is before the first preset time period, the second cluster attribute parameter includes the second IO traffic parameter of the reference storage volume, and step 205 can be implemented by steps 205a to 205b:
[0091] Step 205a: Use a preset prediction algorithm to perform prediction calculation on the second IO traffic parameter to determine the upper threshold of IO traffic prediction.
[0092] In the embodiment of the present application, the preset prediction algorithm can be a trained neural network model algorithm, or can also be a prediction algorithm such as a time series prediction algorithm. In this way, through the preset prediction algorithm, prediction calculation and analysis are performed on the second IO traffic parameter collected for the reference storage volume within the second preset time period, and the upper threshold of IO traffic prediction corresponding to the reference storage volume within the first preset time period after the second preset time period can be obtained.
[0093] Step 205b: Based on the upper threshold of IO traffic prediction and the third IO traffic parameter corresponding to the storage volume to be analyzed in the first IO traffic parameter, determine the abnormal storage volume.
[0094] In an embodiment of the present application, from the first IO traffic parameters, determine the IO traffic parameters corresponding to the storage volume to be analyzed, obtain the third IO traffic parameters, and perform analysis by combining the IO traffic prediction upper limit threshold and the third IO traffic parameters to determine the abnormal storage volume from the storage volume to be analyzed.
[0095] Based on the foregoing embodiment, in other embodiments of the present application, step 205b may be implemented by steps b11 to b12:
[0096] Step b11, determine the residual distribution between the IO traffic prediction upper limit threshold and each IO traffic parameter included in the third IO traffic parameters, and obtain at least one residual distribution.
[0097] Among them, the IO traffic parameters included in the third IO traffic parameters correspond one-to-one with the storage volumes included in the storage volume to be analyzed.
[0098] In an embodiment of the present application, a residual calculation method is used to calculate the residual analysis between each IO traffic parameter included in the third IO traffic parameters and the IO traffic upper limit threshold, and obtain at least one residual analysis. That is to say, for several IO traffic parameters included in the third IO traffic parameters, several corresponding residual distributions are obtained. The residual calculation method may be residual absolute value analysis or residual numerical analysis.
[0099] Exemplarily, when the storage volume to be analyzed includes storage volume 1 and storage volume 2, the corresponding third IO traffic parameters include 2 IO traffic parameters, for example, denoted as IO traffic parameter 1 and IO traffic parameter 2. Correspondingly, calculate the residual distribution between the IO traffic prediction upper limit threshold of storage volume 1 and IO traffic parameter 1, and calculate the residual distribution between the IO traffic prediction upper limit threshold of storage volume 2 and IO traffic parameter 2, and 2 residual distributions can be obtained.
[0100] Step b12, based on at least one residual distribution, determine the abnormal storage volume from the storage volume to be analyzed.
[0101] In an embodiment of the present application, analyze at least one residual distribution, determine at least one target distribution, and determine the storage volume corresponding to at least one target distribution from the storage volume to be analyzed to obtain the abnormal storage volume.
[0102] Based on the foregoing embodiment, in other embodiments of the present application, as shown in Figure 3 After the electronic device executes step 205, it is further configured to execute steps 206 to 207:
[0103] Step 206, based on the second correlation coefficient of the second sub-storage volume included in the abnormal storage volume, sort the second sub-storage volumes included in the abnormal storage volume to obtain a sorting result.
[0104] In the embodiment of the present application, among the n groups of p second correlation coefficients calculated in step a11, determine the second correlation coefficients corresponding to the second sub-storage volumes included in the abnormal storage volume, and analyze and sort the second correlation coefficients corresponding to the second sub-storage volumes included in the abnormal storage volume, so as to sort the second sub-storage volumes included in the abnormal storage volume and obtain a sorting result.
[0105] Step 207: Display the sorting result.
[0106] In the embodiment of the present application, displaying the sorting result can prompt the user about the order of processing the abnormal storage volume, enabling the user to quickly eliminate the faults existing in the storage volume and improving the utilization rate of the target cluster.
[0107] Based on the foregoing embodiment, in other embodiments of the present application, step 206 can be implemented by steps 206a to 206c:
[0108] Step 206a: Obtain the IO traffic parameters of each second sub-storage volume within a third preset time period to obtain fourth IO traffic parameters.
[0109] Among them, the third preset time period includes at least the first preset time period.
[0110] In the embodiment of the present application, the third preset time period is a duration empirical value obtained based on a large number of experiments. Among them, the third preset time period includes at least the first preset time period. In some application scenarios, the user can set the third preset time period according to the actual situation.
[0111] Step 206b: Calculate the eigenvalues of each IO traffic parameter in the fourth IO traffic parameters to obtain at least one eigenvalue.
[0112] In the embodiment of the present application, an eigenvalue calculation method is used to calculate the IO traffic parameters of each second sub-storage volume to obtain the eigenvalue corresponding to each second sub-storage volume.
[0113] Among them, the eigenvalue calculation method can be a mean calculation method or a percentage calculation method, etc.
[0114] Step 206c: Based on at least one eigenvalue and the second correlation coefficient of each second sub-storage volume, sort the second sub-storage volumes included in the abnormal storage volume to obtain a sorting result.
[0115] In the embodiments of the present application, the eigenvalue and the second correlation coefficient of each second sub-storage volume are analyzed and sorted to sort the second sub-storage volumes included in the abnormal storage volume, so as to obtain a sorting result. In the process of analyzing and sorting the eigenvalue and the second correlation coefficient of each second sub-storage volume, a coordinate system with the eigenvalue as the abscissa and the second correlation coefficient as the ordinate can be established to determine the distribution of each second sub-storage volume in the coordinate system. The second sub-storage volumes with larger eigenvalues and larger second correlation coefficients are sorted in the front, and thus the sorting result is obtained. In some application scenarios, the coordinate system can also be established with the second correlation coefficient as the abscissa and the eigenvalue as the ordinate.
[0116] Based on the foregoing embodiments, in other embodiments of the present application, the embodiments of the present application provide an abnormal detection method, and the application scenario of this method can be as Figure 4 shown, including a storage cluster and a monitoring server; wherein, the storage cluster is used to provide cluster services, and the monitoring server is used to perform abnormal detection on the storage cluster. The implementation process of the abnormal detection method by the monitoring server on the storage cluster can be referred to the following steps:
[0117] Step c101: Regularly query the IO response duration of the storage cluster.
[0118] Step c102: If the queried IO response duration is greater than or equal to the IO response threshold, obtain the cluster resource utilization rate, the first IO traffic data of the storage volume, and the IO response duration of the storage cluster within the previous L time periods before the current moment.
[0119] It should be noted that the cluster resource utilization rate, the IO response duration of the storage cluster, and the first IO traffic data of the storage volumes included in the storage cluster obtained by the monitoring server can be preprocessed. Among them, the preprocessing process can be: collecting the cluster resource utilization rate of the storage cluster, the IO response duration of the storage cluster, and the first IO traffic data of all the storage volumes included in the storage cluster at a preset sampling interval to obtain the corresponding sampling data, and then performing smoothing processing and standard dimensionless processing on the sampling data corresponding to the cluster resource utilization rate, the sampling data corresponding to the IO response duration, and the sampling data corresponding to the first IO traffic data, so as to obtain the corresponding cluster resource utilization rate, IO response duration, and first IO traffic data. In some application scenarios, for the first IO traffic data of all the storage volumes included in the storage cluster, it can also be the first IO traffic data of some storage volumes obtained after filtering the first IO traffic data of all the storage volumes with a pre-set high-pass filter and then performing preprocessing. In this way, by using the method of filtering with a high-pass filter, the storage volumes with IO traffic lower than a certain threshold are screened out, effectively reducing the number of storage volumes to be analyzed, improving the processing efficiency, reducing the analysis volume, and saving computing resources.
[0120] During the preprocessing process, the data smoothing method adopted can be, for example, wavelet analysis method, convolution smoothing, etc.; the standard dimensionless processing method can be, for example, the minimum-maximum normalization processing method, the z-value normalization processing method, etc. The cluster resource utilization rate includes the utilization rates corresponding to at least two cluster resources.
[0121] Step c103: Calculate the obtained cluster resource utilization rate and the IO response duration using the Pearson correlation coefficient calculation method to obtain at least one corresponding first Pearson correlation coefficient.
[0122] Step c104: Determine the cluster resources that are strongly correlated with the IO response duration from the at least one first Pearson correlation coefficient obtained by calculation.
[0123] Exemplarily, for a target cluster with a management Internet Protocol (IP) address of 10.96.10x.xxx, the first Pearson correlation coefficient calculated between the compressed CPU utilization rate and the IO response duration is 0.431504, the first Pearson correlation coefficient calculated between the write cache rate and the IO response duration is 0.789922, and the first Pearson correlation coefficient calculated between the CPU utilization rate and the IO response duration is 0.739205.
[0124] The range of the degree of correlation defined in advance for the Pearson correlation coefficient can be exemplarily defined as follows: when the Pearson correlation coefficient is less than 0, the two analysis objects are negatively correlated; when the Pearson correlation coefficient is greater than or equal to 0 and less than 0.4, the two analysis objects are weakly correlated; when the Pearson correlation coefficient is greater than or equal to 0.4 and less than 0.6, the two analysis objects are moderately correlated; when the Pearson correlation coefficient is greater than or equal to 0.6, the two analysis objects are strongly correlated.
[0125] In this way, it can be determined that the cluster resources strongly correlated with the IO response duration are the write cache rate and the CPU utilization rate.
[0126] Step c105: Calculate the cluster resources strongly correlated with the IO response duration and the first IO traffic data of the storage volume using the Pearson correlation coefficient calculation method to obtain at least one corresponding second Pearson correlation coefficient.
[0127] Among them, the first IO traffic data includes at least one of the following types of data: data rate and input / output rate (IO rate). Among them, the input / output rate data further includes read IO rate and write IO rate, and the data rate further includes read data rate and write data rate. When the first IO traffic data includes at least two of the above data, determine the Pearson correlation coefficient between the cluster resources strongly correlated with the IO response duration and each type of data in the first IO traffic data to obtain the second Pearson correlation coefficient.
[0128] Step c106: Among at least one second Pearson correlation coefficient, determine that the cluster resources strongly correlated with the IO response duration and the storage volumes moderately and strongly correlated are the storage volumes to be analyzed.
[0129] Step c107: Obtain the second IO traffic data of the storage volume to be analyzed within the Q time period before the L time period.
[0130] Exemplarily, when the current time is 15:00:00, taking L as 3 hours and Q as 6 hours as an example, the second IO traffic data of the storage volume to be analyzed within the time period from 6:00:00 to 12:00:00 needs to be obtained.
[0131] Step c108: Use the time series prediction algorithm to perform prediction processing on the second IO traffic data to obtain the upper threshold of the IO traffic prediction.
[0132] Step c109: Use the residual data distribution algorithm to calculate the upper threshold of the IO traffic prediction and the first IO traffic data of the storage volume to be analyzed to determine the abnormal storage volume.
[0133] Among them, it can be set to only analyze the part where the residual > 0, or use the absolute value of the residual for analysis, or use the numerical value of the residual for analysis.
[0134] Step c110: Display the abnormal storage volume.
[0135] Further, when the abnormal storage volume includes at least one storage volume, the multiple storage volumes included in the abnormal storage volume can be sorted, and then the sorted multiple storage volumes are displayed. Among them, the process of sorting the multiple storage volumes included in the abnormal storage volume is as follows: obtaining the third IO traffic data of the multiple storage volumes included in the abnormal storage volume during the Z time period, where the Z time period includes the L time period; calculating the eigenvalue corresponding to the third IO traffic data of each storage volume included in the abnormal storage volume during the Z time period, and sorting the multiple storage volumes included in the abnormal storage volume based on the eigenvalue to obtain a sorting result. Further, when sorting the multiple storage volumes included in the abnormal storage volume based on the eigenvalue, the second Pearson correlation coefficient of the multiple storage volumes included in the abnormal storage volume can also be obtained, and sorting is performed from two dimensions based on the second Pearson correlation coefficient and the eigenvalue of each storage volume included in the abnormal storage volume.
[0136] After displaying the sorting result, the above collected data, such as the data obtained in step c102 and the data calculated in each step, can also be persistently processed for subsequent analysis. After displaying the sorting result, operations such as warning processing can also be performed.
[0137] In this way, the range of abnormal detection objects is determined and narrowed by using the Pearson correlation coefficient twice, reducing the analysis quantity, and the time period can be set according to requirements, and subsequent analysis can be performed without a long running time, reducing the order of magnitude of the quantity of data to be checked.
[0138] It should be noted that the descriptions of the same steps and the same content in this embodiment and other embodiments can refer to the descriptions in other embodiments and will not be repeated here.
[0139] The abnormal detection method provided by the embodiment of the present application determines the storage volume to be analyzed by obtaining the first cluster attribute parameter corresponding to the target cluster during the first preset time period, based on the first cluster attribute parameter and the reference storage volume included in the target cluster, then obtains the second cluster attribute parameter of the target cluster during the second preset time period, and finally determines the abnormal storage volume based on the second cluster attribute parameter and the storage volume to be analyzed. In this way, by analyzing the first cluster attribute parameter and the reference storage volume of the target cluster during the first preset time period to determine the storage volume to be analyzed, and then analyzing the storage volume to be analyzed by using the second cluster attribute parameter during the second preset time period to determine the abnormal storage volume, it solves the problem that the current method for analyzing the cause of the abnormal increase in storage response latency caused by the storage capacity needs to rely on a large amount of historical sample data, and proposes a method for determining the abnormal storage capacity causing the abnormal increase in storage response latency. By relying on historical sample data for analysis, the range of abnormal storage capacity is determined, ensuring the accuracy of abnormal analysis.
[0140] Based on the foregoing embodiments, an embodiment of the present application provides an electronic device, which can be applied to Figures 1 to 3 the anomaly detection method provided in the corresponding embodiment, with reference to Figure 5 As shown, the electronic device 3 may include: a memory 31, a processor 32, and a communication bus 33; where:
[0141] The memory 31 is used to store executable instructions;
[0142] The communication bus 33 is used to implement the communication connection between the processor 32 and the memory 31;
[0143] The processor 32 is used to execute the anomaly detection program stored in the memory 31 to implement the following steps:
[0144] In other embodiments of the present application, before the processor executes the step of obtaining the first cluster attribute parameters corresponding to the target cluster within the first preset time period, it is further used to execute the following steps:
[0145] Detect the input / output (IO) response duration of the target cluster;
[0146] Correspondingly, when the processor executes the step of obtaining the first cluster attribute parameters corresponding to the target cluster within the first preset time period, it can be implemented through the following steps:
[0147] If it is detected that the IO response duration exceeds the preset response duration, obtain the first cluster attribute parameters within the first preset time period before the current moment; where the first cluster attribute parameters at least include m storage resource performance parameters of the target cluster, the IO response duration parameter of the target cluster, and the first IO traffic parameter of the reference storage volume.
[0148] In other embodiments of the present application, when the processor executes the step of determining the storage volume to be analyzed based on the first cluster attribute parameters and the reference storage volume included in the target cluster, it can be implemented through the following steps:
[0149] Use the first preset correlation algorithm to calculate the correlation coefficients between each storage resource performance parameter and the IO response duration parameter, and obtain m first correlation coefficients;
[0150] Based on the m first correlation coefficients, determine n target resource performance parameters from the m storage resource performance parameters;
[0151] Based on the n target resource performance parameters and the first IO traffic parameter, determine the storage volume to be analyzed.
[0152] In other embodiments of the present application, when the processor executes the step of referring to a storage volume including p first sub-storage volumes, the first IO traffic parameter includes p sub-traffic parameters, and the sub-traffic parameters correspond to the first sub-storage volumes one by one, and determining the storage volume to be analyzed based on n target resource performance parameters and the first IO traffic parameter, the following steps can be implemented:
[0153] Use a second preset correlation algorithm to calculate the correlation coefficient between each target resource performance parameter and each sub-traffic parameter, and obtain n groups of p second correlation coefficients;
[0154] Based on n groups of p second correlation coefficients, determine the storage volume to be analyzed from the p first sub-storage volumes.
[0155] In other embodiments of the present application, when the processor executes the step that the second preset time period is before the first preset time period, the second cluster attribute parameter includes the second IO traffic parameter of the reference storage volume, and determining the abnormal storage volume based on the second cluster attribute parameter and the storage volume to be analyzed, the following steps can be implemented:
[0156] Use a preset prediction algorithm to perform prediction calculation on the second IO traffic parameter to determine the upper threshold of the IO traffic prediction;
[0157] Based on the upper threshold of the IO traffic prediction and the third IO traffic parameter corresponding to the storage volume to be analyzed in the first IO traffic parameter, determine the abnormal storage volume.
[0158] In other embodiments of the present application, when the processor executes the step of determining the abnormal storage volume based on the upper threshold of the IO traffic prediction and the third IO traffic parameter corresponding to the storage volume to be analyzed in the first IO traffic parameter, the following steps can be implemented:
[0159] Determine the residual distribution between the upper threshold of the IO traffic prediction and each IO traffic parameter included in the third IO traffic parameter, and obtain at least one residual distribution; wherein, the IO traffic parameters included in the third IO traffic parameter correspond to the storage volumes included in the storage volume to be analyzed one by one;
[0160] Based on at least one residual distribution, determine the abnormal storage volume from the storage volume to be analyzed.
[0161] In other embodiments of the present application, after the processor executes the step of determining the abnormal storage volume based on the second cluster attribute parameter and the storage volume to be analyzed, it is further used to execute the following steps:
[0162] Sort the second sub-storage volumes included in the abnormal storage volume based on the second correlation coefficient of the second sub-storage volumes included in the abnormal storage volume to obtain a sorting result;
[0163] Display the sorting result.
[0164] In other embodiments of the present application, when the processor executes the step of sorting the second sub-storage volumes included in the abnormal storage volume based on the second correlation coefficient of the second sub-storage volumes included in the abnormal storage volume to obtain a sorting result, the following steps may be adopted to implement it:
[0165] Obtain the IO traffic parameters of each second sub-storage volume within a third preset time period to obtain fourth IO traffic parameters; wherein, the third preset time period at least includes the first preset time period;
[0166] Calculate the eigenvalues of each IO traffic parameter in the fourth IO traffic parameters to obtain at least one eigenvalue;
[0167] Sort the second sub-storage volumes included in the abnormal storage volume based on at least one eigenvalue and the second correlation coefficient of each second sub-storage volume to obtain a sorting result.
[0168] It should be noted that for the specific implementation process of the steps executed by the processor in this embodiment, reference may be made to Figures 1 to 3 the implementation process in the abnormal detection method provided in the corresponding embodiment, which will not be elaborated here.
[0169] The electronic device provided in the embodiment of the present application, after obtaining the first cluster attribute parameters corresponding to the target cluster within the first preset time period, determines the storage volume to be analyzed based on the first cluster attribute parameters and the reference storage volume included in the target cluster, then obtains the second cluster attribute parameters of the target cluster within the second preset time period, and finally determines the abnormal storage volume based on the second cluster attribute parameters and the storage volume to be analyzed. In this way, by analyzing the first cluster attribute parameters and the reference storage volume of the target cluster within the first preset time period to determine the storage volume to be analyzed, and then analyzing the storage volume to be analyzed with the second cluster attribute parameters within the second preset time period to determine the abnormal storage volume, it solves the problem that the current method for analyzing the cause of the abnormal increase in storage response latency caused by the storage capacity needs to rely on a large amount of historical sample data, and proposes a method for determining the abnormal storage capacity that causes the abnormal increase in storage response latency. By relying on historical sample data for analysis to determine the range of abnormal storage capacity, the accuracy of abnormal analysis is ensured.
[0170] Based on the foregoing embodiments, an embodiment of the present application provides a computer-readable storage medium, which may be simply referred to as a storage medium. The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the implementation process of the method provided in the corresponding embodiment, which will not be elaborated here. Figures 1 to 3
[0171] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories and optical memories, etc.) that contain computer-usable program code.
[0172] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks.
[0173] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one or more of the flows Figure 1 or blocks.
[0174] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks.
[0175] As described above, it is only the preferred embodiment of the present application, and is not used to limit the protection scope of the present application.
Claims
1. An anomaly detection method, the method comprising: Obtaining first cluster attribute parameters corresponding to a target cluster within a first preset time period; Determining a storage volume to be analyzed from the reference storage volumes based on the first cluster attribute parameters and the reference storage volumes included in the target cluster; Obtaining second cluster attribute parameters corresponding to the target cluster within a second preset time period; wherein the first preset time period is different from the second preset time period; Determining an abnormal storage volume from the storage volumes to be analyzed based on the second cluster attribute parameters and the storage volumes to be analyzed; wherein the number of the abnormal storage volumes is less than or equal to the number of the storage volumes to be analyzed.
2. The method according to claim 1, before obtaining the first cluster attribute parameters corresponding to the target cluster within the first preset time period, the method further comprising: Detecting the input / output (IO) response duration of the target cluster; Correspondingly, obtaining the first cluster attribute parameters corresponding to the target cluster within the first preset time period includes: If it is detected that the IO response duration exceeds a preset response duration, obtaining the first cluster attribute parameters within the first preset time period before the current moment; wherein the first cluster attribute parameters at least include m storage resource performance parameters of the target cluster, an IO response duration parameter of the target cluster, and a first IO traffic parameter of the reference storage volume.
3. The method according to claim 2, determining the storage volume to be analyzed based on the first cluster attribute parameters and the reference storage volumes included in the target cluster includes: Calculating a correlation coefficient between each of the storage resource performance parameters and the IO response duration parameter by using a first preset correlation algorithm to obtain m first correlation coefficients; Determining n target resource performance parameters from the m storage resource performance parameters based on the m first correlation coefficients; wherein n is an integer less than or equal to m and greater than or equal to 1; Determining the storage volume to be analyzed based on the n target resource performance parameters and the first IO traffic parameter.
4. The method according to claim 3, the reference storage volume includes p first sub-storage volumes, the first IO traffic parameter includes p sub-traffic parameters, and the sub-traffic parameters correspond to the first sub-storage volumes one by one. Determining the storage volume to be analyzed based on the n target resource performance parameters and the first IO traffic parameter includes: Calculating a correlation coefficient between each of the target resource performance parameters and each of the sub-traffic parameters by using a second preset correlation algorithm to obtain n groups of p second correlation coefficients; wherein p is an integer greater than or equal to 1; Determining the storage volume to be analyzed from the p first sub-storage volumes based on the n groups of p second correlation coefficients.
5. The method according to any one of claims 2 to 4, the second preset time period is before the first preset time period, the second cluster attribute parameter includes a second IO traffic parameter of the reference storage volume, and determining the abnormal storage volume based on the second cluster attribute parameters and the storage volume to be analyzed includes: Use a preset prediction algorithm to perform predictive calculations on the second IO traffic parameter to determine the upper threshold of the IO traffic prediction; Based on the upper threshold of the IO traffic prediction and the third IO traffic parameter corresponding to the storage volume to be analyzed in the first IO traffic parameter, determine the abnormal storage volume.
6. The method according to claim 5, wherein the determining the abnormal storage volume based on the upper threshold of the IO traffic prediction and the third IO traffic parameter corresponding to the storage volume to be analyzed in the first IO traffic parameter includes: Determine the residual distribution between the upper threshold of the IO traffic prediction and each IO traffic parameter included in the third IO traffic parameter to obtain at least one residual distribution; wherein, the IO traffic parameters included in the third IO traffic parameter correspond one-to-one with the storage volumes included in the storage volume to be analyzed; Based on at least one of the residual distributions, determine the abnormal storage volume from the storage volumes to be analyzed.
7. The method according to claim 4, after the determining the abnormal storage volume based on the second cluster attribute parameter and the storage volume to be analyzed, the method further includes: Sort the second sub-storage volumes included in the abnormal storage volume based on the second correlation coefficient of the second sub-storage volumes included in the abnormal storage volume to obtain a sorting result; Display the sorting result.
8. The method according to claim 7, wherein the sorting the second sub-storage volumes included in the abnormal storage volume based on the second correlation coefficient of the second sub-storage volumes included in the abnormal storage volume to obtain a sorting result includes: Obtain the IO traffic parameter of each of the second sub-storage volumes within a third preset time period to obtain a fourth IO traffic parameter; wherein, the third preset time period includes at least the first preset time period; Calculate the eigenvalue of each IO traffic parameter in the fourth IO traffic parameter to obtain at least one eigenvalue; Based on at least one of the eigenvalues and the second correlation coefficient of each of the second sub-storage volumes, sort the second sub-storage volumes included in the abnormal storage volume to obtain the sorting result.
9. An electronic device, the electronic device comprising: A memory, a processor, and a communication bus; wherein: The memory is used to store executable instructions; The communication bus is used to implement the communication connection between the processor and the memory; The processor is used to execute the abnormal detection program stored in the memory to implement the steps of the abnormal detection method according to any one of claims 1 to 8.
10. A storage medium, on which an abnormal detection program is stored, and when the abnormal detection program is executed by a processor, the steps of the abnormal detection method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Script exception detection method and terminal thereof
CN108255710A
Centralized memory monitoring method and device
CN111563022A