Performance monitoring method and system for data storage equipment

By performing single-factor ANOVA analysis of variance and real-time performance indicator data on historical working data, establishing performance prediction models, and dynamically adjusting monitoring strategies, the problem of difficulty in automatically adjusting monitoring strategies in the existing technology is solved, and the accuracy and flexibility of monitoring are improved.

CN120029850AInactive Publication Date: 2025-05-23NANTONG QILU INFORMATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510106798.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-23
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

It is difficult for the prior art to automatically adjust monitoring strategies based on historical data and real-time feedback to adapt to different workloads and performance changes, affecting monitoring effects.

Method used

By retrieving historical working data from the database, performing single-factor variance analysis, obtaining real-time performance indicator data, calculating descriptive statistics, establishing baseline and performance prediction models, monitoring equipment performance in real time, and dynamically adjusting monitoring strategies based on analysis results.

Benefits of technology

Improves early warning accuracy, effectively identify performance abnormalities, reduces false alarms and missed response rates, and ensures that the monitoring strategy matches the current status of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029850A_ABST
    Figure CN120029850A_ABST
Patent Text Reader

Abstract

The invention discloses a data storage equipment performance monitoring method and system, and relates to the technical field of storage equipment performance monitoring, and the method comprises the steps: calling historical working data from a database, carrying out the single-factor variance analysis, outputting a first analysis result, collecting performance index data, calculating descriptive statistic data, building a baseline, and building a first performance prediction model; according to the method, the influence of the load change on the equipment performance is analyzed and judged through the single-factor variance analysis, if the influence is obvious, a baseline and a performance prediction model are established, and performance index data are collected in real time for contrastive analysis. Meanwhile, equipment performance is monitored in real time, a monitoring strategy is dynamically adjusted according to real-time data and a historical trend, early warning accuracy is improved, a machine learning algorithm and an LOF technology are introduced, performance abnormity is recognized more effectively, and the false alarm rate and the missing report rate are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of storage device performance monitoring, and in particular to a data storage device performance monitoring method and system. Background Art

[0002] With the rapid development of information technology, data storage devices are increasingly used in all walks of life. However, performance monitoring of data storage devices has always been a challenge, especially when dealing with large amounts of data and complex workloads. Traditional performance monitoring methods are usually based on static thresholds to determine whether the device is operating normally, but this method has limitations because the performance of the device fluctuates with changes in load. Therefore, a more intelligent and flexible monitoring method is needed to adapt to different workloads and performance changes.

[0003] At present, the Chinese invention patent with application number 202410063153.4 discloses a cache performance test method, device, computer equipment and storage medium of UFS. The method includes the following steps: securely erasing the storage to be tested to keep the storage to be tested in an initial state; configuring the test environment of the storage to be tested; filling the storage to be tested with data and reserving a preset size of available space; performing a large amount of reading and writing on the available space, so that the SLC buffer data is messy; after standing for a set time, the firmware releases the data space of the SLC buffer; using a data model to perform performance test verification on the SLC buffer and output the test results; generating a graphical statistical analysis report based on the test results.

[0004] The above technologies are difficult to automatically adjust monitoring strategies based on historical data and real-time feedback to adapt to different workloads and performance changes, which affects the monitoring effect. Summary of the invention

[0005] The technical problem solved by the present invention is that it is difficult for the above technology to automatically adjust the monitoring strategy according to historical data and real-time feedback to adapt to different workloads and performance changes, which affects the monitoring effect.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0007] A data storage device performance monitoring method comprises the following steps:

[0008] Step S1, retrieving historical working data from a database, performing a one-way variance analysis on historical performance indicator data corresponding to each historical load in the historical working data, and outputting a first analysis result;

[0009] Step S2, collecting real-time performance indicator data from the data storage device according to the first analysis result, the real-time performance indicator data including real-time IOPS, real-time bandwidth, real-time delay and real-time CPU, and storing the real-time performance indicator data in a database;

[0010] Step S3, calculating descriptive statistics data according to the first analysis result, wherein the descriptive statistics data include a standard deviation and a mean value, and establishing a baseline according to the descriptive statistics data;

[0011] Step S4, establishing a first performance prediction model according to the baseline and the corresponding historical working data, establishing a corresponding relationship between the performance prediction model and the historical load, outputting several second performance prediction models, obtaining the real-time load, searching for the second performance prediction model corresponding to the real-time load, obtaining the time from the baseline, outputting the predicted fault data, analyzing the predicted fault data, and outputting the second analysis result;

[0012] Step S5, selecting a monitoring strategy according to the second analysis result, and inputting the monitoring strategy into the control terminal.

[0013] Preferably, the step S1 includes the following sub-steps:

[0014] Step S101, retrieving historical work data from a database, the historical work data including historical load and corresponding historical performance indicator data, the historical performance indicator data including historical IOPS, historical bandwidth, historical delay and historical CPU;

[0015] Step S102: performing a one-way variance analysis on the historical performance indicator data corresponding to each historical load in the historical working data, and outputting a first analysis result.

[0016] Preferably, the one-way ANOVA is:

[0017] Calculate the total internal sum of squares, inter-group sum of squares, intra-group sum of squares, inter-group degrees of freedom, and intra-group degrees of freedom for the historical IOPS, historical bandwidth, historical latency, and historical CPU for each historical load respectively;

[0018] The inter-group mean square is calculated based on the inter-group sum of squares and the inter-group degrees of freedom. The mathematical expression of the inter-group mean square is:

[0019]

[0020] Among them, MSB is the mean square between groups, SSB is the sum of squares between groups, and dfB is the degrees of freedom between groups.

[0021] The within-group mean square is calculated based on the within-group sum of squares and the within-group degrees of freedom. The mathematical expression of the within-group mean square is:

[0022]

[0023] Among them, MSW is the within-group mean square, SSW is the within-group sum of squares, and dfW is the within-group degrees of freedom;

[0024] The F ratio is calculated based on the between-group mean square and the within-group mean square, where the F ratio is the ratio of the between-group mean square to the within-group mean square. The F ratio and the corresponding between-group degrees of freedom and within-group degrees of freedom are input into SPSS, the P value is searched, and a correspondence is established between the P value and the corresponding historical IOPS, historical bandwidth, historical latency, and historical CPU, which is output as the first analysis result.

[0025] Preferably, step S2 includes the following sub-steps:

[0026] Step S201, obtain the first analysis result and analyze it:

[0027] If the P value is greater than or equal to 0.05, it is determined that the load change does not affect the performance and the preset constant standard IOPS threshold, constant standard bandwidth threshold, constant standard delay threshold, constant standard CPU threshold and constant standard acquisition frequency are obtained and input to the control end;

[0028] If the P value is less than 0.05, it is determined that the load change affects the performance and performance indicator data is collected from the data storage device, wherein the real-time performance indicator data includes real-time IOPS, real-time bandwidth, real-time delay and real-time CPU;

[0029] Step S202, storing the real-time performance indicator data in a database.

[0030] Preferably, step S3 includes the following sub-steps:

[0031] Step S301, parsing the first analysis result:

[0032] If the P value is less than 0.05, the descriptive statistics data of the historical IOPS, historical bandwidth, historical latency and historical CPU corresponding to each historical load are calculated respectively, and the descriptive statistics data include the standard deviation and the mean;

[0033] Step S302 , using the mean ± standard deviation as a range to establish baselines corresponding to historical IOPS, historical bandwidth, historical latency, and historical CPU.

[0034] Preferably, step S4 includes the following sub-steps:

[0035] Step S401, establishing a first performance prediction model according to the baseline and the corresponding historical working data, establishing a corresponding relationship between the performance prediction model and the historical load, and outputting them as a plurality of second performance prediction models;

[0036] Step S402, obtaining the real-time load, searching for a second performance prediction model corresponding to the real-time load, obtaining the time from the baseline, and outputting it as predicted fault data;

[0037] Step S403, analyzing the predicted fault data;

[0038] If the time from the baseline is less than the preset first failure time threshold, it is marked as an impending failure;

[0039] If the time from the baseline is at or greater than the first fault time threshold, it is marked as not a fault;

[0040] Impending failure and not-impending failure are output as the second analysis result.

[0041] Preferably, the logic of the first performance prediction model is:

[0042] Input historical performance indicator data, descriptive statistics data and baseline as input objects to the machine learning algorithm, define the neighborhood distance threshold of each data in the historical performance indicator data and descriptive statistics data, and find the k nearest neighbors of each data;

[0043] Calculate the local density of each data in the historical performance indicator data and the descriptive statistics data, where the local density is the average of the reciprocals of the densities of the k nearest neighbors;

[0044] Calculate the LOF value of each data in the historical performance indicator data and the descriptive statistics data, where the LOF value is the ratio of the local density of each data in the historical performance indicator data and the descriptive statistics data to the adjacent local density;

[0045] The LOF threshold is set according to the baseline, and data points exceeding the LOF threshold are marked as abnormal.

[0046] Preferably, step S5 includes the following sub-steps:

[0047] Step S501, selecting a monitoring strategy according to the second analysis result, the monitoring strategy including maintaining the existing monitoring frequency and increasing the existing monitoring frequency, and obtaining a monitoring frequency corresponding to the monitoring strategy, the monitoring frequency including the existing monitoring frequency and the changed monitoring frequency;

[0048] Step S502, input the monitoring frequency to the control end and perform monitoring.

[0049] Preferably, the selection logic of the monitoring strategy is:

[0050] If the second analysis result is that there is no fault, the monitoring strategy is determined to maintain the existing monitoring frequency and obtain the preset existing monitoring frequency;

[0051] If the second analysis result is an impending failure, the monitoring strategy is determined to be to increase the existing monitoring frequency and obtain the time from the baseline, define the time threshold from the baseline and the corresponding monitoring frequency, find the monitoring frequency corresponding to the time from the baseline and output it as the changed monitoring frequency.

[0052] A data storage device performance monitoring system, which is applied to the data storage device performance monitoring method, comprises a historical data analysis module, a real-time data acquisition module, a real-time data analysis module, a prediction model establishment module and a monitoring strategy selection module;

[0053] The historical data analysis module is used to retrieve historical working data from the database, perform a one-way variance analysis on the historical performance indicator data corresponding to each historical load in the historical working data, and output a first analysis result;

[0054] The real-time data acquisition module is used to collect real-time performance indicator data from the data storage device according to the first analysis result, the real-time performance indicator data including real-time IOPS, real-time bandwidth, real-time delay and real-time CPU, and store the real-time performance indicator data in a database;

[0055] The real-time data analysis module is used to calculate descriptive statistics data according to the first analysis result, wherein the descriptive statistics data includes a standard deviation and a mean value, and establish a baseline according to the descriptive statistics data;

[0056] The prediction model establishment module is used to establish a first performance prediction model according to the baseline and the corresponding historical working data, establish a corresponding relationship between the performance prediction model and the historical load, output as a plurality of second performance prediction models, obtain the real-time load, find the second performance prediction model corresponding to the real-time load, obtain the time from the baseline, output as predicted fault data, analyze the predicted fault data, and output a second analysis result;

[0057] The monitoring strategy selection module is used to select a monitoring strategy according to the second analysis result and input the monitoring strategy to the control end.

[0058] Beneficial effects of the present invention: The present invention uses single-factor variance analysis to determine the impact of load changes on equipment performance. If the impact is significant, a baseline and performance prediction model is established, and performance indicator data is collected in real time for comparative analysis. At the same time, the performance of the equipment is monitored in real time, and the monitoring strategy is dynamically adjusted according to real-time data and historical trends to improve the accuracy of early warnings. Machine learning algorithms and LOF technology are introduced to more effectively identify performance anomalies and reduce false positives and false negatives. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 A flowchart of a method for monitoring performance of a data storage device provided by an embodiment of the present invention;

[0060] Figure 2 A basic flow chart of a data storage device performance monitoring system provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0061] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, but not all of the embodiments.

[0062] Example 1, reference Figure 1 , provides a data storage device performance monitoring method, comprising the following steps:

[0063] Step S1, retrieve historical working data from a database, perform a one-way variance analysis on historical performance indicator data corresponding to each historical load in the historical working data, and output a first analysis result.

[0064] Step S2, collecting real-time performance indicator data from the data storage device according to the first analysis result, the real-time performance indicator data including real-time IOPS, real-time bandwidth, real-time delay and real-time CPU, and storing the real-time performance indicator data in a database.

[0065] Step S3, calculating descriptive statistics data according to the first analysis result, wherein the descriptive statistics data include a standard deviation and a mean value, and establishing a baseline according to the descriptive statistics data.

[0066] Step S4, establish a first performance prediction model based on the baseline and the corresponding historical working data, establish a corresponding relationship between the performance prediction model and the historical load, output it as several second performance prediction models, obtain the real-time load, find the second performance prediction model corresponding to the real-time load, obtain the time from the baseline, output it as predicted fault data, analyze the predicted fault data, and output the second analysis result.

[0067] Step S5, selecting a monitoring strategy according to the second analysis result, and inputting the monitoring strategy into the control terminal.

[0068] Step S1 includes the following sub-steps:

[0069] Step S101, retrieve historical work data from a database, the historical work data including historical load and corresponding historical performance indicator data, the historical performance indicator data including historical IOPS, historical bandwidth, historical delay and historical CPU.

[0070] Step S101 accurately retrieves historical work data from the database. These data include historical loads and corresponding historical performance indicator data. The historical performance indicator data specifically covers historical IOPS, historical bandwidth, historical latency, and historical CPU. These data can fully reflect the performance of data storage devices under different loads and provide necessary input data for subsequent one-way analysis of variance.

[0071] Step S102: performing a one-way variance analysis on the historical performance indicator data corresponding to each historical load in the historical working data, and outputting a first analysis result.

[0072] One-way ANOVA is:

[0073] The total intra-group sum of squares, inter-group sum of squares, intra-group sum of squares, inter-group degrees of freedom, and intra-group degrees of freedom of the historical IOPS, historical bandwidth, historical latency, and historical CPU corresponding to each historical load are calculated respectively.

[0074] The inter-group mean square is calculated based on the inter-group sum of squares and the inter-group degrees of freedom. The mathematical expression of the inter-group mean square is:

[0075]

[0076] Among them, MSB is the mean square between groups, SSB is the sum of squares between groups, and dfB is the degrees of freedom between groups.

[0077] The within-group mean square is calculated based on the within-group sum of squares and the within-group degrees of freedom. The mathematical expression of the within-group mean square is:

[0078]

[0079] Among them, MSW is the within-group mean square, SSW is the within-group sum of squares, and dfW is the within-group degrees of freedom.

[0080] The F ratio is calculated based on the between-group mean square and the within-group mean square, where the F ratio is the ratio of the between-group mean square to the within-group mean square. The F ratio and the corresponding between-group degrees of freedom and within-group degrees of freedom are input into SPSS, the P value is searched, and a correspondence is established between the P value and the corresponding historical IOPS, historical bandwidth, historical latency, and historical CPU, which is output as the first analysis result.

[0081] Step S102 calculates the total internal sum of squares, the inter-group sum of squares, the intra-group sum of squares, the inter-group degrees of freedom and the intra-group degrees of freedom of the historical IOPS, historical bandwidth, historical delay and historical CPU corresponding to each historical load, respectively, to provide a basis for the subsequent calculation of the inter-group mean square and the intra-group mean square. The inter-group mean square is calculated based on the inter-group sum of squares and the inter-group degrees of freedom, reflecting the average difference degree of performance indicator data between different loads. The intra-group mean square is calculated based on the intra-group sum of squares and the intra-group degrees of freedom, reflecting the average difference degree of performance indicator data within the same load. The F ratio is calculated based on the inter-group mean square and the intra-group mean square. The larger the F ratio, the more significant the difference in performance indicators between different loads. The F ratio, the corresponding inter-group degrees of freedom and the intra-group degrees of freedom are input into statistical software such as SPSS to find the P value. The P value is used to determine whether there is a significant difference in performance indicators under different loads. A corresponding relationship between the P value and the corresponding historical IOPS, historical bandwidth, historical delay and historical CPU is established, and the output is the first analysis result, which can intuitively display the significance of the difference in performance indicator data under different loads, and provide a basis for the subsequent establishment of a performance prediction model and the selection of a real-time monitoring strategy.

[0082] Step S1 retrieves historical working data from the historical database and performs one-way variance analysis on the data to determine whether there are significant differences in the performance indicators of the data storage device under different historical loads, which provides a basis for the subsequent establishment of a performance prediction model and the selection of a real-time monitoring strategy.

[0083] Step S2 includes the following sub-steps:

[0084] Step S201, obtain the first analysis result and analyze it:

[0085] If the P value is greater than or equal to 0.05, it is determined that the load change does not affect the performance and the preset constant standard IOPS threshold, constant standard bandwidth threshold, constant standard delay threshold, constant standard CPU threshold and constant standard acquisition frequency are obtained and input into the control end.

[0086] If the P value is less than 0.05, it is determined that the load change affects the performance and performance indicator data is collected from the data storage device, where the real-time performance indicator data includes real-time IOPS, real-time bandwidth, real-time delay and real-time CPU.

[0087] Step S201 determines whether the load change affects the performance through the P value, and takes corresponding actions according to the determination result:

[0088] If the P value is greater than or equal to 0.05: it is determined that the load change does not affect the performance. At this time, the preset constant standard threshold and acquisition frequency are used and these values ​​are input to the control end. This can simplify the performance monitoring process and reduce unnecessary resource consumption.

[0089] If the P value is less than 0.05: it is judged that the load change affects the performance. At this time, real-time performance indicator data is collected from the data storage device, including real-time IOPS, real-time bandwidth, real-time latency, and real-time CPU. These data can reflect the current performance status of the system and provide an important basis for subsequent performance analysis and optimization.

[0090] Step S202, storing the real-time performance indicator data in a database.

[0091] Step S202 stores the real-time performance index data in a database. This ensures persistent storage of data, making it easier to analyze and mine performance data later. At the same time, database storage also helps to achieve centralized management and sharing of data, improving the efficiency and value of data use.

[0092] Step S2 makes a preliminary judgment on whether the load change affects the performance, and takes corresponding actions according to the judgment result. If the load change does not affect the performance, the preset constant standard threshold is used; if the load change affects the performance, real-time performance indicator data is collected. This step provides basic data for subsequent performance monitoring and optimization, which helps to ensure the stable operation of the method.

[0093] Step S3 includes the following sub-steps:

[0094] Step S301, parsing the first analysis result:

[0095] If the P value is less than 0.05, the descriptive statistics data of the historical IOPS, historical bandwidth, historical latency and historical CPU corresponding to each historical load are calculated respectively, and the descriptive statistics data include the standard deviation and the mean value.

[0096] Step S301 determines which performance indicators may have significant changes by judging the P value in the first analysis result. A P value less than 0.05 usually means that there is a statistically significant difference, so it is necessary to further analyze these indicators and calculate the standard deviation and average value of each performance indicator under the historical load. These descriptive statistics can reflect the distribution and average level of the indicator data and provide necessary data support for establishing a baseline.

[0097] Step S302 , using the mean ± standard deviation as a range to establish baselines corresponding to historical IOPS, historical bandwidth, historical latency, and historical CPU.

[0098] Based on the standard deviation and average value calculated in step S301, the average value ± standard deviation is used as the range to establish a baseline for indicators such as historical IOPS, historical bandwidth, historical latency, and historical CPU. This baseline range can reflect the possible fluctuation range of these performance indicators under normal load, and the established baseline can provide an important basis for subsequent performance evaluation.

[0099] Step S3 aims to determine whether it is necessary to establish a performance baseline by analyzing the first analysis result, and calculate the descriptive statistics of the historical performance indicators based on this, and then set the baseline range of these indicators. This step provides an important reference for subsequent performance evaluation, anomaly detection or capacity planning.

[0100] Step S4 includes the following sub-steps:

[0101] Step S401, establishing a first performance prediction model according to the baseline and the corresponding historical working data, establishing a corresponding relationship between the performance prediction model and the historical load, and outputting a plurality of second performance prediction models.

[0102] Step S401 uses historical performance indicator data, descriptive statistics data and baselines as inputs, and trains a first performance prediction model through a machine learning algorithm. This model can learn the complex relationship between historical loads and performance indicators. According to the first performance prediction model, for different historical load conditions, several specific second performance prediction models are output. These models can more accurately predict the performance under a specific load, establish a corresponding relationship between the performance prediction model and the historical load, so that the most suitable prediction model can be quickly found for performance prediction under real-time load.

[0103] Step S402, obtain the real-time load, search for the second performance prediction model corresponding to the real-time load, obtain the time from the baseline, and output it as predicted fault data.

[0104] Step S402 monitors the load of the system in real time, obtains the current real-time load data, and searches for the corresponding second performance prediction model: based on the real-time load data, finds the most suitable model in the established second performance prediction model for performance prediction, and calculates the time difference between the real-time load data and the most recent baseline data as part of the predicted fault data.

[0105] Step S403, analyzing the predicted fault data;

[0106] If the time from the baseline is less than the preset first failure time threshold, it is marked as an impending failure.

[0107] If the time from the baseline is at or greater than the first fault time threshold, it is marked as not a fault.

[0108] Impending failure and not-impending failure are output as the second analysis result.

[0109] The logic of the first performance prediction model is:

[0110] The historical performance indicator data, descriptive statistics data and baseline are input into the machine learning algorithm as input objects, the neighborhood distance threshold of each data in the historical performance indicator data and descriptive statistics data is defined, and the k nearest neighbors of each data are found.

[0111] The local density of each data in the historical performance indicator data and the descriptive statistics data is calculated, and the local density is the average value of the reciprocal of the density of the k nearest neighbors.

[0112] The LOF value of each data in the historical performance indicator data and the descriptive statistics data is calculated, and the LOF value is the ratio of the local density of each data in the historical performance indicator data and the descriptive statistics data to the adjacent local density.

[0113] The LOF threshold is set according to the baseline, and data points exceeding the LOF threshold are marked as abnormal.

[0114] Step S403 analyzes the predicted fault data obtained in step S402, including the time from the baseline and possible performance indicator abnormalities, compares the time from the baseline with the preset first fault time threshold, marks it as "imminent fault" or "not about to fail", and outputs the analysis result as the second analysis result to provide a basis for subsequent fault handling or performance optimization.

[0115] Step S4 aims to establish performance prediction models through historical data and baseline information, and use these models to predict performance under real-time load, so as to identify potential fault conditions. Through machine learning algorithms, step S4 can analyze historical performance indicator data, descriptive statistics data and baselines, and establish performance prediction models corresponding to different historical loads, providing strong support for fault prevention and optimization.

[0116] Step S5 includes the following sub-steps:

[0117] Step S501, selecting a monitoring strategy according to the second analysis result, wherein the monitoring strategy includes maintaining the existing monitoring frequency and increasing the existing monitoring frequency, and obtaining the monitoring frequency corresponding to the monitoring strategy, wherein the monitoring frequency includes the existing monitoring frequency and the changed monitoring frequency.

[0118] The selection logic of the monitoring strategy is:

[0119] If the second analysis result is that there is no fault, the monitoring strategy is determined to maintain the existing monitoring frequency and obtain the preset existing monitoring frequency.

[0120] If the second analysis result is an impending failure, the monitoring strategy is determined to be to increase the existing monitoring frequency and obtain the time from the baseline, define the time threshold from the baseline and the corresponding monitoring frequency, find the monitoring frequency corresponding to the time from the baseline and output it as the changed monitoring frequency.

[0121] Step S501 intelligently selects to maintain the existing monitoring frequency or increase the existing monitoring frequency as the monitoring strategy based on the second analysis result, ensuring that the monitoring strategy matches the current state of the system. According to the selected monitoring strategy, the corresponding monitoring frequency is obtained. If the existing monitoring frequency is maintained, the preset existing monitoring frequency is obtained. If the existing monitoring frequency is increased, the changed monitoring frequency that matches the current distance to the baseline time is searched and output based on the time threshold from the baseline and the corresponding monitoring frequency table.

[0122] The monitoring strategy intelligently decides whether to maintain the existing monitoring frequency or increase the monitoring frequency by judging the second analysis result. This decision-making method can ensure that the monitoring strategy matches the actual status and needs of the system. The monitoring frequency is dynamically adjusted according to the time threshold from the baseline and the corresponding monitoring frequency table. When a failure is likely to occur, the monitoring sensitivity and accuracy are improved by increasing the monitoring frequency, thereby more effectively preventing and handling potential problems.

[0123] Step S502, input the monitoring frequency to the control end and perform monitoring.

[0124] Step S502 inputs the monitoring frequency obtained in step S501 to the control end of the system so that the control end performs the monitoring task according to the new monitoring frequency. The control end monitors the system in real time according to the input monitoring frequency to ensure that any potential faults or performance problems can be discovered and handled in time.

[0125] Step S5 aims to dynamically adjust the monitoring strategy based on the second analysis result, so as to more effectively manage performance and prevent potential failures. By selecting an appropriate monitoring strategy, step S5 can ensure that resources are not over-consumed during normal operation, and improve the sensitivity and frequency of monitoring when failures may occur, so as to promptly discover and handle problems. This method of dynamically adjusting the monitoring strategy helps to improve stability and reliability.

[0126] In the first step of the method, a one-way variance analysis is performed on the historical working data to determine whether the load change will affect the performance of the equipment, which provides a basis for the subsequent real-time data collection and analysis, ensuring that more accurate monitoring measures can be taken when the load change significantly affects the performance. If the load change affects the performance, a baseline is established based on the historical data, and a performance prediction model is established based on these baselines. These models can reflect the performance change trend of the equipment under different loads, providing a reliable basis for real-time monitoring. The performance indicator data of the equipment is collected in real time, and these data are compared and analyzed with the prediction model to determine whether the equipment is operating normally. By collecting and analyzing data in real time, abnormal changes in equipment performance can be discovered in time, and the monitoring strategy can be automatically adjusted according to the second analysis result. If the device is about to fail, the monitoring frequency is increased to more closely monitor changes in device performance. If the device is not about to fail, the existing monitoring frequency is maintained to save resources. This adaptive monitoring strategy can be adjusted according to the actual operating status of the device, improving the flexibility and accuracy of monitoring. Compared with the traditional monitoring method based on static thresholds, this method dynamically adjusts the monitoring threshold according to real-time data and historical trends, which can more accurately reflect the performance status of the device and improve the accuracy of the early warning. By introducing advanced technologies such as machine learning algorithms and local outlier factors, it can more effectively identify abnormal changes in device performance and reduce false alarm and missed alarm rates.

[0127] Example 2, reference Figure 2 , provides a data storage device performance monitoring system, including a historical data analysis module, a real-time data acquisition module, a real-time data analysis module, a prediction model building module and a monitoring strategy selection module.

[0128] The historical data analysis module is used to retrieve historical working data from a database, perform a one-way variance analysis on historical performance indicator data corresponding to each historical load in the historical working data, and output a first analysis result.

[0129] The historical data analysis module retrieves historical work data from the database and performs one-way variance analysis to identify the differences and significance of performance indicator data under different historical loads, outputs the first analysis results, and provides a decision-making basis for subsequent modules.

[0130] The real-time data acquisition module is used to collect real-time performance indicator data from the data storage device according to the first analysis result, wherein the real-time performance indicator data includes real-time IOPS, real-time bandwidth, real-time delay and real-time CPU, and store the real-time performance indicator data in a database.

[0131] The real-time data acquisition module collects real-time performance indicator data from the data storage device in a targeted manner according to the first analysis result, and stores the real-time data in the database to provide updates and supplements for historical data analysis and performance prediction.

[0132] The real-time data analysis module is used to calculate descriptive statistics data according to the first analysis result, wherein the descriptive statistics data includes a standard deviation and a mean value, and to establish a baseline according to the descriptive statistics data.

[0133] The real-time data analysis module calculates the descriptive statistics of real-time performance indicator data to evaluate the stability and consistency of the current state, establishes a baseline based on the descriptive statistics data, and provides a reference standard for subsequent performance prediction and anomaly detection.

[0134] The prediction model establishment module is used to establish a first performance prediction model based on the baseline and the corresponding historical working data, establish a corresponding relationship between the performance prediction model and the historical load, output several second performance prediction models, obtain the real-time load, find the second performance prediction model corresponding to the real-time load, obtain the time from the baseline, output it as predicted fault data, analyze the predicted fault data, and output the second analysis result.

[0135] The prediction model building module uses the baseline and historical working data to build a first performance prediction model, and generates multiple second performance prediction models corresponding to different historical loads. It searches for the corresponding prediction model according to the real-time load, combines the time from the baseline, outputs the predicted fault data, analyzes the predicted fault data, and outputs the second analysis result to indicate whether there is a potential failure risk.

[0136] The monitoring strategy selection module is used to select a monitoring strategy according to the second analysis result and input the monitoring strategy to the control end.

[0137] The monitoring strategy selection module intelligently selects a monitoring strategy based on the second analysis result, and inputs the selected monitoring strategy to the control end to adjust the monitoring level and response mechanism to ensure stability and reliability.

[0138] It should be understood by those skilled in the art that the embodiments of the present invention can be provided as methods, systems or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program codes. Among them, the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, referred to as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, referred to as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, referred to as EPROM), programmable read-only memory (Programmable Read-Only Memory, referred to as PROM), read-only memory (Read-Only Memory, referred to as ROM), magnetic memory, flash memory, disk or optical disk. These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0139] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A data storage device performance monitoring method, characterized in that: The steps include: Step S1, retrieving historical working data from a database, performing a one-way variance analysis on historical performance indicator data corresponding to each historical load in the historical working data, and outputting a first analysis result; Step S2, collecting real-time performance indicator data from the data storage device according to the first analysis result, the real-time performance indicator data including real-time IOPS, real-time bandwidth, real-time delay and real-time CPU, and storing the real-time performance indicator data in a database; Step S3, calculating descriptive statistics data according to the first analysis result, wherein the descriptive statistics data include a standard deviation and a mean value, and establishing a baseline according to the descriptive statistics data; Step S4, establishing a first performance prediction model according to the baseline and the corresponding historical working data, establishing a corresponding relationship between the performance prediction model and the historical load, outputting several second performance prediction models, obtaining the real-time load, searching for the second performance prediction model corresponding to the real-time load, obtaining the time from the baseline, outputting the predicted fault data, analyzing the predicted fault data, and outputting the second analysis result; Step S5, selecting a monitoring strategy according to the second analysis result, and inputting the monitoring strategy into the control terminal.

2. A data storage device performance monitoring method as claimed in claim 1, characterized in that: The step S1 includes the following sub-steps: Step S101, retrieving historical work data from a database, the historical work data including historical load and corresponding historical performance indicator data, the historical performance indicator data including historical IOPS, historical bandwidth, historical delay and historical CPU; Step S102: performing a one-way variance analysis on the historical performance indicator data corresponding to each historical load in the historical working data, and outputting a first analysis result.

3. A data storage device performance monitoring method as claimed in claim 2, characterized in that: The one-way ANOVA is: Calculate the total internal sum of squares, inter-group sum of squares, intra-group sum of squares, inter-group degrees of freedom, and intra-group degrees of freedom for each historical load, respectively; The inter-group mean square is calculated based on the inter-group sum of squares and the inter-group degrees of freedom. The mathematical expression of the inter-group mean square is: Among them, MSB is the mean square between groups, SSB is the sum of squares between groups, and dfB is the degrees of freedom between groups; The within-group mean square is calculated based on the within-group sum of squares and the within-group degrees of freedom. The mathematical expression of the within-group mean square is: Among them, MSW is the within-group mean square, SSW is the within-group sum of squares, and dfW is the within-group degrees of freedom; The F ratio is calculated based on the between-group mean square and the within-group mean square, where the F ratio is the ratio of the between-group mean square to the within-group mean square. The F ratio and the corresponding between-group degrees of freedom and within-group degrees of freedom are input into SPSS, the P value is searched, and a correspondence is established between the P value and the corresponding historical IOPS, historical bandwidth, historical latency, and historical CPU, which is output as the first analysis result.

4. A data storage device performance monitoring method as claimed in claim 3, characterized in that: The step S2 includes the following sub-steps: Step S201, obtain the first analysis result and analyze it: If the P value is greater than or equal to 0.05, it is determined that the load change does not affect the performance and the preset constant standard IOPS threshold, constant standard bandwidth threshold, constant standard delay threshold, constant standard CPU threshold and constant standard acquisition frequency are obtained and input to the control end; If the P value is less than 0.05, it is determined that the load change affects the performance and performance indicator data is collected from the data storage device, wherein the real-time performance indicator data includes real-time IOPS, real-time bandwidth, real-time delay and real-time CPU; Step S202, storing the real-time performance indicator data in a database.

5. A data storage device performance monitoring method as claimed in claim 4, characterized in that: The step S3 includes the following sub-steps: Step S301, parsing the first analysis result: If the P value is less than 0.05, the descriptive statistics data of the historical IOPS, historical bandwidth, historical latency and historical CPU corresponding to each historical load are calculated respectively, and the descriptive statistics data include the standard deviation and the mean; Step S302 , using the mean ± standard deviation as a range to establish baselines corresponding to historical IOPS, historical bandwidth, historical latency, and historical CPU.

6. A data storage device performance monitoring method as claimed in claim 5, characterized in that: The step S4 includes the following sub-steps: Step S401, establishing a first performance prediction model according to the baseline and the corresponding historical working data, establishing a corresponding relationship between the performance prediction model and the historical load, and outputting them as a plurality of second performance prediction models; Step S402, obtaining the real-time load, searching for a second performance prediction model corresponding to the real-time load, obtaining the time from the baseline, and outputting it as predicted fault data; Step S403, analyzing the predicted fault data; If the time from the baseline is less than the preset first failure time threshold, it is marked as an impending failure; If the time from the baseline is at or greater than the first fault time threshold, it is marked as not a fault; Impending failure and not-impending failure are output as the second analysis result.

7. A data storage device performance monitoring method as claimed in claim 6, characterized in that: The logic of the first performance prediction model is: Input historical performance indicator data, descriptive statistics data and baseline as input objects to the machine learning algorithm, define the neighborhood distance threshold of each data in the historical performance indicator data and descriptive statistics data, and find the k nearest neighbors of each data; Calculate the local density of each data in the historical performance indicator data and the descriptive statistics data, where the local density is the average of the reciprocals of the densities of the k nearest neighbors; Calculate the LOF value of each data in the historical performance indicator data and the descriptive statistics data, where the LOF value is the ratio of the local density of each data in the historical performance indicator data and the descriptive statistics data to the adjacent local density; The LOF threshold is set according to the baseline, and data points exceeding the LOF threshold are marked as abnormal.

8. A data storage device performance monitoring method as claimed in claim 7, characterized in that: The step S5 includes the following sub-steps: Step S501, selecting a monitoring strategy according to the second analysis result, the monitoring strategy including maintaining the existing monitoring frequency and increasing the existing monitoring frequency, and obtaining a monitoring frequency corresponding to the monitoring strategy, the monitoring frequency including the existing monitoring frequency and the changed monitoring frequency; Step S502, input the monitoring frequency to the control end and perform monitoring.

9. A data storage device performance monitoring method as claimed in claim 8, characterized in that: The selection logic of the monitoring strategy is: If the second analysis result is that there is no fault, the monitoring strategy is determined to maintain the existing monitoring frequency and obtain the preset existing monitoring frequency; If the second analysis result is an impending failure, the monitoring strategy is determined to be to increase the existing monitoring frequency and obtain the time from the baseline, define the time threshold from the baseline and the corresponding monitoring frequency, find the monitoring frequency corresponding to the time from the baseline and output it as the changed monitoring frequency.

10. A data storage device performance monitoring system, applied to a data storage device performance monitoring method as claimed in any one of claims 1 to 8, characterized in that: It includes historical data analysis module, real-time data collection module, real-time data analysis module, prediction model building module and monitoring strategy selection module; The historical data analysis module is used to retrieve historical working data from the database, perform a one-way variance analysis on the historical performance indicator data corresponding to each historical load in the historical working data, and output a first analysis result; The real-time data acquisition module is used to collect real-time performance indicator data from the data storage device according to the first analysis result, the real-time performance indicator data including real-time IOPS, real-time bandwidth, real-time delay and real-time CPU, and store the real-time performance indicator data in a database; The real-time data analysis module is used to calculate descriptive statistics data according to the first analysis result, wherein the descriptive statistics data includes a standard deviation and a mean value, and establish a baseline according to the descriptive statistics data; The prediction model establishment module is used to establish a first performance prediction model according to the baseline and the corresponding historical working data, establish a corresponding relationship between the performance prediction model and the historical load, output as a plurality of second performance prediction models, obtain the real-time load, find the second performance prediction model corresponding to the real-time load, obtain the time from the baseline, output as predicted fault data, analyze the predicted fault data, and output a second analysis result; The monitoring strategy selection module is used to select a monitoring strategy according to the second analysis result and input the monitoring strategy to the control end.

Citation Information

Patent Citations

  • UFS cache performance test method and device, computer equipment and storage medium

    CN117891402A

Cited By

  • Network card management method and system, computer equipment, storage medium and program product

    CN120710873A

  • Network card management method and system, computer device, storage medium and program product

    CN120710873B