Multi-factor Cloud Service Storage Device Error Prediction
Through multi-factor analysis and cost-sensitive machine learning model, the failure probability of storage devices in cloud storage systems is predicted, which solves the problem of difficult to deal with gray failures in traditional methods, realizes early prediction and optimizes the use of storage devices, and improves the availability and economic benefits of cloud services.
Patent Information
- Application Number
- CN201880095070.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2018-06-29
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2038-06-29
AI Technical Summary
The prior art is difficult to predict subtle errors in storage devices in cloud storage systems early, resulting in service interruptions and unnecessary capacity losses at high cost. Traditional methods cannot effectively deal with gray failures and system-level signal changes.
Proactive repair and migration is achieved by combining multi-factor analysis of storage devices and system-level signals using cost-sensitive machine learning models to predict the failure probability of storage devices and identify devices with high failure risk through feature selection and ranking models.
Improve the availability of cloud services, reduce service interruptions, reduce costs due to wrong devices, and optimize storage device usage through early prediction and migration, improving system reliability and economic benefits.
Smart Images

Figure CN112771504B_ABST
Abstract
Description
Technical Field
[0001] The embodiments described herein generally relate to computer disk error detection and, in some embodiments, more particularly to multi-factor storage device error prediction. Background Art
[0002] Cloud service providers maintain data storage, computing, and networking services for entities external to the service provider organization. The data storage can include a cluster of storage devices (e.g., hard disk drives, solid state drives, non-volatile memories, etc.). The computing services can include a cluster of computing nodes hosting virtual computer devices, each virtual computer device using a portion of the computing hardware available to the computing node or cluster. The networking services can include a virtual network infrastructure for interconnecting the virtual computing devices. The operating system and data of a given virtual computing device can be located on one or more storage devices that are distributed across the data storage system. If a storage device storing the operating system or data of a virtual computing device encounters an error, the virtual computing device may experience an unexpected shutdown or may encounter operational problems (e.g., loss of service, slow response time, data loss, errors, etc.). Detecting and replacing a storage device that encounters an error can enable the cloud service provider to reposition the resources of the virtual computing device to a different storage device to mitigate service interruptions.
[0003] Self-Monitoring, Analysis and Reporting Technology (S.M.A.R.T.) is a monitoring system included in certain types of storage devices. S.M.A.R.T. can monitor a storage device and can provide notification of storage device errors. S.M.A.R.T. errors can provide an indication of a failing drive. However, by the time the S.M.A.R.T. error report has been triggered, the virtual computing device using the storage device may already have experienced a service interruption. Summary of the Invention
[0004] The embodiments described herein relate to computer disk error detection and, in particular, to multi-factor storage device error prediction.
[0005] In a first aspect of the present disclosure, a system for storage device error prediction is provided. The system includes: at least one processor; and a memory including instructions that, when executed by the at least one processor, cause the at least one processor to perform operations to: obtain a set of storage device metrics and a set of computing system metrics; generate a feature set including the set of storage device metrics and the set of computing system metrics; perform verification of the features in the feature set by evaluating a validation training dataset using the features in the feature set; create a modified feature set including the verified features in the feature set; create a storage device failure model using the modified feature set, wherein the storage device failure model determines the probability that a given storage device may fail; determine a storage device rating range by minimizing the misclassification cost of the storage device; and identify a set of storage devices to generate an indication of storage devices having a high probability of failure, wherein the set of storage devices includes a plurality of storage devices within the storage device rating range, and wherein the storage devices in the set of storage devices are ranked based on an evaluation of the storage devices using the storage device failure model.
[0006] In a second aspect of the present disclosure, a method for storage device error prediction is provided. The method includes: obtaining a set of storage device metrics and a set of computing system metrics; generating a feature set including the set of storage device metrics and the set of computing system metrics; performing verification of the features in the feature set by evaluating a validation training dataset using the features in the feature set; creating a modified feature set including the verified features in the feature set; creating a storage device failure model using the modified feature set, wherein the storage device failure model determines the probability that a given storage device may fail; determining a storage device rating range by minimizing the misclassification cost of the storage device; and identifying a set of storage devices to generate an indication of storage devices having a high probability of failure, wherein the set of storage devices includes a plurality of storage devices within the storage device rating range, and wherein the storage devices in the set of storage devices are ranked based on an evaluation of the storage devices using the storage device failure model.
[0007] In a third aspect of the present disclosure, at least one non-transitory machine-readable storage medium is provided, including instructions for storing device error prediction, which when executed by at least one processor, cause the at least one processor to perform operations to: obtain a set of storage device metrics and a set of computing system metrics; generate a feature set including the set of storage device metrics and the set of computing system metrics, where the set of storage device metrics includes self-monitoring, analysis, and reporting technology signals from storage devices in a cloud computing storage system, and the set of computing system metrics includes system-level signals from corresponding virtual machines, where operating system data resides on storage devices in the cloud computing storage system; perform verification of features in the feature set by evaluating a validation training data set using features in the feature set; create a modified feature set including the verified features in the feature set; create a storage device failure model using the modified feature set, where the storage device failure model determines the probability that a given storage device may fail; determine a storage device rating range by minimizing the misclassification cost of the storage device, where misclassifying a storage device as having a high failure probability is identified as a first cost and misclassifying a storage device as not having a high failure probability is identified as a second cost, where the storage device rating range is the number that results in the lowest sum of the number of misclassified storage devices multiplied by each of the first cost and the second cost; and identify a set of storage devices to produce an indication of storage devices having a high failure probability, where the set of storage devices includes a plurality of storage devices within the storage device rating range, and where the storage devices in the set of storage devices are ranked based on the evaluation of the storage devices using the storage device failure model. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] In the drawings, which are not necessarily to scale, like numerals may describe like components in different views. Like numerals with different letter suffixes may represent different instances of like components. The drawings generally illustrate, by way of example and not limitation, the various embodiments discussed in this document.
[0009] Figure 1 is a block diagram of an example of an environment and system for multi-factor cloud service storage device error prediction according to one embodiment.
[0010] Figure 2 is a flowchart of an example of a process for multi-factor cloud service storage device error prediction according to one embodiment.
[0011] Figure 3 is a flowchart of an example of a process for feature selection for multi-factor cloud service storage device error prediction according to one embodiment.
[0012] Figure 4 is a flowchart of an example of a method for multi-factor cloud service storage device error prediction according to one embodiment.
[0013] Figure 5 It is a block diagram showing an example of a machine on which one or more embodiments can be implemented. Detailed implementation
[0014] High service availability is crucial for cloud systems. Typical cloud systems use a large number of physical storage devices. As used herein, the term storage device refers to any persistent storage medium used in a cloud storage system. Storage devices can include, for example, hard disk drives (HDDs), solid state drives (SSDs), non-volatile memories (e.g., conforming to the NVM (NVMe) standard, the Non-Volatile Memory Host Controller Interface Specification (NVMHCIS) standard, etc.).
[0015] Storage device errors can be a major cause for a cloud service provider to be unable to use the service. Storage device errors (e.g., sector errors, latency errors, etc.) can be regarded as a form of grey failure, which may be a subtle failure that is difficult to detect even when it causes application errors. The systems and techniques discussed herein can reduce service interruptions due to storage device errors by proactively predicting storage device errors before they cause more serious damage to the cloud system. The ability to predict failing storage devices enables real-time migration of existing virtual machines (VMs) (e.g., virtual computing devices, etc.) and the allocation of new virtual machines to healthy storage devices.
[0016] Proactively transferring workloads to healthy storage devices can also improve service availability. To establish an accurate online prediction model, both storage device-level sensor data and system-level signals are evaluated to identify storage devices that are likely to fail. A cost-sensitive rank-based machine learning model is used, which can learn the characteristics of previously failed storage devices and rank the current storage devices based on the likelihood of the current storage device having errors recently. The solutions discussed herein outperform traditional methods of detecting storage device errors by including system-level signals. Individual system-level signals may not indicate a failure by themselves. However, the combination of system-level signals and storage device-level sensor data can provide an indication of the likelihood of a storage device failing. System-level factors and storage device-level factors can be evaluated against a failure prediction model. The failure prediction model can represent a multi-factor analysis model that can determine the probability that a given storage device will fail. The collection and analysis of system-level and storage device-level data can produce a more accurate prediction of storage device failures compared to single-factor analysis that includes either storage device-level data or system-level data.
[0017] In recent years, software applications are increasingly being deployed as online services on cloud computing platforms. Millions of users globally can use cloud service platforms on a 24 / 7 / 365 basis. As a result, high availability has become essential for cloud-based service providers. Although many cloud service providers aim for high service availability (e.g., 99.999% uptime, etc.), services can fail and cause great dissatisfaction among users and revenue losses. For example, according to a study of data from 63 data center organizations in the United States, the average cost of downtime has steadily increased from $505,502 per hour in 2010 to $740,357 per hour in 2016.
[0018] There are several traditional storage device error detection solutions available. For example, some existing methods have attempted to train a prediction model based on historical storage device failure data and use the trained model to predict whether a storage device will fail in the near future (e.g., whether the storage device will operate). Then proactive remedial measures can be taken, such as replacing storage devices prone to failure. However, the prediction model is mainly built using S.M.A.R.T. data, which is storage device-level sensor data provided by firmware embedded in the storage device drive.
[0019] These existing methods can focus on predicting complete storage device failures (e.g., the storage device is operational / non-operational). However, in a cloud environment, before a complete storage device failure, upper-layer services may already be affected by storage device errors (e.g., experiencing latency errors, timeout errors, sector errors, etc.). These symptoms can include, for example, file operation errors, VMs not responding to communication requests, etc. These subtle failures may not trigger the storage device error detection system to perform a quick and deterministic detection. If no measures are taken, more serious problems may occur or a service interruption may happen. Early prediction of disk failures using the solutions discussed in this article can allow for proactive repairs, including, for example, error-aware VM allocation (e.g., allocating VMs to healthier storage devices), real-time VM migration (e.g., moving VMs from a failing storage device to a healthy storage device), etc. This can allow for the discontinuation of the use of a storage device before an error that may cause a service interruption occurs.
[0020] Storage device errors can be reflected by system-level signals such as operating system (OS) events. A prediction algorithm can be used to rank the health of storage devices in a cloud storage system that contains both S.M.A.R.T. data and system-level signals. Machine learning algorithms can be used to train a prediction model using historical data. The output is a model for predicting storage devices that are likely to fail in the short term. The prediction model can be used to rank all storage devices according to the degree of error proneness of each storage device, so that the cloud service computing system can allocate new VMs and migrate existing VMs to the storage devices ranked as the healthiest under cost and capacity constraints.
[0021] Predicting storage device errors in a cloud service storage system poses challenges. Data imbalance in the cloud storage system makes prediction more difficult. For example, out of 1,000,000 storage devices per day, 300 may fail. The cost of removing healthy storage devices from the system is high and may result in unnecessary capacity loss. Therefore, a cost-sensitive ranking model is used to address this challenge. Storage devices are ranked according to their error proneness, and faulty drives are identified by minimizing the total cost. Using a cost-sensitive ranking model, the r storage devices that are most error-prone can be identified, rather than classifying error-prone storage devices as faulty storage devices. In this way, error-prone storage devices can be removed from the cloud storage system in a cost-effective manner that more closely represents the actual expected failure rate.
[0022] Some features, especially system-level signals, may be time-sensitive (e.g., the value changes sharply continuously over time) or environment-sensitive (e.g., due to the constantly changing cloud environment, its data distribution will change significantly). Models built using these unstable features may lead to good results during cross-validation (e.g., randomly dividing data into training and test sets), but perform poorly during real-world online prediction (e.g., dividing data into training and test sets by time). To address this challenge, unique feature selection techniques are used to perform systematic feature engineering to select stable and predictable features.
[0023] Compared with traditional storage device error detection solutions, the systems and techniques disclosed herein provide more robust and earlier storage device error detection and mitigation. A multi-factor assessment is performed on storage devices in a cloud storage system. A unique feature selection model is used to evaluate both system-level factors and storage device-level factors to select storage device failure prediction features. A cost-sensitive ranking model is used to rank storage devices according to their error proneness. The advantage of this solution is that error-prone storage devices can be identified as early as possible and removed from the cloud storage system before the storage devices cause system downtime. By using cost-sensitive ranking, error-prone storage devices can be removed from the cloud storage system at a rate close to the expected actual failure rate, which can minimize the cost of early removal of storage devices by cloud storage providers.
[0024] Figure 1 FIG. 4 is a block diagram of an example of an environment 100 and a system 120 for multi-factor cloud service storage device error prediction according to one embodiment. The environment 100 may include a cloud service infrastructure 105 that includes a cloud storage system 110 (e.g., a storage area network (SAN), a hyper-converged computing system, a redundant array of inexpensive disks (RAID) array, etc.), the cloud storage system 110 including a plurality of storage devices that store other data including operating system (OS) data and virtual machines (VMs) (e.g., virtual computing devices that share the physical computing hardware of a host computing node, etc.) 115. The cloud storage system 110 and the virtual machines 115 may be communicatively coupled (e.g., via a wired network, a wireless network, a cellular network, a shared bus, etc.) to the system 120. In one example, the system 120 may be a multi-factor storage device error detection engine. The system 120 may include various components such as a metric collector 125, a feature set generator 130, a feature selector 135, a model generator 140, a comparator 145, one or more databases 150, and a storage manager 155.
[0025] The cloud storage system 110 may contain up to hundreds of millions of storage devices serving various services and applications. The storage devices may be used in various types of clusters, such as, for example, clusters for data storage and clusters for cloud applications. The data storage clusters may use redundancy mechanisms such as redundant arrays of inexpensive disks (RAID) technology, which can increase the fault tolerance of the storage devices. The cloud application clusters may host a large number of virtual machines 115, and it should be understood that storage device errors may cause unwanted interruptions to the services and applications hosted by the virtual machines 115. The techniques discussed herein can reduce interruptions by predicting storage device errors before the storage devices cause service failures.
[0026] Service interruptions can lead to revenue loss and user dissatisfaction. Therefore, service providers make every effort to improve service availability. For example, service providers can seek to increase reliability from "four nines" (e.g., 99.99%) to "five nines" (e.g., 99.999%), and then to "six nines" (e.g., 99.9999%). Storage devices are one of the most frequently failing components in the cloud service infrastructure 105 and are thus an important focus for improving reliability.
[0027] Automatically predicting the occurrence of storage device failures can allow cloud service providers to avoid storage device errors that cause system failures (e.g., impacts on the operation of virtual machines 115, etc.). In this way, proactive measures such as storage device replacement can be taken. Traditional storage device error prediction solutions can use Self-Monitoring, Analysis and Reporting Technology (S.M.A.R.T.) data to monitor the internal attributes of individual storage devices to build a failure prediction model.
[0028] However, before a storage device completely fails, the storage device may have started reporting errors. There are various storage device errors, such as storage device partition errors (e.g., storage device volumes and volume sizes become abnormal), latency errors (e.g., an unexpectedly long delay between a request for data and the return of the data), timeout errors (e.g., exceeding a predefined storage device timeout value), and sector errors (e.g., individual sectors on the storage device become unavailable), etc. Storage device failures can be detected by conventional system failure detection mechanisms. However, these conventional mechanisms typically employ overly simplistic failure models in which the storage device is either operating or failed. Such conventional mechanisms are insufficient to handle storage device errors because they may manifest as subtle gray failures.
[0029] Storage device errors are common and can affect the normal operation of upper-layer applications and may cause unexpected downtime of the VM 115. Symptoms can include I / O request timeouts, the VM 115 or container not responding to communication requests, etc. If no measures are taken, more serious problems may occur, even service interruptions. Therefore, it is important to capture and predict storage device errors before VM 115 errors occur.
[0030] The metric collector 125 can obtain a set of storage device metrics from storage devices in the cloud storage system 110 and a set of computing system metrics from the virtual machines 115. The metric collector 125 can store the obtained metrics in the database(s) 150 (e.g., arranged by the relationship between the metrics and the physical storage devices). Two types of data are collected: storage device metrics (e.g., S.M.A.R.T. data, etc.) and computing system metrics (e.g., system-level signals, etc.). For example, S.M.A.R.T. data can be obtained from the monitoring firmware of each storage device, which allows the storage device to report data about its internal activities. Table 1 provides some examples of S.M.A.R.T. characteristics.
[0031]
[0032] Table 1
[0033] In the cloud system, there are also various system-level events that can be collected periodically (e.g., every hour, etc.). Many of these system-level events (e.g., OS events, file system operation errors, unexpected telemetry loss, etc.) are early signals of storage device errors. Table 2 gives descriptions of some system-level signals. In one example, the set of computing system metrics includes system-level signals from virtual machines, where the operating system data resides on storage devices in the cloud computing storage system. For example, FileSystemError is an event caused by a storage device-related error that can be traced back to bad sectors or damaged storage device integrity. These computing system metrics can correspond to the storage devices containing the data of the virtual machines 115 and can be included in the computing system metrics.
[0034]
[0035]
[0036] Table 2
[0037] The feature set generator 130 can use the set of storage device metrics and the set of computing system metrics to generate a feature set. The feature set can be stored in the database(s) 150. In addition to the features directly identified from the raw data, some statistical features can also be calculated, such as:
[0038] Diff: The change of the feature value over time can help distinguish storage device errors. Given a time window w, the Diff of feature x at timestamp t is defined as follows:
[0039] Diff(x, t, w) = x(t) - x(t - w)
[0040] Sigma: Sigma calculates the variance of the attribute values over a specific period. Given a time window w, the Sigma of attribute x at timestamp t is defined as: Sigma(x, t, w) = E[(X - μ) 2 , where X = (xt-w, xt-w-1,..., xt) and
[0041] Bin: Bin calculates the sum of the attribute values in window w as follows:
[0042]
[0043] Three different window sizes (e.g., 3, 5, 7) can be used when calculating Diff, Bin, and Sigma. Multiple features (e.g., 457 features, etc.) can be identified from storage device metrics and computing system metrics. However, not all features can distinguish healthy and faulty storage devices, especially in the case of online prediction. Therefore, the feature selector 135 can perform verification of the features in the feature set by using the feature evaluation of the verification training data set.
[0044] In one example, the verification training data set can be divided into a training data set and a test data set by time. The training data set can be used to train a prediction model. The reference accuracy result can be calculated by predicting the results in the test data set by using the prediction model. The features in the feature set can be removed, and the prediction model can be retrained without the features in the feature set. The feature accuracy result can be calculated by predicting the results in the test data set by using the retrained prediction model. The features in the feature set can also be removed from the test data set. If the reference accuracy result is greater than the feature set accuracy result, the features in the feature set can be verified. In other words, if the prediction model predicts more accurately without the feature compared to the model with the feature, the feature will be removed from the feature set because it is determined that the feature is not a predictive feature.
[0045] The feature set generator 130 can create a modified feature set based on the verification. It has been proven that the feature selection process is very useful for selecting relevant features for constructing a machine learning model. Existing feature selection methods are divided into two major categories: statistical indicators (e.g., chi-square, mutual information, etc.) and machine learning-based methods (e.g., random forest, etc.). Due to the existence of time-sensitive and environment-sensitive features, traditional feature selection methods may not be able to obtain good prediction performance. The information carried by these features is highly correlated with the training period, but may not be applicable to predicting samples in the next period. These represent non-predictive features, which means they have no predictive ability in online prediction.
[0046] The model generator 140 can use the modified feature set to create a storage device failure model. The storage device failure model can represent the probability that a given storage device may fail. After features have been collected from historical data, a storage device failure model (e.g., a prediction model) is then constructed to predict the error propensity of the storage device in the coming days. The prediction problem is formulated as a ranking problem rather than a classification problem. That is, instead of simply determining whether a storage device is faulty, the storage devices are ranked according to their error propensity. The ranking method alleviates the problem of extremely imbalanced failure data because it is insensitive to class imbalance.
[0047] To train the ranking model, historical failure data about the storage devices is obtained and the storage devices are ranked according to their relative failure times (e.g., the number of days between data collections and the detection of the first error). The concept of "learning to rank" is adopted, which can automatically learn from a large amount of data to optimize the ranking model to minimize a loss function. The FastTree algorithm (which is a form of "Multiplicative Additive Regression Tree" (MART) gradient boosting algorithm) can be used to construct each regression tree in a step-by-step manner (e.g., this is a decision tree that has a scalar value in its leaves).
[0048] The comparator 145 can determine the rating range of the storage devices by minimizing the misclassification cost of the storage devices. To improve service availability, new VMs 115 can be assigned to healthier storage devices (e.g., the storage devices ranked as least likely to error, etc.), so that these VMs 115 are less likely to encounter storage device errors in the near future. To achieve this, faulty and healthy storage devices are identified based on the probability of failure of faulty and healthy storage devices. Since most storage devices are healthy and only a small fraction of storage devices are faulty, the top r selected results returned by the ranking model are taken as faulty storage devices.
[0049] In one example, a first cost of misclassifying a storage device as having a high probability of failure and a second cost of misclassifying a storage device as not having a high probability of failure can be identified. The storage device rating range can be the number that results in the lowest sum of the number of misclassified storage devices multiplied by each of the first cost and the second cost. For example, the best top r storage devices are selected such that they can minimize the total misclassification cost:
[0050] cost = Cost 1 * FP r + Cost 2 * FN r ,
[0051] where FP r and FN rThey are respectively the numbers of false positives and false negatives in the first r prediction results. Cost1 is the cost of misidentifying a healthy storage device as faulty, which includes the cost of unnecessary real-time migration from a "faulty" storage device to a healthy storage device. The migration process incurs a non-negligible cost and reduces the capacity of the cloud system. Cost2 is the cost of failing to identify a faulty storage device.
[0052] The values of Cost1 and Cost2 can be determined empirically by experts in the product team. In one example, due to concerns about the migration cost of VM115 and cloud capacity, Cost1 may be much higher than Cost2 (for example, the value of precision is greater than the value of recall). The ratio between Cost1 and Cost2 can be set by domain experts to, for example, 3:1. The numbers of false positives and false negatives are estimated from the false positive rate and false negative rate obtained from historical data. The optimal r value is determined by minimizing the total misclassification cost. The first r storage devices are the predicted faulty storage devices, which are high-risk storage devices, and thus the VMs 115 hosted on them should be migrated out.
[0053] Comparator 145 can identify a set of storage devices to be marked as having a high probability of failure. The set of storage devices can include multiple storage devices equal to the rated range of the storage devices. The storage devices in the set can be ranked based on an assessment of the storage devices using a storage device failure model.
[0054] The ranked storage devices can be marked as retired (e.g., replaced, removed, data migration, etc.) by the storage manager 155 and can be avoided when generating new VMs 115 in the cloud service infrastructure 105. Additionally or alternatively, the VM115 data can be migrated from the marked storage devices to healthy storage devices. The healthy storage devices can be identified as a series of storage devices ranked as having the lowest probability of failure. In one example, the number of identified healthy storage devices can be equal to the number of marked storage devices, and their capacity can be equal to the data stored by the marked storage devices, etc. In one example, the data can be migrated out of the marked storage devices based on their rankings.
[0055] In one example, the healthy storage devices can be identified based on an assessment of the healthy storage devices using a storage device failure model. The data of the virtual machines residing on the member storage devices in the set can be determined, and the data of the virtual machines can be migrated from the members of the storage devices to the healthy storage devices.
[0056] In another example, healthy storage devices can be identified based on an assessment of storage devices using a storage device failure model. A request to create a new virtual machine 115 can be received, and data for the virtual machine can be created on a healthy storage device rather than on the set of storage devices.
[0057] The prediction model can be updated periodically to capture changes occurring in the cloud service infrastructure 105. For example, storage device metrics and computing system metrics can be obtained daily, and storage devices can be ranked and labeled daily. Feature selection can also be updated periodically when new features are identified or historical data is updated to indicate additional features that can predict storage device failures. Thus, as new types of storage devices are added to the cloud storage system 110 and as the cloud service infrastructure 105 evolves, the prediction model can evolve. Additionally, as new technologies become available for managing the storage of VMs 115, the cost calculation function can be adjusted to allow for variable changes in the competing costs of false positive and false negative detection.
[0058] Figure 2 FIG. shows a flowchart of an example of a process 200 for multi-factor cloud service storage device error prediction according to one embodiment. Process 200 can provide features as Figure 1 described.
[0059] The systems and techniques discussed herein predict the error propensity of storage devices based on an analysis of historical data. The ability to predict storage device errors can help improve service availability by proactively allocating VMs to healthier storage devices rather than failing storage devices, and by proactively migrating VMs from predicted failing storage devices to healthy storage devices. Machine learning techniques are used to build a prediction model based on historical storage device error data, and then the model is used to predict the likelihood that a storage device will experience an error in the near future. There are several technical challenges in designing a storage device error prediction model for a large cloud service system:
[0060] (a) Extreme data imbalance: For a large cloud service system, only three out of every ten thousand storage devices may fail per day. The unbalanced storage device failure rate poses difficulties in training classification models. With this unbalanced data, naive classification models may attempt to classify all storage devices as healthy because, in this way, the likelihood of making an incorrect guess is lowest. Some methods may apply data rebalancing techniques (such as oversampling and undersampling techniques) to attempt to address this challenge. These methods may help improve recall, but may simultaneously introduce a large number of false positives, which may lead to a decrease in precision. Removing incorrectly detected storage devices may reduce capacity and may be costly due to causing unnecessary VM migrations.
[0061] (b) Online Prediction: Traditional solutions may solve the prediction problem in a cross - validation manner. However, cross - validation may not be the most effective solution for evaluating the storage device error prediction model. In cross - validation, the dataset can be randomly divided into a training set and a test set. Thus, the training set may contain some future data, while the test set may contain some past data. However, when it comes to online prediction (e.g., using historical data to train the model and predict future states), there is no time overlap between the training and test data.
[0062] In storage device error prediction, some data (especially system - level signals) are time - sensitive (e.g., their values change rapidly over time) or environment - sensitive (e.g., their data distribution may change due to the changing cloud environment). For example, if the storage device rack undergoes an environmental change due to voltage instability or an OS upgrade, all storage devices on it will experience changes. With cross - validation, environment - specific knowledge can be propagated to both the training set and the test set. The knowledge learned from the training set can be applied to the test set. Therefore, in order to construct an effective prediction model in practice, online prediction is used instead of cross - validation. Future knowledge should not be known during prediction.
[0063] Cloud Disk Error Prediction (CDEF) can improve service availability by predicting storage device errors. Process 200 shows an overview of CDEF. First, historical data on faulty and healthy storage devices is collected. Storage device error labels 210 are obtained through root - cause analysis of service problems by on - site engineers. Feature data includes S.M.A.R.T. data 205 and system - level signals 215.
[0064] CDEF addresses some challenges in storage device error detection in cloud systems by incorporating a feature engineering process 220. Features can be identified at operation 225. Features can include features based on the raw data from S.M.A.R.T. data 205 and system - level signals 215 corresponding to the error labels 210.
[0065] Then the features are processed by a feature selection process 230 to select stable and predictive features. A ranking model 235 is generated and used to improve the accuracy of cost - sensitive online prediction. At operation 230, stable and predictive features are selected for training. Based on the selected features, a cost - sensitive ranking model 235 is constructed, which ranks the storage devices.
[0066] The top r storage devices 245 that can identify the misclassification cost of the predicted faulty storage devices can be minimized. For example, the top one hundred most error-prone storage devices can be identified as faulty 250. The remaining storage devices can be unmarked or can be marked as healthy 255. When new data 240 is evaluated by the ranking model 235, the number of storage devices identified as faulty and the characteristics of the storage devices identified as faulty may change.
[0067] Figure 3 FIG. shows a flowchart of an example of a process 300 for feature selection for multi-factor cloud service storage device error prediction according to one embodiment. The process 300 can provide as Figure 1 and Figure 2 described features.
[0068] Some features such as SeekTimePerformance can be non-predictive features. The feature values of healthy storage devices in the training set over time and the feature values of faulty storage devices in the training set can prove that the average feature value of healthy storage devices is lower than the average feature value of faulty storage devices. However, this is not the case. The feature values of healthy and faulty storage devices in the test set over time can show that the average feature value of healthy storage devices is higher than the average feature value of faulty storage devices. Therefore, the behavior of this feature is unstable. Therefore, this feature is considered a non-predictive feature and is not suitable for online prediction. In comparison, predictive features such as ReallocatedSectors can prove to have stable behavior - in both the training and test sets, the values of healthy storage devices are close to zero and the values of faulty storage devices increase over time.
[0069] To select stable and predictive features, a feature selection process 300 is performed to prune out features that will perform poorly in prediction. The goal of the process 300 is to simulate online prediction on the training set.
[0070] Using a feature set F including features (f1, f2,,,, f m ) to obtain training data (TR) (e.g., at operation 305). The training set is divided into two parts over time: one part for training and the other part for validation (e.g., at operation 310). Evaluate each feature f in the feature set F i (e.g., at operation 315).
[0071] Use a model including this feature to calculate an accuracy result (e.g., at operation 320). And use a model not including this feature to calculate an accuracy result (e.g., at operation 325). If the performance on the validation set becomes better after deleting a feature, then the feature is deleted (e.g., at operation 330). Evaluate features until the number of remaining features is less than θ% of the total number of features (e.g., as determined at decision 335). In one example, by default, θ is set to be equal to 10%, which means that if the number of remaining features is less than 10%, the pruning process will stop. If the remaining features are above the threshold, evaluate other features (e.g., at operation 315). If the remaining features are below or equal to the threshold, return the modified feature set (e.g., at operation 340). Then, rescale the range of all selected features using zero-mean normalization as follows: x zero - mean = x - mean(X). The modified feature set includes features determined to be difficult to predict storage device failures.
[0072] Figure 4 Shows an example of a method 400 for multi-factor cloud service storage device error prediction according to one embodiment. Method 400 can provide features as Figures 1 to 3 described.
[0073] A set of storage device metrics and a set of computing system metrics can be obtained (e.g., by a metrics collector 125 as Figure 1 described) (e.g., at operation 405). In one example, the set of storage device metrics can include Self-Monitoring, Analysis and Reporting Technology (S.M.A.R.T.) signals from storage devices in a cloud computing storage system. In one example, the set of computing system metrics includes system-level signals from virtual machines where operating system data resides on a storage device in the cloud computing storage system.
[0074] The set of storage device metrics and the set of computing system metrics can be used to generate a feature set (e.g., by a feature set generator 130 as Figure 1 described) (e.g., at operation 410). In one example, the feature set can include statistical features for calculating statistical values for a time window included in a dataset.
[0075] The members of the feature set can be verified by evaluating a validation training dataset using the members of the feature set (e.g., by a Figure 1The described feature selector 135) (e.g., at operation 415). In one example, the validation training data set can be divided into a training data set and a test data set over time. The training data set can be used to train a prediction model. The reference accuracy result can be calculated by predicting results in the test data set using the prediction model. Members of the feature set can be removed, and the prediction model can be retrained without the members of the feature set. In the case where members of the feature set have been removed from the test data set, the feature accuracy result can be calculated by predicting results in the test data set using the retrained prediction model. If the reference accuracy result is greater than the feature accuracy result, the members of the feature set can be validated.
[0076] A modified feature set can be created based on the validation (e.g., by the feature set generator 130) as described (e.g., at operation 420). The modified feature set can be used to create a storage device failure model (e.g., by the model generator 140) as described (e.g., at operation 425). The storage device failure model can represent the probability that a given storage device may fail. Figure 1 Figure 1 Figure 1 The storage device failure model can represent the probability that a given storage device may fail.
[0077] The storage device rating range can be determined by minimizing the misclassification cost of the storage device (e.g., by the comparator 145) as described (e.g., at operation 430). In one example, a first cost of misclassifying the storage device as having a high failure probability and a second cost of misclassifying the storage device as not having a high failure probability can be identified, and the storage device rating range can be the number that results in the lowest sum of the number of misclassified storage devices multiplied by each of the first cost and the second cost. Figure 1 Figure 1
[0078] A group of storage devices can be identified as being marked as having a high failure probability. The group of storage devices can include a number of storage devices equal to the storage device rating range. The storage devices in the group of storage devices can be ranked based on the evaluation of the storage devices using the storage device failure model.
[0079] In one example, healthy storage devices can be identified based on the evaluation of healthy storage devices using the storage device failure model. The data of the virtual machines residing on the member storage devices in the group of storage devices can be determined, and the data of the virtual machines from the members of the storage devices can be migrated to the healthy storage devices.
[0080] In another example, healthy storage devices can be identified based on the evaluation of healthy storage devices using the storage device failure model. A request to create a new virtual machine can be received, and the data of the virtual machine can be created on the healthy storage device instead of on the group of storage devices.
[0081] Figure 5 FIG. 500 is a block diagram of an example machine on which any one or more of the techniques (e.g., methods) discussed herein may be performed. In alternative embodiments, machine 500 may operate as a stand-alone device or may be connected (e.g., networked) to other machines. In a networked deployment, machine 500 may operate in a server-client network environment as a server machine, a client machine, or both. In one example, machine 500 may act as a peer machine in a peer-to-peer (P2P) (or other distributed) network environment. Machine 500 may be a personal computer (PC), a tablet computer, a set-top box (STB), a personal digital assistant (PDA), a mobile phone, a network appliance, a network router, a switch or bridge, or any machine capable of executing instructions (sequentially or otherwise) that specify actions to be taken by that machine. Further, while only a single machine is illustrated, the term "machine" shall also be taken to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methods discussed herein, such as cloud computing, software as a service (SaaS), and other computer cluster configurations.
[0082] As described herein, an example may include or be operated on by logic or a plurality of components or mechanisms. A circuit set is a collection of circuits implemented in a tangible entity that includes hardware (e.g., simple circuits, gates, logic, etc.). Circuit set membership may be flexible over time and underlying hardware variability. A circuit set includes members that may individually or in combination perform a specified operation when operating. In one example, the hardware of a circuit set may be immutably designed to perform a particular operation (e.g., hardwired). In one example, the hardware of a circuit set may include physically coupled components (e.g., execution units, transistors, simple circuits, etc.) that include a computer-readable medium that is physically modified (e.g., magnetic, electrical, movable placement of invariant mass particles, etc.) to encode instructions for a particular operation. When physically coupling the components, the underlying electrical properties of the hardware components change, such as from an insulator to a conductor and vice versa. The instructions cause the embedded hardware (e.g., execution unit or loading mechanism) to create members of the circuit set through variable connections to perform portions of a particular operation when operating. Thus, when the device is operating, the computer-readable medium is communicatively coupled to other components of the members of the circuit set. In one example, any physical component may be used in more than one member of more than one circuit set. For example, under operation, an execution unit may be used in a first circuit of a first circuit set at one point in time and, at a different time, may be reused by a second circuit in the first circuit set or by a third circuit in a second circuit set.
[0083] A machine (e.g., a computer system) 500 can include a hardware processor 502 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), a hardware processor core, or any combination thereof), a main memory 504, and a static memory 506, some or all of which may communicate with each other via an interconnecting link (e.g., a bus) 508. The machine 500 can also include a display unit 510, an alphanumeric input device 512 (e.g., a keyboard), and a user interface (UI) navigation device 514 (e.g., a mouse). In one example, the display unit 510, the input device 512, and the UI navigation device 514 can be a touch screen display. The machine 500 can additionally include a storage device (e.g., a drive unit) 516, a signal generation device 518 (e.g., a speaker), a network interface device 520, and one or more sensors 521, such as a global positioning system (GPS) sensor, a compass, an accelerometer, or other sensors. The machine 500 can include an output controller 528, such as a serial (e.g., universal serial bus (USB)), parallel, or other wired or wireless (e.g., infrared (IR), near field communication (NFC), etc.) connection, to communicate with or control one or more peripheral devices (e.g., a printer, a card reader, etc.).
[0084] The storage device 516 can include a machine-readable medium 522 having stored thereon a set or multiple sets of data structures or instructions 524 (e.g., software) embodied or utilized by any one or more of the techniques or functions described herein. During execution of the instructions 524 by the machine 500, the instructions 524 can also reside, in whole or at least in part, within the main memory 504, within the static memory 506, or within the hardware processor 502. In one example, one or any combination of the hardware processor 502, the main memory 504, the static memory 506, or the storage device 516 can constitute a machine-readable medium.
[0085] Although the machine-readable medium 522 is shown as a single medium, the term "machine-readable medium" can include a single medium or multiple media (e.g., a centralized or distributed database and / or associated caches and servers) configured to store one or more instructions 524.
[0086] The term "machine-readable medium" can include any medium that can store, encode, or carry instructions executable by a machine 500 and cause the machine 500 to perform any one or more of the techniques of the present disclosure or can store, encode, or carry data structures used by or associated with such instructions. Non-limiting examples of machine-readable media can include solid-state memory and optical and magnetic media. In one example, a machine-readable medium can exclude transitory propagated signals (e.g., non-transitory machine-readable media). Specific examples of non-volatile machine-readable media can include: non-volatile memory such as semiconductor storage devices (e.g., electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM)) and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; CD-ROM and DVD-ROM disks; etc.
[0087] In one example, the machine-readable medium can include storage devices in a cloud service platform (e.g., cloud service infrastructure 105, etc.). In one example, the storage devices can include hard disk drives (HDDs), solid state drives (SSDs), non-volatile memory (e.g., compliant with NVM (NVMe) standards, Non-Volatile Memory Host Controller Interface Specification (NVMHCIS) standards, etc.).
[0088] Instructions 524 can also be transmitted or received via a network interface device 520 over a communication network 526 using a transmission medium using any one of a variety of transmission protocols (e.g., Frame Relay, Internet Protocol (IP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), etc.). Example communication networks can include local area networks (LANs), wide area networks (WANs), packet data networks (e.g., the Internet), mobile telephone networks (e.g., cellular networks), plain old telephone (POTS) networks, and wireless data networks (e.g., the Institute of Electrical and Electronics Engineers (IEEE) 802.11 standard series known as and the Institute of Electrical and Electronics Engineers (IEEE) 802.11 standard series known as the IEEE 802.16 standard series), the IEEE 802.15.4 standard series, peer-to-peer (P2P) networks, the 3rd Generation Partnership Project (3GPP) standards for 4G and 5G wireless communications, including: the 3GPP Long Term Evolution (LTE) standard series, the 3GPP LTE-Advanced standard series, the 3GPP LTE Advanced Pro standard series, the 3GPP New Radio (NR) series of standards, and so on. In one example, the network interface device 520 may include one or more physical jacks (e.g., Ethernet, coaxial, or phone jacks) or one or more antennas to connect to the communication network 526. In one example, the network interface device 520 may include multiple antennas to perform wireless communication using at least one of single-input multiple-output (SIMO), multiple-input multiple-output (MIMO), or multiple-input single-output (MISO) techniques. The term "transmission medium" shall be taken to include any non-transitory medium that is capable of storing, encoding, or carrying instructions executed by the machine 500, and includes digital or analog communication signals or other non-transitory media that facilitate the communication of such software.
[0089] Example 1 is a system for proactive storage device error prediction, the system including: at least one processor; and a memory including instructions that, when executed by the at least one processor, cause the at least one processor to perform operations to: obtain a set of storage device metrics and a set of computing system metrics; generate a feature set using the set of storage device metrics and the set of computing system metrics; perform verification of features in the feature set by evaluating a validation training dataset using features in the feature set; create a modified feature set including the verified features in the feature set; create a storage device failure model using the modified feature set, wherein the storage device failure model determines the probability that a given storage device may fail; determine a storage device rating range by minimization of misclassification cost of the storage device; and identify a set of storage devices to produce an indication of storage devices having a high probability of failure, wherein the set of storage devices includes multiple storage devices within the storage device rating range, and wherein the storage devices in the set of storage devices are ranked based on an evaluation of the storage devices using the storage device failure model.
[0090] In Example 2, the subject matter of Example 1 includes that the memory further includes instructions for: identifying the healthy storage devices based on an evaluation of the healthy storage devices using the storage device failure model; determining data of virtual machines resident on member storage devices in the set of storage devices; and migrating the data of the virtual machines from the member storage devices to the healthy storage devices.
[0091] In Example 3, the subject matter of Examples 1-2 includes that the memory further includes instructions for: identifying healthy storage devices based on an evaluation of the storage device using the storage device failure model; receiving a request to create a new virtual machine; and creating data for the new virtual machine on the healthy storage device in place of the storage device in the set of storage devices.
[0092] In Example 4, the subject matter of Examples 1-3 includes that the instructions for performing verification of the features in the feature set further include instructions for: partitioning the verification training dataset into a training dataset and a test dataset over time; training a prediction model using the training dataset; calculating a reference accuracy result by predicting results in the test dataset using the prediction model; removing a feature from the feature set and retraining the prediction model without the feature in the feature set; calculating a feature accuracy result by predicting results in the test dataset using the retrained prediction model, wherein the feature in the feature set has been removed from the test dataset; and verifying the feature in the feature set if the reference accuracy result is greater than the feature accuracy result.
[0093] In Example 5, the subject matter of Examples 1-4 includes that the set of storage device metrics includes Self-Monitoring, Analysis and Reporting Technology (S.M.A.R.T.) signals from storage devices in a cloud computing storage system.
[0094] In Example 6, the subject matter of Examples 1-5 includes: the set of computing system metrics includes system-level signals from respective virtual machines, where operating system data resides on storage devices in a cloud computing storage system.
[0095] In Example 7, the subject matter of Examples 1-6 includes that the instructions for determining the storage device rating range further include instructions for: identifying a first cost of misclassifying a storage device as having a high failure probability and a second cost of misclassifying the storage device as not having a high failure probability, where the storage device rating range is the number that results in the lowest sum of the number of misclassified storage devices multiplied by each of the first cost and the second cost.
[0096] In Example 8, the subject matter of Examples 1-7 includes that the feature set includes statistical features for calculating statistical values for time windows included in a dataset.
[0097] Example 9 is at least one machine-readable storage medium including instructions for proactive storage device error prediction, which when executed by at least one processor, cause the at least one processor to perform operations to: obtain a set of storage device metrics and a set of computing system metrics; generate a feature set using the set of storage device metrics and the set of computing system metrics; perform verification of features in the feature set by evaluating a validation training data set using features in the feature set; create a modified feature set including the verified features in the feature set; create a storage device failure model using the modified feature set, wherein the storage device failure model determines a probability that a given storage device may fail; determine a storage device rating range by minimizing a misclassification cost of the storage device; and identify a set of storage devices to produce an indication of storage devices having a high probability of failure, wherein the set of storage devices includes a plurality of storage devices within the storage device rating range, and wherein storage devices in the set of storage devices are ranked based on an evaluation of the storage devices using the storage device failure model.
[0098] In Example 10, the subject matter of Example 9 includes instructions for: identifying the healthy storage device based on an evaluation of the healthy storage device using the storage device failure model; determining data of a virtual machine resident on a member storage device in the set of storage devices; and migrating the data of the virtual machine from the member storage device to the healthy storage device.
[0099] In Example 11, the subject matter of Examples 9 - 10 includes instructions for: identifying a healthy storage device based on an evaluation of the storage device using the storage device failure model; receiving a request to create a new virtual machine; and creating data of the new virtual machine on the healthy storage device instead of a storage device in the set of storage devices.
[0100] In Example 12, the subject matter of Examples 9 - 11 includes, wherein the instructions for performing verification of features in the feature set further include instructions for: partitioning the validation training data set into a training data set and a test data set over time; training a prediction model using the training data set; calculating a reference accuracy result by predicting results in the test data set using the prediction model; removing a feature from the feature set and retraining the prediction model without the feature in the feature set; calculating a feature accuracy result by predicting results in the test data set using the retrained prediction model, wherein the feature in the feature set has been removed from the test data set; and verifying the feature in the feature set if the reference accuracy result is greater than the feature accuracy result.
[0101] In Example 13, the subject matter of Examples 9 - 12 includes, wherein the set of storage device metrics includes Self - Monitoring, Analysis and Reporting Technology (S.M.A.R.T.) signals from storage devices in a cloud computing storage system.
[0102] In Example 14, the subject matter of Examples 9 - 13 includes: wherein the set of computing system metrics includes system - level signals from respective virtual machines, where operating system data resides on storage devices in a cloud computing storage system.
[0103] In Example 15, the subject matter of Examples 9 - 14 includes, wherein the instructions for determining the storage device rating range further include instructions for: identifying a first cost of misclassifying a storage device as having a high probability of failure and a second cost of misclassifying the storage device as not having a high probability of failure, where the storage device rating range is the number that results in the lowest sum of the number of misclassified storage devices multiplied by each of the first cost and the second cost.
[0104] In Example 16, the subject matter of Examples 9 - 15 includes, wherein the feature set includes statistical features for calculating statistical values for time windows included in a dataset.
[0105] Example 17 is a method for proactive storage device error prediction, the method comprising: obtaining a set of storage device metrics and a set of computing system metrics; generating a feature set using the set of storage device metrics and the set of computing system metrics; performing verification of the features in the feature set by evaluating a validation training dataset using the features in the feature set; creating a modified feature set including the verified features in the feature set; creating a storage device failure model using the modified feature set, where the storage device failure model determines the probability that a given storage device may fail; determining a storage device rating range by minimizing the cost of misclassifying storage devices; and identifying a set of storage devices to produce an indication of storage devices having a high probability of failure, where the set of storage devices includes a plurality of storage devices within the storage device rating range, and where the storage devices in the set of storage devices are ranked based on an evaluation of the storage devices using the storage device failure model.
[0106] In Example 18, the subject matter of Example 17 includes: identifying the healthy storage devices based on an evaluation of healthy storage devices using the storage device failure model; determining data of virtual machines residing on member storage devices in the set of storage devices; and migrating the data of the virtual machines from the member storage devices to the healthy storage devices.
[0107] In Example 19, the subject matter of Examples 17 - 18 includes: identifying healthy storage devices based on an evaluation of the storage devices using the storage device failure model; receiving a request to create a new virtual machine; and creating the new virtual machine on the healthy storage devices instead of the storage devices in the set of storage devices.
[0108] In Example 20, the subject matter of Examples 17 - 19 includes, wherein performing the verification of the features in the feature set further includes: dividing the verification training data set into a training data set and a test data set over time; training a prediction model using the training data set; calculating a reference accuracy result by predicting results in the test data set using the prediction model; removing a feature from the feature set and retraining the prediction model without the feature in the feature set; calculating a feature accuracy result by predicting results in the test data set using the retrained prediction model, wherein the feature in the feature set has been removed from the test data set; and verifying the feature in the feature set if the reference accuracy result is greater than the feature accuracy result.
[0109] In Example 21, the subject matter of Examples 17 - 20 includes, wherein the set of storage device metrics includes Self - Monitoring, Analysis and Reporting Technology (S.M.A.R.T.) signals from storage devices in a cloud computing storage system.
[0110] In Example 22, the subject matter of Examples 17 - 21 includes, wherein the set of computing system metrics includes system - level signals from respective virtual machines, where operating system data resides on storage devices in a cloud computing storage system.
[0111] In Example 23, the subject matter of Examples 17 - 22 includes, wherein determining the storage device rating range further includes: identifying a first cost of misclassifying a storage device as having a high failure probability and a second cost of misclassifying the storage device as not having a high failure probability, wherein the storage device rating range is the number that results in the lowest sum of the number of misclassified storage devices multiplied by each of the first cost and the second cost.
[0112] In Example 24, the subject matter of Examples 17 - 23 includes, wherein the feature set includes statistical features for calculating statistical values for time windows included in a data set.
[0113] Example 25 is at least one machine - readable medium including instructions that, when executed by a processing circuit system, cause the processing circuit system to perform operations for implementing any one of Examples 1 - 24.
[0114] Example 26 is an apparatus including a module for implementing any one of Examples 1 to 24.
[0115] Example 27 is a system implementing any one of Examples 1 to 24.
[0116] Example 28 is a method implementing any one of Examples 1 to 24.
[0117] The detailed description above includes references to the accompanying drawings, which form a part of the detailed description. The drawings illustrate, by way of example, specific embodiments that may be practiced. Such embodiments are also referred to herein as "examples". Such examples may include other elements in addition to those shown or described. However, the inventors also contemplate examples in which only the elements shown or described are provided. In addition, the inventors also contemplate examples using any combination or arrangement of the elements (or aspects thereof) shown or described for a particular example (or one or more aspects thereof) or for other examples (or one or more aspects thereof) shown or described herein.
[0118] All publications, patents, and patent documents cited herein are hereby incorporated by reference in their entirety as if individually incorporated by reference. If there is an inconsistency in usage between this document and the documents incorporated by reference, the usage in the incorporated (multiple) references shall be regarded as supplementary to the usage in this document; for irreconcilable inconsistencies, the usage in this document shall prevail.
[0119] In this document, the terms "a" or "an" are used as in patent documents to include one or more than one, independent of any other instances or usages of "at least one" or "one or more". In this document, unless otherwise indicated, the term "or" is used to mean a non-exclusive or, such that "A or B" includes "A but not B", "B but not A", and "A and B". In the appended claims, the terms "including" and "in which" are used as the ordinary English equivalents of the respective terms "comprising" and "wherein". Also, in the following claims, the terms "including" and "comprising" are open-ended, i.e., a system, device, article, or process that includes elements other than those listed after such term in the claim is still considered to fall within the scope of the claim. Further, in the appended claims, the terms "first", "second", "third", etc. are used only as labels and are not intended to impose a numerical requirement on their objects.
[0120] The foregoing description is intended to be illustrative and not restrictive. For example, the above examples (or one or more aspects thereof) may be used in combination with each other. After reviewing the foregoing description, other embodiments may be used by, for example, those of ordinary skill in the art. The "Abstract" is provided to enable the reader to quickly ascertain the nature of the technical disclosure and is submitted on the understanding that it will not be used to interpret or limit the scope or meaning of the claims. Additionally, in the foregoing "Detailed Description", various features may be grouped together to streamline the disclosure. This should not be construed as intending that unclaimed disclosed features are essential to any claim. On the contrary, the inventive subject matter may lie in less than all of the features of a particular disclosed embodiment. Accordingly, the appended claims are incorporated into the "Detailed Description", where each claim stands on its own as a separate embodiment. The scope of an embodiment should be determined with reference to the appended claims and the full scope of equivalents to which such claims are entitled.
Claims
1. A system for predicting storage device errors, the system comprising: At least one processor; And A memory including instructions that, when executed by the at least one processor, cause the at least one processor to perform operations to: Obtain a set of storage device metrics and a set of computing system metrics, wherein the set of storage device metrics includes Self-Monitoring, Analysis and Reporting Technology (SMART) signals from storage devices in a cloud computing storage system, and the set of computing system metrics includes system-level signals from corresponding virtual machines, wherein operating system data resides on storage devices in the cloud computing storage system; Generate a feature set including the set of storage device metrics and the set of computing system metrics; Perform verification of the features in the feature set by evaluating a validation training dataset using the features in the feature set; Create a modified feature set including the verified features in the feature set; Create a storage device failure model using the modified feature set, wherein the storage device failure model determines the probability that a given storage device may fail; Determine a storage device rating range by minimizing the misclassification cost of storage devices, wherein misclassifying a storage device as having a high failure probability is identified as a first cost and misclassifying the storage device as not having a high failure probability is identified as a second cost, wherein the storage device rating range is the number that results in the lowest sum of the number of misclassified storage devices multiplied by each of the first cost and the second cost; And Identify a set of storage devices to produce an indication of storage devices having a high failure probability, wherein the set of storage devices includes a plurality of storage devices within the storage device rating range, and wherein the storage devices in the set of storage devices are ranked based on an evaluation of the storage devices using the storage device failure model.
2. The system according to claim 1, wherein the memory further includes instructions for: Identifying the healthy storage devices based on an evaluation of the healthy storage devices using the storage device failure model; Determining data of virtual machines residing on storage devices in the set of storage devices; And Migrating the data of the virtual machines from the storage devices in the set of storage devices to the healthy storage devices.
3. The system according to claim 1, wherein the memory further includes instructions for: Identifying healthy storage devices based on an evaluation of the storage devices using the storage device failure model; Receiving a request to create a new virtual machine; and Creating data of the new virtual machine on the healthy storage devices instead of the storage devices in the set of storage devices.
4. The system according to claim 1, wherein the instructions for performing verification of the features in the feature set further include instructions for: Partitioning the validation training dataset into a training dataset and a test dataset over time; Training a prediction model using the training dataset; Calculating a reference accuracy result by predicting results in the test dataset using the prediction model; Remove the features from the feature set and retrain the prediction model without the features in the feature set; Calculate a feature accuracy result by predicting results in the test data set using the retrained prediction model, where the features in the feature set have been removed from the test data set; and If the reference accuracy result is greater than the feature accuracy result, verify the features in the feature set.
5. The system according to claim 1, wherein the feature set includes statistical features for calculating statistical values for time windows included in the validation training data set.
6. A method for storage device error prediction, the method comprising: Obtain a set of storage device metrics and a set of computing system metrics; Generate a feature set including the set of storage device metrics and the set of computing system metrics, where the set of storage device metrics includes Self-Monitoring, Analysis and Reporting Technology signals from storage devices in a cloud computing storage system, and the set of computing system metrics includes system-level signals from corresponding virtual machines, where operating system data resides on storage devices in the cloud computing storage system; Perform verification of the features in the feature set by evaluating a validation training data set using the features in the feature set; Create a modified feature set including the verified features in the feature set; Create a storage device failure model using the modified feature set, where the storage device failure model determines the probability that a given storage device may fail; Determine a storage device rating range by minimizing the misclassification cost of the storage device, where misclassifying the storage device as having a high failure probability is identified as a first cost and misclassifying the storage device as not having a high failure probability is identified as a second cost, where the storage device rating range is the number that results in the lowest sum of the number of misclassified storage devices multiplied by each of the first cost and the second cost; Identify a set of storage devices to produce an indication of storage devices having a high failure probability, where the set of storage devices includes a plurality of storage devices within the storage device rating range, and where the storage devices in the set of storage devices are ranked based on an evaluation of the storage devices using the storage device failure model.
7. The method according to claim 6, further comprising: Identify the healthy storage devices based on an evaluation of healthy storage devices using the storage device failure model; Determine the data of the virtual machines residing on the storage devices in the set of storage devices; And Migrate the data of the virtual machines from the storage devices in the set of storage devices to the healthy storage devices.
8. The method according to claim 6, further comprising: Identify healthy storage devices based on an evaluation of the storage devices using the storage device failure model; Receive a request to create a new virtual machine; And Create the data of the new virtual machine on the healthy storage device instead of the storage device in the set of storage devices.
9. The method according to claim 6, wherein performing the verification of the features in the feature set further comprises: Dividing the verification training data set into a training data set and a test data set by time; Training a prediction model using the training data set; Calculating a reference accuracy result by predicting results in the test data set using the prediction model; Removing the features in the feature set and retraining the prediction model without the features in the feature set; Calculating a feature accuracy result by predicting results in the test data set using the retrained prediction model, wherein the features in the feature set have been removed from the test data set; and Verifying the features in the feature set if the reference accuracy result is greater than the feature accuracy result.
10. The method according to claim 6, wherein the feature set includes statistical features for calculating statistical values for time windows included in the verification training data set.
11. At least one non-transitory machine-readable storage medium, comprising instructions for storing device error prediction, which when executed by at least one processor, cause the at least one processor to perform operations to: Obtain a set of storage device metrics and a set of computing system metrics; Generate a feature set including the set of storage device metrics and the set of computing system metrics, wherein the set of storage device metrics includes self-monitoring, analysis, and reporting technology signals from storage devices in a cloud computing storage system, and the set of computing system metrics includes system-level signals from corresponding virtual machines, wherein operating system data resides on a storage device in the cloud computing storage system; Perform verification of the features in the feature set by evaluating a verification training data set using the features in the feature set; Create a modified feature set including the verified features in the feature set; Create a storage device failure model using the modified feature set, wherein the storage device failure model determines the probability that a given storage device may fail; Determine a storage device rating range by minimizing the misclassification cost of the storage device, wherein misclassifying the storage device as having a high failure probability is identified as a first cost and misclassifying the storage device as not having a high failure probability is identified as a second cost, wherein the storage device rating range is the number that results in the lowest sum of the number of misclassified storage devices multiplied by each of the first cost and the second cost; And Identify a set of storage devices to generate an indication of storage devices having a high failure probability, wherein the set of storage devices includes a plurality of storage devices within the storage device rating range, and wherein the storage devices in the set of storage devices are ranked based on an evaluation of the storage devices using the storage device failure model.
12. The at least one non-transitory machine-readable storage medium according to claim 11, further comprising instructions which, when executed by at least one processor, cause the at least one processor to perform operations to: Identify the healthy storage device based on an assessment of a healthy storage device using the storage device failure model; Determine data of a virtual machine resident on a storage device in the set of storage devices; And Migrate the data of the virtual machine from the storage device in the set of storage devices to the healthy storage device.
13. The at least one non-transitory machine-readable storage medium according to claim 11, further comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations to: Identify a healthy storage device based on an assessment of the storage device using the storage device failure model; Receive a request to create a new virtual machine; And Create data of the new virtual machine on the healthy storage device in place of the storage device in the set of storage devices.
14. The at least one non-transitory machine-readable storage medium according to claim 11, further comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations to: Partition the validation training data set into a training data set and a test data set over time; Train a prediction model using the training data set; Calculate a reference accuracy result by predicting results in the test data set using the prediction model; Remove a feature from the feature set and retrain the prediction model without the feature in the feature set; Calculate a feature accuracy result by predicting results in the test data set using the retrained prediction model, wherein the feature in the feature set has been removed from the test data set; and Validate the feature in the feature set if the reference accuracy result is greater than the feature accuracy result.
15. The at least one non-transitory machine-readable storage medium according to claim 11, wherein the feature set includes statistical features for calculating statistical values for time windows included in the validation training data set.
Citation Information
Patent Citations
Time sequence classification early warning method for storage device
CN108052528A