Method, device, equipment, medium and program for predicting recovery time

By acquiring and normalizing the feature values of the storage device, and using the deep neural network model to predict the recovery time of the streamlined pool, the problem that manual experience cannot accurately determine the recovery time and improve the service reliability of the storage device.

CN120104394BActive Publication Date: 2025-08-08INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510591973.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-08-08
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

In the prior art, the recovery time of the storage device streamlined pool cannot be accurately determined through manual experience, resulting in the storage device being unable to recover on time, affecting service reliability.

Method used

By obtaining the initial eigenvalues of multiple target features, calculating the feature mean and standard deviation, and normalizing the process, the target deep neural network model is used to predict the recovery time of the streamlined pool.

Benefits of technology

No manual prediction is required, which improves the reliability of the service provided by the storage device and the accuracy of predictions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104394B_ABST
    Figure CN120104394B_ABST
Patent Text Reader

Abstract

This application discloses a recovery time prediction method, apparatus, device, medium, and program, relating to the field of server repair technology. The method comprises obtaining initial feature values corresponding to multiple target features, where the target features are features that affect the recovery time of a target thin pool in a target storage device; obtaining feature means and standard deviations corresponding to each target feature; normalizing the multiple initial feature values based on the feature means and standard deviations corresponding to each target feature to obtain multiple target feature values; and predicting the recovery time of the thin pool based on the multiple target feature values using a target deep neural network model to obtain a target recovery time. This method can improve the reliability of services provided by storage devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of server repair technology, and in particular to a method, apparatus, device, medium, and program for predicting recovery time. Background Art

[0002] T3 recovery restores system configurations and data after all cluster nodes on a storage device fail. After a T3 recovery, a thin pool recovery is required to safely redeploy and bring the storage device back online. The duration of thin pool recovery varies depending on the storage device configuration. Long recovery times can impact subsequent storage device operations. Based on experience and the storage device configuration, staff can determine the duration of thin pool recovery and schedule the storage device recovery accordingly.

[0003] However, it is impossible to accurately determine the recovery time of the thin pool through manual experience, resulting in the storage device being unable to resume work on time, making the reliability of the service provided by the storage device low. Summary of the Invention

[0004] The present application provides a method, apparatus, device, medium, and program for predicting recovery time, in order to at least solve the problem of low reliability of services provided by storage devices in related technologies.

[0005] This application provides a method for predicting recovery time, including:

[0006] Obtaining initial feature values corresponding to multiple target features, where the target features are features that affect the recovery time of a target thin pool in a target storage device;

[0007] Get the feature mean and standard deviation corresponding to each target feature;

[0008] According to the feature mean and standard deviation corresponding to each target feature, multiple initial feature values are normalized to obtain multiple target feature values;

[0009] The target recovery time of the target thin pool is predicted through a target deep neural network model and based on multiple target feature values to obtain the target recovery time.

[0010] The present application provides a recovery time prediction device, comprising: a first acquisition module, a second acquisition module, a normalization processing module and a prediction module, wherein:

[0011] The first acquisition module is used to obtain initial feature values corresponding to multiple target features, where the target features are features that affect the recovery time of a target thin pool in a target storage device;

[0012] The second acquisition module is used to obtain the feature mean and standard deviation corresponding to each target feature;

[0013] The normalization processing module is used to normalize multiple initial eigenvalues according to the feature mean and standard deviation corresponding to each target feature to obtain multiple target eigenvalues;

[0014] The prediction module is used to predict the recovery time of the target thin pool through a target deep neural network model and based on multiple target feature values to obtain a target recovery time.

[0015] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned recovery time prediction methods when executing the computer program.

[0016] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned recovery time prediction methods are implemented.

[0017] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned recovery time prediction methods when executed by a processor.

[0018] Through the recovery time prediction method, device, equipment, medium and program provided in this application, it is possible to obtain initial feature values corresponding to multiple target features, obtain the feature mean and standard deviation corresponding to each target feature, normalize the multiple target feature values according to the feature mean and standard deviation corresponding to each target feature, and obtain multiple target feature values. Through the target deep neural network model, the thin pool recovery time is predicted based on the multiple target feature values to obtain the target recovery time. There is no need for manual recovery time prediction, which can improve the reliability of services provided by storage devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0020] Figure 1 A schematic diagram of an application scenario provided in an embodiment of the present application;

[0021] Figure 2 A flowchart of a recovery time prediction method provided in an embodiment of the present application;

[0022] Figure 3 A schematic diagram of a process for determining target features provided in an embodiment of the present application;

[0023] Figure 4 A flowchart of a model training method provided in an embodiment of the present application;

[0024] Figure 5 A schematic diagram of the architecture of a deep neural network model provided in an embodiment of the present application;

[0025] Figure 6 A schematic diagram of the architecture of a model configuration provided in an embodiment of the present application;

[0026] Figure 7 A schematic diagram of the structure of a recovery time prediction device provided in an embodiment of the present application;

[0027] Figure 8 This is a schematic diagram of the structure of an electronic device provided in this application. DETAILED DESCRIPTION

[0028] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0029] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0030] If all cluster nodes on a storage device fail, you can use T3 to restore the storage device's system configuration and data. T3 recovery can rebuild the system using saved configuration data and XML backups. However, T3 recovery only retrieves the latest XML file from the quorum disk for system recovery and cannot restore all volume data.

[0031] The specific steps for T3 recovery can be as follows:

[0032] Step 1: Create a storage cluster and ensure normal operation. Step 2: Generate a configuration backup file: Run the mcsconfig backup command on the configuration node to generate a backup file. This file will be used for cluster data recovery. Step 3: Simulate a T3 failure trigger: Simultaneously remove the power modules or controllers on all nodes to cause a failure, or run the reboot command on all nodes to simulate removing a controller. Step 4: Node recovery and status reset: Reinsert the power modules (or controllers) on all nodes and wait for the nodes to boot and the operating system to load. Wait for the node status to change from 878 to 578, then run the mtop leavecluster –force command to forcibly leave the cluster, placing all nodes in the candidate state and waiting to rejoin the cluster. Step 5: Prepare the T3 recovery environment: Run the mtop t3recovery -prepare command, check the command execution status using the mtinq lscmdstatus command, and wait for the command to complete. Step 6: Cluster recovery: Run the mtop t3recovery -execute command, check the command execution status using the mtinq lscmdstatus command, and wait for the command to complete. Step 7. Repair the thin pool: Run the recovervdiskbypool mdiskgrp_id command to repair the thin pools one by one.

[0033] After T3 recovery, all data stored on the storage device can be further restored through thin pool repair. This process locates, repairs, or reconstructs damaged data blocks, effectively filling data gaps and restoring data to its pre-failure state by leveraging redundant backups or distributed storage. This allows the storage device to be safely redeployed and brought online.

[0034] In related technologies, the duration of thin pool recovery varies depending on the configuration of the storage device. A long recovery time can affect subsequent storage device operations. Based on experience and the storage device configuration, personnel can determine the recovery time for thin pool recovery and schedule the storage device recovery accordingly. However, manual experience alone cannot accurately determine the recovery time for thin pool recovery, resulting in storage devices being unable to resume operations on time and lowering the reliability of the services provided by the storage devices.

[0035] The recovery time prediction method provided in the embodiments of the present application can obtain initial feature values corresponding to multiple target features, obtain feature means and standard deviations corresponding to each target feature, normalize the multiple target feature values according to the feature means and standard deviations corresponding to each target feature, obtain multiple target feature values, and predict the recovery time of the thin pool based on the multiple target feature values through a target deep neural network model to obtain the target recovery time. This eliminates the need for manual recovery time prediction and can improve the reliability of services provided by the storage device.

[0036] Figure 1 This is a schematic diagram of an application scenario provided by an embodiment of this application. Figure 1 , including a storage device 101 and a detection device 102. A storage system can be installed in storage device 101 to provide storage services. Storage device 101 can manage multiple cluster nodes through the storage system, which are used to store data. The storage resources of the multiple cluster nodes in storage device 101 can be managed through a resource pool. A thin pool is a virtualized storage resource pool that uses thin provisioning technology to dynamically allocate physical space and relies on metadata to track usage.

[0037] The configuration information of the storage device 101 may include the number of processor cores, processor main frequency, memory size and memory main frequency, etc. The configuration information of the thin pool corresponding to the storage device 101 may include pool type, volume type, number of volumes, storage pool capacity, number of storage pools, amount of written data, compression ratio, deduplication ratio and write mode, etc.

[0038] The detection device 102 may obtain initial feature values corresponding to multiple target features from the storage device 101 , and obtain a target recovery time for thin pool recovery by processing the multiple initial feature values.

[0039] Hereinafter, various embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. Note that, in the accompanying drawings, the same reference numerals are given to components having substantially the same or similar structures and functions, and repeated descriptions thereof will be omitted.

[0040] Figure 2 This is a flowchart of a recovery time prediction method provided in an embodiment of the present application. Figure 2 , the method may include:

[0041] S201. Obtain initial feature values corresponding to multiple target features.

[0042] The target characteristic may be a characteristic that affects the recovery time of the target thin pool in the target storage device.

[0043] When recovering a thin pool of a storage device, both the storage device configuration information and the thin pool configuration information may affect the recovery time of the thin pool. However, some features have little or no impact on the recovery time. It is only necessary to obtain the initial feature values corresponding to multiple target features from the storage device configuration information and the thin pool configuration information.

[0044] The multiple target characteristics may include storage device characteristics and thin pool characteristics. Storage device characteristics may include the number of processor cores, processor frequency, memory size, and memory frequency. Storage thin pool characteristics may include storage pool capacity, number of storage pools, written data volume, compression ratio, deduplication ratio, and write mode.

[0045] For example, assuming that multiple target features are the number of processor cores, memory size, storage pool capacity, number of storage pools, amount of written data, and compression ratio, then the number of processor cores can be obtained as 16, memory size as 32G, storage pool capacity as 5T, number of storage pools as 10, amount of written data as 20T, and compression ratio as 2.5 times compression in the configuration information of the storage device and the configuration information of the thin pool.

[0046] The thin pool serves as a virtualization layer. The configuration information of the thin pool is concentrated in the metadata managed by the software. The configuration information of the storage device is hardware parameters, which are directly managed by the underlying firmware and controller.

[0047] S202: Obtain the feature mean and standard deviation corresponding to each target feature.

[0048] In some possible embodiments, for any target feature, the feature mean may be obtained by averaging multiple historical feature values corresponding to the target feature during the training phase, and the standard deviation may be the standard deviation of multiple historical feature values corresponding to the target feature during the training phase.

[0049] In some possible embodiments, the feature mean and standard deviation corresponding to each target feature may be preset.

[0050] In some possible embodiments, for any target feature, multiple recorded feature values of the storage device at different stages may be recorded, and the feature mean and standard deviation corresponding to the target feature may be determined based on the multiple recorded feature values.

[0051] S203 : Normalize the multiple initial eigenvalues according to the feature mean and standard deviation corresponding to each target feature to obtain multiple target eigenvalues.

[0052] In some possible embodiments, for any target feature, the difference between the initial eigenvalue corresponding to the target feature and the feature mean can be determined as the first difference corresponding to the target feature; and the ratio of the first difference corresponding to the target feature to the standard deviation can be determined as the target feature value corresponding to the target feature.

[0053] Please refer to the following formula:

[0054]

[0055] in, is the initial eigenvalue, is the characteristic mean, is the standard deviation, is the target feature value.

[0056] In an embodiment of the present application, normalization processing can be performed to ensure that the eigenvalues of the input target deep neural network model maintain the same dimension as the target deep neural network model during the training phase. After eliminating the dimensional difference, the accuracy of the prediction can be improved.

[0057] S204: predicting the recovery duration of the target thin pool using a target deep neural network model and based on multiple target feature values to obtain a target recovery duration.

[0058] The target deep neural network model is a prediction model built based on a deep neural network.

[0059] A deep neural network (DNN) is a complex model composed of multiple layers of neurons. A DNN usually includes an input layer, a hidden layer, and an output layer.

[0060] The input layer can receive raw data, which can be multiple target feature values.

[0061] The hidden layer can be responsible for extracting high-order features of the data through multi-layer nonlinear transformations. The number of hidden layers and the number of neurons in the hidden layers determine the depth and complexity of the model. For example, if there are 3 hidden layers, the number of nodes in each hidden layer can be 16, 32, and 64 respectively.

[0062] The output layer can provide the final prediction result, which can be the target recovery time.

[0063] Each neuron in the hidden layer can receive the output of all neurons in the previous layer, calculate the input according to the weight and activation function, and pass the result to the next layer of neurons.

[0064] In this application, when predicting the recovery time of the thin pool, by stacking multiple such hidden layers, the DNN can gradually extract high-level features in the data that affect the recovery time of the thin pool, thereby achieving efficient learning and processing of complex data.

[0065] The target deep neural network model can be a solidified deep neural network model after training with training samples. By directly inputting multiple target feature values into the target deep neural network model, the target prediction duration can be obtained.

[0066] The recovery time prediction method provided in the embodiment of the present application can obtain initial feature values corresponding to multiple target features, normalize the multiple target feature values according to the feature mean and standard deviation corresponding to each target feature, and obtain multiple target feature values. Through the target deep neural network model, the thin pool recovery time is predicted based on the multiple target feature values, and the target recovery time can be obtained. There is no need for manual recovery time prediction, which can improve the reliability of services provided by the storage device.

[0067] Before training the deep neural network model using a sample set, target features that affect the recovery time of the thin pool can be identified from multiple configuration features to reduce the number of features input into the initial neural network model. This can shorten the training time, reduce the impact of invalid features, and improve prediction accuracy.

[0068] Next, combine Figure 3 , further illustrating the execution process of determining target features provided in the embodiment of the present application.

[0069] Figure 3 This is a flow chart of determining target features provided in an embodiment of the present application. Figure 3 , the method may include:

[0070] S301. Acquire multiple configuration features.

[0071] The multiple configuration characteristics may include at least one first configuration characteristic of the target storage device and at least one second configuration characteristic corresponding to the target thin pool.

[0072] Among them, at least one first configuration feature can be the number of processor cores, processor main frequency, memory size and memory main frequency, etc., and at least one second configuration feature can be the storage pool capacity, number of storage pools, amount of written data, compression rate, deduplication rate and write mode, etc.

[0073] S302: Perform correlation analysis on the multiple configuration features and the target thin pool recovery duration to obtain multiple target features.

[0074] Correlation analysis methods can include Pearson product-distance correlation coefficient, non-parametric correlation analysis, principal component analysis, multiple regression analysis, and grey clustering and weight assignment method.

[0075] In the embodiment of the present application, the correlation between multiple configuration features and the recovery time can be analyzed using the Pearson product-moment correlation coefficient.

[0076] Specifically, multiple feature value groups corresponding to multiple configuration features can be obtained, and the feature value groups can include multiple configuration feature values corresponding to the configuration features and the historical recovery duration corresponding to each configuration feature value; for any configuration feature, the correlation between the configuration feature and the recovery duration is determined based on the multiple configuration feature values and the multiple historical recovery durations corresponding to the configuration feature; the configuration feature with a correlation greater than a preset correlation is determined as the target feature to obtain multiple target features.

[0077] For example, assuming there are five configuration features, namely configuration features 1-5, configuration feature 1 can be the number of sub-processor cores, configuration feature 2 can be the memory size, configuration feature 3 can be the storage pool capacity, configuration feature 4 can be the number of storage pools, and configuration feature 5 can be the amount of written data. Multiple configuration feature values corresponding to each configuration feature and the historical recovery duration corresponding to each configuration feature value can be obtained. This can be seen in Table 1. Based on the multiple configuration feature values and multiple historical recovery durations corresponding to the configuration feature, the correlation between the configuration feature and the recovery duration can be determined, that is, configuration feature 1 corresponds to correlation 1, configuration feature 2 corresponds to correlation 2, configuration feature 3 corresponds to correlation 3, configuration feature 4 corresponds to correlation 4, and configuration feature 5 corresponds to correlation 5.

[0078] Table 1

[0079] For any configuration feature, the correlation can be calculated as follows: determine the first mean of multiple configuration feature values and the second mean of multiple historical recovery time lengths; determine the covariance and standard deviation corresponding to the configuration feature based on the first mean and the second mean; and determine the ratio of the covariance and the standard deviation as the correlation corresponding to the configuration feature.

[0080] Specifically, please refer to the following formula:

[0081] in, is the correlation, the number of multiple configuration feature values and multiple historical recovery durations is , is the first mean of multiple configuration eigenvalues, It is the second average of multiple historical recovery times. is the eigenvalue of the i-th configuration, is the i-th historical recovery time, is the covariance corresponding to the configuration feature, is the standard deviation corresponding to the configuration feature.

[0082] The range of correlation is If the correlation is greater than 0, the configuration feature is positively correlated with the recovery time. If the correlation is less than 0, the configuration feature is negatively correlated with the recovery time. The closer the absolute value of the correlation is to 1, the closer the configuration feature is to the recovery time. If the correlation is 0, it indicates that there is no correlation between the configuration feature and the recovery time.

[0083] For example, assuming that the preset correlation can be 0.5, assuming that there are 5 configuration features in total, the correlation 1 corresponding to configuration feature 1 is 0.3, the correlation corresponding to configuration feature 2 is 0.6, the correlation corresponding to configuration feature 3 is -0.8, the correlation corresponding to configuration feature 4 is -0.2, and the correlation corresponding to configuration feature 5 is 0.8, then configuration feature 2, configuration feature 3 and configuration feature 5 can be determined as target features.

[0084] It is worth noting that the preset relevance and target configuration provided in the embodiments of the present application are merely examples. The preset relevance and target configuration can be determined according to actual conditions and are not specifically limited here.

[0085] The recovery time prediction method provided in the embodiments of the present application can identify target features related to recovery time from among multiple configuration features through correlation analysis, thereby improving the accuracy of recovery time prediction. Furthermore, correlation analysis using the Pearson product-distance correlation coefficient can identify features that are negatively and positively correlated with recovery time, thereby improving the accuracy of feature correlation analysis.

[0086] Before predicting the recovery time of the thin pool of the target storage device in a targeted manner, it is necessary to train the initial deep neural network model to obtain the target deep neural network model. Figure 4 , the execution process of the model training provided in the embodiment of the present application is explained.

[0087] Figure 4 This is a flow chart of a model training method provided in an embodiment of the present application. Figure 4 , the method may include:

[0088] S401. Acquire multiple target samples.

[0089] The target sample may include multiple target historical feature values and the historical recovery time of the target sample annotation;

[0090] Because multiple initial samples are obtained from different storage devices and at different times, there is a lack of uniformity in the measurements between the multiple initial samples. Therefore, the multiple initial samples need to be normalized. The specific normalization process is as follows:

[0091] Acquire multiple initial samples, each including multiple initial historical feature values corresponding to multiple target features and historical recovery durations annotated with the initial samples. For any target feature, obtain multiple initial historical feature values corresponding to the target feature from the multiple initial samples. Determine the feature mean and standard deviation corresponding to the target feature based on the multiple initial historical feature values. Normalize the multiple initial historical feature values based on the feature mean and standard deviation to obtain multiple target historical feature values. Obtain multiple target samples based on the multiple target historical feature values corresponding to each initial sample and the historical recovery duration annotated with each initial sample.

[0092] The specific normalization execution process can be found in the normalization description of the execution process of predicting the recovery time in the above embodiment, which will not be repeated here.

[0093] The multiple target samples may include multiple training samples and multiple test samples. The multiple training samples are used to train the initial deep neural network model, and the multiple test samples are used to test the trained initial deep neural network model.

[0094] S402: Train the initial deep neural network model based on multiple training samples until a preset convergence condition is reached, and determine the deep neural network model that reaches the preset convergence condition as the neural network model to be tested.

[0095] The initial deep neural network model includes an input layer, multiple hidden layers, and an output layer. The number of neurons in the input layer is the number of multiple target historical feature values. Since the prediction target is the recovery time of the streamlined pool, the number of neurons in the output layer is 1.

[0096] The activation function in the hidden layer of the initial deep neural network model can be a rectified linear unit (for example, ReLU). ReLU can introduce nonlinearity, alleviate gradient vanishing, be computationally efficient, and enable sparse activation. The activation function in the output layer can be a linear activation function (for example, Linear). Linear can directly output continuous values, avoiding truncation or compression of the output range, and is suitable for regression tasks.

[0097] The initial deep neural network may be trained for a preset number of training rounds using multiple training samples. The preset number of training rounds may be a maximum number of training rounds for training the multiple training samples during the training process.

[0098] A preset training round (epoch) can fully train the model using all training samples once. For example, if there are 10,00 samples in total, when the model traverses these 10,00 samples, one epoch is completed.

[0099] The preset convergence condition can minimize the loss function for the target.

[0100] The target loss function is the sum of the preset product and the initial loss function. The preset product is the product of the regularization term and the regularization coefficient. The target loss function can be found in the following formula:

[0101] in, is the target loss function, is the initial loss function, is the preset product, is the regularization coefficient, is the regularization term, is the initial weight of the jth hidden layer of the initial deep neural network model, is the number of hidden layers of the initial deep neural network model, is an integer greater than 1, and the regularization term is used to indicate the sum of the squares of all initial weights; For historical restoration time, To predict the repair time, is a preset difference used to control the limit of the error size. When the error is less than δ, the mean squared error (MSE) is used, and when the error is greater than δ, the mean absolute error (MAE) is used. is the mean square error, is the mean absolute error.

[0102] The larger the value of the regularization coefficient, the stronger the constraint on the weight, and the range is generally between 0.001 and 0.1.

[0103] The initial loss function is the Huber Loss function, which combines MSE and MAE. MSE can be used for small errors, and MAE can be used for large errors. When dealing with regression problems, it can both smooth the training process and reduce the impact of outliers.

[0104] In addition, through L2 regularization, the sum of the squares of the weight parameters is added to the initial loss function to limit the weight amplitude, which can make the model parameters smoothly distributed, avoid the occurrence of extreme values, and prevent the model from being overly dependent on a single feature (for example, the number of processor cores).

[0105] The initial deep neural network model may include an initial weight matrix, where the initial weight matrix is used to indicate initial weights of all hidden layers in the initial deep neural network.

[0106] When training the initial deep neural network model based on multiple training samples, the initial weight matrix is updated through the optimization function; wherein the optimization function is the difference between the first weight matrix and the gradient update term, the first weight matrix is the initial weight matrix minus the product of the weight attenuation coefficient and the initial weight matrix, and the gradient update term is determined by the target loss function value and the learning rate.

[0107] The optimization function can be the AdamW optimizer, and the formula can be seen as follows:

[0108] in, is the target weight matrix; is the first weight matrix, is the weight attenuation coefficient, is the initial weight matrix; is the gradient update term, used to adjust the model weight; is the learning rate, which is used to control the parameter update step size; is the bias-corrected first-order momentum estimate, is the bias-corrected second-order momentum estimate; It is a numerical stability term, a preset minimum constant; To indicate the training, To indicate the training, is an integer greater than or equal to 1.

[0109] The learning rate can be adjusted using cosine decay, a learning rate optimization method used in deep learning training to help models converge. This method gradually reduces the learning rate over the training cycle, using a larger learning rate early in training to more quickly approach the global minimum and a smaller learning rate for fine-tuning later in the training cycle. This improves model stability and generalization.

[0110] In the embodiment of the present application, the attenuation term can be separated from the update of the first-order moment and the second-order moment by optimizing the function, making the optimization process more stable.

[0111] Early stopping can monitor the performance of the validation set and terminate training early when the model starts to overfit, thereby saving parameters when the model is in the best generalization state and avoiding ineffective subsequent training.

[0112] Specifically, if there are multiple consecutive training rounds with the same target loss function value, the training of the deep neural network model is stopped, and the deep neural network model with the smallest target loss function value is determined as the neural network model to be tested.

[0113] The target loss function can be used as a validation set. If the target loss function value does not improve within multiple consecutive epochs (for example, the number of epochs is 10 and the target loss function value does not improve), the training is stopped. The model weights with the minimum target loss function value can be automatically retained.

[0114] In this application, by stopping training early, invalid training can be avoided, the training efficiency can be improved, and overfitting can be prevented to ensure the generalization ability of the model.

[0115] The initial deep neural network model includes multiple initial hidden layers, and the number of the multiple initial hidden layers can be indivual.

[0116] In deep neural network models, Dropout is a regularization technique that prevents the model from becoming overly dependent on certain features or neurons by randomly dropping (i.e., temporarily shutting down) some neurons, thereby improving the generalization ability and robustness of the model.

[0117] When an initial deep neural network model is trained based on multiple training samples, when multiple abstract feature data of an initial hidden layer are input into a next initial hidden layer, multiple selected feature data and at least one data to be processed are determined from the multiple abstract feature data, and the at least one data to be processed is set to zero.

[0118] For example, during the training process, the proportion of randomly "closing" some neurons can be set to 20%. Assuming that there are 100 abstract feature data in total, the number of multiple selected feature data is 80, and the number of at least one data to be processed is 20.

[0119] Figure 5 This is a schematic diagram of the architecture of a deep neural network model provided in an embodiment of the present application. Figure 5The deep neural network model includes an input layer, a first hidden layer, a second hidden layer, a third hidden layer, and an output layer. The input layer may include 128 neurons, the first hidden layer may include 256 neurons, the second hidden layer may include 128 neurons, the third hidden layer may include 64 neurons, and the output layer may include 1 neuron. The activation function uses Reluctant Unit (ReLU). When the 256 abstract feature data corresponding to the 256 neurons in the first hidden layer are input into the second hidden layer, the probability of random zeroing (dropout) is 0.3, meaning that 30% of the abstract feature data can be randomly zeroed. When the 128 abstract feature data corresponding to the 128 neurons in the second hidden layer are input into the third hidden layer, the probability of random zeroing (dropout) is 0.2, meaning that 20% of the abstract feature data can be randomly zeroed.

[0120] Figure 6 This is a schematic diagram of the architecture of a model configuration provided in the embodiment of this application. Figure 6 When configuring the initial deep neural network model, the Huber loss function was used, the AdamW optimizer was used, the cosine annealing learning rate schedule was used, the L2 regularization with dropout was used, and the number of early stopping attempts was 10. The batch size was set to 64, the number of epochs was set to 100, and the validation set monitoring was either MSE or MAE.

[0121] S403: Use multiple test samples to test the neural network model to be tested, and determine the deep neural network model that meets the test requirements as the target deep neural network model.

[0122] In some embodiments, the prediction error between the predicted recovery time and the historical recovery time can be calculated using MSE to verify whether the neural network model to be tested will meet the test requirements.

[0123] MSE can measure the squared average of the prediction error between the predicted repair time and the historical repair time. The smaller the MSE, the more accurate the model prediction.

[0124] In some embodiments, the prediction error between the predicted recovery time and the historical recovery time can be calculated by MAE to verify whether the neural network model to be tested will meet the test requirements.

[0125] MAE can measure the average difference in prediction error between the predicted repair time and the historical repair time. The smaller the MAE, the more accurate the model prediction.

[0126] The test requirement may be that the MSE value or the MAE value is less than or equal to a preset threshold.

[0127] The recovery time prediction method provided in the embodiment of the present application can train an initial neural network model through the historical feature values of the storage device and the corresponding historical recovery time to obtain a target deep neural network model. When determining the recovery time of the thin pool, the target features of the storage device can be directly processed through the target deep neural network model to obtain the target recovery time, thereby improving the efficiency of the target recovery time prediction. In addition, there is no need for manual empirical prediction, which can improve the accuracy of the prediction and thereby improve the reliability of the services provided by the storage device.

[0128] Figure 7 This is a schematic diagram of the structure of a recovery time prediction device provided in an embodiment of the present application. Figure 7 The recovery time prediction device 700 may include a first acquisition module 701, a second acquisition module 702, a normalization processing module 703 and a prediction module 704, wherein:

[0129] The first acquisition module 701 is used to obtain initial feature values corresponding to multiple target features, where the target features are features that affect the recovery time of a target thin pool in a target storage device;

[0130] The second acquisition module 702 is used to obtain the feature mean and standard deviation corresponding to each target feature;

[0131] The normalization processing module 703 is used to perform normalization processing on the multiple initial eigenvalues according to the feature mean and standard deviation corresponding to each target feature to obtain multiple target eigenvalues;

[0132] The prediction module 704 is configured to predict the recovery duration of the target thin pool by using a target deep neural network model and according to a plurality of target feature values to obtain a target recovery duration.

[0133] In one possible embodiment, the recovery time prediction device 700 further includes a third acquisition module and a correlation analysis module, wherein:

[0134] The third acquisition module is used to acquire a plurality of configuration features, where the plurality of configuration features include at least one first configuration feature of the target storage device and at least one second configuration feature corresponding to the target thin pool;

[0135] The correlation analysis module is used to perform correlation analysis on a plurality of configuration features and a recovery duration of a target thin pool to obtain a plurality of target features.

[0136] In one possible embodiment, the correlation analysis module is specifically configured to:

[0137] Acquire multiple feature value groups corresponding to multiple configuration features, where the feature value groups include multiple configuration feature values corresponding to the configuration features and historical recovery durations corresponding to the configuration feature values;

[0138] For any configuration feature, determine the correlation between the configuration feature and the recovery duration based on multiple configuration feature values corresponding to the configuration feature and multiple historical recovery durations;

[0139] The configuration features whose absolute values of the correlations are greater than the preset correlations are determined as target features to obtain a plurality of target features.

[0140] In one possible embodiment, the correlation analysis module is specifically configured to:

[0141] Determining a first mean of a plurality of configuration characteristic values and a second mean of a plurality of historical recovery durations;

[0142] Determine the covariance and standard deviation corresponding to the configuration feature based on the first mean and the second mean;

[0143] The ratio of the covariance to the standard deviation is determined as the correlation corresponding to the configuration features.

[0144] In one possible embodiment, the normalization processing module 703 is specifically configured to:

[0145] For any target feature, the difference between the initial eigenvalue and the feature mean corresponding to the target feature is determined as the first difference corresponding to the target feature;

[0146] For any target feature, the ratio of the first difference value corresponding to the target feature to the standard deviation is determined as the target feature value corresponding to the target feature.

[0147] In one possible embodiment, the recovery time prediction device 700 further includes a fourth acquisition module, a training module, and a testing module:

[0148] The fourth acquisition module is used to acquire multiple target samples, the multiple target samples include multiple training samples and multiple test samples, the target samples include multiple target historical feature values, and historical recovery time of target sample annotations;

[0149] The training module is used to train the initial deep neural network model based on multiple training samples until a preset convergence condition is reached, and determine the deep neural network model that reaches the preset convergence condition as the neural network model to be tested;

[0150] The testing module is used to test the neural network model to be tested using multiple test samples, and determine the deep neural network model that meets the test requirements as the target deep neural network model.

[0151] In one possible embodiment, the preset convergence condition is that the target loss function is minimized; the target loss function is the sum of a preset product and an initial loss function, and the preset product is the product of the regularization term and the regularization coefficient.

[0152] In one possible embodiment, the initial loss function includes a mean square error and a mean absolute error. The mean square error is used when the difference between the historical recovery time and the predicted recovery time is greater than or equal to a preset difference, and the mean absolute error is used when the difference between the historical recovery time and the predicted recovery time is less than a preset difference.

[0153] In one possible embodiment, the initial deep neural network model includes an initial weight matrix; when the initial deep neural network model is trained based on multiple training samples, the initial weight matrix is updated by an optimization function;

[0154] Among them, the optimization function is the difference between the first weight matrix and the gradient update term. The first weight matrix is the initial weight matrix minus the product of the weight attenuation coefficient and the initial weight matrix. The gradient update term is determined by the target loss function value and the learning rate.

[0155] In one possible embodiment, the initial deep neural network is trained for a preset number of training rounds using the plurality of training samples, where the preset number of training rounds is the maximum number of training rounds for training the plurality of training samples during the training process;

[0156] Among them, if there are multiple consecutive training rounds with the same target loss function value, the training of the deep neural network model is stopped, and the deep neural network model with the smallest target loss function value is determined as the neural network model to be tested.

[0157] In one possible embodiment, the initial deep neural network model includes multiple initial hidden layers; when the initial deep neural network model is trained based on multiple training samples, when multiple abstract feature data of the initial hidden layer are input into the next initial hidden layer, multiple selected feature data and at least one to-be-processed data are determined from the multiple abstract feature data, and the at least one to-be-processed data is set to zero.

[0158] In one possible embodiment, the fourth acquisition module is specifically configured to:

[0159] Acquire multiple initial samples, where the initial samples include multiple initial historical feature values corresponding to multiple target features, and historical recovery durations annotated with the initial samples;

[0160] For any target feature, multiple initial historical feature values corresponding to the target feature are obtained from multiple initial samples; the feature mean and standard deviation corresponding to the target feature are determined based on the multiple initial historical feature values; and the multiple initial historical feature values are normalized based on the feature mean and standard deviation to obtain multiple target historical feature values.

[0161] Multiple target samples are obtained based on multiple target historical feature values corresponding to each initial sample and the historical recovery time marked for each initial sample.

[0162] For the description of the features in the embodiment corresponding to the recovery time prediction device, reference can be made to the relevant description of the embodiment corresponding to the recovery time prediction method, which will not be repeated here.

[0163] Figure 8 This is a schematic diagram of the structure of an electronic device provided by this application. Figure 8 As shown, the electronic device 800 provided in this embodiment includes: at least one processor 801 and a memory 802. Optionally, the electronic device 800 further includes a communication component 803. The processor 801, the memory 802 and the communication component 803 are connected via a bus.

[0164] During the specific implementation process, at least one processor 801 executes the computer-executable instructions stored in the memory 802, so that the at least one processor 801 executes the above-mentioned embodiment of the method for predicting the recovery time.

[0165] The specific implementation process of the processor 801 can be found in the above method embodiment. Its implementation principle and technical effects are similar and will not be repeated here in this embodiment.

[0166] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the application may be directly executed by a hardware processor or by a combination of hardware and software modules within the processor.

[0167] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage.

[0168] A bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, and control buses. For ease of illustration, the buses in the drawings of this application are not limited to just one bus or just one type of bus.

[0169] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned recovery duration prediction embodiments when running.

[0170] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0171] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned recovery time prediction method embodiments are implemented.

[0172] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps in any of the above-mentioned recovery time prediction method embodiments.

[0173] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0174] The above is a detailed introduction to a memory scheduling method provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A method for predicting recovery time, characterized in that: include: Acquire a plurality of configuration features, where the plurality of configuration features include at least one first configuration feature of a target storage device and at least one second configuration feature corresponding to a target thin pool; Acquire multiple feature value groups corresponding to the multiple configuration features, where the feature value groups include multiple configuration feature values corresponding to the configuration features and historical recovery durations corresponding to the configuration feature values; For any configuration feature, determining the correlation between the configuration feature and the recovery duration according to multiple configuration feature values corresponding to the configuration feature and multiple historical recovery durations; Determining the configuration features whose absolute values of the correlations are greater than a preset correlation as target features to obtain multiple target features; Obtaining initial feature values corresponding to the plurality of target features, where the target features are features that affect a target thin pool recovery time in a target storage device; Get the feature mean and standard deviation corresponding to each target feature; Normalizing the multiple initial eigenvalues according to the feature means and standard deviations corresponding to the target features to obtain multiple target eigenvalues; The recovery duration of the target thin pool is predicted by a target deep neural network model and according to the multiple target feature values to obtain a target recovery duration.

2. The method according to claim 1, characterized in that Determining, based on multiple configuration feature values corresponding to the configuration feature and multiple historical recovery durations, a correlation between the configuration feature and the recovery duration, including: Determining a first average of the plurality of configuration characteristic values and a second average of the plurality of historical recovery durations; Determining a covariance and a standard deviation corresponding to the configuration feature based on the first mean and the second mean; The ratio of the covariance to the standard deviation is determined as the correlation corresponding to the configuration feature.

3. The method according to claim 1, characterized in that Normalizing the multiple initial eigenvalues according to the feature means and standard deviations corresponding to the target features to obtain multiple target eigenvalues, including: For any target feature, the difference between the initial feature value and the feature mean corresponding to the target feature is determined as the first difference corresponding to the target feature; For any target feature, the ratio of the first difference value corresponding to the target feature to the standard deviation is determined as the target feature value corresponding to the target feature.

4. The method according to any one of claims 1 to 3, characterized in that The training steps of the target deep neural network model include: Acquire multiple target samples, the multiple target samples including multiple training samples and multiple test samples, the target samples including multiple target historical feature values and historical recovery time annotated by the target samples; Training the initial deep neural network model based on multiple training samples until a preset convergence condition is reached, and determining the deep neural network model that reaches the preset convergence condition as the neural network model to be tested; The neural network model to be tested is tested using the multiple test samples, and a deep neural network model that meets the test requirements is determined as a target deep neural network model.

5. The method according to claim 4, characterized in that The preset convergence condition is that the target loss function reaches a minimum; the target loss function is the sum of a preset product and an initial loss function, and the preset product is the product of a regularization term and a regularization coefficient.

6. The method according to claim 5, characterized in that The initial loss function includes the mean square error and the mean absolute error. The mean square error is used when the difference between the historical recovery time and the predicted recovery time is greater than or equal to the preset difference. The mean absolute error is used when the difference between the historical recovery time and the predicted recovery time is less than the preset difference.

7. The method according to claim 6, characterized in that The initial deep neural network model includes an initial weight matrix; when the initial deep neural network model is trained based on multiple training samples, the initial weight matrix is updated by an optimization function; The optimization function is the difference between a first weight matrix and a gradient update term, the first weight matrix is the product of the initial weight matrix minus the weight attenuation coefficient and the initial weight matrix, and the gradient update term is determined by the target loss function value and the learning rate.

8. The method according to claim 4, characterized in that Performing a preset number of training rounds on the initial deep neural network using the multiple training samples, where the preset number of training rounds is a maximum number of training rounds performed on the multiple training samples during a training process; Among them, if there are multiple consecutive training rounds with the same target loss function value, the training of the deep neural network model is stopped, and the deep neural network model with the smallest target loss function value is determined as the neural network model to be tested.

9. The method according to claim 4, characterized in that The initial deep neural network model includes multiple initial hidden layers; when the initial deep neural network model is trained based on multiple training samples, when multiple abstract feature data of the initial hidden layer are input into the next initial hidden layer, multiple selected feature data and at least one to-be-processed data are determined from the multiple abstract feature data, and the at least one to-be-processed data is set to zero.

10. The method according to claim 4, characterized in that Acquire multiple target samples, including: Acquire multiple initial samples, where the initial samples include multiple initial historical feature values corresponding to multiple target features, and historical recovery durations annotated with the initial samples; For any target feature, obtain multiple initial historical feature values corresponding to the target feature from the multiple initial samples; determine a feature mean and a standard deviation corresponding to the target feature based on the multiple initial historical feature values; and perform normalization processing on the multiple initial historical feature values based on the feature mean and standard deviation to obtain multiple target historical feature values; The multiple target samples are obtained according to the multiple target historical feature values corresponding to the initial samples and the historical recovery time marked on the initial samples.

11. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the method for predicting the recovery time as claimed in any one of claims 1 to 10 when executing the computer program.

12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the method for predicting the recovery time according to any one of claims 1 to 10 are implemented.

13. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method for predicting the recovery time according to any one of claims 1 to 10 are implemented.

Citation Information

Patent Citations

  • Prediction model generation method and device for repair duration of simplified pool and prediction method

    CN116340853A

  • Time length prediction method of target time period, equipment and medium

    CN119378756A