Recovery duration prediction method, device, equipment, medium and program

The deep neural network model predicts the recovery time of streamlined pools, which solves the problem of difficulty in accurate predictions by manual prediction and improves the service reliability of storage devices.

CN120104394AActive Publication Date: 2025-06-06INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510591973.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-06-06
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

In the prior art, it is difficult to accurately determine the recovery time of streamlined pool recovery through manual experience, resulting in the inability to restore work on time, reducing the reliability of the storage device providing services.

Method used

The deep neural network model is used to predict the recovery time of the streamlined pool. By obtaining the initial feature values, feature mean and standard deviations of multiple target features, normalizing the process, and then using the deep neural network model for prediction.

Benefits of technology

There is no need to manually predict the recovery time, which improves the reliability of storage device service and ensures that storage devices can be restored and operated online on time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104394A_ABST
    Figure CN120104394A_ABST
Patent Text Reader

Abstract

The invention discloses a recovery duration prediction method and device, equipment, a medium and a program, and relates to the technical field of server recovery, initial feature values corresponding to a plurality of target features are obtained, and the target features are features affecting the recovery duration of a target simplified pool in target storage equipment; obtaining a feature mean value and a standard deviation corresponding to each target feature; according to the feature mean value and the standard deviation corresponding to each target feature, performing normalization processing on the plurality of initial feature values to obtain a plurality of target feature values; and through the target deep neural network model, predicting the recovery duration of the simplified pool according to the plurality of target feature values to obtain a target recovery duration. The service providing reliability of the storage device can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of server repair technology, and in particular to a method, device, equipment, medium and program for predicting recovery time. Background Art

[0002] T3 recovery can restore system configuration and data after all cluster nodes of the storage device fail. After T3 recovery, the storage device needs to be safely redeployed and put online through thin pool recovery. The recovery time of thin pool recovery varies depending on the configuration of the storage device. If the recovery time is long, it will affect the subsequent work of the storage device. The staff can determine the recovery time of thin pool recovery based on work experience and the configuration of the storage device, and then arrange the recovery work of the storage device.

[0003] However, it is impossible to accurately determine the recovery time of the thin pool through manual experience, resulting in the storage device being unable to resume work on time, making the reliability of the service provided by the storage device low. Summary of the invention

[0004] The present application provides a recovery time prediction method, apparatus, device, medium and program to at least solve the problem of low reliability of services provided by storage devices in related technologies.

[0005] The present application provides a method for predicting recovery time, including:

[0006] Acquire initial feature values ​​corresponding to a plurality of target features, where the target features are features that affect the recovery time of a target thin pool in a target storage device;

[0007] Get the feature mean and standard deviation corresponding to each target feature;

[0008] According to the feature mean and standard deviation corresponding to each target feature, multiple initial feature values ​​are normalized to obtain multiple target feature values;

[0009] The target recovery time of the target thin pool is predicted through a target deep neural network model according to multiple target feature values ​​to obtain the target recovery time.

[0010] The present application provides a recovery duration prediction device, comprising: a first acquisition module, a second acquisition module, a normalization processing module and a prediction module, wherein:

[0011] The first acquisition module is used to acquire initial feature values ​​corresponding to a plurality of target features, where the target features are features that affect the recovery time of a target thin pool in a target storage device;

[0012] The second acquisition module is used to obtain the feature mean and standard deviation corresponding to each target feature;

[0013] The normalization processing module is used to perform normalization processing on multiple initial eigenvalues ​​according to the feature mean and standard deviation corresponding to each target feature to obtain multiple target eigenvalues;

[0014] The prediction module is used to predict the recovery time of the target thin pool through a target deep neural network model and according to multiple target feature values ​​to obtain a target recovery time.

[0015] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned recovery time prediction methods when executing the computer program.

[0016] The present application also provides a computer-readable storage medium, in which a computer program is stored, wherein when the computer program is executed by a processor, the steps of any of the above-mentioned recovery time prediction methods are implemented.

[0017] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned recovery time prediction methods when the computer program is executed by a processor.

[0018] Through the recovery time prediction method, device, equipment, medium and program provided in the present application, it is possible to obtain initial feature values ​​corresponding to multiple target features, obtain feature means and standard deviations corresponding to each target feature, normalize multiple target feature values ​​according to the feature means and standard deviations corresponding to each target feature, obtain multiple target feature values, use a target deep neural network model, and predict the recovery time of the thin pool based on the multiple target feature values ​​to obtain the target recovery time. There is no need for manual recovery time prediction, and the reliability of services provided by the storage device can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0020] Figure 1 A schematic diagram of an application scenario provided in an embodiment of the present application;

[0021] Figure 2 A flowchart of a recovery time prediction method provided in an embodiment of the present application;

[0022] Figure 3 A schematic diagram of a process for determining target features provided in an embodiment of the present application;

[0023] Figure 4 A flowchart of a model training method provided in an embodiment of the present application;

[0024] Figure 5 A schematic diagram of the architecture of a deep neural network model provided in an embodiment of the present application;

[0025] Figure 6 A schematic diagram of the architecture of a model configuration provided in an embodiment of the present application;

[0026] Figure 7 A schematic diagram of the structure of a recovery time prediction device provided in an embodiment of the present application;

[0027] Figure 8 A schematic diagram of the structure of an electronic device provided in this application. DETAILED DESCRIPTION

[0028] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0029] It should be noted that, in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, and are not used to describe a specific order or sequence.

[0030] In the event that all cluster nodes in the storage device fail, the system configuration and data of the storage device can be restored through T3. T3 recovery can rebuild the system through the saved configuration data and XML backup. However, T3 recovery can only obtain the latest XML file in the arbitration disk for system recovery, and cannot restore all volume data.

[0031] The specific steps of T3 recovery can be as follows:

[0032] Step 1: Create a storage cluster and ensure normal operation; Step 2: Generate a configuration backup file: Execute mcsconfig backup on the configuration node to generate a backup file, and use this file to restore cluster data; Step 3: Simulate T3 fault trigger: Unplug all node power modules or controllers at the same time to create a fault, or execute the reboot command on all nodes at the same time to simulate the controller unplugging; Step 4: Node recovery and status reset: Plug back all node power modules (or controllers) and wait for the node to start until the operating system is loaded. Wait for the node status to change from 878->578, then execute the mtop leavecluster–force command to force exit the cluster, set all nodes to candidate status, and wait to rejoin the cluster. Step 5: Prepare the T3 recovery environment: Execute the mtop t3recovery -prepare command, check the command execution status through mtinq lscmdstatus, and wait for the command to be executed; Step 6: Cluster recovery process: Execute the mtop t3recovery -execute command, check the command execution status through mtinq lscmdstatus, and wait for the command to be executed. Step 7. Repair the thin pool: Execute recovervdiskbypool mdiskgrp_id to repair the thin pools one by one.

[0033] After T3 recovery, all data stored in the storage device can be further recovered through thin pool repair. The thin pool repair process can locate, repair or rebuild damaged data blocks, effectively fill data gaps by taking advantage of redundant backup or distributed storage, and restore data to the state before the failure. The storage device can then be safely redeployed and put online.

[0034] In the related art, the recovery time of the thin pool recovery varies due to the configuration of the storage device. If the recovery time is long, it will affect the subsequent work of the storage device. The staff can determine the recovery time of the thin pool recovery based on the configuration of the storage device according to their work experience, and then arrange the recovery work of the storage device. However, it is impossible to accurately determine the recovery time of the thin pool recovery through manual experience, resulting in the inability of the storage device to resume work on time, making the reliability of the service provided by the storage device low.

[0035] The recovery time prediction method provided in the embodiment of the present application can obtain initial feature values ​​corresponding to multiple target features, obtain feature means and standard deviations corresponding to each target feature, normalize multiple target feature values ​​according to the feature means and standard deviations corresponding to each target feature, obtain multiple target feature values, and predict the recovery time of the thin pool based on the multiple target feature values ​​through a target deep neural network model to obtain the target recovery time. There is no need to manually predict the recovery time, which can improve the reliability of services provided by the storage device.

[0036] Figure 1 A schematic diagram of an application scenario provided by an embodiment of the present application. Figure 1 , including a storage device 101 and a detection device 102. A storage system can be installed in the storage device 101, and storage services are provided through the storage system. The storage device 101 can manage multiple cluster nodes through the storage system, and the cluster nodes are used to store data. The storage resources of the corresponding multiple cluster nodes in the storage device 101 can be managed through a resource pool. A thin pool is a virtualized storage resource pool that uses thin configuration technology, dynamically allocates physical space, and relies on metadata to track usage.

[0037] The configuration information of the storage device 101 may include the number of processor cores, processor main frequency, memory size and memory main frequency, etc. The configuration information of the thin pool corresponding to the storage device 101 may include pool type, volume type, number of volumes, storage pool capacity, number of storage pools, amount of written data, compression rate, deduplication rate and write mode, etc.

[0038] The detection device 102 may obtain initial feature values ​​corresponding to multiple target features in the storage device 101 , and obtain a target recovery duration of the thin pool by processing the multiple initial feature values.

[0039] Hereinafter, various embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. It should be noted that in the accompanying drawings, the same reference numerals are given to components having substantially the same or similar structures and functions, and repeated descriptions thereof will be omitted.

[0040] Figure 2 A flowchart of a recovery time prediction method provided in an embodiment of the present application. Figure 2 , the method may include:

[0041] S201. Obtain initial feature values ​​corresponding to multiple target features.

[0042] The target feature may be a feature that affects the recovery time of the target thin pool in the target storage device.

[0043] When recovering a thin pool of a storage device, the configuration information of the storage device and the configuration information of the thin pool may affect the recovery time of the thin pool recovery. However, some features have little effect on the recovery time or even no effect. It is only necessary to obtain initial feature values ​​corresponding to multiple target features from the configuration information of the storage device and the configuration information of the thin pool.

[0044] The multiple target features may include features of storage devices and features of thin pools. The features of storage devices may include the number of processor cores, processor frequency, memory size, and memory frequency, and the features of storage thin pools may include storage pool capacity, number of storage pools, amount of written data, compression rate, deduplication rate, and write mode, etc.

[0045] For example, assuming that multiple target features are the number of processor cores, memory size, storage pool capacity, number of storage pools, amount of written data, and compression rate, then the number of processor cores can be obtained as 16, the memory size is 32G, the storage pool capacity is 5T, the number of storage pools is 10, the amount of written data is 20T, and the compression rate is 2.5 times compression in the configuration information of the storage device and the configuration information of the thin pool.

[0046] The thin pool serves as a virtualization layer. The configuration information of the thin pool is concentrated in the metadata managed by the software. The configuration information of the storage device is the hardware parameters, which are directly managed by the underlying firmware and controller.

[0047] S202: Obtain the feature mean and standard deviation corresponding to each target feature.

[0048] In some possible embodiments, for any target feature, the feature mean may be obtained by averaging multiple historical feature values ​​corresponding to the target feature during the training phase, and the standard deviation may be the standard deviation of multiple historical feature values ​​corresponding to the target feature during the training phase.

[0049] In some possible embodiments, the feature mean and standard deviation corresponding to each target feature may be preset.

[0050] In some possible embodiments, for any target feature, multiple recorded feature values ​​of the storage device at different stages may be recorded, and the feature mean and standard deviation corresponding to the target feature may be determined based on the multiple recorded feature values.

[0051] S203. Normalize the multiple initial eigenvalues ​​according to the feature mean and standard deviation corresponding to each target feature to obtain multiple target eigenvalues.

[0052] In some possible embodiments, for any target feature, the difference between an initial feature value corresponding to the target feature and the feature mean can be determined as a first difference corresponding to the target feature; and the ratio of the first difference corresponding to the target feature to the standard deviation can be determined as a target feature value corresponding to the target feature.

[0053] See the following formula:

[0054]

[0055] in, is the initial eigenvalue, is the feature mean, is the standard deviation, is the target feature value.

[0056] In an embodiment of the present application, normalization processing can be used to ensure that the eigenvalues ​​of the input target deep neural network model maintain the same dimension as the target deep neural network model during the training phase. After eliminating the dimensional difference, the accuracy of the prediction can be improved.

[0057] S204. Predicting the recovery duration of the target thin pool by using a target deep neural network model and according to multiple target feature values ​​to obtain a target recovery duration.

[0058] The target deep neural network model is a prediction model built based on a deep neural network.

[0059] A deep neural network (DNN) is a complex model composed of multiple layers of neurons. DNN usually includes an input layer, a hidden layer, and an output layer.

[0060] The input layer can receive raw data, which can be multiple target feature values.

[0061] The hidden layer can be responsible for extracting high-order features of the data through multi-layer nonlinear transformations, where the number of hidden layers and the number of neurons in the hidden layers determine the depth and complexity of the model. For example, if there are 3 hidden layers, the number of nodes in each hidden layer can be 16, 32, and 64, respectively.

[0062] The output layer can provide the final prediction result, which can be the target recovery time.

[0063] Each neuron in the hidden layer can receive the output of all neurons in the previous layer, calculate the input according to the weight and activation function, and pass the result to the next layer of neurons.

[0064] In the present application, when predicting the recovery time of the streamlined pool, by stacking multiple such hidden layers, the DNN can gradually extract high-level features in the data that affect the recovery time of the streamlined pool, thereby achieving efficient learning and processing of complex data.

[0065] The target deep neural network model can be a solidified deep neural network model after being trained with training samples. By directly inputting multiple target feature values ​​into the target deep neural network model, the target prediction duration can be obtained.

[0066] The recovery time prediction method provided in the embodiment of the present application can obtain initial feature values ​​corresponding to multiple target features, and normalize the multiple target feature values ​​according to the feature mean and standard deviation corresponding to each target feature to obtain multiple target feature values. Through the target deep neural network model, the thin pool recovery time is predicted based on the multiple target feature values, and the target recovery time can be obtained. There is no need to manually predict the recovery time, which can improve the reliability of services provided by the storage device.

[0067] Before training the deep neural network model through the sample set, the target features that affect the recovery time of the streamlined pool can be determined from multiple configuration features to reduce the number of features input into the initial neural network model, which can reduce the training time, reduce the impact of invalid features, and improve the accuracy of prediction.

[0068] Next, combine Figure 3 , further illustrating the execution process of determining target features provided in the embodiment of the present application.

[0069] Figure 3 A schematic diagram of a process for determining target features provided in an embodiment of the present application. Figure 3 , the method may include:

[0070] S301. Acquire multiple configuration features.

[0071] The multiple configuration features may include at least one first configuration feature of the target storage device and at least one second configuration feature corresponding to the target thin pool.

[0072] Among them, at least one first configuration feature can be the number of processor cores, processor main frequency, memory size and memory main frequency, etc., and at least one second configuration feature can be storage pool capacity, number of storage pools, amount of written data, compression rate, deduplication rate and write mode, etc.

[0073] S302: Perform correlation analysis on multiple configuration features and recovery duration of a target thin pool to obtain multiple target features.

[0074] Correlation analysis methods can include Pearson product-distance correlation coefficient, nonparametric correlation analysis, principal component analysis, multiple regression analysis, and grey clustering and weight assignment method.

[0075] In the embodiment of the present application, the correlation analysis between multiple configuration features and the recovery time can be performed using the Pearson product-moment correlation coefficient.

[0076] Specifically, multiple feature value groups corresponding to multiple configuration features can be obtained, and the feature value groups may include multiple configuration feature values ​​corresponding to the configuration features and historical recovery durations corresponding to each configuration feature value; for any configuration feature, the correlation between the configuration feature and the recovery duration is determined based on the multiple configuration feature values ​​and the multiple historical recovery durations corresponding to the configuration feature; the configuration feature having a correlation greater than a preset correlation is determined as a target feature to obtain multiple target features.

[0077] For example, assuming that there are 5 configuration features, namely configuration features 1-5, configuration feature 1 can be the number of sub-processor cores, configuration feature 2 can be the memory size, configuration feature 3 can be the storage pool capacity, configuration feature 4 can be the number of storage pools, and configuration feature 5 can be the amount of written data. Multiple configuration feature values ​​corresponding to each configuration feature and the historical recovery duration corresponding to each configuration feature value can be obtained. See Table 1. According to the multiple configuration feature values ​​and multiple historical recovery durations corresponding to the configuration feature, the correlation between the configuration feature and the recovery duration can be determined, that is, configuration feature 1 corresponds to correlation 1, configuration feature 2 corresponds to correlation 2, configuration feature 3 corresponds to correlation 3, configuration feature 4 corresponds to correlation 4, and configuration feature 5 corresponds to correlation 5.

[0078] Table 1

[0079] For any configuration feature, the correlation can be determined in the following manner: determine a first mean of multiple configuration feature values ​​and a second mean of multiple historical recovery durations; determine the covariance and standard deviation corresponding to the configuration feature based on the first mean and the second mean; and determine the ratio of the covariance and the standard deviation as the correlation corresponding to the configuration feature.

[0080] Specifically, please refer to the following formula:

[0081] in, is the correlation, the number of multiple configuration feature values ​​and multiple historical recovery durations is , is the first mean of multiple configuration eigenvalues, It is the second average of multiple historical recovery times. is the eigenvalue of the i-th configuration, is the i-th historical recovery time, is the covariance corresponding to the configuration feature, is the standard deviation corresponding to the configuration feature.

[0082] The range of correlation is If the correlation is greater than 0, the configuration feature is positively correlated with the recovery time. If the correlation is less than 0, the configuration is negatively correlated with the recovery time. The closer the absolute value of the correlation is to 1, the closer the configuration feature is to the recovery time. If the correlation is 0, it means that there is no correlation between the configuration feature and the recovery time.

[0083] For example, assuming that the preset correlation can be 0.5, assuming that there are 5 configuration features in total, the correlation corresponding to configuration feature 1 is 0.3, the correlation corresponding to configuration feature 2 is 0.6, the correlation corresponding to configuration feature 3 is -0.8, the correlation corresponding to configuration feature 4 is -0.2, and the correlation corresponding to configuration feature 5 is 0.8, then configuration feature 2, configuration feature 3 and configuration feature 5 can be determined as target features.

[0084] It is worth noting that the preset relevance and target configuration provided in the embodiments of the present application are merely examples, and the preset relevance and target configuration can be determined according to actual conditions and are not specifically limited here.

[0085] The recovery time prediction method provided in the embodiment of the present application can determine the target feature related to the recovery time among multiple configuration features through correlation analysis, which can improve the accuracy of predicting the recovery time. In addition, through correlation analysis using the Pearson product-distance correlation coefficient, features that are negatively correlated and positively correlated with the recovery time can be analyzed, which can improve the accuracy of feature correlation analysis.

[0086] Before predicting the recovery time of the thin pool of the target storage device in a targeted manner, it is necessary to train the initial deep neural network model to obtain the target deep neural network model. Figure 4 , the execution process of the model training provided in the embodiment of the present application is explained.

[0087] Figure 4 A flow chart of a model training method provided in an embodiment of the present application. Figure 4 , the method may include:

[0088] S401. Acquire multiple target samples.

[0089] The target sample may include multiple target historical feature values ​​and the historical recovery time annotated by the target sample;

[0090] Since multiple initial samples are obtained from different storage devices and at different times, there is a situation where the measurements of the multiple initial samples are not uniform. It is necessary to normalize the multiple initial samples. The specific normalization process can be seen as follows:

[0091] Obtain multiple initial samples, the initial samples include multiple initial historical feature values ​​corresponding to multiple target features, and the historical recovery time annotated by the initial samples; for any target feature, obtain multiple initial historical feature values ​​corresponding to the target feature in multiple initial samples; determine the feature mean and standard deviation corresponding to the target feature according to the multiple initial historical feature values; normalize the multiple initial historical feature values ​​according to the feature mean and standard deviation to obtain multiple target historical feature values. Obtain multiple target samples according to the multiple target historical feature values ​​corresponding to each initial sample and the historical recovery time annotated by each initial sample.

[0092] The specific normalization execution process can refer to the normalization description of the execution process of predicting the recovery time in the above embodiment, which will not be repeated here.

[0093] The multiple target samples may include multiple training samples and multiple test samples. The multiple training samples are used to train the initial deep neural network model, and the multiple test samples are used to test the trained initial deep neural network model.

[0094] S402: Train the initial deep neural network model based on multiple training samples until a preset convergence condition is reached, and determine the deep neural network model that reaches the preset convergence condition as the neural network model to be tested.

[0095] The initial deep neural network model includes an input layer, multiple hidden layers and an output layer, wherein the number of neurons in the input layer is the number of multiple target historical feature values, and since the prediction target is the recovery time of the streamlined pool, the number of neurons in the output layer is 1.

[0096] The activation function in the hidden layer of the initial deep neural network model can be a modified linear unit (for example, ReLU), which can introduce nonlinearity, alleviate gradient vanishing, be computationally efficient, and perform sparse activation, etc. The activation function in the output layer can be a linear activation function (for example, Linear), which can directly output continuous values, avoid truncation or compression of the output range, and meet the requirements of regression tasks.

[0097] The initial deep neural network may be trained for a preset number of training rounds using multiple training samples, and the preset number of training rounds may be a maximum number of training rounds for training the multiple training samples during the training process.

[0098] A preset training round (epoch) can be used to fully train the model with all training samples. For example, if there are 10,00 samples in total, when the model traverses these 10,00 samples, 1 epoch is completed.

[0099] The preset convergence condition can minimize the loss function for the target.

[0100] The target loss function is the sum of the preset product and the initial loss function. The preset product is the product of the regularization term and the regularization coefficient. The target loss function can be found in the following formula:

[0101] in, is the target loss function, is the initial loss function, is the preset product, is the regularization coefficient, is the regularization term, is the initial weight of the jth hidden layer of the initial deep neural network model, is the number of hidden layers of the initial deep neural network model, is an integer greater than 1, and the regularization term is used to indicate the sum of squares of all initial weights; For historical restoration time, To predict the repair time, is a preset difference used to control the limit of the error size. When the error is less than δ, the mean square error (MSE) is used, and when the error is greater than δ, the mean absolute error (MAE) is used. is the mean square error, is the mean absolute error.

[0102] The larger the value of the regularization coefficient, the stronger the constraint on the weight, and the range is generally between 0.001 and 0.1.

[0103] The initial loss function is the Huber Loss function, which combines MSE and MAE. MSE can be used for small errors, and MAE can be used for large errors. When dealing with regression problems, it can both smooth the training process and reduce the impact of outliers.

[0104] In addition, by using L2 regularization, adding the sum of squares of weight parameters to the initial loss function and limiting the weight amplitude, the model parameters can be smoothly distributed, extreme values ​​can be avoided, and the model can be prevented from being overly dependent on a single feature (for example, the number of processor cores).

[0105] The initial deep neural network model may include an initial weight matrix, where the initial weight matrix is ​​used to indicate initial weights of all hidden layers in the initial deep neural network.

[0106] When the initial deep neural network model is trained based on multiple training samples, the initial weight matrix is ​​updated by an optimization function; wherein the optimization function is the difference between the first weight matrix and the gradient update term, the first weight matrix is ​​the initial weight matrix minus the product of the weight attenuation coefficient and the initial weight matrix, and the gradient update term is determined by the target loss function value and the learning rate.

[0107] The optimization function can be the AdamW optimizer, and the formula can be seen as follows:

[0108] in, is the target weight matrix; is the first weight matrix, is the weight decay coefficient, is the initial weight matrix; is the gradient update term, used to adjust the model weights; is the learning rate, which is used to control the parameter update step size; is the bias-corrected first-order momentum estimate, is the bias-corrected second-order momentum estimate; It is a numerical stability term, a preset minimum constant; To indicate the training, To indicate the training, is an integer greater than or equal to 1.

[0109] Among them, the learning rate can be scheduled through the cosine annealing (Cosine Decay) learning rate. The cosine annealing learning rate is a method for optimizing the learning rate and can be used in deep learning training to help the model converge better. The cosine annealing learning rate adjustment method gradually reduces the learning rate during the training cycle, uses a larger learning rate in the early stage of training to approach the global minimum faster, and uses a smaller learning rate for fine-tuning in the later stage, which can improve the stability and generalization ability of the model.

[0110] In the embodiment of the present application, the attenuation term can be separated from the update of the first-order moment and the second-order moment by optimizing the function, so that the optimization process is more stable.

[0111] Early stopping can monitor the performance of the validation set and terminate training early when the model starts to overfit, thereby saving parameters when the model is in the best generalization state and avoiding invalid subsequent training.

[0112] Specifically, if there are multiple consecutive training rounds with the same target loss function value, the training of the deep neural network model is stopped, and the deep neural network model with the smallest target loss function value is determined as the neural network model to be tested.

[0113] The target loss function can be used as a validation set. If the target loss function value does not improve within multiple consecutive epochs (for example, the number of epochs is 10 and the target loss function value does not improve), the training is stopped, and the model weights when the target loss function value is the smallest can be automatically retained.

[0114] In this application, by stopping the training in advance, invalid training can be avoided, the training efficiency can be improved, and overfitting can be prevented to ensure the generalization ability of the model.

[0115] The initial deep neural network model includes multiple initial hidden layers, and the number of the multiple initial hidden layers can be indivual.

[0116] In deep neural network models, Dropout is a regularization technique that prevents the model from becoming overly dependent on certain features or neurons by randomly discarding (i.e., temporarily shutting down) a portion of neurons, thereby improving the generalization ability and robustness of the model.

[0117] When an initial deep neural network model is trained based on multiple training samples, when multiple abstract feature data of an initial hidden layer are input into a next initial hidden layer, multiple selected feature data and at least one data to be processed are determined from the multiple abstract feature data, and at least one data to be processed is set to zero.

[0118] For example, during the training process, the proportion of randomly "closing" some neurons can be set to 20%. Assuming that there are 100 abstract feature data in total, then the number of multiple selected feature data is 80, and the number of at least one data to be processed is 20.

[0119] Figure 5 This is a schematic diagram of the architecture of a deep neural network model provided in an embodiment of the present application. Figure 5The deep neural network model includes an input layer, a first hidden layer, a second hidden layer, a third hidden layer and an output layer, wherein the input layer may include 128 neurons, the first hidden layer may include 256 neurons, the second hidden layer may include 128 neurons, the third hidden layer may include 64 neurons, and the output layer may include 1 neuron. The activation function uses ReLU. When the 256 abstract feature data corresponding to the 256 neurons in the first hidden layer are input into the second hidden layer, the probability of random zeroing (Dropout) is 0.3, that is, 30% of the abstract feature data can be randomly zeroed. When the 128 abstract feature data corresponding to the 128 neurons in the second hidden layer are input into the third hidden layer, the probability of random zeroing (Dropout) is 0.2, that is, 20% of the abstract feature data can be randomly zeroed.

[0120] Figure 6 This is a schematic diagram of a model configuration architecture provided in an embodiment of the present application. Figure 6 When configuring the initial deep neural network model, the loss function is Huber Loss, the optimizer is AdamW, the learning rate schedule is cosine annealing, the regularization is L2 regularization and random zeroing (Dropout), and the number of early stops is 10. The batch size is 64, the preset training epochs is 100, and the validation set monitoring is MSE or MAE.

[0121] S403, using multiple test samples to test the neural network model to be tested, and determining the deep neural network model that meets the test requirements as the target deep neural network model.

[0122] In some embodiments, the prediction error between the predicted recovery time and the historical recovery time can be calculated through MSE to verify whether the neural network model to be tested will meet the test requirements.

[0123] MSE can measure the square average of the prediction error between the predicted repair time and the historical repair time. The smaller it is, the more accurate the model prediction is.

[0124] In some embodiments, the prediction error between the predicted recovery time and the historical recovery time can be calculated by MAE to verify whether the neural network model to be tested will meet the test requirements.

[0125] MAE can measure the average difference in prediction error between the predicted repair time and the historical repair time. The smaller it is, the more accurate the model prediction is.

[0126] The test requirement may be that the MSE value or the MAE value is less than or equal to a preset threshold.

[0127] The recovery time prediction method provided in the embodiment of the present application can train the initial neural network model through the historical feature values ​​of the storage device and the corresponding historical recovery time to obtain the target deep neural network model. When determining the recovery time of the thin pool, the target features of the storage device can be directly processed through the target deep neural network model to obtain the target recovery time, thereby improving the efficiency of the target recovery time prediction. In addition, there is no need for manual empirical prediction, which can improve the accuracy of the prediction, thereby improving the reliability of the service provided by the storage device.

[0128] Figure 7 This is a schematic diagram of the structure of a recovery time prediction device provided in an embodiment of the present application. Figure 7 The recovery time prediction device 700 may include a first acquisition module 701, a second acquisition module 702, a normalization processing module 703 and a prediction module 704, wherein:

[0129] The first acquisition module 701 is used to acquire initial feature values ​​corresponding to a plurality of target features, where the target features are features that affect the recovery time of a target thin pool in a target storage device;

[0130] The second acquisition module 702 is used to obtain the feature mean and standard deviation corresponding to each target feature;

[0131] The normalization processing module 703 is used to perform normalization processing on multiple initial feature values ​​according to the feature mean and standard deviation corresponding to each target feature to obtain multiple target feature values;

[0132] The prediction module 704 is used to predict the recovery duration of the target thin pool through a target deep neural network model and according to multiple target feature values ​​to obtain a target recovery duration.

[0133] In a possible embodiment, the recovery time prediction device 700 further includes a third acquisition module and a correlation analysis module, wherein:

[0134] The third acquisition module is used to acquire multiple configuration features, where the multiple configuration features include at least one first configuration feature of the target storage device and at least one second configuration feature corresponding to the target thin pool;

[0135] The correlation analysis module is used to perform correlation analysis on a plurality of configuration features and the recovery duration of a target thin pool to obtain a plurality of target features.

[0136] In a possible embodiment, the correlation analysis module is specifically used for:

[0137] Acquire multiple feature value groups corresponding to the multiple configuration features, where the feature value groups include multiple configuration feature values ​​corresponding to the configuration features and historical recovery durations corresponding to the configuration feature values;

[0138] For any configuration feature, determining the correlation between the configuration feature and the recovery duration according to multiple configuration feature values ​​corresponding to the configuration feature and multiple historical recovery durations;

[0139] The configuration features whose absolute values ​​of correlation are greater than a preset correlation are determined as target features to obtain multiple target features.

[0140] In a possible embodiment, the correlation analysis module is specifically used for:

[0141] Determining a first mean value of a plurality of configuration characteristic values ​​and a second mean value of a plurality of historical recovery durations;

[0142] Determine the covariance and standard deviation corresponding to the configuration feature according to the first mean and the second mean;

[0143] The ratio of the covariance and the standard deviation is determined as the correlation corresponding to the configuration features.

[0144] In a possible embodiment, the normalization processing module 703 is specifically used for:

[0145] For any target feature, the difference between the initial feature value and the feature mean corresponding to the target feature is determined as the first difference corresponding to the target feature;

[0146] For any target feature, the ratio of the first difference value corresponding to the target feature to the standard deviation is determined as the target feature value corresponding to the target feature.

[0147] In a possible embodiment, the recovery time prediction device 700 further includes a fourth acquisition module, a training module and a testing module:

[0148] The fourth acquisition module is used to acquire multiple target samples, the multiple target samples include multiple training samples and multiple test samples, the target samples include multiple target historical feature values, and historical recovery time annotated by the target samples;

[0149] The training module is used to train the initial deep neural network model based on multiple training samples until a preset convergence condition is reached, and determine the deep neural network model that reaches the preset convergence condition as the neural network model to be tested;

[0150] The testing module is used to test the neural network model to be tested using multiple test samples, and determine the deep neural network model that meets the test requirements as the target deep neural network model.

[0151] In one possible embodiment, the preset convergence condition is that the target loss function is minimized; the target loss function is the sum of a preset product and an initial loss function, and the preset product is the product of a regularization term and a regularization coefficient.

[0152] In one possible embodiment, the initial loss function includes a mean square error and a mean absolute error. The mean square error is used when the difference between the historical recovery time and the predicted recovery time is greater than or equal to a preset difference, and the mean absolute error is used when the difference between the historical recovery time and the predicted recovery time is less than a preset difference.

[0153] In one possible embodiment, the initial deep neural network model includes an initial weight matrix; when the initial deep neural network model is trained based on multiple training samples, the initial weight matrix is ​​updated by an optimization function;

[0154] Among them, the optimization function is the difference between the first weight matrix and the gradient update term, the first weight matrix is ​​the initial weight matrix minus the product of the weight attenuation coefficient and the initial weight matrix, and the gradient update term is determined by the target loss function value and the learning rate.

[0155] In one possible embodiment, the initial deep neural network is trained for a preset number of training rounds using a plurality of training samples, where the preset number of training rounds is a maximum number of training rounds for training the plurality of training samples during the training process;

[0156] Among them, if there are multiple consecutive training rounds corresponding to the same target loss function value, the training of the deep neural network model is stopped, and the deep neural network model with the smallest target loss function value is determined as the neural network model to be tested.

[0157] In one possible embodiment, the initial deep neural network model includes multiple initial hidden layers; when the initial deep neural network model is trained based on multiple training samples, when multiple abstract feature data of the initial hidden layer are input into the next initial hidden layer, multiple selected feature data and at least one data to be processed are determined from the multiple abstract feature data, and at least one data to be processed is set to zero.

[0158] In a possible embodiment, the fourth acquisition module is specifically used for:

[0159] Acquire multiple initial samples, where the initial samples include multiple initial historical feature values ​​corresponding to multiple target features and historical recovery time annotated by the initial samples;

[0160] For any target feature, multiple initial historical feature values ​​corresponding to the target feature are obtained from multiple initial samples; the feature mean and standard deviation corresponding to the target feature are determined based on the multiple initial historical feature values; the multiple initial historical feature values ​​are normalized based on the feature mean and standard deviation to obtain multiple target historical feature values.

[0161] According to the multiple target historical feature values ​​corresponding to each initial sample and the historical recovery time annotated by each initial sample, multiple target samples are obtained.

[0162] For the description of the features in the embodiment corresponding to the recovery time prediction device, reference may be made to the relevant description of the embodiment corresponding to the recovery time prediction method, which will not be described in detail here.

[0163] Figure 8 This is a schematic diagram of the structure of an electronic device provided in this application. Figure 8 As shown, the electronic device 800 provided in this embodiment includes: at least one processor 801 and a memory 802. Optionally, the electronic device 800 further includes a communication component 803. The processor 801, the memory 802 and the communication component 803 are connected via a bus.

[0164] In a specific implementation process, at least one processor 801 executes the computer-executable instructions stored in the memory 802, so that at least one processor 801 executes the above-mentioned recovery time prediction method embodiment.

[0165] The specific implementation process of the processor 801 can be found in the above method embodiment, and its implementation principle and technical effect are similar, so this embodiment will not be repeated here.

[0166] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in the application may be directly implemented as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor.

[0167] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (NVM), such as at least one disk storage.

[0168] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus in the drawings of this application is not limited to only one bus or one type of bus.

[0169] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned recovery duration prediction embodiments when running.

[0170] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0171] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any one of the above-mentioned recovery time prediction method embodiments are implemented.

[0172] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, the non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned recovery time prediction method embodiments are implemented.

[0173] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0174] The above is a detailed introduction to a memory scheduling method provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A method for predicting recovery time, characterized in that: include: Acquire initial feature values ​​corresponding to a plurality of target features, where the target features are features that affect the recovery time of a target thin pool in a target storage device; Get the feature mean and standard deviation corresponding to each target feature; Normalizing the multiple initial eigenvalues ​​according to the feature means and standard deviations corresponding to the target features to obtain multiple target eigenvalues; The recovery duration of the target thin pool is predicted by using a target deep neural network model and according to the multiple target feature values ​​to obtain a target recovery duration.

2. The method according to claim 1, characterized in that Before obtaining the initial feature values ​​corresponding to the multiple target features, the method further includes: Acquire a plurality of configuration features, where the plurality of configuration features include at least one first configuration feature of a target storage device and at least one second configuration feature corresponding to the target thin pool; A correlation analysis is performed between the multiple configuration features and the recovery duration of the target thin pool to obtain the multiple target features.

3. The method according to claim 2, characterized in that The plurality of configuration features are correlated with the recovery duration of the target thin pool to obtain the plurality of target features, including: Acquire multiple feature value groups corresponding to the multiple configuration features, the feature value groups including multiple configuration feature values ​​corresponding to the configuration features and historical recovery durations corresponding to the configuration feature values; For any configuration feature, determining the correlation between the configuration feature and the recovery duration according to multiple configuration feature values ​​and multiple historical recovery durations corresponding to the configuration feature; The configuration features whose absolute values ​​of the correlation are greater than the preset correlation are determined as target features to obtain the multiple target features.

4. The method according to claim 3, characterized in that Determining the correlation between the configuration feature and the recovery duration according to the multiple configuration feature values ​​and the multiple historical recovery durations corresponding to the configuration feature includes: Determining a first mean value of the plurality of configuration characteristic values ​​and a second mean value of the plurality of historical recovery durations; Determining a covariance and a standard deviation corresponding to the configuration feature according to the first mean and the second mean; The ratio of the covariance to the standard deviation is determined as the correlation corresponding to the configuration feature.

5. The method according to claim 1, characterized in that: According to the feature mean and standard deviation corresponding to each target feature, multiple initial feature values ​​are normalized to obtain multiple target feature values, including: For any target feature, the difference between the initial feature value and the feature mean corresponding to the target feature is determined as the first difference corresponding to the target feature; For any target feature, a ratio of the first difference corresponding to the target feature to the standard deviation is determined as the target feature value corresponding to the target feature.

6. The method according to any one of claims 1 to 5, characterized in that: The training steps of the target deep neural network model include: Acquire a plurality of target samples, the plurality of target samples comprising a plurality of training samples and a plurality of test samples, the target samples comprising a plurality of target historical feature values, and historical recovery durations annotated by the target samples; Training the initial deep neural network model based on multiple training samples until a preset convergence condition is reached, and determining the deep neural network model that reaches the preset convergence condition as the neural network model to be tested; The neural network model to be tested is tested using the multiple test samples, and a deep neural network model that meets the test requirements is determined as a target deep neural network model.

7. The method according to claim 6, characterized in that The preset convergence condition is that the target loss function is minimized; the target loss function is the sum of a preset product and an initial loss function, and the preset product is the product of a regularization term and a regularization coefficient.

8. The method according to claim 7, characterized in that The initial loss function includes the mean square error and the mean absolute error. The mean square error is used when the difference between the historical recovery time and the predicted recovery time is greater than or equal to the preset difference, and the mean absolute error is used when the difference between the historical recovery time and the predicted recovery time is less than the preset difference.

9. The method according to claim 7, characterized in that: The initial deep neural network model includes an initial weight matrix; when the initial deep neural network model is trained based on multiple training samples, the initial weight matrix is ​​updated by an optimization function; Among them, the optimization function is the difference result of the first weight matrix and the gradient update term, the first weight matrix is ​​the initial weight matrix minus the product of the weight attenuation coefficient and the initial weight matrix, and the gradient update term is determined by the target loss function value and the learning rate.

10. The method according to claim 6, characterized in that Using the multiple training samples, the initial deep neural network is trained for a preset number of training rounds, where the preset number of training rounds is the maximum number of training rounds for training the multiple training samples during the training process; Among them, if there are multiple consecutive training rounds corresponding to the same target loss function value, the training of the deep neural network model is stopped, and the deep neural network model with the smallest target loss function value is determined as the neural network model to be tested.

11. The method according to claim 6, characterized in that The initial deep neural network model includes multiple initial hidden layers; when the initial deep neural network model is trained based on multiple training samples, when multiple abstract feature data of the initial hidden layer are input into the next initial hidden layer, multiple selected feature data and at least one data to be processed are determined from the multiple abstract feature data, and the at least one data to be processed is set to zero.

12. The method according to claim 6, characterized in that Get multiple target samples, including: Acquire a plurality of initial samples, wherein the initial samples include a plurality of initial historical feature values ​​corresponding to a plurality of target features, and a historical recovery time annotated with the initial samples; For any target feature, obtain multiple initial historical feature values ​​corresponding to the target feature from the multiple initial samples; determine a feature mean and a standard deviation corresponding to the target feature according to the multiple initial historical feature values; and perform normalization processing on the multiple initial historical feature values ​​according to the feature mean and the standard deviation to obtain multiple target historical feature values; The multiple target samples are obtained according to the multiple target historical feature values ​​corresponding to the initial samples and the historical recovery time marked by the initial samples.

13. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the method for predicting the recovery time as claimed in any one of claims 1 to 12 when executing the computer program.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the method for predicting the recovery time as claimed in any one of claims 1 to 12.

15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method for predicting the recovery time length according to any one of claims 1 to 12 are implemented.

Citation Information

Patent Citations

  • Prediction model generation method and device for repair duration of simplified pool and prediction method

    CN116340853A

  • Time length prediction method of target time period, equipment and medium

    CN119378756A

  • Prediction method and apparatus for faulty GPU, electronic device and storage medium

    US20240118984A1

Cited By

  • Fault repairing method and device for full flash memory storage system

    CN121029472A

  • Method and apparatus for failure recovery of all-flash storage system

    CN121029472B