Disk life monitoring method and system based on LSTM model

Through the disk life monitoring method based on the LSTM model, the feature matrix and life score are constructed, and the disk life prediction model is trained in combination with the optimization algorithm, the problem of unintuitive disk life monitoring and low prediction accuracy in the existing technology is solved, and the quantification and accurate prediction of disk life are achieved.

CN120234226BActive Publication Date: 2025-08-26YUNNAN YUANXIN TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510730566.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-08-26
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

Existing disk life monitoring methods cannot provide intuitive quantitative indicators, and ignore the differences in the use of disks in different users in the time dimension, resulting in low prediction accuracy.

Method used

The disk life monitoring method based on the LSTM model is adopted, and the disk life prediction model is trained in combination with the optimization algorithm to achieve quantification and accurate prediction of disk life.

Benefits of technology

Eliminates the difference in using disks in the time dimension of different users, improves the quantization ability and prediction accuracy of disk life, and enhances the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234226B_ABST
    Figure CN120234226B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of disk life monitoring, and discloses a disk life monitoring method and system based on an LSTM model. The disk life monitoring method based on the LSTM model comprises the following steps: step S101, constructing a feature matrix; step S102, obtaining a disk life score through manual labeling; step S103, training a disk life prediction model using an optimization algorithm using the feature matrix and the disk life score; step S104, obtaining the disk life score through the trained disk life prediction model, and if it is determined that the disk life score is less than a preset threshold, reminding the user to back up the disk data in time; the present invention constructs disk usage data into a feature matrix, and establishes a nonlinear mapping relationship between the feature matrix and the disk life score through the disk life prediction model, thereby eliminating differences in disk usage by different users in the time dimension and realizing the quantification of disk life.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of disk life monitoring, and more specifically, to a disk life monitoring method and system based on an LSTM model. Background Art

[0002] With the rapid growth of data storage demand, disks, as key storage media in information systems, have a lifespan that directly affects the reliability of data storage. Therefore, disk life monitoring has become an important means to ensure secure data storage.

[0003] Currently, disk lifespan monitoring primarily relies on SMART (Self-Monitoring Analysis and Reporting Technology) technology, which reflects disk health by providing internal operating status indicators (such as power-on duration, power-on times, number of remapped sectors, and number of uncorrectable sectors). Furthermore, some approaches use these indicators to predict disk lifespan using regression models (such as random forest regression and ridge regression). However, these methods have the following drawbacks: 1. Numerous internal operating status indicators exist, and different disk manufacturers define different thresholds for these indicators, resulting in less intuitive monitoring results and difficulty providing users with clear quantitative indicators of disk lifespan. 2. Predicting disk lifespan based solely on internal operating status indicators ignores the temporal variability of disk usage across different users, such as read / write frequency, access load, and usage time, resulting in low disk lifespan prediction accuracy. Summary of the Invention

[0004] The present invention provides a disk life monitoring method and system based on an LSTM model to solve the technical problems in the above-mentioned background technology.

[0005] The present invention provides a disk life monitoring method based on an LSTM model, comprising the following steps:

[0006] Step S101: collecting disk usage data through a sliding window, and constructing a feature matrix through time division and normalization processing;

[0007] The feature matrix consists of M rows and N columns, where M represents the length of the sliding window, N represents the number of time periods, and the element value in the mth row and nth column is represented by the normalized usage data of the disk in the nth time period at the mth time point, where 1≤m≤M and 1≤n≤N.

[0008] Usage data includes: number of reads, number of writes, total number of files operated, average file size per operation, and average duration per operation.

[0009] Step S102: moving the sliding window backward for a preset period of time, collecting internal operating status indicators of the disk, and obtaining a disk life score through manual labeling;

[0010] Internal operating status indicators include: power-on time, power-on times, number of remapped sectors, number of uncorrectable sectors, and temperature;

[0011] The lifespan score ranges from 0 to 1 and is positively correlated with the remaining lifespan of the disk.

[0012] Step S103: Using the feature matrix and the disk life score as sample data and sample labels for training a disk life prediction model, and completing the training of the disk life prediction model through an optimization algorithm;

[0013] Step S104 , obtaining a disk life score through the trained disk life prediction model, and judging that the score is less than a preset threshold, reminding the user to back up the disk data in time.

[0014] Furthermore, the length of the sliding window, the step size of the sliding window, the number of time periods, the length of the time periods, the preset time period, and the preset threshold are all custom parameters.

[0015] Furthermore, the calculation formula for normalization processing is as follows:

[0016] ;

[0017] in and Respectively represent the disk usage data before and after normalization at the mth time point and the nth period. and They represent the maximum and minimum values ​​of disk usage data in the nth period at the mth time point.

[0018] Furthermore, the disk life prediction model consists of M first hidden layers, 1 second hidden layer, and 1 first classifier;

[0019] Each first hidden layer includes N first units, the nth first unit of the mth first hidden layer inputs the element value of the mth row and nth column of the feature matrix, and outputs a first update vector, wherein the number of dimensions of the first update vector is a custom parameter;

[0020] The second hidden layer includes M second units, the mth second unit inputs the first update vector output by the Nth first unit of the mth first hidden layer, and outputs a second update vector;

[0021] The first classifier inputs the second update vector output by the Mth second unit of the second hidden layer, and the category space of the first classifier represents the life score of the disk, wherein the number of dimensions of the second update vector is a custom parameter.

[0022] Furthermore, the calculation formula of the nth first unit of the mth first hidden layer includes:

[0023] ;

[0024] in and They represent the first update vectors of the nth and n-1th first unit outputs of the mth first hidden layer, Represents the element value of the mth row and nth column of the feature matrix of the nth first unit input of the mth first hidden layer, 、 、 and They represent the attenuation vector, storage vector, first weight parameter and first bias parameter of the nth first unit of the mth first hidden layer respectively. Mish represents the Mish activation function. represents point-by-point multiplication;

[0025] ;

[0026] in Represents the element value of the mth row and fth column of the feature matrix of the fth first unit input of the mth first hidden layer, and denote the second weight parameter and the second bias parameter of the nth first unit of the mth first hidden layer, respectively; Swish denotes the Swish activation function; and || denotes the concatenation operation.

[0027] ;

[0028] in and They respectively represent the third weight parameter and the third bias parameter of the nth first unit of the mth first hidden layer, Hea represents the Heaviside function, and tanh represents the tanh activation function.

[0029] Furthermore, the second units of the second hidden layer are all constructed based on LSTM units, the first classifier is constructed based on a multilayer perceptron, and the corresponding activation function is a Sigmoid activation function.

[0030] Furthermore, the second update vector output by the Mth second unit of the second hidden layer is input into the second, third, fourth, fifth and sixth classifiers respectively, and the corresponding category spaces are power-on duration, power-on times, number of remapped sectors, number of uncorrectable sectors and temperature, respectively. A preliminary score is first calculated according to the internal operating status indicators, and then the disk life score is obtained through manual labeling based on the preliminary score. The preliminary score is first obtained by weighted summing the internal operating status indicators, and then the value range of the preliminary score is controlled between 0 and 1 through the Sigmoid activation function.

[0031] Furthermore, the disk life prediction model is trained by optimizing the algorithm, including the following steps:

[0032] Step S201, randomly generating weight parameters and bias parameters of a disk life prediction model as codes of individuals in an initialization population of an optimization algorithm;

[0033] Step S202: The disk life prediction model applies the encoding of the individual and calculates the loss value of all individuals in the initialization population through the loss function;

[0034] The loss function is represented by the mean square error between the output value of the disk life prediction model at the preset number of training iterations and the sample label, where the preset number of training iterations is a custom parameter;

[0035] Step S203: Generate an iteration coefficient for the current iteration number, and determine if it is greater than or equal to a preset iteration coefficient threshold, then proceed to step S204; otherwise, proceed to step S205;

[0036] The calculation formula of the iteration coefficient F is as follows:

[0037] ;

[0038] Where t indicates the current number of iterations is t, and the starting value of the current number of iterations is 0. represents the maximum number of iterations, Represents the first random number in the range of 0 to 1, Represents a second random number with a value range between -2 and 2, where the preset iteration coefficient threshold is a custom parameter;

[0039] Step S204: Generate a random number between 0 and 1 for each individual in the initialized population, and determine if the random number is greater than or equal to a first threshold. Then, update the code of the individual using the first exploration strategy; otherwise, update the code of the individual using the second exploration strategy.

[0040] Step S205: Generate a random number between 0 and 1 for each individual in the initialized population, and determine if the random number is greater than or equal to a second threshold. If so, update the code of the individual using the first development strategy; otherwise, update the code of the individual using the second development strategy.

[0041] Step S206: If the current number of iterations is incremented by 1 and the iteration termination condition is determined to be satisfied, the code of the individual with the smallest loss value in the initialized population is applied to complete the training of the disk life prediction model; otherwise, the process returns to step S202 to continue the training.

[0042] The iteration termination condition is that the current number of iterations is greater than or equal to the maximum number of iterations or the loss values ​​of more than 1 / 3 of the individuals in the initialized population are less than the preset loss threshold, where the preset loss threshold is a custom parameter.

[0043] Furthermore, the calculation formula of the optimization algorithm includes:

[0044] The calculation formula of the first exploration strategy is as follows:

[0045] ;

[0046] Where 1≤i≤H, i≠j, H represents the total number of individuals in the initialization population and is a custom parameter. represents the code of the i-th individual whose current iteration number is t+1, and They represent the codes of the i-th and j-th individuals with the current iteration number t, Represents a third random number ranging from 0 to 1;

[0047] The calculation formula of the second exploration strategy is as follows:

[0048] ;

[0049] in Indicates the code of the individual with the smallest loss value in the initialization population with the current iteration number t, represents a fourth random number ranging from 0 to 2;

[0050] The calculation formula for the first development strategy includes:

[0051] ;

[0052] ;

[0053] ;

[0054] in and They represent the codes of the first intermediate individual and the second intermediate individual of the current iteration number t, Indicates the fifth random number ranging from 0 to 1;

[0055] The calculation formula for the second development strategy is as follows:

[0056] ;

[0057] Where round represents the rounding function.

[0058] The present invention provides a disk life monitoring system based on an LSTM model, comprising:

[0059] A feature matrix construction module is used to collect disk usage data through a sliding window and construct a feature matrix through time division and normalization processing;

[0060] The annotation module is used to move the sliding window backward by a preset time period, collect the internal operating status indicators of the disk, and obtain the disk life score through manual annotation;

[0061] A training module, which uses the feature matrix and the disk life score as sample data and sample labels for training the disk life prediction model, and completes the training of the disk life prediction model through an optimization algorithm;

[0062] The threshold judgment module is used to obtain the disk life score through the trained disk life prediction model, and if it is judged to be less than the preset threshold, it reminds the user to back up the disk data in time.

[0063] The beneficial effects of the present invention are as follows: the present invention constructs the disk usage data into a feature matrix, and establishes a nonlinear mapping relationship between it and the disk life score through a disk life prediction model, thereby eliminating the differences in disk usage among different users in the time dimension and realizing the quantification of disk life. In addition, the present invention solves the optimal solution of the parameters in the disk life prediction model through an optimization algorithm, which can improve the prediction accuracy and generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 It is a flow chart of the disk life monitoring method based on the LSTM model of the present invention;

[0065] Figure 2 This is a flow chart of the present invention for completing the training of a disk life prediction model through an optimization algorithm;

[0066] Figure 3 Schematic diagram of the disk life monitoring system based on the LSTM model of the present invention.

[0067] In the figure: feature matrix construction module 301, labeling module 302, training module 303, threshold judgment module 304. DETAILED DESCRIPTION

[0068] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed solely to enable those skilled in the art to better understand and implement the subject matter described herein, and that the functions and arrangements of the elements discussed may be varied without departing from the scope of this specification. Various examples may omit, substitute, or add various processes or components as needed. In addition, features described with respect to some examples may also be combined in other examples.

[0069] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in one or more embodiments of the present invention should have the usual meanings understood by people with ordinary skills in the field to which the present invention belongs. The "first", "second" and similar words used in one or more embodiments of the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprising" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0070] like Figures 1 to 3 As shown in FIG, the disk life monitoring method based on the LSTM model includes the following steps:

[0071] Step S101: collecting disk usage data through a sliding window, and constructing a feature matrix through time division and normalization processing;

[0072] The feature matrix consists of M rows and N columns, where M represents the length of the sliding window, N represents the number of time periods, and the element value in the mth row and nth column is represented by the normalized usage data of the disk in the nth time period at the mth time point, where 1≤m≤M and 1≤n≤N.

[0073] Usage data includes: number of reads, number of writes, total number of files operated, average file size per operation, and average duration per operation.

[0074] Step S102, moving the sliding window backward for a preset period of time, collecting internal operating status indicators of the disk, and obtaining a disk life score through manual labeling;

[0075] Internal operating status indicators include: power-on time, power-on times, number of remapped sectors, number of uncorrectable sectors, and temperature;

[0076] The lifespan score ranges from 0 to 1 and is positively correlated with the remaining lifespan of the disk.

[0077] Step S103: Using the feature matrix and the disk life score as sample data and sample labels for training a disk life prediction model, and completing the training of the disk life prediction model through an optimization algorithm;

[0078] Step S104 , obtaining a disk life score through the trained disk life prediction model, and judging that the score is less than a preset threshold, reminding the user to back up the disk data in time.

[0079] In one embodiment of the present invention, the length of the sliding window, the step size of the sliding window, the number of time periods, the duration of the time period, the preset time period, and the preset threshold are all custom parameters. For example, the length of the sliding window is set to 14 days, the step size of the sliding window is set to 1 day, the number of time periods is set to 6, the duration of the time period is set to 4 hours, the preset time period is set to 3 days, and the preset threshold is set to 0.1. According to the above example, the feature matrix includes 14 rows and 6 columns, and the preset threshold represents the critical value at which the disk is about to fail.

[0080] In one embodiment of the present invention, the calculation formula for normalization processing is as follows:

[0081] ;

[0082] in and Respectively represent the disk usage data before and after normalization at the mth time point and the nth period. and They represent the maximum and minimum values ​​of disk usage data in the nth period at the mth time point.

[0083] It should be noted that normalization can also be performed using the z-score method. Normalization can eliminate dimensional differences and, to a certain extent, accelerate the convergence of the disk life prediction model, prevent gradient explosion or gradient disappearance, and enhance the generalization ability of the model.

[0084] In one embodiment of the present invention, the disk life prediction model consists of M first hidden layers, 1 second hidden layer and 1 first classifier;

[0085] Each first hidden layer includes N first units, the nth first unit of the mth first hidden layer inputs the element value of the mth row and nth column of the feature matrix, and outputs a first update vector, wherein the number of dimensions of the first update vector is a custom parameter, for example, the number of dimensions of the first update vector is set to 8;

[0086] The second hidden layer includes M second units, the mth second unit inputs the first update vector output by the Nth first unit of the mth first hidden layer, and outputs a second update vector;

[0087] The first classifier inputs the second update vector output by the Mth second unit of the second hidden layer, and the category space of the first classifier represents the life score of the disk, wherein the number of dimensions of the second update vector is a custom parameter, for example, the number of dimensions of the second update vector is set to 16.

[0088] It should be noted that the M first hidden layers share weight parameters and bias parameters. The number of first hidden layers is determined by the length of the sliding window (the number of rows in the feature matrix). The first hidden layer is a "hot-swappable" design and can be adaptively adjusted according to the length of the sliding window. In addition, parallel computing can speed up the overall calculation speed of the disk life prediction model.

[0089] In one embodiment of the present invention, the calculation formula of the nth first unit of the mth first hidden layer includes:

[0090] ;

[0091] in and They represent the first update vectors of the nth and n-1th first unit outputs of the mth first hidden layer, Represents the element value of the mth row and nth column of the feature matrix of the nth first unit input of the mth first hidden layer, 、 、 and They represent the attenuation vector, storage vector, first weight parameter and first bias parameter of the nth first unit of the mth first hidden layer respectively. Mish represents the Mish activation function. represents point-by-point multiplication;

[0092] It should be noted that, according to the above example, the number of dimensions of the attenuation vector and the storage vector are both designed to be 8, and The number of dimensions is 5 (i.e., the normalized usage data), then the first weight parameter needs to be designed as a 5×8 matrix, and the first bias parameter needs to be designed as a 1×8 vector;

[0093] ;

[0094] in Represents the element value of the mth row and fth column of the feature matrix of the fth first unit input of the mth first hidden layer, and denote the second weight parameter and the second bias parameter of the nth first unit of the mth first hidden layer, respectively; Swish denotes the Swish activation function; and || denotes the concatenation operation.

[0095] It should be noted that, according to the above example, and The concatenation result is a 1×13 (8+5) size vector, so the second weight parameter needs to be designed as a 13×8 size matrix, and the second bias parameter needs to be designed as a 1×8 size vector;

[0096] ;

[0097] in and They represent the third weight parameter and the third bias parameter of the nth first unit of the mth first hidden layer, Hea represents the Heaviside function, and tanh represents the tanh activation function;

[0098] It should be noted that, according to the above example, the third weight parameter is the same as the first weight parameter, and the third bias parameter is the same as the first bias parameter.

[0099] It should be noted that the design of the first unit in the present invention utilizes the concept of HMM (Hidden Markov Model), that is, the current output is related to the output of the previous time step and the current input, and the redundant gating mechanism in LSTM is removed, thereby improving the calculation speed of the disk life prediction model. In addition, the design of the decay vector in the present invention also integrates the influence of all time steps before the current time step, and the Heaviside function of the storage vector can achieve the effect of "information screening" (that is, controlling the element value to a binary representation of 0 or 1).

[0100] In one embodiment of the present invention, the second units of the second hidden layer are all constructed based on LSTM (long short-term memory recurrent neural network) units, the first classifier is constructed based on a multi-layer perceptron, and the corresponding activation function is a Sigmoid activation function, wherein LSTM is a conventional technical means, and the specific calculation formula is not repeated here.

[0101] In one embodiment of the present invention, the second update vector output by the Mth second unit of the second hidden layer is input into the second, third, fourth, fifth and sixth classifiers respectively, and the corresponding category spaces are power-on duration, power-on times, number of remapped sectors, number of uncorrectable sectors and temperature, respectively. A preliminary score is first calculated based on the internal operating status indicators, and then the life score of the disk is obtained through manual labeling based on the preliminary score, wherein the preliminary score is first obtained by weighted summing the internal operating status indicators, and then the value range of the preliminary score is controlled between 0 and 1 through the Sigmoid activation function.

[0102] It should be noted that the value range of the preliminary score is the same as that of the life score. Although manual labeling has higher accuracy, the workload is huge. This semi-automatic labeling method can significantly improve labeling efficiency.

[0103] In one embodiment of the present invention, Figure 2 As shown in the figure, the training of the disk life prediction model is completed through the optimization algorithm, which includes the following steps:

[0104] Step S201, randomly generating weight parameters and bias parameters of a disk life prediction model as codes of individuals in an initialization population of an optimization algorithm;

[0105] Step S202: The disk life prediction model applies the encoding of the individual and calculates the loss value of all individuals in the initialization population through the loss function;

[0106] The loss function is represented by the mean square error between the value output by the disk life prediction model at the preset number of training iterations and the sample label, where the preset number of training iterations is a custom parameter, for example, the preset number of training iterations is set to 100;

[0107] Step S203: Generate an iteration coefficient for the current iteration number, and determine if it is greater than or equal to a preset iteration coefficient threshold, then proceed to step S204; otherwise, proceed to step S205;

[0108] The calculation formula of the iteration coefficient F is as follows:

[0109] ;

[0110] Where t indicates the current number of iterations is t, and the starting value of the current number of iterations is 0. represents the maximum number of iterations, Represents the first random number in the range of 0 to 1, represents a second random number in the range of -2 to 2, where the preset iteration coefficient threshold is a custom parameter, for example, the preset iteration coefficient threshold is set to 1;

[0111] Step S204: Generate a random number between 0 and 1 for each individual in the initialized population, and determine if the random number is greater than or equal to a first threshold. Then, update the code of the individual using the first exploration strategy; otherwise, update the code of the individual using the second exploration strategy.

[0112] Step S205: Generate a random number between 0 and 1 for each individual in the initialized population, and determine if the random number is greater than or equal to a second threshold. If so, update the code of the individual using the first development strategy; otherwise, update the code of the individual using the second development strategy.

[0113] Step S206: If the current number of iterations is incremented by 1 and the iteration termination condition is determined to be satisfied, the code of the individual with the smallest loss value in the initialized population is applied to complete the training of the disk life prediction model; otherwise, the process returns to step S202 to continue the training.

[0114] The iteration termination condition is that the current number of iterations is greater than or equal to the maximum number of iterations or the loss values ​​of more than 1 / 3 of the individuals in the initialized population are less than the preset loss threshold, where the preset loss threshold is a custom parameter, for example, the preset loss threshold is set to 0.01.

[0115] In one embodiment of the present invention, the calculation formula of the optimization algorithm includes:

[0116] The calculation formula of the first exploration strategy is as follows:

[0117] ;

[0118] Where 1≤i≤H, i≠j, H represents the total number of individuals in the initial population and is a custom parameter. For example, H is set to 10. represents the code of the i-th individual whose current iteration number is t+1, and They represent the codes of the i-th and j-th individuals with the current iteration number t, Represents a third random number ranging from 0 to 1;

[0119] The calculation formula of the second exploration strategy is as follows:

[0120] ;

[0121] in Indicates the code of the individual with the smallest loss value in the initialization population with the current iteration number t, represents a fourth random number ranging from 0 to 2;

[0122] The calculation formula for the first development strategy includes:

[0123] ;

[0124] ;

[0125] ;

[0126] in and They represent the codes of the first intermediate individual and the second intermediate individual of the current iteration number t, Indicates the fifth random number ranging from 0 to 1;

[0127] The calculation formula for the second development strategy is as follows:

[0128] ;

[0129] Where round represents the rounding function.

[0130] In one embodiment of the present invention, Figure 3 As shown in the figure, the disk life monitoring system based on the LSTM model includes:

[0131] A feature matrix construction module 301 is used to collect disk usage data through a sliding window and construct a feature matrix through time division and normalization processing;

[0132] A labeling module 302 is used to move the sliding window backward for a preset period of time, collect internal operating status indicators of the disk, and obtain a disk life score through manual labeling;

[0133] A training module 303 is configured to use the feature matrix and the disk life score as sample data and sample labels for training a disk life prediction model, and complete the training of the disk life prediction model through an optimization algorithm;

[0134] The threshold judgment module 304 is used to obtain the disk life score through the trained disk life prediction model, and if it is judged to be less than a preset threshold, it reminds the user to back up the disk data in time.

[0135] The above describes the embodiments of this embodiment, but this embodiment is not limited to the above specific implementation methods. The above specific implementation methods are merely illustrative and not restrictive. Ordinary technicians in this field can also make many forms based on the inspiration of this embodiment, all of which are protected by this embodiment.

Claims

1. A disk life monitoring method based on the LSTM model is characterized by: The following steps are involved: Step S101: collecting disk usage data through a sliding window, and constructing a feature matrix through time division and normalization processing; The feature matrix consists of M rows and N columns, where M represents the length of the sliding window, N represents the number of time periods, and the element value in the mth row and nth column is represented by the normalized usage data of the disk in the nth time period at the mth time point, where 1≤m≤M and 1≤n≤N. Usage data includes: number of reads, number of writes, total number of files operated, average file size per operation, and average duration per operation. Step S102: moving the sliding window backward for a preset period of time, collecting internal operating status indicators of the disk, and obtaining a disk life score through manual labeling; Internal operating status indicators include: power-on time, power-on times, number of remapped sectors, number of uncorrectable sectors, and temperature; The lifespan score ranges from 0 to 1 and is positively correlated with the remaining lifespan of the disk. Step S103: Using the feature matrix and the disk life score as sample data and sample labels for training a disk life prediction model, and completing the training of the disk life prediction model through an optimization algorithm; Step S104: Obtain a disk lifespan score using the trained disk lifespan prediction model, and if it is determined to be less than a preset threshold, the user is reminded to back up the disk data in a timely manner. The disk life prediction model consists of M first hidden layers, 1 second hidden layer and 1 first classifier; Each first hidden layer includes N first units, the nth first unit of the mth first hidden layer inputs the element value of the mth row and nth column of the feature matrix, and outputs a first update vector, wherein the number of dimensions of the first update vector is a custom parameter; The second hidden layer includes M second units, the mth second unit inputs the first update vector output by the Nth first unit of the mth first hidden layer, and outputs a second update vector; The first classifier inputs a second update vector output by the Mth second unit of the second hidden layer, the category space of the first classifier represents the life score of the disk, wherein the number of dimensions of the second update vector is a custom parameter; The calculation formula for the nth first unit of the mth first hidden layer includes: ; in and They represent the first update vectors of the nth and n-1th first unit outputs of the mth first hidden layer, Represents the element value of the mth row and nth column of the feature matrix of the nth first unit input of the mth first hidden layer, 、 、 and They represent the attenuation vector, storage vector, first weight parameter and first bias parameter of the nth first unit of the mth first hidden layer respectively. Mish represents the Mish activation function. represents point-by-point multiplication; ; in Represents the element value of the mth row and fth column of the feature matrix of the fth first unit input of the mth first hidden layer, and denote the second weight parameter and the second bias parameter of the nth first unit of the mth first hidden layer, respectively; Swish denotes the Swish activation function; and || denotes the concatenation operation. ; in and They respectively represent the third weight parameter and the third bias parameter of the nth first unit of the mth first hidden layer, Hea represents the Heaviside function, and tanh represents the tanh activation function.

2. The disk life monitoring method based on the LSTM model according to claim 1 is characterized in that: The length of the sliding window, the step size of the sliding window, the number of time periods, the duration of the time period, the preset time period, and the preset threshold are all custom parameters.

3. The disk life monitoring method based on the LSTM model according to claim 1 is characterized in that: The calculation formula for normalization is as follows: ; in and Respectively represent the disk usage data before and after normalization at the mth time point and the nth period. and They represent the maximum and minimum values ​​of disk usage data in the nth period at the mth time point.

4. The disk life monitoring method based on the LSTM model according to claim 1 is characterized in that: The second units of the second hidden layer are all constructed based on LSTM units, the first classifier is constructed based on multi-layer perceptron, and the corresponding activation function is Sigmoid activation function.

5. The disk life monitoring method based on the LSTM model according to claim 1 is characterized in that: The second update vector output by the Mth second unit of the second hidden layer is input into the second, third, fourth, fifth and sixth classifiers respectively. The corresponding category spaces are power-on duration, power-on times, number of remapped sectors, number of uncorrectable sectors and temperature, respectively. A preliminary score is first calculated based on the internal operating status indicators, and then the disk life score is obtained through manual labeling based on the preliminary score. The preliminary score is first obtained by weighted summation of the internal operating status indicators, and then the value range of the preliminary score is controlled between 0 and 1 through the Sigmoid activation function.

6. The disk life monitoring method based on the LSTM model according to claim 1 is characterized in that: The disk life prediction model is trained using an optimization algorithm, which includes the following steps: Step S201, randomly generating weight parameters and bias parameters of a disk life prediction model as codes of individuals in an initialization population of an optimization algorithm; Step S202: The disk life prediction model applies the encoding of the individual and calculates the loss value of all individuals in the initialization population through the loss function; The loss function is represented by the mean square error between the output value of the disk life prediction model at the preset number of training iterations and the sample label, where the preset number of training iterations is a custom parameter; Step S203: Generate an iteration coefficient for the current iteration number, and determine if it is greater than or equal to a preset iteration coefficient threshold, then proceed to step S204; otherwise, proceed to step S205; The calculation formula of the iteration coefficient F is as follows: ; Where t indicates the current number of iterations is t, and the starting value of the current number of iterations is 0. represents the maximum number of iterations, Represents the first random number in the range of 0 to 1, Represents a second random number with a value range between -2 and 2, where the preset iteration coefficient threshold is a custom parameter; Step S204: Generate a random number between 0 and 1 for each individual in the initialized population, and determine if the random number is greater than or equal to a first threshold. Then, update the code of the individual using the first exploration strategy; otherwise, update the code of the individual using the second exploration strategy. Step S205: Generate a random number between 0 and 1 for each individual in the initialized population, and determine if the random number is greater than or equal to a second threshold. If so, update the code of the individual using the first development strategy; otherwise, update the code of the individual using the second development strategy. Step S206: If the current number of iterations is incremented by 1 and the iteration termination condition is determined to be satisfied, the code of the individual with the smallest loss value in the initialized population is applied to complete the training of the disk life prediction model; otherwise, the process returns to step S202 to continue the training. The iteration termination condition is that the current number of iterations is greater than or equal to the maximum number of iterations or the loss values ​​of more than 1 / 3 of the individuals in the initialized population are less than the preset loss threshold, where the preset loss threshold is a custom parameter.

7. The disk life monitoring method based on the LSTM model according to claim 6 is characterized in that: The calculation formula of the optimization algorithm includes: The calculation formula of the first exploration strategy is as follows: ; Where 1≤i≤H, i≠j, H represents the total number of individuals in the initialization population and is a custom parameter. represents the code of the i-th individual whose current iteration number is t+1, and They represent the codes of the i-th and j-th individuals with the current iteration number t, Represents a third random number ranging from 0 to 1; The calculation formula of the second exploration strategy is as follows: ; in Indicates the code of the individual with the smallest loss value in the initialization population with the current iteration number t, represents a fourth random number ranging from 0 to 2; The calculation formula for the first development strategy includes: ; ; ; in and They represent the codes of the first intermediate individual and the second intermediate individual of the current iteration number t, Indicates the fifth random number ranging from 0 to 1; The calculation formula for the second development strategy is as follows: ; Where round represents the rounding function.

8. The disk life monitoring system based on LSTM model is characterized by: The method for monitoring disk life based on the LSTM model according to any one of claims 1 to 7 is implemented, comprising: A feature matrix construction module is used to collect disk usage data through a sliding window and construct a feature matrix through time division and normalization processing; The annotation module is used to move the sliding window backward by a preset time period, collect the internal operating status indicators of the disk, and obtain the disk life score through manual annotation; A training module, which uses the feature matrix and the disk life score as sample data and sample labels for training the disk life prediction model, and completes the training of the disk life prediction model through an optimization algorithm; The threshold judgment module is used to obtain the disk life score through the trained disk life prediction model, and if it is judged to be less than the preset threshold, it reminds the user to back up the disk data in time.

Citation Information

Patent Citations

  • A method and system for predicting the life of a solid-state drive based on big data

    CN119739351A