Solid state disk remaining life prediction method, memory and computer program product
By fusing multidimensional health indicators and using a pre-trained network model, the remaining lifespan of solid-state drives is dynamically predicted, solving the problem of inaccurate prediction in existing technologies and achieving higher prediction accuracy and system reliability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHENGDU BIWIN STORAGE TECHNOLOGY CO LTD
- Filing Date
- 2025-12-03
- Publication Date
- 2026-05-05
AI Technical Summary
Current technologies rely on a single-dimensional indicator to predict the remaining lifespan of solid-state drives (SSDs), leading to inaccurate predictions, failure to replace them in a timely manner, and potential data loss and decreased system stability.
By acquiring multidimensional health indicators, a pre-trained convolutional autoencoder network is used to fuse the indicators, and a pre-defined degradation model and gated attention unit network are combined to dynamically predict the remaining lifespan of solid-state drives.
It improves the accuracy and precision of solid-state drive (SSD) remaining life prediction, reduces human intervention, enhances the ability to characterize degradation trends, and ensures the security and reliability of storage systems.
Smart Images

Figure CN121979702A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of storage technology, and in particular to a method for predicting the remaining lifespan of a solid-state drive, a memory, and a computer program product. Background Technology
[0002] Solid-state drives (SSDs), with their low latency and high throughput, have become core storage components in high-performance computing clusters, distributed storage systems (such as OST (Object Storage Target) / MDT (Metadata Target) nodes in the Lustre parallel file system), and cloud computing data centers. However, if SSDs are not inspected and replaced in a timely manner when they reach the end of their lifespan, especially for enterprise customers with high storage demands, it can lead to irreversible data loss, decreased storage system stability, and business downtime, resulting in significant financial and security losses. Therefore, monitoring degradation, building a Health Indicator (HI), and predicting the Remaining Useful Life (RUL) are crucial for ensuring the security and reliability of SSDs. However, current methods for predicting the remaining lifespan of SSDs rely on single-dimensional metrics from log data (such as write / erase cycles), leading to inaccurate predictions.
[0003] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention
[0004] The main objective of this application is to provide a method for predicting the remaining lifespan of a solid-state drive (SSD), a memory, and a computer program product, aiming to solve the technical problem of how to improve the accuracy of predicting the remaining lifespan of an SSD.
[0005] To achieve the above objectives, this application proposes a method for predicting the remaining lifespan of a solid-state drive (SSD). The method includes: To obtain the first multidimensional health indicator of the lifespan of the associated solid-state drive, where the first multidimensional... Health indicators include multiple different dimensions of health indicators; A pre-trained convolutional autoencoder network is used to perform index fusion processing on the first multidimensional health index to obtain the first fused health index. The fault threshold corresponding to the first fusion health index is determined using a preset degradation model; Using a pre-trained gated attention unit network, the remaining lifespan of the solid-state drive is predicted based on the first fused health metric and the fault threshold.
[0006] In one embodiment, the step of determining the fault threshold corresponding to the first fused health indicator using a preset degradation model includes: Based on the first fusion health indicator, at least one degradation model is selected from multiple preset degradation models; In response to selecting a degradation model, the first fused health index is input into the selected degradation model to predict the failure threshold; In response to selecting multiple degradation models, the information criteria information corresponding to the selected multiple degradation models is determined. The degradation model corresponding to the smallest information criterion information among the selected multiple degradation models is taken as the target degradation model. The first fused health index is input into the target degradation model to predict and obtain the fault threshold.
[0007] In one embodiment, the step of selecting at least one degradation model from a plurality of preset degradation models based on a first fusion health index includes: Determine the trend chart of the first integrated health indicator, which includes the first time point to the second time point, where the first time point is shorter than the second time point; Based on the first trend information corresponding to the first integrated health indicator represented by the trend chart, at least one degradation model is selected from multiple preset degradation models.
[0008] In one embodiment, the degradation model includes a linear distribution model, an exponential distribution model, and a power-law distribution model. The step of selecting at least one degradation model from multiple preset degradation models based on the first trend information corresponding to the first fusion health indicator represented by the trend chart includes at least one of the following: In response to the first trend information, and to characterize the second trend information that the first fusion health indicator shows a stable linear decline over time, a linear distribution model is selected from the linear distribution model, exponential distribution model, and power law distribution model. In response to the first trend information, to characterize the third trend information of the accelerated decline in the later stage of the first integrated health indicator, the exponential distribution model is selected from the linear distribution model, exponential distribution model and power law distribution model. In response to the first trend information, and to characterize the fourth trend information that represents the nonlinear degradation trend caused by sudden load, the power law distribution model is selected from the linear distribution model, the exponential distribution model, and the power law distribution model. In response to first trend information including at least two of first trend information, second trend information and third trend information, at least two degenerate models are selected from linear distribution models, exponential distribution models and power law distribution models.
[0009] In one embodiment, the step of selecting at least two degenerate models from a linear distribution model, an exponential distribution model, and a power-law distribution model in response to the first trend information including at least two of the first trend information, second trend information, and third trend information includes: In response to the first trend information including at least two of the first trend information, the second trend information, and the third trend information, the boundary point between different trend information in the trend graph is determined; In response to the boundary point including a first boundary point, for a first fusion health indicator belonging to the range from the first time point to the first boundary point, the step of selecting at least one degradation model from a plurality of preset degradation models based on the first fusion health indicator is executed; for a first fusion health indicator belonging to the range from the first boundary point to the second time point, the step of selecting at least one degradation model from a plurality of preset degradation models based on the first fusion health indicator is executed. In response to the boundary points including two different second and third boundary points, for the first fusion health indicator belonging to the range from the first time point to the second boundary point, the step of selecting at least one degradation model from multiple preset degradation models based on the first fusion health indicator is executed; for the first fusion health indicator belonging to the range from the second to the third boundary point, the step of selecting at least one degradation model from multiple preset degradation models based on the first fusion health indicator is executed; for the first fusion health indicator belonging to the range from the third boundary point to the second time point, the step of selecting at least one degradation model from multiple preset degradation models based on the first fusion health indicator is executed.
[0010] In one embodiment, the step of predicting the remaining lifespan of a solid-state drive (SSD) based on a first fused health metric and a fault threshold using a pre-trained gated attention unit network includes: Construct a first Hankel matrix that includes the first fusion health metric for the first time period; The first Hankel matrix is input into a pre-trained gated attention network for prediction, and the first prediction result is output. In response to the first prediction result being less than the fault threshold, the first prediction result is used as the first fusion health indicator for the next time step and updated to the first Hankel matrix. Based on the updated first Hankel matrix, the step of inputting the first Hankel matrix into the pre-trained gated attention network for prediction is performed until the latest obtained first prediction result is detected to be greater than or equal to the fault threshold. Based on the latest first prediction result, the number of predictions to be made by the pre-trained gating attention network is determined, and the sampling time corresponding to the first fused health indicator is determined. The sampling time includes the sampling duration and sampling interval time that characterize the sampling of the first multidimensional health indicator. The sampling time is updated based on the number of predictions to obtain the remaining lifespan of the solid-state drive.
[0011] In one embodiment, the method for predicting the remaining lifespan of a solid-state drive further includes at least one of the following: The pre-trained convolutional autoencoder network is trained using a pre-set first training dataset until the first training cutoff condition is met, resulting in a pre-trained convolutional autoencoder network. The training data in the first training dataset consists of a second multidimensional health indicator and a second fusion health indicator corresponding to the second multidimensional health indicator. The pre-defined gated attention unit network is trained using a pre-defined second training dataset until the second training cutoff condition is met, resulting in a pre-trained gated attention unit network. The training data in the second training dataset includes at least one third fusion health metric.
[0012] In one embodiment, the convolutional autoencoder network includes a convolutional encoder and a deconvolutional decoder. The steps of training a pre-defined convolutional autoencoder network using a pre-defined first training dataset until a first training cutoff condition is met, to obtain a pre-trained convolutional autoencoder network, include: The second multidimensional health indicator is input into a preset convolutional autoencoder network, and the data features of the second multidimensional health indicator are extracted according to the convolutional encoder to obtain the encoded features. The fourth fusion health index is obtained by reconstructing the encoded features based on the deconvolution decoder. The first loss function is used to calculate the loss function value of the fourth fusion health index and the second fusion health index, and the first loss function value is obtained. The convolutional autoencoder network is then updated in reverse according to the first loss function value until the first training cutoff condition is reached, and the pre-trained convolutional autoencoder network is obtained.
[0013] In one embodiment, the convolutional encoder includes a convolutional layer, a first activation function, a Dropout layer, a pooling layer, and a fully connected layer. The steps for extracting data features from the second multidimensional health indicator using a convolutional encoder to obtain encoded features include: Based on the convolutional layer, each one-dimensional health indicator in the second multidimensional health indicator is convolved to obtain the first convolution result; The activation process is performed on each first convolution result according to the first activation function, and the first convolution result after activation is randomly deactivated according to the Dropout layer. Then, the first convolution result after random deactivation is downsampled according to the pooling layer to obtain the downsampled first convolution result. The encoded features are obtained by performing a full connection on the downsampled first convolution results based on the fully connected layer.
[0014] In one embodiment, the deconvolution decoder includes a deconvolution layer, a second activation function, and an upsampling layer. The steps for reconstructing the encoded features using a deconvolutional decoder to obtain the fourth fused health indicator include: The encoded features are upsampled using the upsampling layer to obtain sampled data; The sampled data is deconvolutionally activated based on the deconvolution layer and the second activation function to obtain the fourth fused health index.
[0015] In one embodiment, the gated attention unit network includes a reset gate, an update gate, and an attention gate. The steps of training a pre-defined gated attention unit network using a pre-defined second training dataset until a second training cutoff condition is met, to obtain a pre-trained gated attention unit network, include: Construct a second Hankel matrix that includes at least one third fusion health indicator, and input the second Hankel matrix and the historical information of the solid-state drive into a preset gated attention unit network, wherein the historical information includes the previous hidden state corresponding to the third fusion health indicator; Using the reset gate, the output information of the reset gate is determined based on the second Hankel matrix and historical information; The update gate output information is determined based on the second Hankel matrix and historical information. The candidate hidden state is determined based on the reset door output information, the updated door output information, and historical information; Using attention gates, the attention distribution matrix is determined based on the output information of the reset gate and the output information of the update gate; The sixth fusion health indicator is determined based on the candidate hidden state, reset gate information, updated gate information, and attention distribution matrix; The gated attention network is updated in reverse according to the preset second loss function and the sixth fusion health index until the second training cutoff condition is met, thus obtaining the pre-trained gated attention network.
[0016] In one embodiment, the health metrics include at least one of the following: total written data volume, daily drive write volume, write amplification factor, Nand write volume, Host write volume, P / E cycles, percentage of lifetimes, number of bad blocks, wear index, available remaining space, uncorrectable error count, write error count, erase error count, correctable error count, cyclic redundancy check error count, sequential read / write bandwidth performance, random read / write IOPS performance, random read / write latency performance, and random read / write QoS performance.
[0017] In addition, to achieve the above objectives, this application also proposes a memory, which includes a main control chip and a storage chip. The storage chip stores a computer program, and the main control chip can execute the computer program to implement the steps of the solid-state drive remaining life prediction method described above.
[0018] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the solid-state drive remaining life prediction method as described above.
[0019] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the solid-state drive remaining life prediction method as described above.
[0020] One or more technical solutions proposed in this application have at least the following technical effects: In this embodiment, a first multi-dimensional health indicator related to the lifespan of the solid-state drive (SSD) is obtained. This first multi-dimensional health indicator includes multiple health indicators of different dimensions. This allows for the use of multiple health indicators reflecting the SSD's degradation trend as the first multi-dimensional health indicator, providing sufficient data for the preliminary step of SSD lifespan prediction and avoiding the inaccurate prediction of remaining SSD lifespan using a single-dimensional indicator. Furthermore, a pre-trained convolutional autoencoder network is used to perform indicator fusion processing on the first multi-dimensional health indicator to obtain a first fused health indicator. This enables unsupervised multi-dimensional feature fusion of multiple health indicators of different dimensions through the convolutional autoencoder network, reducing manual intervention while improving the representation of SSD degradation trends by the health indicators. Finally, a preset degradation model is used to determine the fault threshold corresponding to the first fused health indicator, allowing for dynamic and reasonable setting of fault thresholds for different dimensions of health indicators. Furthermore, it utilizes a pre-trained gated attention unit network to predict the remaining lifespan of the solid-state drive (SSD) based on the first fused health metric and fault threshold. In turn, the gated attention unit network, which combines gated recurrent units and attention mechanisms, can effectively capture long-term trends and enhance dynamic feature acquisition capabilities, thereby improving the accuracy and precision of SSD remaining lifespan prediction. Attached Figure Description
[0021] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0022] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a schematic diagram of the first process of an embodiment of the solid-state drive remaining lifespan prediction method of this application; Figure 2 This is a flowchart illustrating an embodiment of the solid-state drive remaining lifespan prediction method of this application; Figure 3 This is a schematic diagram of the second process in an embodiment of the solid-state drive remaining lifespan prediction method of this application; Figure 4 This is a schematic diagram of the third process in an embodiment of the solid-state drive remaining lifespan prediction method of this application; Figure 5 This is a schematic diagram of the CAE network architecture in the solid-state drive remaining life prediction method of this application; Figure 6 This is a schematic diagram of the GRU network architecture in the solid-state drive remaining life prediction method of this application; Figure 7 This is a schematic diagram of the GAU network architecture in the solid-state drive remaining life prediction method of this application; Figure 8 This is a schematic diagram of the memory structure involved in the solid-state drive remaining lifespan prediction method in the embodiments of this application.
[0024] Explanation of icon numbers: C. Join operation; M. Product operation; A. Summation operation; Tanh function; Sigmoid function Input data; The previous hidden state; , Reset the output of the door; Update the gate output; 1. Current hidden state; 400. Memory; 410. Main control chip; 420. Storage chip.
[0025] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0026] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0027] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0028] As a high-performance storage medium, solid-state drives (SSDs) have become core storage components in high-performance computing clusters, distributed storage systems (such as OST / MDT nodes in the Lustre parallel file system), and cloud computing data centers due to their low latency and high throughput. However, if SSDs are not inspected and replaced in a timely manner when they reach the end of their lifespan, especially for enterprise customers with high storage demands, it can lead to irreversible data loss, decreased storage system stability, and business downtime, resulting in significant financial and security losses. Therefore, monitoring degradation, building a Health Indicator (HI), and predicting their Remaining Useful Life (RUL) are crucial for ensuring the security and reliability of SSDs.
[0029] However, current methods for predicting the lifetime of SSDs (RUL) have some shortcomings. For example, the traditional Hierarchical Index (HI) suffers from weak degradation characterization capabilities due to its singular nature. It relies on single-dimensional indicators from log data and lacks collaborative modeling of physical degradation mechanisms (such as electron migration effects and oxide layer fatigue in NAND flash memory) and user behavior characteristics (such as write amplification factor and load fluctuations). Furthermore, the reliance on human experience in integrated HI leads to high time costs. Although it can extract multi-dimensional indicators reflecting SSD degradation, including UCEr (Uncorrectable Error Rate), PER (Program Error Ratio), EEr (Erase Error Ratio), CER (Correctable Error Ratio), wear status, CRC (Cyclic Redundancy Check Error Count), and performance jitter, and fit them into a lifetime prediction model function using a binomial formula, the high degree of human intervention results in a high time cost for integrated HI. Furthermore, some HI fusion methods rely on labeled data to generate degradation indicators, but in real-world scenarios, SSD failure samples are scarce, limiting the model's generalization ability. For example, the RUL prediction model's insufficient ability to acquire long-term trends and dynamic features leads to low prediction accuracy.
[0030] In summary, in the current field of RUL prediction for SSD, the degradation characterization capabilities of multi-dimensional fusion HI need to be strengthened, the time cost of manual intervention needs to be reduced, and the accuracy of RUL prediction needs to be improved.
[0031] Therefore, in this embodiment, by acquiring a first multi-dimensional health indicator related to the lifespan of the solid-state drive (SSD), and this first multi-dimensional health indicator includes multiple health indicators of different dimensions, it is possible to use multiple health indicators of different dimensions that can reflect the degradation trend of the SSD as the first multi-dimensional health indicator. This provides sufficient data for the preliminary step of SSD lifespan prediction and avoids the inaccurate prediction of the remaining lifespan of the SSD when using a single-dimensional indicator. Furthermore, a pre-trained convolutional autoencoder network is used to perform indicator fusion processing on the first multi-dimensional health indicator to obtain a first fused health indicator. This enables unsupervised multi-dimensional feature fusion of multiple health indicators of different dimensions through the convolutional autoencoder network, reducing manual intervention while improving the representation of the SSD degradation trend by the health indicators. Finally, a preset degradation model is used to determine the fault threshold corresponding to the first fused health indicator, allowing for dynamic and reasonable setting of fault thresholds for different dimensions of health indicators. Furthermore, it utilizes a pre-trained gated attention unit network to predict the remaining lifespan of the solid-state drive (SSD) based on the first fused health metric and fault threshold. In turn, the gated attention unit network, which combines gated recurrent units and attention mechanisms, can effectively capture long-term trends and enhance dynamic feature acquisition capabilities, thereby improving the accuracy and precision of SSD remaining lifespan prediction.
[0032] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication and program execution functions, such as a memory, computer, mobile phone, etc., or an electronic device capable of realizing the above functions.
[0033] Reference Figure 1 , Figure 1 This paper illustrates a first flowchart of a solid-state drive (SSD) remaining lifespan prediction method provided in an embodiment of this application. In this embodiment, the SSD remaining lifespan prediction method includes steps S10-S40: Step S10: Obtain the first multidimensional health indicator of the associated solid-state drive lifespan, wherein the first multidimensional health indicator includes multiple health indicators of different dimensions. Optionally, the solid-state drive (SSD) can be monitored during operation to obtain health indicators reflecting SSD degradation information. These indicators can be used as health indicators related to the lifespan of the SSD. Multiple health indicators of different dimensions can be obtained and preprocessed, such as outlier removal, timing smoothing interpolation, absolute value taking, batch standardization, and normalization, to obtain multi-dimensional health indicators of the SSD, which can then be used as the first multi-dimensional health indicator.
[0034] Optionally, health indicators can be time-series data reflecting the degradation trend of solid-state drives, and can be presented in the form of text data or image data.
[0035] Optionally, health metrics include at least one of the following: total written data, daily drive writes, write amplification factor, Nand writes, Host writes, P / E cycles (Program / Erase Cycle Count), percentage of lifetimes, number of bad blocks, wear index, available remaining space, uncorrectable error count, write error count, erase error count, correctable error count, cyclic redundancy check error count, sequential read / write bandwidth performance, random read / write IOPS performance, random read / write latency performance, and random read / write QoS performance.
[0036] Optionally, when collecting the first multidimensional health indicator, a fixed time interval can be set as the collection time interval (e.g., 5 minutes), and multiple health indicators of different dimensions can be collected once at each collection time interval.
[0037] Optionally, during the operation of the SSD, it can be monitored using tools such as SMART, nvme-cli, and fio to collect multiple health indicators from different dimensions. These health indicators can be queried using management commands such as smartctl -a / dev / $dev (applicable to both NVMe and SATA SSDs), nvme smart-log / dev / $dev, and nvme intel smart-log-add / dev / $dev (NVMe SSD).
[0038] Optionally, for health indicators in terms of performance dimensions, the fio tool can be used to perform 128k sequential read / write bandwidth performance tests, 4k random read / write IOPS performance tests, 4k random read / write latency performance and QoS performance tests, and after the tests are completed, the log data of each test item (such as the result.log result) can be viewed to obtain health indicators in terms of performance dimensions at different time points.
[0039] Optionally, performance testing may include at least one of the following: (1) Sequential write to the entire disk twice. (2) Sequential read bandwidth test for 10 minutes. (3) Sequential bandwidth test for 10 minutes. (4) Random write to the entire disk twice (testing steady-state performance). (5) Random read IOPS test for 10 minutes. (6) Random write IOPS test for 10 minutes. (7) Random read latency & QoS test for 10 minutes. (8) Random write latency & QoS test for 10 minutes.
[0040] Optionally, after collecting various health indicators from multiple dimensions, preprocessing operations such as outlier removal, time-series smoothing interpolation, absolute value taking, batch standardization, and normalization can be performed to obtain a first multidimensional health indicator that reflects the degradation trend of the solid-state drive.
[0041] Optionally, outlier removal can filter out values that differ significantly from the degradation trend of solid-state drives, thereby improving the usability of the health indicator; time-series smoothing interpolation can make each health indicator form a time series data that is easy to process; taking absolute values is to unify the changing trends of different indicators, making data analysis easier; batch standardization and normalization are to unify the units of measurement, making data processing easier.
[0042] Optionally, the formula for batch standardization can be the following formula (I), and the formula for normalization can be the following formula (II).
[0043] Formula (1); Formula (II); Optionally, Represents the raw sequence data of various health indicators. This represents data after batch standardization of health indicators. This represents data after normalization of health indicators. This represents the mean of a data sample (such as a health indicator). The standard deviation of the data sample is represented by... and These represent the maximum and minimum values of each sampling period (e.g., the maximum and minimum values of multiple health indicators within the same sampling period).
[0044] Step S20: Use a pre-trained convolutional autoencoder network to perform index fusion processing on the first multidimensional health index to obtain the first fused health index; Optionally, the pre-trained convolutional autoencoder network can be a pre-trained convolutional autoencoder (CAE) network.
[0045] Alternatively, historical data (such as known multidimensional health indicators, like the second multidimensional health indicator) can be used to train the convolutional autoencoder network.
[0046] Optionally, a convolutional autoencoder network can combine the feature extraction function of a convolutional neural network with the unsupervised feature reconstruction function of an autoencoder to achieve unsupervised fusion of input data.
[0047] Optionally, the first multidimensional health indicator can be input into a pre-trained convolutional autoencoder network for unsupervised fusion, thereby realizing the fusion processing of multiple health indicators of different dimensions and obtaining the first fused health indicator after fusion processing.
[0048] Alternatively, a multi-dimensional first multidimensional health indicator can be converted into a one-dimensional first fused health indicator using a pre-trained convolutional autoencoder network.
[0049] Step S30: Determine the fault threshold corresponding to the first fusion health indicator using a preset degradation model; Alternatively, the fault threshold can be determined by using users' experience values and historical data statistics, or by constructing a statistical degradation model.
[0050] Optionally, the degradation model may include a linear distribution model, an exponential distribution model, and a power-law distribution model.
[0051] Optionally, the linear distribution model can be as shown in Formula (III), the exponential distribution model can be as shown in Formula (IV), and the power law distribution model can be as shown in Formula (V).
[0052] = + t+ , Formula (3); = Formula (IV); = Formula (5); in, This is a linear distribution model; t is a time point (e.g., the current time point). The initial health status (intercept), such as the first fusion health indicator; The degradation rate (slope); The random error follows a normal distribution with a mean of 0 and a variance of . . It is an exponential distribution model; This represents the initial health state, such as the first integrated health indicator; is the decay rate, and is greater than 0; 'a' is a time point, such as 't'; It is random noise. For a power-law distribution model, This is a scaling parameter (the scaling factor for the initial state of the solid-state drive). The power exponent ( >0 determines the rate of degradation of the solid-state drive (SSD). This is the time offset to prevent divergence at t=0.
[0053] Optionally, a suitable degradation model can be selected from multiple degradation models to determine the fault threshold based on the trend performance corresponding to different first fusion health indicators.
[0054] Optionally, the first fused health indicator can be input into a preset degradation model to output a fault threshold.
[0055] Step S40: Using a pre-trained gated attention unit network, predict the remaining lifespan of the solid-state drive based on the first fused health indicator and the fault threshold.
[0056] Optionally, the pre-trained gated attention network can be a pre-trained gated attention unit (GAU) network.
[0057] Optionally, when training the gated attention unit network model, the root mean square error can be used as the loss function for training the gated attention unit network, and a real-time recurrent learning algorithm can be used to train the gated attention unit network model.
[0058] Optionally, the remaining lifespan of the solid-state drive can also be predicted based on the first fused health index according to the GRU. In this embodiment, only the GAU is used as an example.
[0059] Optionally, step S40, which involves using a pre-trained gated attention unit network to predict the remaining lifespan of the solid-state drive based on the first fused health indicator and the fault threshold, includes steps a10-a60.
[0060] Step a10: Construct a first Hankel matrix containing a first fusion health metric for a first time range; Optionally, the first time range may include multiple time points. These time points can be the time nodes where health indicators from different dimensions are collected and fused to obtain the first fused health indicator. For example, it may include the time point where the first multidimensional health indicator was collected most recently and fused to obtain the first fused health indicator (such as the current time point). Furthermore, there may be multiple first fused health indicators within the first time range, such as five.
[0061] Optionally, multiple health indicators of different dimensions at multiple different time points can be collected, and the first multidimensional health indicators at each different time point can be fused using a pre-trained convolutional autoencoder network to obtain the first fused health indicators at each time point. Then, a Hankel matrix can be constructed from the first fused health indicators at each time point to obtain the first Hankel matrix.
[0062] Step a20: Input the first Hankel matrix into the pre-trained gated attention network for prediction, and output the first prediction result; Step a30: In response to the first prediction result being less than the fault threshold, the first prediction result is used as the first fusion health indicator for the next time step and updated to the first Hankel matrix. Optionally, the first Hankel matrix can be input into a pre-trained gated attention network for prediction, and the output can be the first prediction result, which can be the first fusion health index of the next time step, i.e. the first fusion health index of the next sampling period.
[0063] Optionally, the first fused health indicator can be compared with the fault threshold determined by the degradation model. If the first fused health indicator is greater than or equal to the fault threshold, it can be determined that the solid-state drive may be faulty at the current moment, and the corresponding alarm information can be output.
[0064] Optionally, if the first fusion health index is less than the fault threshold, the pre-trained gated attention unit network can continue to make predictions. The first prediction result is used as the first fusion health index in the next time step to update the first Hankel matrix. For example, the first prediction result can be added to the last position in the first Hankel matrix, and the first position in the first Hankel matrix can be deleted or not deleted to obtain the updated first Hankel matrix.
[0065] Step a40: Based on the updated first Hankel matrix, perform the step of inputting the first Hankel matrix into the pre-trained gated attention network for prediction until the latest obtained first prediction result is detected to be greater than or equal to the fault threshold. Optionally, the updated first Hankel matrix can be input into the pre-trained gated attention network for prediction again until the prediction result is greater than or equal to the fault threshold. At this point, the gated attention network can be stopped for prediction processing.
[0066] Optionally, when updating the first Hankel matrix, the fault threshold can also be updated synchronously based on the first prediction result. For example, the first prediction result can be used as the first fusion health indicator for the next time step, and the degradation model corresponding to the first fusion health indicator for the next time step can be selected from among the degradation models. The fault threshold can be determined based on the selected degradation model. That is, step S30 can be executed based on the first prediction result.
[0067] Optionally, each time a prediction is made using the pre-trained gated attention network, the predicted result is compared with the latest fault threshold until the latest predicted result is detected to be greater than or equal to the latest fault threshold.
[0068] For example, if the first Hankel matrix is X= The first prediction result is Then, the first prediction result is updated to the first Hankel matrix, and the updated first Hankel matrix is X= The updated first Hankel matrix is then input into the pre-trained gated attention network for prediction, and training is performed m times. The output at the nth prediction time is then: .in, This can represent a pre-trained gated attention network, equivalent to a mapping function. m and n are positive integers.
[0069] Step a50: Based on the latest first prediction result, determine the number of predictions to be made by the pre-trained gating attention network, and determine the sampling time corresponding to the first fused health indicator. The sampling time includes the sampling duration and sampling interval time that characterize the sampling of the first multidimensional health indicator. Optionally, the remaining lifespan of the solid-state drive can be predicted when the first prediction result obtained by the pre-trained gated attention unit network is greater than or equal to the fault threshold, such as when the latest first prediction result is greater than or equal to the fault threshold.
[0070] Optionally, the number of predictions made by the pre-trained gated attention network within the current sampling period can be determined, and the sampling duration and sampling interval representing the first multidimensional health indicator can be determined.
[0071] Optionally, the sampling duration can be the collection time for a single multidimensional health indicator collection, i.e., the time it takes to collect the first multidimensional health indicator within a sampling period. The sampling interval can be the time interval between adjacent times corresponding to the first multidimensional health indicator, such as a pre-set fixed time interval or a collection time interval.
[0072] Alternatively, the sum of the sampling duration and the sampling interval can be used as the sampling time.
[0073] Step a60: Update the sampling time based on the number of predictions to obtain the remaining lifespan of the solid-state drive.
[0074] Optionally, the product of the number of predictions and the sampling time can be calculated and used as the remaining lifespan of the solid-state drive, which is the time required for the first multidimensional health indicator at the current moment to reach the failure threshold.
[0075] Alternatively, the remaining lifespan of the solid-state drive can be calculated using the following formula (vi).
[0076] RUL= *( + Formula (VI); RUL represents the remaining lifespan of the solid-state drive. To predict the number of times, and These represent the sampling duration and the sampling interval, respectively.
[0077] In this embodiment, the first fused health index is converted into a Hankel matrix and input into a pre-trained gated attention network. The gated attention network is then used for prediction until the latest first prediction result is greater than or equal to the fault threshold. The remaining lifespan of the solid-state drive is determined based on the number of predictions made by the gated attention network and the sampling time corresponding to the first fused health index. This effectively captures the long-term trend of the first fused health index and enhances the dynamic feature acquisition capability, thereby improving the prediction accuracy of the remaining lifespan of the solid-state drive.
[0078] In addition, to aid in understanding the principle and process of predicting the remaining lifespan of the solid-state drive in this embodiment, an example is provided below.
[0079] For example, such as Figure 2 As shown, the SSD test environment is first prepared, for example, an enterprise-grade SSD working normally as a data disk on a server system. Then, data preprocessing is performed, which involves collecting health indicators from multiple dimensions and performing preprocessing operations to obtain the first multi-dimensional health indicators. For example, 23 important health indicators reflecting SSD degradation information are monitored through Linux commands to form an original feature set, which is then divided into a training set and a test set. Health indicators can include total written data, daily drive write volume, write amplification factor, Nand write volume, Host write volume, P / E cycles, percentage of used life, number of bad blocks, wear index, available remaining space, uncorrectable error count, write error count, erase error count, correctable error count, cyclic redundancy check error count, sequential read / write bandwidth performance, random read / write IOPS performance, random read / write latency performance, and random read / write QoS performance.
[0080] Then, the CAE network is used to generate SSD-HI, that is, the first multidimensional health index is fused using a pre-trained convolutional autoencoder network to obtain the first fused health index. For example, the CAE network is constructed using programming software, and the network structure is as follows: input layer - encoder (6 convolutional layers) - dropout layer - max pooling layer - fully connected layer - upsampling layer - decoder (6 deconvolutional layers) - output layer. The pre-processed training set data is input into the CAE for training, and the test set data is input into the trained CAE. The data is fused to generate SSD-HI, thus obtaining the pre-trained convolutional autoencoder network for subsequent RUL prediction.
[0081] Then, the GAU network is used for remaining lifespan prediction. This involves using a pre-trained gated attention unit network to predict the remaining lifespan of the solid-state drive (SSD) based on a first fused health metric and a fault threshold. For example, the GAU network can be constructed using programming software, with the network structure consisting of: input unit - reset gate - update gate - attention gate - output unit. The first fused health metric generated by the CAE network in the previous step is input into the GAU network to predict the remaining lifespan of the SSD.
[0082] Then, when the remaining lifespan of the solid-state drive (SSD) is about to expire, a fault warning is triggered. For example, programming software is used to execute the fault warning triggering steps. Based on the trend information of the first fused health indicator, the corresponding degradation model is identified, and the target degradation model (including degradation models for a single time period or multiple time periods) is selected according to the AIC and BIC information criteria. Combined with the graded warning mechanism, a fault threshold is determined, and when the fault threshold is exceeded, a fault warning is triggered, prompting the manufacturer that the SSD's lifespan is about to expire and it needs to be replaced.
[0083] In this embodiment, a first multi-dimensional health indicator related to the lifespan of the solid-state drive (SSD) is obtained. This first multi-dimensional health indicator includes multiple health indicators of different dimensions. This allows for the use of multiple health indicators reflecting the SSD's degradation trend as the first multi-dimensional health indicator, providing sufficient data for the preliminary step of SSD lifespan prediction and avoiding the inaccurate prediction of remaining SSD lifespan using a single-dimensional indicator. Furthermore, a pre-trained convolutional autoencoder network is used to perform indicator fusion processing on the first multi-dimensional health indicator to obtain a first fused health indicator. This enables unsupervised multi-dimensional feature fusion of multiple health indicators of different dimensions through the convolutional autoencoder network, reducing manual intervention while improving the representation of SSD degradation trends by the health indicators. Finally, a preset degradation model is used to determine the fault threshold corresponding to the first fused health indicator, allowing for dynamic and reasonable setting of fault thresholds for different dimensions of health indicators. Furthermore, it utilizes a pre-trained gated attention unit network to predict the remaining lifespan of the solid-state drive (SSD) based on the first fused health metric and fault threshold. In turn, the gated attention unit network, which combines gated recurrent units and attention mechanisms, can effectively capture long-term trends and enhance dynamic feature acquisition capabilities, thereby improving the accuracy and precision of SSD remaining lifespan prediction.
[0084] Reference Figure 3 , Figure 3 This illustration shows a second flowchart of the solid-state drive remaining lifespan prediction method provided in this application embodiment. Contents identical or similar to those in the above embodiments can be referred to the above description and will not be repeated hereafter. In step S30, the step of determining the fault threshold corresponding to the first fused health index using a preset degradation model includes steps S31 to S33: Step S31: Select at least one degradation model from multiple preset degradation models based on the first fusion health index; Optionally, multiple degradation models can be set in advance, and one or more degradation models can be selected from multiple degradation models based on the first fusion health indicator at the current moment, such as linear distribution model, exponential distribution model and power law distribution model.
[0085] Optionally, step S31, which involves selecting at least one degradation model from among multiple preset degradation models based on the first fusion health index, includes steps b10-b20.
[0086] Step b10: Determine the trend chart of the first fused health indicator, which includes the data from the first time point to the second time point. Step b20: Based on the first trend information corresponding to the first fusion health indicator represented by the trend chart, select at least one degradation model from multiple preset degradation models.
[0087] Optionally, the start and end times of the first time range can be designated as the first and second time points, respectively. The first time point can be a first historical time point, such as the time point 10 minutes before the current time point, and the second time point can be the current time point.
[0088] Optionally, the first and second time points can be the start and end time points for collecting health indicators in multiple different dimensions. For example, health indicators can be collected in five sampling cycles between the first and second time points to obtain five first multidimensional health indicators at five different time points. These indicators can then be fused to obtain five first fused health indicators. A trend chart of these five first fused health indicators over time can then be constructed.
[0089] Optionally, a trend chart corresponding to the first integrated health indicator can be constructed. This trend chart can represent the trend information of the first integrated health indicator changing over time. For example, when the first time point is a historical time node and the second time point is the current time node, it can be the trend information of the first integrated health indicator changing over time within the range from the first historical time node to the current time node, and use it as the first trend information.
[0090] Optionally, at least one degradation model can be selected from multiple degradation models based on the first trend information, and the failure threshold can be predicted based on the selected degradation model.
[0091] Optionally, different degradation models are selected for different first trend information.
[0092] In this embodiment, at least one degradation model is selected from multiple preset degradation models based on the first trend information corresponding to the first fused health indicator represented by the trend chart. Then, the fault threshold is determined based on the selected degradation model. This ensures that the selected degradation model threshold is closely related to the actual first fused health indicator, making the subsequently determined fault threshold more accurate and effective, thereby indirectly improving the accuracy of the prediction of the remaining lifespan of the solid-state drive.
[0093] Optionally, step b20, which involves selecting at least one degradation model from a plurality of preset degradation models based on the first trend information corresponding to the first fusion health indicator represented by the trend graph, includes at least one of the following steps c10-c40.
[0094] Step c10: In response to the first trend information, to characterize the second trend information that the first fused health indicator shows a stable linear decline over time, a linear distribution model is selected from the linear distribution model, exponential distribution model, and power law distribution model. Optionally, image recognition can be performed on the trend graph to determine the first trend information of the first fused health indicator. When the first trend information is a trend information that represents the first fused health indicator showing a stable linear decline over time, the first trend information can be used as the second trend information. A linear distribution model can be selected from linear distribution models, exponential distribution models, and power law distribution models, and then the fault threshold can be predicted based on the selected linear distribution model.
[0095] Step c20: In response to the first trend information, to characterize the third trend information of the accelerated decline in the later stage of the first fused health indicator, the exponential distribution model is selected from the linear distribution model, the exponential distribution model, and the power law distribution model. Optionally, when the first trend information is the trend information that characterizes the accelerated decline of the first fused health indicator in the later stage, the first trend information can be used as the third trend information, and the exponential distribution model can be selected from the linear distribution model, exponential distribution model and power law distribution model, and then the fault threshold can be predicted based on the selected exponential distribution model.
[0096] Step c30: In response to the first trend information, and to select the power law distribution model from the linear distribution model, exponential distribution model, and power law distribution model for the fourth trend information characterizing the nonlinear degradation trend caused by the sudden load. Optionally, when the first trend information is a characterization of the nonlinear degradation trend caused by sudden load, the first trend information can be used as the fourth trend information, and the power law distribution model can be selected from the linear distribution model, the exponential distribution model and the power law distribution model, and then the fault threshold can be predicted based on the selected power law distribution model.
[0097] Step c40, in response to the first trend information including at least two of the second, third, and fourth trend information, select at least two degenerate models from the linear distribution model, the exponential distribution model, and the power-law distribution model.
[0098] Optionally, since the first trend information may involve a combination of multiple trend information, such as initial stability, mid-term fluctuations, and rapid decline at the end, multiple degradation models can be selected to jointly predict the fault threshold.
[0099] Optionally, when the first trend information includes second, third, and fourth trend information, a linear distribution model, an exponential distribution model, and a power-law distribution model can be selected to jointly predict the fault threshold. (This is repeated three times in the original text.)
[0100] In this embodiment, by selecting different degradation models in response to different trend information, it can be ensured that the selected degradation model threshold is closely related to the actual first fusion health index, making the subsequently determined fault threshold more accurate and effective, thereby indirectly improving the accuracy of the prediction of the remaining lifespan of the solid-state drive.
[0101] Optionally, in step c40, in response to the first trend information including at least two of the second, third, and fourth trend information, the step of selecting at least two degenerate models from the linear distribution model, the exponential distribution model, and the power-law distribution model includes steps d10-d30.
[0102] Step d10: In response to the first trend information including at least two of the second trend information, the third trend information, and the fourth trend information, determine the dividing point between different trend information in the trend graph; Optionally, when the first trend information includes at least two of the second, third, and fourth trend information, different degradation models can be selected based on the trend information for different time periods to predict the fault threshold for different time periods.
[0103] Optionally, the dividing point between different trend information in the trend chart can be determined. For example, if the first trend information represented in the trend chart is followed by the second trend information and the third trend information in chronological order, then the time interval between the second trend information and the third trend information in the trend chart can be determined and used as the dividing point.
[0104] Optionally, the CUSUM algorithm (cumulative sum algorithm) can be used to detect and determine the dividing point in the trend chart, or a large model can be used to perform image recognition on the trend chart to determine the dividing point. Other methods can also be used, and no restrictions are placed here.
[0105] Step d20: In response to the boundary point including a first boundary point, for the first fusion health index belonging to the range from the first time point to the first boundary point, perform the step of selecting at least one degradation model from a plurality of preset degradation models based on the first fusion health index; for the first fusion health index belonging to the range from the first boundary point to the second time point, perform the step of selecting at least one degradation model from a plurality of preset degradation models based on the first fusion health index. Optionally, when the first trend information includes two different trend information, the dividing point between the two different trend information can be determined and used as the first dividing point.
[0106] Optionally, since the first trend information is the trend information representing the change of the first fusion health index over time within the time range from the first time point to the second time point, the first fusion health index within the time range from the first time point to the first time point and the first fusion health index within the time range from the first time point to the second time point can be determined after the first boundary point is determined.
[0107] Optionally, the trend information of the first integrated health indicator changing over time (hereinafter referred to as the first sub-trend information) can be determined within the time range from the first time point to the first boundary point, and the trend information of the first integrated health indicator changing over time (hereinafter referred to as the second sub-trend information) can be determined within the time range from the first boundary point to the second time point.
[0108] Optionally, the same operation can then be performed on the first sub-trend information and the second sub-trend information respectively, i.e., selecting their respective corresponding degradation models.
[0109] For example, if the first sub-trend information indicates that the first fused health indicator shows a stable linear decline over time, a linear distribution model can be selected to predict the fault threshold within the time range from the first time point to the first cutoff point. If the second sub-trend information indicates that the first fused health indicator declines more rapidly in the later stages, an exponential distribution model can be selected to predict the fault threshold within the time range from the first cutoff point to the second time point.
[0110] Step d30: In response to the boundary point including two different second boundary points and a third boundary point, for the first fusion health indicator belonging to the range from the first time point to the second boundary point, the step of selecting at least one degradation model from multiple preset degradation models based on the first fusion health indicator is executed; for the first fusion health indicator belonging to the range from the second boundary point to the third boundary point, the step of selecting at least one degradation model from multiple preset degradation models based on the first fusion health indicator is executed; for the first fusion health indicator belonging to the range from the third boundary point to the second time point, the step of selecting at least one degradation model from multiple preset degradation models based on the first fusion health indicator is executed.
[0111] Optionally, when the first trend information includes three different trend information, the dividing point between these three different trend information can be determined and used as the second dividing point and the third dividing point.
[0112] Optionally, since the first trend information is the trend information representing the change of the first fusion health index over time within the time range from the first time point to the second time point, after determining the second and third boundary points, the first fusion health index within the time range from the first time point to the second boundary point, the first fusion health index within the time range from the second boundary point to the third boundary point, and the first fusion health index within the time range from the third boundary point to the second time point can be determined.
[0113] Optionally, the trend information of the first integrated health indicator changing over time during the period from the first time point to the second boundary point (hereinafter referred to as the third sub-trend information), the trend information of the first integrated health indicator changing over time during the period from the second boundary point to the third boundary point (hereinafter referred to as the fourth sub-trend information), and the trend information of the first integrated health indicator changing over time during the period from the third boundary point to the second time point (hereinafter referred to as the fifth sub-trend information) can be determined.
[0114] Optionally, the same operation can then be performed on the third, fourth, and fifth sub-trend information respectively, i.e., selecting their respective corresponding degradation models.
[0115] For example, if the third sub-trend information indicates that the first fused health indicator shows a stable linear decline over time, a linear distribution model can be selected to predict the fault threshold within the time range from the first time point to the second boundary point. If the fourth sub-trend information indicates that the first fused health indicator accelerates its decline in the later stages, an exponential distribution model can be selected to predict the fault threshold within the time range from the second boundary point to the third boundary point. If the fifth sub-trend information indicates a nonlinear degradation trend caused by sudden load, a power-law distribution model can be selected to predict the fault threshold within the time range from the third boundary point to the second time point.
[0116] For example, a piecewise regression model can be used to match the first fused health index for fault threshold prediction, and the corresponding formula can be shown in the following formula (VII).
[0117] Formula (VII); in, It could be the second dividing point. It can be the third dividing point, and 0 to... During this period, a linear distribution trend was observed (for example, the first fusion health index showed a stable linear decline over time). to During this period, an exponential distribution trend was observed (for example, the first fusion health index declined more rapidly in the later stages). The current time point exhibits a power-law distribution trend (such as a nonlinear degradation trend caused by sudden load), and 0 can be taken as the first time point and the current time point as the second time point.
[0118] Alternatively, the dividing point can be determined by the following formulas (viii), (ix), and (x), such as the second dividing point. and the third dividing point .
[0119] Formula (8); Formula (IX); Formula (10); in, Indicates the cumulative sum; Indicates the lower cumulative sum; This represents the t-th observation (e.g., the first fused health indicator), where t is a positive integer; k represents the baseline mean, such as the rate of degradation of the first fusion health indicator in the current stable phase; k1 represents the allowable deviation, such as... / 2. When or If the threshold is exceeded, a change point is identified, and the time point corresponding to that change point is used as the dividing point.
[0120] In this embodiment, when the first trend information includes multiple different trend information, the boundary point between different trend information can be determined. Based on the boundary point at different time periods, and according to their respective trend information, an appropriate degradation model can be selected. This ensures that the selected degradation model threshold is closely related to the actual first fused health indicator, making the subsequently determined fault threshold more accurate and effective, thereby indirectly improving the accuracy of the prediction of the remaining lifespan of the solid-state drive.
[0121] Step S32: In response to selecting a degradation model, the first fused health index is input into the selected degradation model to predict and obtain the fault threshold; Optionally, when the number of degradation models selected based on the first fusion health indicator is one degradation model, the failure threshold can be obtained by directly predicting based on that degradation model.
[0122] Step S33: In response to selecting multiple degradation models, determine the information criteria information corresponding to the selected multiple degradation models, take the degradation model corresponding to the smallest information criteria information among the selected multiple degradation models as the target degradation model, and input the first fused health index into the target degradation model to predict and obtain the fault threshold.
[0123] Optionally, the information criteria include AIC (Akaike Information Criterion) and BIC (Bayesian Information Criterion). AIC can balance goodness of fit and parameter complexity, while BIC can emphasize model simplicity.
[0124] Optionally, when there are multiple degradation models to choose from, such as for the time node where the boundary point is located, two degradation models can be selected. In this case, the multiple degradation models can be filtered again to obtain the target degradation model, and the target degradation model can be used for prediction to obtain the fault threshold.
[0125] Optionally, the AIC and BIC corresponding to the selected multiple degradation models can be calculated using the following formulas (xi) and (xii), and used as information criterion information.
[0126] Formula AIC = 2K - 2ln(L) (XI); BIC = ln(n)K - 2ln(L) Formula (XII); Where K represents the number of parameters, L represents the likelihood function value, and n is the sample size (e.g., the number of first fusion health values).
[0127] Optionally, the information criteria information corresponding to each of the selected multiple degradation models can be calculated according to formulas (xi) and (xii), and the degradation model corresponding to the smallest information criteria information can be selected as the target degradation model, such as the degradation model with the smallest AIC and / or BIC.
[0128] Optionally, when the first fused health indicator is input into the target degradation model to predict the fault threshold, a graded early warning method can also be used to determine the fault threshold.
[0129] Optionally, the first fused health index from the first time point to the second time point can be abstracted into a finite discrete state (such as the stable period, transition period and exhaustion period), and the next finite discrete state is required to depend on the current finite discrete state. The corresponding transition probability can be reflected by the transition matrix, which can be the following formula (XIII).
[0130] Formula (XIII); Where P is the transition matrix; It can represent the probability of transitioning from state i to state j, where state i and state j can be different finite discrete states, and i and j are integers not less than 0.
[0131] Optionally, the transition probability of the first fusion health indicator having a state transition can be determined based on the transition matrix, and whether to reselect a new degradation model for fault threshold prediction can be determined based on the transition probability. For example, if the transition probability is greater than a preset probability threshold (e.g., 70%), a new degradation model can be reselected to predict the fault threshold.
[0132] Alternatively, the fault threshold can be expressed as different degradation models. The corresponding formula can be found in the following formula (XIV).
[0133] Formula (XIV); in, At the fault threshold time, This is the fault threshold.
[0134] In this embodiment, by selecting a degradation model based on the first fused health indicator, and when there are multiple selected degradation models, a target degradation model can be selected based on information criteria, and a fault threshold can be predicted based on the target degradation model. This ensures that the selected target degradation model threshold is closely related to the actual first fused health indicator, making the subsequently determined fault threshold more accurate and effective, thereby indirectly improving the accuracy of the prediction of the remaining lifespan of the solid-state drive.
[0135] Reference Figure 4 , Figure 4 This illustration shows a third flowchart of the solid-state drive (SSD) remaining lifespan prediction method provided in this application embodiment. Contents identical or similar to those in the above embodiments can be referred to the above description and will not be repeated hereafter. The SSD remaining lifespan prediction method includes at least one of steps S01-S02: Step S01: Use the preset first training dataset to train the preset convolutional autoencoder network until the first training cutoff condition is reached to obtain the pre-trained convolutional autoencoder network. The training data in the first training dataset are the second multidimensional health index and the second fusion health index corresponding to the second multidimensional health index. Optionally, the convolutional autoencoder network includes a convolutional encoder and a deconvolutional decoder.
[0136] Optionally, the convolutional encoder includes a convolutional layer, a first activation function, a dropout layer, a pooling layer, and a fully connected layer.
[0137] Optionally, the deconvolution decoder includes a deconvolution layer, a second activation function, and an upsampling layer.
[0138] Optionally, such as Figure 5 As shown, a convolutional autoencoder network can include two parts: an encoder (i.e., a convolutional encoder) and a decoder (i.e., a deconvolutional decoder). The encoder can include multiple convolutional layers and pooling layers, while the decoder includes upsampling layers and deconvolutional layers. Furthermore, the loss function MSE used by the convolutional autoencoder network can be represented in the form of reconstruction error.
[0139] Optionally, the function formula corresponding to the convolution encoder can be formula (XV).
[0140] Ec=pool( Formula (XV); Where X represents the input data, such as the second multidimensional health indicator, the first multidimensional health indicator, etc.; Ec represents the encoded features, such as encoded features; This represents the i-th convolutional kernel in the convolutional layer; This represents the i-th bias term; This represents the convolution operation; Indicates the activation function; The expression represents the summation operation; pool() represents the pooling operation.
[0141] Optionally, the function formula corresponding to the convolutional decoder can be formula (XVI).
[0142] =upsp( Formula (XVI); in, This indicates reconstructed data, such as the fourth integrated health indicator; This represents the i-th deconvolution kernel in the deconvolution layer; This represents the i-th bias term; The expression indicates a deconvolution operation; upsp() indicates an upsampling operation.
[0143] Optionally, a training dataset including the correspondence between the second multidimensional health indicator and the second fused health indicator can be constructed in advance and used as the first training dataset.
[0144] Optionally, both the second multidimensional health indicator and the second fused health indicator can be historical data. The second multidimensional health indicator can be a health indicator that includes multiple dimensions, and the second fused health indicator can be a one-dimensional health indicator obtained by fusing the second multidimensional health indicator.
[0145] For example, the first training dataset It can include multiple health indicators from different dimensions. =[ , ].
[0146] in, It can be a health indicator representing the total amount of data written; It can be a health indicator representing the daily drive write volume; It can be a health indicator representing the amplification factor. It can be a health indicator representing the amount of Nand writes; It can be a health indicator representing the amount of writes to the host; It can be a health indicator representing the number of P / E wear cycles; It can be a health indicator representing the percentage of lifespan already served; It can be a health indicator representing the number of bad blocks; It can be a health indicator representing the degree of wear and tear; It can be a health indicator representing the available remaining space; It can be a health indicator representing the count of uncorrectable errors; It can be a health indicator representing the write error count; It can be a health indicator representing the erase error count; It can be a health indicator representing a correctable error count; It can be a health indicator representing the count of cyclic redundancy check errors; It can be a health metric characterizing sequential read bandwidth performance; It can be a health metric characterizing random write bandwidth IOPS performance; It can be a health metric characterizing random read IOPS performance; It can be a health metric characterizing random write IOPS performance; It can be a health indicator that characterizes random read latency performance; It can be a health metric characterizing random write latency performance; It can be a health metric characterizing random read QoS performance; It is a health metric characterizing random write QoS performance.
[0147] Optionally, if the text length corresponding to each health indicator in the first training dataset is 1000, the specific network structure of the convolutional autoencoder network can be referred to Table 1 below.
[0148] Table 1:
[0149] Optionally, step S01, which involves training a pre-set convolutional autoencoder network using a pre-set first training dataset until a first training cutoff condition is met, to obtain a pre-trained convolutional autoencoder network, includes steps e10-e30.
[0150] Step e10: Input the second multidimensional health indicator into the preset convolutional autoencoder network, and extract data features from the second multidimensional health indicator according to the convolutional encoder to obtain the encoded features; Optionally, a second multidimensional health indicator can be selected from the first training dataset and input into a pre-defined convolutional autoencoder network for model training. The convolutional encoder can be used to extract data features from multiple health indicators of different dimensions, and the data features can be fused to obtain encoded features.
[0151] Optionally, step e10, which involves extracting data features from the second multidimensional health indicator based on the convolutional encoder to obtain encoded features, includes steps f10-f30.
[0152] Step f10: Perform convolution processing on each one-dimensional health indicator in the second multidimensional health indicator according to the convolution layer to obtain each first convolution result; Step f20: Activate each first convolution result according to the first activation function, randomly deactivate each first convolution result after activation according to the Dropout layer, and then downsample each first convolution result after random deactivation according to the pooling layer to obtain each first convolution result after downsampling. Step f30: Perform a full connection on the downsampled first convolution results based on the fully connected layer to obtain the encoded features.
[0153] Optionally, the second multidimensional health index input can be processed using the encoder formula to obtain the encoded features, and the operations of steps f10-f30 can be performed using the following formula (xvii).
[0154] SSD-HI Tr =FC(pool(Dropout( n=1,...,6 Formula (XVII); in, It can represent health indicators after CAE fusion, such as encoded features; FC(.) represents a fully connected operation; pool(.) represents a pooling operation; Dropout(.) represents a Dropout operation (such as random deactivation). (.) represents an activation function (such as the first activation function); This represents the summation operation; This represents the input data for the nth convolutional layer; This represents the convolution operation; This represents the i-th convolutional kernel in the n-th convolutional layer; This represents the i-th bias term in the n-th convolutional layer.
[0155] In this embodiment, the CAE network is trained using the first training dataset. During model training, the CAE network's convolutional encoder is used for encoding to obtain encoded features, thereby ensuring the effectiveness of the obtained encoded features and realizing the effective training of the CAE network model.
[0156] Step e20: The encoded features are reconstructed based on the deconvolution decoder to obtain the fourth fusion health index; Optionally, after encoding features using a convolutional encoder, these features can be input into a deconvolutional decoder for reconstruction, and the resulting reconstructed data can be used as a fourth fusion health indicator.
[0157] Optionally, step e20, which involves reconstructing the encoded features based on the deconvolution decoder to obtain the fourth fused health index, includes steps h10-h20.
[0158] Step h10: Upsample the encoded features based on the upsampling layer to obtain sampled data; Step h20: Perform deconvolution activation processing on the sampled data based on the deconvolution layer and the second activation function to obtain the fourth fused health index.
[0159] Optionally, in the deconvolution decoder, the encoded features can be processed according to formula (xvii). That is, the encoded features are used as input data of the deconvolution decoder. They can be upsampled in the upsampling layer to obtain sampled data, and then deconvolution processing is performed in the deconvolution layer. Since there are multiple deconvolution layers, the result of deconvolution processing of each deconvolution layer can be activated and activated using a pre-set second activation function. Finally, the fourth fused health index is output.
[0160] In this embodiment, the CAE network is trained using the first training dataset, and the model is trained based on the convolutional encoder and deconvolutional decoder, thereby ensuring the effective training of the CAE network.
[0161] Step e30: Calculate the loss function value of the fourth fusion health index and the second fusion health index according to the preset first loss function to obtain the first loss function value, and update the convolutional autoencoder network in reverse according to the first loss function value until the first training cutoff condition is reached to obtain the pre-trained convolutional autoencoder network.
[0162] Optionally, the first loss function can be represented in the form of reconstruction error, which can represent the Euclidean distance between the second and fourth fused health indicators.
[0163] Optionally, the formula for the first loss function can be shown in formula (XVIII).
[0164] Re= Formula (18); Where n represents the amount of data, This represents the i-th reconstructed data, such as the fourth integrated health indicator; Represents the i-th original data, such as the second fused health indicator; Re represents the reconstruction error, i.e., the value of the first loss function; i is a positive integer.
[0165] Optionally, the first training cutoff condition can be a pre-set training termination condition, such as the first loss function value being minimized, or the preset number of training iterations being reached.
[0166] Optionally, the convolutional autoencoder network can be trained iteratively multiple times using the first training dataset until the first training cutoff condition is met, resulting in a trained convolutional autoencoder network. A first test set can be constructed to test the trained convolutional autoencoder network. If the test meets the requirements, the trained convolutional autoencoder network can be used as a pre-trained convolutional autoencoder network.
[0167] In this embodiment, the CAE network is trained using the first training dataset. During model training, a convolutional encoder, a convolutional decoder, and a first loss function are combined to train the model until the first training cutoff condition is met, resulting in a pre-trained convolutional autoencoder network. This network facilitates subsequent prediction of the remaining lifespan of the solid-state drive (SSD) based on the pre-trained convolutional autoencoder network, thereby improving the accuracy of the predicted SSD lifespan.
[0168] Step S02: Use the preset second training dataset to train the preset gated attention unit network until the second training cutoff condition is met, and obtain the pre-trained gated attention unit network. The training data in the second training dataset includes at least one third fusion health indicator.
[0169] Optionally, the Gated Attention Unit (GAU) network can be a variation of the GRU network. The GAU seamlessly integrates the attention mechanism into the original GRU network structure, thereby focusing on the outputs of the reset and update gates. (See also...) Figure 6 As shown, a GRU network can include a reset gate and an update gate, with the input being the previous hidden state. and current input data For example, the first Hankel matrix or the second Hankel matrix, the output is In a GRU network, connection operations (C), product operations (M), and summation operations (A) can be performed, and two different activation functions, namely the Sigmoid function, can be used. and the Tanh function , and will and These are used as the outputs of the reset gate and the update gate, respectively.
[0170] Optionally, the function formula corresponding to the GRU network can be shown in formulas (19) to (22).
[0171] Formula (19); Formula (20); Formula (21); Formula (22); in, and These represent the outputs of the reset gate and the update gate, respectively. Represents the Sigmoid function; This represents the Tanh function; This represents the input data at the current moment; This indicates the current temporary hidden state; Indicates the current hidden state; , , , , , Both represent the weight matrix in the calculation process; , , Both represent bias matrices in the calculation process; ⊙ represents the dot product; This represents historical information and can include the previous hidden state. .
[0172] Optionally, the gating attention unit network includes a reset gate, an update gate, and an attention gate.
[0173] For reference Figure 7 As shown, the GAU network can include a reset gate, an update gate, and an attention gate, with the input being the previous hidden state. and current input data For example, the first Hankel matrix or the second Hankel matrix, the output is In a GAU network, connection operations (C), product operations (M), and summation operations (A) can be performed, and two different activation functions, namely the sigmoid function, can be used. and the Tanh function , and will and These are used as the outputs of the reset gate and the update gate, respectively.
[0174] Optionally, the function formula corresponding to the GAU network can be shown in formulas (23) to (29) below.
[0175] Formula (23); Formula (24); Formula (25); ( Formula (26); ( Formula (27); Formula (28); Formula (29); in, Indicates to and The level of attention; Indicates whether the output information is positive or negative; Indicates to and The overall attention distribution of the fusion , , , Both represent the weight matrix in the calculation process; , Both represent the bias matrix in the calculation process.
[0176] Optionally, in step S02, the step of training the preset gated attention network using the preset second training dataset until the second training cutoff condition is met, and obtaining the pre-trained gated attention network, includes steps i10-i70.
[0177] Step i10: Construct a second Hankel matrix including at least one third fusion health indicator, and input the second Hankel matrix and the historical information of the solid-state drive into a preset gated attention unit network, wherein the historical information includes the previous hidden state corresponding to the third fusion health indicator; Step i20: Using the reset gate, determine the output information of the reset gate based on the second Hankel matrix and historical information; Step i30: Using the update gate, determine the output information of the update gate based on the second Hankel matrix and historical information; Step i40: Determine the candidate hidden state based on the reset door output information, the updated door output information, and the historical information; Step i50: Using the attention gate, determine the attention distribution matrix based on the reset gate output information and the update gate output information; Step i60: Determine the sixth fusion health indicator based on the candidate hidden state, reset gate information, update gate information, and attention distribution matrix; Step i70: The gated attention network is updated in reverse according to the preset second loss function and the sixth fusion health index until the second training cutoff condition is met, and the pre-trained gated attention network is obtained.
[0178] Optionally, a Hankel matrix can be constructed based on the second training dataset to obtain a second Hankel matrix.
[0179] Optionally, if there are multiple one-dimensional fusion health indicators V, such as V= Then the first k elements in the first V can be used as the training set (i.e., the second training dataset), and the remaining elements can be used as the second test set, where n and k are positive integers greater than 1, and n is greater than k.
[0180] Optionally, if the second training dataset = Then, the second Hankel matrix constructed based on the second training dataset can be X= Where i represents the number of neurons in the input layer of the gated attention unit network, and the second Hankel matrix... It can be represented as = e is a positive integer.
[0181] Alternatively, the gated attention unit network can be used as a mapping function, taking the first i vectors of the second Hankel matrix X as its input and the last vector as its input. As its output, it is shown in the following formula (30).
[0182] Formula (30); Furthermore, a real-time recurrent learning algorithm can be used to train the gating attention unit network until the second training cutoff condition is met, thus obtaining a pre-trained gating attention unit network.
[0183] Optionally, the second training cutoff condition can be a pre-set training termination condition, such as the second loss function value being minimized, or the preset number of training iterations being reached.
[0184] Optionally, the gated attention network can be trained iteratively multiple times using the second training dataset until the second training cutoff condition is met, resulting in a trained gated attention network. A second test set can be constructed to test the trained gated attention network. If the test meets the requirements, the trained gated attention network can be used as a pre-trained convolutional autoencoder network.
[0185] In this embodiment, since the gated attention unit network includes a reset gate, an update gate, and an attention gate, the gated attention unit can be trained using the second training dataset and the second loss function until the second training cutoff condition is met, thus obtaining a pre-trained gated attention unit network. This facilitates subsequent prediction of the remaining lifespan of the solid-state drive based on the pre-trained gated attention unit network, thereby improving the accuracy of the predicted remaining lifespan of the solid-state drive.
[0186] Optionally, refer to the following: Figure 8 This is a schematic diagram of a memory 400 provided in an embodiment of this application. The memory 400 includes a main control chip 410 and a storage chip 420. The main control chip 410 and the storage chip 420 are electrically connected directly or indirectly to realize data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines.
[0187] The storage chip 420 stores a computer program that can be executed by the main control chip 410. The main control chip is used to read / write the data or computer program stored in the storage chip 420 and perform corresponding functions. For example, when the computer program stored in the storage chip 420 is executed by the main control chip 410, the solid-state drive remaining life prediction method disclosed in the above embodiments can be implemented. It should be understood that, Figure 8 The structure shown is only a schematic diagram of the memory 400. The memory 400 may also include components such as memory 400. Figure 8 The more or fewer components shown, or having the same Figure 8 The different configurations shown. Figure 8 The components shown can be implemented using hardware, software, or a combination thereof.
[0188] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the solid-state drive remaining life prediction method in the above embodiments.
[0189] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0190] The aforementioned computer-readable storage medium may be contained within a memory or may exist independently without being assembled into a memory.
[0191] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by the memory, cause the memory to: To obtain the first multidimensional health indicator of the lifespan of the associated solid-state drive, where the first multidimensional... Health indicators include multiple different dimensions of health indicators; A pre-trained convolutional autoencoder network is used to perform index fusion processing on the first multidimensional health index to obtain the first fused health index. The fault threshold corresponding to the first fusion health index is determined using a preset degradation model; Using a pre-trained gated attention unit network, the remaining lifespan of the solid-state drive is predicted based on the first fused health metric and the fault threshold.
[0192] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0193] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0194] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0195] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described solid-state drive (SSD) remaining lifespan prediction method, thereby solving the technical problem of how to improve the accuracy of SSD remaining lifespan prediction. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the SSD remaining lifespan prediction method provided in the above embodiments, and will not be repeated here.
[0196] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the solid-state drive remaining life prediction method as described above.
[0197] The computer program product provided in this application solves the technical problem of how to improve the accuracy of predicting the remaining lifespan of solid-state drives (SSDs). Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the SSD remaining lifespan prediction method provided in the above embodiments, and will not be repeated here.
[0198] The above are only some embodiments of this application and do not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A method for predicting the remaining lifespan of a solid-state drive, characterized in that, The method for predicting the remaining lifespan of a solid-state drive includes: Obtain the first multidimensional health indicator of the associated solid-state drive's lifespan, wherein the first multidimensional... Health indicators include multiple different dimensions of health indicators; The first multidimensional health indicator is fused using a pre-trained convolutional autoencoder network to obtain the first fused health indicator. The fault threshold corresponding to the first fused health index is determined using a preset degradation model; Using a pre-trained gated attention unit network, the remaining lifespan of the solid-state drive is predicted based on the first fused health indicator and the fault threshold.
2. The method for predicting the remaining lifespan of a solid-state drive as described in claim 1, characterized in that, The step of determining the fault threshold corresponding to the first fused health index using a preset degradation model includes: Based on the first fused health index, at least one degradation model is selected from multiple preset degradation models; In response to selecting a degradation model, the first fused health index is input into the selected degradation model to predict and obtain the fault threshold; In response to selecting multiple degradation models, the information criteria information corresponding to the selected multiple degradation models is determined. The degradation model corresponding to the smallest information criteria information among the selected multiple degradation models is taken as the target degradation model, and the first fused health index is input into the target degradation model to predict and obtain the fault threshold.
3. The method for predicting the remaining lifespan of a solid-state drive as described in claim 2, characterized in that, The step of selecting at least one degradation model from multiple preset degradation models based on the first fused health index includes: Determine the trend chart of the first integrated health indicator, which includes the first time point to the second time point, where the first time point is shorter than the second time point; Based on the first trend information corresponding to the first fused health indicator represented by the trend chart, at least one degradation model is selected from multiple preset degradation models.
4. The method for predicting the remaining lifespan of a solid-state drive as described in claim 3, characterized in that, The degradation models include linear distribution models, exponential distribution models, and power-law distribution models. The step of selecting at least one degradation model from multiple preset degradation models based on the first trend information corresponding to the first fused health indicator represented by the trend chart includes at least one of the following: In response to the first trend information, and to characterize the second trend information that the first fused health indicator shows a stable linear decline over time, the linear distribution model is selected from the linear distribution model, the exponential distribution model, and the power law distribution model. In response to the first trend information, to characterize the third trend information of the accelerated decline of the first fused health indicator in the later stage, the exponential distribution model is selected from the linear distribution model, exponential distribution model and power law distribution model. In response to the first trend information, a fourth trend information is provided to characterize the nonlinear degradation trend caused by the sudden load, in which the power law distribution model is selected from the linear distribution model, the exponential distribution model, and the power law distribution model. In response to the first trend information including at least two of the first trend information, the second trend information, and the third trend information, at least two degenerate models are selected from the linear distribution model, the exponential distribution model, and the power law distribution model.
5. The method for predicting the remaining lifespan of a solid-state drive as described in claim 4, characterized in that, The step of selecting at least two degenerate models from the linear distribution model, exponential distribution model, and power law distribution model in response to the first trend information including at least two of the first trend information, second trend information, and third trend information includes: In response to the fact that the first trend information includes at least two of the first trend information, the second trend information, and the third trend information, the boundary point between different trend information in the trend graph is determined; In response to the boundary point including a first boundary point, for a first fusion health indicator belonging to the range from the first time point to the first boundary point, the step of selecting at least one degradation model from a plurality of preset degradation models based on the first fusion health indicator is executed; for a first fusion health indicator belonging to the range from the first boundary point to the second time point, the step of selecting at least one degradation model from a plurality of preset degradation models based on the first fusion health indicator is executed. In response to the boundary point including two different second boundary points and a third boundary point, for the first fusion health indicator belonging to the range from the first time point to the second boundary point, the step of selecting at least one degradation model from multiple preset degradation models based on the first fusion health indicator is executed; for the first fusion health indicator belonging to the range from the second boundary point to the third boundary point, the step of selecting at least one degradation model from multiple preset degradation models based on the first fusion health indicator is executed; for the first fusion health indicator belonging to the range from the third boundary point to the second time point, the step of selecting at least one degradation model from multiple preset degradation models based on the first fusion health indicator is executed.
6. The method for predicting the remaining lifespan of a solid-state drive as described in claim 1, characterized in that, The step of predicting the remaining lifespan of the solid-state drive using a pre-trained gated attention unit network based on the first fused health metric and the fault threshold includes: Construct a first Hankel matrix that includes the first fusion health metric for the first time period; The first Hankel matrix is input into a pre-trained gated attention network for prediction, and the first prediction result is output. In response to the first prediction result being less than the fault threshold, the first prediction result is used as the first fusion health indicator for the next time step and updated to the first Hankel matrix. Based on the updated first Hankel matrix, the step of inputting the first Hankel matrix into the pre-trained gated attention network for prediction is performed until the latest obtained first prediction result is detected to be greater than or equal to the fault threshold. Based on the latest first prediction result, the number of predictions to be made by the pre-trained gating attention network is determined, and the sampling time corresponding to the first fused health indicator is determined, wherein the sampling time includes the sampling duration and sampling interval time characterizing the sampling of the first multidimensional health indicator. The sampling time is updated based on the number of predictions to obtain the remaining lifespan of the solid-state drive.
7. The method for predicting the remaining lifespan of a solid-state drive as described in claim 1, characterized in that, The method for predicting the remaining lifespan of a solid-state drive further includes at least one of the following: The pre-trained convolutional autoencoder network is trained using a pre-set first training dataset until the first training cutoff condition is met, resulting in a pre-trained convolutional autoencoder network. The training data in the first training dataset consists of a second multidimensional health indicator and a second fusion health indicator corresponding to the second multidimensional health indicator. The pre-defined gated attention unit network is trained using a pre-defined second training dataset until the second training cutoff condition is met, resulting in a pre-trained gated attention unit network. The training data in the second training dataset includes at least one third fusion health metric.
8. The method for predicting the remaining lifespan of a solid-state drive as described in claim 7, characterized in that, The convolutional autoencoder network includes a convolutional encoder and a deconvolutional decoder. The step of training a pre-defined convolutional autoencoder network using a pre-defined first training dataset until a first training cutoff condition is met, to obtain a pre-trained convolutional autoencoder network, includes: The second multidimensional health indicator is input into a preset convolutional autoencoder network, and the data features of the second multidimensional health indicator are extracted according to the convolutional encoder to obtain the encoded features. The encoded features are reconstructed using the deconvolution decoder to obtain the fourth fused health index; The first loss function is used to calculate the loss function values of the fourth fusion health index and the second fusion health index, and the first loss function value is obtained. The convolutional autoencoder network is then updated in reverse according to the first loss function value until the first training cutoff condition is reached, thus obtaining the pre-trained convolutional autoencoder network.
9. The method for predicting the remaining lifespan of a solid-state drive as described in claim 8, characterized in that, The convolutional encoder includes a convolutional layer, a first activation function, a Dropout layer, a pooling layer, and a fully connected layer. The step of extracting data features from the second multidimensional health indicator based on the convolutional encoder to obtain encoded features includes: Based on the convolutional layer, each one-dimensional health indicator in the second multidimensional health indicator is convolved to obtain each first convolution result; The activation process is performed on each of the first convolution results according to the first activation function, and the first convolution results after activation are randomly deactivated according to the Dropout layer. Then, the first convolution results after random deactivation are downsampled according to the pooling layer to obtain the downsampled first convolution results. The fully connected layer is used to perform a full connection on each of the downsampled first convolution results to obtain the encoded features.
10. The method for predicting the remaining lifespan of a solid-state drive as described in claim 9, characterized in that, The deconvolutional decoder includes a deconvolutional layer, a second activation function, and an upsampling layer. The step of reconstructing the encoded features based on the deconvolution decoder to obtain the fourth fused health indicator includes: The encoded features are upsampled according to the upsampling layer to obtain sampled data; The sampled data is subjected to deconvolution activation processing based on the deconvolution layer and the second activation function to obtain the fourth fused health index.
11. The method for predicting the remaining lifespan of a solid-state drive as described in claim 7, characterized in that, The gated attention unit network includes a reset gate, an update gate, and an attention gate. The step of training a pre-defined gated attention unit network using a pre-defined second training dataset until a second training cutoff condition is met, to obtain a pre-trained gated attention unit network, includes: A second Hankel matrix is constructed, including at least one of the third fusion health indicators, and the second Hankel matrix and the historical information of the solid-state drive are input into a preset gated attention unit network, wherein the historical information includes the previous hidden state corresponding to the third fusion health indicator; Using the reset gate, the reset gate output information is determined based on the second Hankel matrix and the historical information; Using the update gate, the update gate output information is determined based on the second Hankel matrix and the historical information; The candidate hidden state is determined based on the reset gate output information, the update gate output information, and the historical information; Using the attention gate, an attention distribution matrix is determined based on the reset gate output information and the update gate output information; The sixth fusion health indicator is determined based on the candidate hidden state, the reset gate information, the update gate information, and the attention distribution matrix; The gated attention network is updated in reverse according to the preset second loss function and the sixth fusion health index until the second training cutoff condition is reached, thus obtaining the pre-trained gated attention network.
12. The method for predicting the remaining lifespan of a solid-state drive as described in any one of claims 1-11, characterized in that, The health metrics include at least one of the following: total written data, daily drive write volume, write amplification factor, Nand write volume, Host write volume, P / E cycles, percentage of lifetimes, number of bad blocks, wear index, available remaining space, uncorrectable error count, write error count, erase error count, correctable error count, cyclic redundancy check error count, sequential read / write bandwidth performance, random read / write IOPS performance, random read / write latency performance, and random read / write QoS performance.
13. A memory, characterized in that, The memory includes a main control chip and a storage chip. The storage chip stores a computer program, and the main control chip can execute the computer program to implement the solid-state drive remaining life prediction method according to any one of claims 1-12.
14. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the solid-state drive remaining life prediction method as described in any one of claims 1-12.