Hard disk failure prediction method, lightweight prediction model construction method, and electronic device
By deploying a lightweight prediction model on the hard drive backplane and using a convolutional long short-term memory network for hard drive failure prediction, the problems of slow fault response and strong dependence in existing technologies are solved, achieving more efficient hard drive failure prediction and independent health assessment, and improving the stability and intelligence of the system.
Patent Information
- Application Number
- CN202511243503.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-09-02
AI Technical Summary
Existing technologies for hard drive failure prediction rely on static thresholds and centralized processing, lacking edge autonomy and multi-parameter trend analysis capabilities. This results in slow fault response, untimely warnings, and a high risk of false alarms and missed alarms. It is also unable to effectively identify early hard drive anomalies and thermal interference, leading to strong system dependence and poor timeliness. In particular, predictions fail when edge devices are offline or central nodes fail.
A lightweight prediction model is deployed on the hard drive backplane. By collecting sensor datasets from multiple hard drives, a temporal-space structure is constructed after preprocessing. A lightweight convolutional long short-term memory network is used for multidimensional array modeling to output hard drive failure prediction results. It supports independent health assessment and trend modeling.
It significantly improves the ability to detect hard drive failures in advance and enhances local autonomy, supports more granular and higher-precision predictive maintenance, reduces the risk of data loss and maintenance costs, and improves system stability and intelligence.
Smart Images

Figure CN120803829B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of servers, in particular to a hard disk fault prediction method, a lightweight prediction model construction method and an electronic device. BACKGROUND
[0002] Hard disk health monitoring mechanisms such as self-monitoring, analysis and reporting technology generally rely on static thresholds and centralized processing, which are difficult to effectively identify the evolution trend of parameters and the coupling relationship between multiple parameters, react slowly to sudden failures, and are prone to false positives or false negatives under complex working conditions, and cannot capture early "soft anomalies" and systematic degradation characteristics such as thermal interference and aging between multiple hard disks. At the same time, alarms and judgments are heavily dependent on host systems or central servers, lack autonomous capabilities on the edge side, and cannot achieve independent health assessment, trend modeling and proactive warning at the hard disk backplane level, resulting in high delay in fault response, strong system dependency, and problems such as prediction failure, poor timeliness, low stability and high network pressure when edge devices are offline, bandwidth is limited or central nodes fail.
[0003] Therefore, in view of the shortcomings of the prior art, the present application provides a hard disk fault prediction method. SUMMARY
[0004] The present application provides a hard disk fault prediction method, a lightweight prediction model construction method and an electronic device to at least solve the problem of slow response to complex failures, delayed warning and false positives and false negatives in related technologies due to reliance on static thresholds and centralized processing, lack of edge autonomy and multi-parameter trend analysis capabilities.
[0005] The present application provides a hard disk fault prediction method, which is suitable for a hard disk backplane connected to multiple hard disks, and the method comprises: collecting a sensor data set of sensors corresponding to the multiple hard disks, and preprocessing the sensor data set; constructing a time-space structure, converting the preprocessed sensor data set into a multi-dimensional array; inputting the multi-dimensional array into a lightweight prediction model, obtaining a prediction state corresponding to the multiple hard disks through a convolutional long short-term memory network layer in the lightweight prediction model; processing the prediction state corresponding to the multiple hard disks through a prediction layer in the lightweight prediction model, and outputting a fault prediction result of the multiple hard disks.
[0006] The application provides a lightweight prediction model construction method, which comprises the following steps: obtaining an initial prediction model according to a training data set; replacing standard convolution in a convolutional long short-term memory network layer in the initial prediction model with deep separable convolution to obtain a first convolutional long short-term memory network layer; training the first convolutional long short-term memory network layer according to a preset gating mechanism and a loss function, removing output channels with gating variables less than a preset value and fine-tuning to obtain a first prediction model; determining a quantization strategy according to an application scenario; quantizing the first prediction model according to the quantization strategy to obtain a second prediction model; and evaluating the second prediction model, and obtaining a lightweight prediction model in response to the second prediction model passing the evaluation.
[0007] The application also provides an electronic device, comprising a memory for storing a computer program and a processor for executing the computer program to implement the steps of any of the hard disk failure prediction methods.
[0008] The application also provides a computer readable storage medium, which stores a computer program, wherein the computer program is executed by a processor to implement the steps of any of the hard disk failure prediction methods.
[0009] The application also provides a computer program product comprising a computer program, wherein the computer program is executed by a processor to implement the steps of any of the hard disk failure prediction methods.
[0010] According to the application, the sensor data sets of the plurality of hard disks are collected, the sensor data sets are preprocessed, the time-space structure is constructed, the preprocessed sensor data sets are converted into multi-dimensional arrays, the multi-dimensional arrays are input into the lightweight prediction model, the prediction states of the plurality of hard disks are obtained through the convolutional long short-term memory network layer in the lightweight prediction model, and the prediction states of the plurality of hard disks are processed through the prediction layer in the lightweight prediction model to output the failure prediction results of the plurality of hard disks. Therefore, the lightweight prediction model is deployed on the hard disk backboard, the time-space joint modeling of the multi-channel sensor data and the intelligent prediction of the hard disk health trend are realized, the advanced perception ability and the local autonomy level of the server system to the hard disk failure are significantly improved, the system can independently run under abnormal conditions such as main system failure or network interruption, more fine-grained and higher-precision predictive maintenance is supported, the risk of data loss and the operation and maintenance cost are effectively reduced, and the stability, the intelligent degree and the safety performance of the system are comprehensively improved. BRIEF DESCRIPTION OF DRAWINGS
[0011] In order to more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. Obviously, the drawings described below are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort based on these drawings.
[0012] Figure 1 A flowchart of a hard disk failure prediction method provided by the embodiments of the present application is shown in FIG. 1.
[0013] Figure 2 A system architecture diagram of a hard disk failure prediction method provided by the embodiments of the present application is shown in FIG. 2.
[0014] Figure 3 A flowchart of a prediction model data processing of a hard disk failure prediction method provided by the embodiments of the present application is shown in FIG. 3.
[0015] Figure 4 A structural block diagram of a hard disk failure prediction device provided by the embodiments of the present application is shown in FIG. 4.
[0016] Figure 5 An internal structure diagram of an electronic device provided by the embodiments of the present application is shown in FIG. 5. DETAILED DESCRIPTION
[0017] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort fall within the protection scope of the present application.
[0018] It should be noted that, in the description of the present application, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that the processes, methods, articles or devices comprising a series of elements not only include those elements, but also include other elements not explicitly listed, or further include the elements inherent to such processes, methods, articles or devices. The terms "first", "second" and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.
[0019] It should be noted that the terms "S1", "S2" and the like are only used for the purpose of describing the steps and do not particularly indicate the order or sequence, nor are they used to limit the present application, but are merely used for the convenience of describing the method of the present application, and cannot be understood as indicating the sequence of the steps. In addition, the technical solutions of various embodiments can be combined with each other, but it must be based on the realization of a person skilled in the art, when the combination of technical solutions appears contradictory or unachievable, it should be considered that the combination of technical solutions does not exist, nor is it within the protection scope required by the present application.
[0020] In order for those skilled in the art to better understand the present application, the present application will be further described in detail below in conjunction with the drawings and specific embodiments.
[0021] The embodiments of the present application provide a hard disk failure prediction method, which is suitable for a hard disk backplane, the hard disk backplane is connected with a plurality of hard disks, and the method is described in detail in combination with an execution flow of the hard disk failure prediction method.
[0022] Here, the hard disk backplane is a key component for connecting a plurality of hard disks in a server or a storage device, often integrating SATA (Serial Advanced Technology Attachment), SAS (Serial Attached SCSI), PCIe (Peripheral Component Interconnect Express) and the like interfaces at the same time, and communicating with a BMC (Baseboard Management Controller) through an I2C (Inter-Integrated Circuit) / SMBus (System Management Bus) line to realize hard disk temperature monitoring, state indication and other management functions.
[0023] Specifically, the hardware of the hard disk backplane can include a plurality of hard disk interfaces (such as SATA / SAS / PCIe), a plurality of sensor modules (temperature, current, voltage, vibration, etc.), a lightweight prediction model, a local storage unit (storing recent sensor data and model parameters), a communication interface (I2C / SMBus, communicating with the mainboard BMC), an indication / alarm unit (LED lamp, buzzer, etc.) and a power management and protection circuit.
[0024] S101: Collecting a sensor data set of a plurality of hard disks corresponding to the sensors, and preprocessing the sensor data set.
[0025] Here, the sensors can include temperature sensors, Hall current sensors, voltage sampling points, and vibration MEMS (Micro-Electro-Mechanical Systems) sensors. Among them, the temperature sensor can be selected from NTC (Negative Temperature Coefficient) thermosensitive or digital thermometer (such as TMP75).
[0026] Each hard disk corresponds to one or more sensors for collecting different data.
[0027] The preprocessed sensor data set includes different sensor data corresponding to a plurality of hard disks.
[0028] The preprocessing can include denoising, missing compensation, outlier removal, normalization, and the like.
[0029] Specifically, the collection module accesses the MCU module through an ADC (Analog-to-Digital Converter) or I2C / SPI (Serial Peripheral Interface) mode.
[0030] In one embodiment, the sampling frequency can also be determined according to user configuration, for example, 1Hz~0.1Hz.
[0031] S102: Construct a time-space structure to convert the preprocessed sensor data set into a multi-dimensional array.
[0032] Here, the spatial structure can be a spatial layout corresponding to the physical layout of the hard disk backboard and the plurality of hard disks.
[0033] Here, the time structure can be a sliding window determined according to user requirements.
[0034] Each row in the multi-dimensional array represents a time point, and each column represents a certain sensor value of a hard disk.
[0035] Specifically, the sensor data set can be preprocessed by the MCU.
[0036] S103: Input the multi-dimensional array into a lightweight prediction model to obtain the predicted state of the plurality of hard disks through the convolutional long short-term memory network layer in the lightweight prediction model.
[0037] Here, ConvLSTM (Convolutional Long Short-Term Memory Network) is a deep learning model combining convolutional neural networks and long short-term memory networks.
[0038] wherein the hidden state represents a state calculated by the ConvLSTM, is an abstract representation formed after the network understands, filters and compresses all historical frames / snapshots (from t=0 to t) up to the current time t, captures the spatio-temporal dependence (such as time series patterns and spatial correlations), and is used for subsequent prediction or passed to the next time step.
[0039] Here, light weight refers to a process of reducing the size of a deep learning model, reducing its computational complexity, reducing memory occupation and energy consumption, so as to make it more suitable for deployment and operation on resource-constrained devices, without losing the performance (such as accuracy) of the model as much as possible.
[0040] wherein the light weight method can include model pruning, knowledge distillation, parameter quantization, light weight network design, low rank decomposition, and model sharing and compression.
[0041] In one embodiment, the light weight prediction model supports updating the model weights or parameters through the I²C channel.
[0042] S104: Process the prediction states corresponding to the plurality of hard disks through the prediction layer in the light weight prediction model, and output the failure prediction results of the plurality of hard disks.
[0043] Here, the failure prediction result can include the next time prediction value and / or health score of the plurality of hard disks.
[0044] In one embodiment, it can also include a gated heat map and channel contribution, combined with the failure prediction result, to locate the abnormal hard disk.
[0045] Specifically, the prediction model outputs the health score and trend label of the plurality of hard disks every preset time length (such as ten minutes), triggers the prompt of the indication unit according to the health score (MCU controls the indication unit to display green / orange / red corresponding to the health, warning, and failure states), and reports the health state, current score, and abnormal log to the server motherboard through the SMBus.
[0046] In one embodiment, the hard disk backplane can communicate with the server motherboard or BMC system in real time through the I2C, SMBus, and the like, to realize the uploading of the backplane prediction result, the synchronization of abnormal record, the management of state information, and the like. Support multi-backplane device address configuration and dynamic recognition mechanism, suitable for batch deployment in high-density server scenarios. Among them, it can support the realization of multi-backplane communication address recognition in the EEPROM address mapping mode.
[0047] It should be noted that the application realizes the spatio-temporal joint modeling of multi-channel sensor data and the intelligent prediction of hard disk health trends by deploying a lightweight prediction model on the hard disk backplane, significantly improves the early perception ability and local autonomy level of the server system to hard disk failure, can not only run independently in abnormal conditions such as main system failure or network interruption, but also supports more fine-grained and high-precision predictive maintenance, effectively reduces the risk of data loss and operation and maintenance cost, and comprehensively improves the stability, intelligent degree and security performance of the system.
[0048] In some embodiments, a plurality of sensor data sets corresponding to sensors of a plurality of hard disks are collected, and the sensor data sets are preprocessed, including:
[0049] According to a preset data range, an abnormal value in the sensor data set is determined, the abnormal value is converted into a missing value, and a first sensor data set is obtained;
[0050] According to the length of the missing segment, the missing value in the first sensor data set is processed, and a second sensor data set is obtained;
[0051] The second sensor data set is denoised by three-point median filtering combined with exponential moving average, and a third sensor data set is obtained;
[0052] The third sensor data set is normalized to obtain a preprocessed sensor data set.
[0053] Here, the three-point median filtering is a nonlinear digital filtering technique used to filter out sharp, bursty impulse noise and outliers in signals.
[0054] Here, the exponential moving average is a data processing method commonly used for time series smoothing, which performs weighted averaging on historical data, so that the weight of recent data is larger and the weight of long-term data is exponentially decayed.
[0055] For example, the preset data range can include ranges corresponding to multiple data, such as temperature range and current range, etc. The temperature range can be -40℃-125℃.
[0056] Wherein, the normalization can include standardization z=(x-μ) / σ or robust standardization.
[0057] Specifically, the three-point median filtering is performed on each point in the second sensor data set to obtain a median result, and the median result is taken as an input to perform exponential moving average to obtain a denoising output value.
[0058] In this way, the original data is converted into high-quality and high-credibility analysis base data.
[0059] In some embodiments, the missing values in the first sensor data set are processed according to the length of the missing segment to obtain a second sensor data set, including:
[0060] In response to the length of the missing segment being less than or equal to a preset length, the missing values in the missing segment are filled by linear interpolation.
[0061] In response to the length of the missing segment being greater than the preset length, the mask of the missing values in the missing segment is set to zero.
[0062] The preset length can be three consecutive points, and less than or equal to three consecutive points is classified as a short missing, and greater than three consecutive points is classified as a long missing. Specifically, for a short missing, the mask is updated to valid by linear interpolation filling according to the boundary valid point; for a long missing, the data is processed by zero, and the mask is kept invalid.
[0063] In one embodiment, when the missing segment is located at the beginning of the data sequence: find the first valid data point in the sequence, fill all the starting boundary missing points with the valid value, and if the entire sequence is invalid, keep it as 0; when the missing segment is located at the end of the data sequence: find the last valid data point in the sequence, fill all the ending boundary missing points with the valid value, and if the entire sequence is invalid, keep it as 0.
[0064] In this way, the data authenticity can be maintained, and false data trends introduced by excessive interpolation can be avoided.
[0065] In some embodiments, a time-space structure is constructed to convert the preprocessed sensor data set into a multi-dimensional array, including:
[0066] A two-dimensional space grid is constructed according to the layout of the hard disk backplane and the plurality of hard disks;
[0067] An index mapping table is created according to the two-dimensional space grid;
[0068] A vacancy mask is determined according to the layout of the plurality of hard disks;
[0069] The sensor data corresponding to a plurality of time points is obtained, the sensor data corresponding to the plurality of time points is mapped on the two-dimensional space grid according to the index mapping table and the vacancy mask, and an image sequence is obtained according to the time sequence.
[0070] The image sequence is intercepted by a sliding window to obtain a multi-dimensional array.
[0071] Here, the sliding window analyzes or extracts features of local data by moving a fixed length window frame by frame on a sequence.
[0072] The step size and window size of the sliding window are determined according to user requirements.
[0073] wherein the two-dimensional spatial grid can be a 3*4 grid.
[0074] Specifically, the multi-dimensional array can be [T, H, W, C], wherein C is the number of sensing channels (for example: temperature + current = 2), T is the length of the time window (for example: 72 time points), H and W are consistent with the two-dimensional spatial grid of the actual backboard physical layout.
[0075] Specifically, the plurality of hard disk slots are mapped to a two-dimensional grid (such as 3*4) consistent with the physical layout of the backboard, an index mapping table and a vacancy mask are generated, and then the sliding window is divided to obtain a multi-dimensional array.
[0076] For example, assuming that sampling is performed every 5 minutes, the last 6 hours (72 points) are used to form an input multi-dimensional array, the backboard has a layout of three rows and four columns, and temperature and current information are collected, then the multi-dimensional array is [72, 3, 4, 2].
[0077] In this way, the spatial and temporal dependencies of the sensing data can be preserved.
[0078] In some specific embodiments, the predicted state corresponding to the plurality of hard disks is obtained by a lightweight convolutional long short-term memory network layer in the prediction model, including:
[0079] The multi-dimensional array, the hidden state of the previous time step, and the cell state of the previous time step are taken as inputs, and a first depth separable convolution and a second depth separable convolution are performed to obtain an input gate and a forget gate;
[0080] The multi-dimensional array, the hidden state of the previous time step, the input gate, and the forget gate are taken as inputs, and a third depth separable convolution is performed to calculate a candidate cell state and an updated cell state;
[0081] The multi-dimensional array, the hidden state of the previous time step, and the updated cell state are taken as inputs, and a fourth depth separable convolution is performed to obtain an output gate;
[0082] The updated cell state and the output gate are taken as inputs, and an activation function is performed to obtain the predicted state corresponding to the plurality of hard disks.
[0083] The depth separable convolution ensures that the gated memory is recursively in time and can capture long-term / short-term trends (such as temperature rise slope and periodic vibration).
[0084] The number of input channels is determined by the number of sensing data of each hard disk, such as current and temperature, and the number of hidden channels is determined by the computing power.
[0085] For example, the lightweight convolutional long short-term memory network layer processing process is as follows:
[0086] ;
[0087] ;
[0088] ;
[0089] ;
[0090] ;
[0091] wherein, in the above formula, σ represents a sigmoid activation function, tanh(·) represents a hyperbolic tangent activation function, X t represents an input of a current time step, H t-1 represents a hidden state of a previous time step, C t-1 represents a memory cell of the previous time step, f t , i t , C t , o t respectively represent a forget gate, an input gate, a memory cell and an output gate, represents a Hadamard product.
[0092] wherein, in the gating, W is a local convolution kernel (such as 3x3), H t and C t are spatial feature maps, and have sharing and receptive field expansion with adjacent hard disks. The processing process of the convolutional long short-term memory network layer includes local neighborhood (3x3 / 5x5 / empty convolution), and can encode spatial dependencies such as thermal coupling and current linkage between hard disks. Meanwhile, the memory cell C t accumulates across time, and the forget gate f t and the input gate i t adaptively retain long-term information or introduce new evidence, and can capture slow temperature rise, periodic fluctuations or sudden jumps.
[0093] For example, when a hard disk in (2, 3) is identified, and the temperature slowly rises due to blocked air ducts, the gate corresponding to the current hard disk will spread the abnormal signal to the neighborhood gate, prompting a potential "heat island".
[0094] In this way, the time sequence change trend of the hard disk running state and the spatial correlation characteristics between different hard disks can be effectively identified, and the fault prediction accuracy and advance amount are significantly improved.
[0095] In some embodiments, the convolutional long short-term memory network layer includes a first convolutional long short-term memory network layer and a second convolutional long short-term memory network layer, and the plurality of hard disks corresponding predicted states are obtained by the convolutional long short-term memory network layer in the lightweight prediction model, and the method further includes:
[0096] According to the multi-dimensional array, the first hidden state corresponding to the plurality of hard disks is obtained by the first convolutional long short-term memory network layer;
[0097] According to the first hidden state, the predicted state corresponding to the plurality of hard disks is obtained by the second convolutional long short-term memory network layer.
[0098] Among them, the first convolutional long short-term memory network layer, the second convolutional long short-term memory network layer and the prediction layer are included in the prediction model.
[0099] In this way, the receptive field can be expanded by stacking the convolutional long short-term memory network layer.
[0100] In some embodiments, the predicted state corresponding to the plurality of hard disks is processed by the prediction layer in the lightweight prediction model, and the failure prediction result of the plurality of hard disks is output, including:
[0101] According to the predicted state corresponding to the plurality of hard disks, the historical data and the current data, the health score of the plurality of hard disks is obtained;
[0102] According to the health score of the plurality of hard disks, the failure prediction result of the plurality of hard disks is output.
[0103] Among them, the health score ranges from 0 to 100, and is classified according to the threshold, greater than or equal to 80 is classified as healthy, greater than or equal to 60 and less than 80 is classified as sub-healthy, greater than or equal to 40 and less than 60 is classified as early warning, and less than 40 is classified as failure.
[0104] Specifically, the variance of the prediction sequence is obtained according to the predicted state, the historical slope is obtained according to the historical data, the prediction error is obtained according to the current data, the abnormality degree of the hard disk is calculated according to the variance of the prediction sequence, the historical slope and the prediction error, and the health score is obtained by normalizing the abnormality degree.
[0105] In one embodiment, when the health score of the current hard disk is in the early warning state, the adjacent hard disk of the current hard disk is identified, and when it is judged that the abnormality of the current hard disk affects the health score of the adjacent hard disk, the health score of the current hard disk is reduced and enters the early warning state.
[0106] For example, assuming that the system configures 12 hard disks, the monitoring points include temperature and current, the sampling is performed once per minute, the inference process of the MCU is performed once every 10 minutes, the scores of all the hard disks are calculated, and it is detected that the temperature of the hard disks 3 and 7 suddenly rises and the current tends to be unstable, the system reduces the scores of the hard disks 3 and 7 to “sub-health”, the slot LEDs of the hard disks 3 and 7 change to orange, and the BMC records are uploaded for system operation and maintenance analysis.
[0107] In this way, the health score mechanism based on trend judgment is realized, early warning and hierarchical prompting are supported, an active and predictable maintenance path is provided for the server, and the system availability and fault defense capability are significantly improved.
[0108] In some specific embodiments, according to the predicted states, historical data and current data of the plurality of hard disks, an expression of the health scores of the plurality of hard disks is specifically as follows:
[0109] ;
[0110] wherein, is an absolute error between the predicted value and the actual value at time t, α is a weight coefficient, is a predicted value sequence for K future time steps, is a variance of the predicted sequence, β is a weight coefficient, is an actual observation value at a past k time point, is a slope obtained after linear fitting of the historical sequence, γ is a weight coefficient.
[0111] Specifically, wherein, is a prediction error term, α is a weight coefficient, and controls the importance of the term; is a prediction uncertainty term, β is a weight coefficient, and controls the influence of future prediction fluctuations; is a historical trend term, γ is a weight coefficient, and controls the influence of the trend term.
[0112] In some specific embodiments, the method further comprises:
[0113] determining a fault trend and an influence range according to the fault prediction results of the plurality of hard disks;
[0114] dynamically adjusting a control operation according to the fault trend and the influence range.
[0115] The control operation can include a fan, data migration, or a dispatching strategy.
[0116] In this way, predictive maintenance is realized.
[0117] Embodiments of the present application also provide a construction method of a lightweight prediction model, and the method is described in detail in combination with an execution process of the construction method of the lightweight prediction model.
[0118] In some embodiments, the method comprises:
[0119] obtaining an initial prediction model according to a training data set;
[0120] replacing standard convolution in a convolutional long short-term memory network layer in the initial prediction model with a depthwise separable convolution to obtain a first convolutional long short-term memory network layer;
[0121] training the first convolutional long short-term memory network layer according to a preset gating mechanism and a loss function, removing output channels with gating variables less than a preset value and fine-tuning to obtain a first prediction model;
[0122] determining a quantization strategy according to an application scenario;
[0123] quantizing the first prediction model according to the quantization strategy to obtain a second prediction model;
[0124] evaluating the second prediction model, and obtaining a lightweight prediction model in response to the second prediction model passing the evaluation.
[0125] Here, the depthwise separable convolution significantly reduces the amount of calculation and model parameters while maintaining similar performance of the standard convolution. The depthwise separable convolution decomposes the standard convolution into two steps, depthwise convolution and pointwise convolution.
[0126] Here, the gating mechanism is to add a trainable gating variable g (scalar) to each output channel in the first convolutional long short-term memory network layer.
[0127] Here, the loss function is the loss function of the original task.
[0128] Specifically, by training the first convolutional long short-term memory network layer according to the preset gating mechanism and the loss function, during the training process, the gating variables will be optimized, and most of the gating variables will tend to 0, indicating that the corresponding channel is not important. After the training is completed, all channels are sorted according to the absolute value of the gating variable, a clipping threshold or a target clipping rate (such as cutting off the smallest 30% of channels) is set, and the convolution kernel weight corresponding to the channel with a gating variable less than the threshold and the corresponding channel dimension in the next layer input are removed.
[0129] Here, the quantization strategy can include post-training quantization and quantization-aware training. Among them, the post-training quantization includes weight-only quantization and full integer quantization.
[0130] Among them, the application scenario can include model accuracy sensitivity, hardware acceleration demand, development time cost, data preparation difficulty, model structure complexity and output layer protection demand, etc.
[0131] Specifically, the parameters in the application scenario form a decision matrix, and the quantization strategy is determined according to the decision matrix.
[0132] Specifically, the quantization perception training includes: inserting a Fake Quantization node (simulating quantization error) in the pruned model, fine-tuning using original training data (or part), small learning rate, and after training, converting the FakeQuant node to a real int8 operation to generate a lightweight prediction model.
[0133] Among them, the evaluation can include evaluating the accuracy (such as Top-1 / Top-5 accuracy, mAP), parameter quantity, calculation quantity (FLOPs) and inference speed of the model.
[0134] In one embodiment, the prediction layer in the lightweight prediction model includes a prediction head, and different hard disks are predicted through the same prediction head.
[0135] In one embodiment, a lightweight channel attention (such as a Squeeze-and-Excitation module) or a spatial attention mechanism can be introduced, combined with a convolutional long short-term memory network layer, so that the prediction model can focus on more important features with fewer channels and parameters, thereby reducing the total parameter quantity while maintaining performance.
[0136] In this way, the lightweight prediction model can adapt to resource-constrained chip environments, support local storage, circular buffer, power-off protection and other mechanisms to ensure the continuity and safety of model inference, and balance accuracy and efficiency.
[0137] In one embodiment, Figure 2 The system architecture in the embodiment of the present application is shown in FIG. 1. Figure 2 As shown in FIG. 1, the system architecture in the present application includes: hard disk slot group HDD1~HDDn, sensor acquisition module, prediction module, local buffer storage, I2C communication interface, indicator light / buzzer control unit.
[0138] Specifically, the sensor acquisition module is used to acquire multi-point temperature, current and vibration data.
[0139] Specifically, the prediction module is used for data preprocessing, model inference and scoring, and risk level judgment.
[0140] Specifically, the local buffer storage is used for ring storage of N hours of original data.
[0141] Specifically, the I2C communication interface is used to connect the BMC.
[0142] In one embodiment, Figure 3 The flowchart of the prediction model data processing in the embodiment of the present application is shown in FIG. 2.Figure 3 As shown, the flow of the prediction model data processing in the present application includes: S1: sensor data sampling (multi-channel); S2: data normalization and windowing (time series formatting); S3: constructing a tensor (i.e., a multi-dimensional array) input lightweight prediction model; S4: model inference; S5: output score (health degree 0-100); S6: determine alarm level; S7: local control feedback and state reporting.
[0143] It should be understood that, although Figures 1-3 The steps in the flowchart of the method of the present application are shown in sequence according to the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other orders. Moreover, Figures 1-3 At least a part of the steps in the method of the present application can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily sequential, but can be executed in rotation or alternation with other steps or sub-steps or stages of other steps.
[0144] From the above description of the embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and the necessary general hardware platform, of course, it can also be realized by hardware, but in many cases the former is a better embodiment.
[0145] The embodiments of the present application also provide a hard disk failure prediction device, which is suitable for a hard disk backboard, the hard disk backboard is connected with a plurality of hard disks, and the device comprises: a first processing module 401, which is used to collect a sensor data set of sensors corresponding to the plurality of hard disks and pre-process the sensor data set; a second processing module 402, which is used to construct a time-space structure and convert the pre-processed sensor data set into a multi-dimensional array; a third processing module 403, which is used to input the multi-dimensional array into a lightweight prediction model and obtain a prediction state corresponding to the plurality of hard disks through a convolution long short-term memory network layer in the lightweight prediction model; and a fourth processing module 404, which is used to process the prediction state corresponding to the plurality of hard disks through a prediction layer in the lightweight prediction model and output a failure prediction result of the plurality of hard disks.
[0146] As a preferred implementation, in the embodiment of the present application, the first processing module 401 is specifically configured to: determine the abnormal values in the sensor data set according to a preset data range, convert the abnormal values into missing values, and obtain a first sensor data set; process the missing values in the first sensor data set according to the length of the missing segment, and obtain a second sensor data set; perform denoising processing on the second sensor data set by combining three-point median filtering with exponential moving average, and obtain a third sensor data set; and perform normalization on the third sensor data set, and obtain the preprocessed sensor data set.
[0147] As a preferred implementation, in the embodiment of the present application, the first processing module 401 is specifically further configured to: in response to the length of the missing segment being less than or equal to a preset length, fill the missing values in the missing segment by linear interpolation; and in response to the length of the missing segment being greater than the preset length, set the mask of the missing values in the missing segment to zero.
[0148] As a preferred implementation, in the embodiment of the present application, the second processing module 402 is specifically configured to: construct a two-dimensional space grid according to the layout of the hard disk backboard and the plurality of hard disks; create an index mapping table according to the two-dimensional space grid; determine a vacancy mask according to the layout of the plurality of hard disks; obtain sensor data corresponding to a plurality of time points, map the sensor data corresponding to the plurality of time points on the two-dimensional space grid according to the index mapping table and the vacancy mask, and arrange according to time sequence to obtain an image sequence; and obtain a multi-dimensional array by sliding window interception of the image sequence.
[0149] As a preferred implementation, in the embodiment of the present application, the third processing module 403 is specifically configured to: take the multi-dimensional array, the hidden state of the previous time step, and the cell state of the previous time step as inputs, obtain an input gate and a forget gate through a first depth separable convolution and a second depth separable convolution; take the multi-dimensional array, the hidden state of the previous time step, the input gate, and the forget gate as inputs, calculate a candidate cell state and an updated cell state through a third depth separable convolution; take the multi-dimensional array, the hidden state of the previous time step, and the updated cell state as inputs, obtain an output gate through a fourth depth separable convolution; and take the updated cell state and the output gate as inputs, obtain the predicted state corresponding to the plurality of hard disks through an activation function.
[0150] As a preferred implementation, in the embodiment of the present application, the convolutional long short-term memory network layer includes a first convolutional long short-term memory network layer and a second convolutional long short-term memory network layer, and the third processing module 403 is specifically further configured to: obtain the first hidden state corresponding to the plurality of hard disks through the first convolutional long short-term memory network layer according to the multi-dimensional array; and obtain the predicted state corresponding to the plurality of hard disks through the second convolutional long short-term memory network layer according to the first hidden state.
[0151] As a preferred implementation, in the embodiment of the present application, the fourth processing module 404 is specifically configured to: obtain health scores of the plurality of hard disks according to the predicted states, the historical data and the current data corresponding to the plurality of hard disks; and output the failure prediction results of the plurality of hard disks according to the health scores of the plurality of hard disks.
[0152] As a preferred implementation, in the embodiment of the present application, the fourth processing module is specifically further configured to:
[0153] ;
[0154] wherein, is the absolute error between the predicted value and the actual value at time t, α is a weight coefficient, is the predicted value sequence for the future K time steps, is the variance of the predicted sequence, β is a weight coefficient, is the actual observation value at the past k time points, is the slope obtained after linear fitting of the historical sequence, γ is a weight coefficient.
[0155] Embodiments of the present application also provide a hard disk failure prediction device, the device comprising: a first construction module configured to obtain an initial prediction model according to a training data set; a second construction module configured to replace a standard convolution in a convolutional long short-term memory network layer in the initial prediction model with a depth separable convolution to obtain a first convolutional long short-term memory network layer; a third construction module configured to train the first convolutional long short-term memory network layer according to a preset gating mechanism and a loss function, remove output channels with gating variables less than a preset value and fine-tune to obtain a first prediction model; a fourth construction module configured to determine a quantization strategy according to an application scenario; a fifth construction module configured to quantize the first prediction model according to the quantization strategy to obtain a second prediction model; and a sixth construction module configured to evaluate the second prediction model, and obtain a lightweight prediction model in response to the second prediction model passing the evaluation.
[0156] The features of the embodiments of the hard disk failure prediction device can be referred to the related descriptions of the embodiments of the hard disk failure prediction method, which will not be repeated here.
[0157] Embodiments of the present application also provide an electronic device, which can be a terminal, and the internal structure diagram thereof can be as shown in Figure 5As shown in the figure. The electronic device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium, an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The network interface of the electronic device is used to communicate with the external terminal through the network connection. The computer program is executed by the processor to implement a hard disk failure prediction method. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the electronic device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the electronic device, or an external keyboard, touchpad or mouse, etc.
[0158] Those skilled in the art can understand that, Figure 5 The skilled in the art can understand that,
[0159] In one embodiment, a computer device is provided, including a memory, a processor and a computer program stored on the memory and executable on the processor, and the processor implements the following steps when executing the computer program: S1: collecting a plurality of sensor data sets corresponding to the sensors of the hard disks, and preprocessing the sensor data sets; S2: constructing a time-space structure, and converting the preprocessed sensor data sets into a multi-dimensional array; S3: inputting the multi-dimensional array into a lightweight prediction model, and obtaining a plurality of predicted states corresponding to the hard disks through a convolutional long short-term memory network layer in the lightweight prediction model; S4: processing the plurality of predicted states corresponding to the hard disks through a prediction layer in the lightweight prediction model, and outputting a failure prediction result of the plurality of hard disks.
[0160] In one embodiment, the processor further implements the following steps when executing the computer program: determining an abnormal value in the sensor data set according to a preset data range, converting the abnormal value into a missing value, and obtaining a first sensor data set; processing the missing value in the first sensor data set according to the length of the missing segment, and obtaining a second sensor data set; performing denoising processing on the second sensor data set through three-point median filtering combined with exponential moving average, and obtaining a third sensor data set; normalizing the third sensor data set, and obtaining the preprocessed sensor data set.
[0161] In one embodiment, the processor, when executing the computer program, also implements the following steps: in response to the length of the missing segment being less than or equal to the preset length, filling the missing values in the missing segment by linear interpolation; and in response to the length of the missing segment being greater than the preset length, setting a mask of the missing values in the missing segment to zero.
[0162] In one embodiment, the processor, when executing the computer program, also implements the following steps: constructing a two-dimensional spatial grid according to the layout of the hard disk backplane and the plurality of hard disks; creating an index mapping table according to the two-dimensional spatial grid; determining a vacancy mask according to the layout of the plurality of hard disks; obtaining sensing data corresponding to a plurality of time points, mapping the sensing data corresponding to the plurality of time points on the two-dimensional spatial grid according to the index mapping table and the vacancy mask, and arranging the sensing data in time sequence to obtain an image sequence; and obtaining a multi-dimensional array by sliding window cutting of the image sequence.
[0163] In one embodiment, the processor, when executing the computer program, also implements the following steps: taking the multi-dimensional array, a hidden state of a previous time step, and a cell state of the previous time step as inputs, obtaining an input gate and a forget gate through a first depth separable convolution and a second depth separable convolution; taking the multi-dimensional array, the hidden state of the previous time step, the input gate, and the forget gate as inputs, calculating a candidate cell state and an updated cell state through a third depth separable convolution; taking the multi-dimensional array, the hidden state of the previous time step, and the updated cell state as inputs, obtaining an output gate through a fourth depth separable convolution; and taking the updated cell state and the output gate as inputs, obtaining a predicted state corresponding to the plurality of hard disks through an activation function.
[0164] In one embodiment, the processor, when executing the computer program, also implements the following steps: obtaining a first hidden state corresponding to the plurality of hard disks through a first convolutional long short-term memory network layer according to the multi-dimensional array; and obtaining a predicted state corresponding to the plurality of hard disks through a second convolutional long short-term memory network layer according to the first hidden state.
[0165] In one embodiment, the processor, when executing the computer program, also implements the following steps: obtaining a health score of the plurality of hard disks according to the predicted state corresponding to the plurality of hard disks, historical data, and current data; and outputting a failure prediction result of the plurality of hard disks according to the health score of the plurality of hard disks.
[0166] In one embodiment, the processor, when executing the computer program, also implements the following steps:
[0167] ;
[0168] wherein, is an absolute error between a predicted value and an actual value at time t, and α is a weight coefficient, is a predicted value sequence for K future time steps, wherein, for predicting the variance of the sequence, β is a weight coefficient, is an actual observation value at a past k time point, is a slope obtained after linear fitting of the historical sequence, and γ is a weight coefficient.
[0169] In an embodiment, a computer device is also provided, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor implements the following steps when executing the computer program: S1: obtaining an initial prediction model according to a training data set; S2: replacing a standard convolution in a convolutional long short-term memory network layer in the initial prediction model with a depth separable convolution to obtain a first convolutional long short-term memory network layer; S3: training the first convolutional long short-term memory network layer according to a preset gating mechanism and a loss function, removing an output channel with a gating variable less than a preset value and fine-tuning to obtain a first prediction model; S4: determining a quantization strategy according to an application scenario; S5: quantizing the first prediction model according to the quantization strategy to obtain a second prediction model; and S6: evaluating the second prediction model, and obtaining a lightweight prediction model in response to the second prediction model passing the evaluation.
[0170] In an embodiment, a computer readable storage medium is provided, having a computer program stored thereon, and the computer program is executable by a processor to implement the following steps: S1: collecting a sensor data set of sensors corresponding to a plurality of hard disks, and preprocessing the sensor data set; S2: constructing a time-space structure, and converting the preprocessed sensor data set into a multi-dimensional array; S3: inputting the multi-dimensional array into a lightweight prediction model, and obtaining a predicted state of the plurality of hard disks through a convolutional long short-term memory network layer in the lightweight prediction model; and S4: processing the predicted state of the plurality of hard disks through a prediction layer in the lightweight prediction model, and outputting a failure prediction result of the plurality of hard disks.
[0171] In an embodiment, the computer program is executable by the processor to further implement the following steps: determining an outlier in the sensor data set according to a preset data range, and converting the outlier into a missing value to obtain a first sensor data set; processing the missing value in the first sensor data set according to a length of a missing segment to obtain a second sensor data set; performing denoising processing on the second sensor data set through three-point median filtering combined with exponential moving average to obtain a third sensor data set; and normalizing the third sensor data set to obtain the preprocessed sensor data set.
[0172] In an embodiment, the computer program is executable by the processor to further implement the following steps: in response to the length of the missing segment being less than or equal to a preset length, filling the missing value in the missing segment through linear interpolation; and in response to the length of the missing segment being greater than the preset length, setting a mask of the missing value in the missing segment to zero.
[0173] In one embodiment, the computer program is further implemented to perform the following steps when executed by the processor: constructing a two-dimensional spatial grid according to the layout of the hard disk backplane and the plurality of hard disks; creating an index mapping table according to the two-dimensional spatial grid; determining a vacancy mask according to the layout of the plurality of hard disks; obtaining sensing data corresponding to a plurality of time points, mapping the sensing data corresponding to the plurality of time points on the two-dimensional spatial grid according to the index mapping table and the vacancy mask, and obtaining an image sequence according to a time sequence arrangement; and obtaining a multi-dimensional array by intercepting the image sequence through a sliding window.
[0174] In one embodiment, the computer program is further implemented to perform the following steps when executed by the processor: taking the multi-dimensional array, a hidden state of a previous time step, and a cell state of the previous time step as inputs, obtaining an input gate and a forget gate through a first deep separable convolution and a second deep separable convolution; taking the multi-dimensional array, the hidden state of the previous time step, the input gate, and the forget gate as inputs, calculating a candidate cell state and an updated cell state through a third deep separable convolution; taking the multi-dimensional array, the hidden state of the previous time step, and the updated cell state as inputs, obtaining an output gate through a fourth deep separable convolution; and taking the updated cell state and the output gate as inputs, obtaining a predicted state corresponding to the plurality of hard disks through an activation function.
[0175] In one embodiment, the computer program is further implemented to perform the following steps when executed by the processor: obtaining a first hidden state corresponding to the plurality of hard disks through a first convolutional long short-term memory network layer according to the multi-dimensional array; and obtaining a predicted state corresponding to the plurality of hard disks through a second convolutional long short-term memory network layer according to the first hidden state.
[0176] In one embodiment, the computer program is further implemented to perform the following steps when executed by the processor: obtaining a health score of the plurality of hard disks according to the predicted state corresponding to the plurality of hard disks, historical data, and current data; and outputting a failure prediction result of the plurality of hard disks according to the health score of the plurality of hard disks.
[0177] In one embodiment, the computer program is further implemented to perform the following steps when executed by the processor:
[0178] ;
[0179] wherein, is an absolute error between a predicted value and an actual value at a time t, a is a weight coefficient, is a predicted value sequence for K future time steps, is a variance of the predicted sequence, β is a weight coefficient, is an actual observation value at a past k time point, is a slope obtained after linear fitting of a historical sequence, γ is a weight coefficient.
[0180] In one embodiment, a computer readable storage medium is also provided, and the computer program is stored on the computer readable storage medium, and the computer program is executed by a processor to implement the following steps: S1: obtaining an initial prediction model according to a training data set; S2: replacing a standard convolution in a convolutional long short-term memory network layer in the initial prediction model with a depth separable convolution to obtain a first convolutional long short-term memory network layer; S3: training the first convolutional long short-term memory network layer according to a preset gating mechanism and a loss function, removing an output channel with a gating variable less than a preset value and fine-tuning to obtain a first prediction model; S4: determining a quantization strategy according to an application scenario; S5: quantizing the first prediction model according to the quantization strategy to obtain a second prediction model; and S6: evaluating the second prediction model, and obtaining a lightweight prediction model in response to the second prediction model passing the evaluation.
[0181] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware, and the computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, the computer program can include the processes of the above-mentioned embodiments. Any reference to memory, storage, database or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memory. The non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. The volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct RAM bus dynamic RAM (DRDRAM) and memory bus dynamic RAM (RDRAM) and the like.
[0182] Any combination of the technical features in the above embodiments can be made, and in order to make the description concise, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist, it should be considered as the scope of the present disclosure.
[0183] The above embodiments only express several implementation ways of the present application, and the description is specific and detailed, but it should not be understood as a limitation to the scope of the patent. It should be pointed out that, for ordinary skilled in the art, without departing from the concept of the present application, several modifications and improvements can be made, which are all within the scope of protection of the present application.
Claims
1. A hard disk failure prediction method, characterized in that, Applicable to a hard drive backplane, wherein the hard drive backplane connects to multiple hard drives, the method includes: Collect sensor datasets corresponding to multiple hard drives, and preprocess the sensor datasets. Construct a time-space structure to convert the preprocessed sensor dataset into a multidimensional array; The multidimensional array is input into the lightweight prediction model, and the predicted state corresponding to the multiple hard disks is obtained through the convolutional long short-term memory network layer in the lightweight prediction model. The prediction layer in the lightweight prediction model processes the prediction status of multiple hard drives and outputs the failure prediction results of the multiple hard drives. The construction of the time-space structure, which converts the preprocessed sensor dataset into a multidimensional array, includes: Based on the layout of the hard drive backplane and the plurality of hard drives, a two-dimensional spatial grid is constructed; Create an index mapping table based on the two-dimensional spatial grid; Determine the empty space mask based on the layout of the multiple hard drives; Acquire sensor data corresponding to multiple time points, map the sensor data corresponding to the multiple time points onto the two-dimensional spatial grid according to the index mapping table and the empty space mask, and arrange them in chronological order to obtain an image sequence; The multidimensional array is obtained by cropping the image sequence using a sliding window; The lightweight prediction model construction methods include: Based on the training dataset, an initial prediction model is obtained; The standard convolutions in the convolutional long short-term memory network layer of the initial prediction model are replaced with depthwise separable convolutions to obtain the first convolutional long short-term memory network layer. According to the preset gating mechanism and loss function, the first convolutional long short-term memory network layer is trained, the output channels with gating variables less than the preset value are removed and fine-tuned to obtain the first prediction model. The gating mechanism is to add a trainable gating variable to each output channel in the first convolutional long short-term memory network layer. Based on the application scenario, a quantization strategy is determined, wherein the quantization strategy includes post-training quantization and quantization-aware training, and the post-training quantization includes weight-only quantization and all-integer quantization. According to the quantization strategy, the first prediction model is quantized to obtain the second prediction model; The second prediction model is evaluated, and in response to the second prediction model passing the evaluation, the lightweight prediction model is obtained.
2. The hard disk failure prediction method according to claim 1, characterized in that, The process of collecting sensor datasets from sensors corresponding to multiple hard drives and preprocessing the sensor datasets includes: Based on a preset data range, outliers in the sensor dataset are determined, and the outliers are converted into missing values to obtain the first sensor dataset. Based on the length of the missing segment, the missing values in the first sensor dataset are processed to obtain the second sensor dataset. The second sensor dataset is denoised by combining three-point median filtering with exponential moving average to obtain the third sensor dataset. The third sensor dataset is normalized to obtain a preprocessed sensor dataset.
3. The hard disk failure prediction method according to claim 2, characterized in that, The step of processing the missing values in the first sensor dataset according to the length of the missing segment to obtain the second sensor dataset includes: When the length of the missing segment is less than or equal to a preset length, the missing value in the missing segment is filled by linear interpolation. When the length of the missing segment is greater than the preset length, the mask of the missing value in the missing segment is set to zero.
4. The hard disk failure prediction method according to claim 1, characterized in that, The step of obtaining the predicted state corresponding to the multiple hard drives through the convolutional long short-term memory network layer in the lightweight prediction model includes: The multidimensional array, the hidden state of the previous time step, and the cell state of the previous time step are used as inputs. The input gate and the forget gate are obtained by passing the first depthwise separable convolution and the second depthwise separable convolution. Using the multidimensional array, the hidden state of the previous time step, the input gate, and the forget gate as input, the candidate cell state and the updated cell state are calculated through a third depthwise separable convolution. The multidimensional array, the hidden state of the previous time step, and the updated cell state are used as inputs, and the output gate is obtained through a fourth depthwise separable convolution. Using the updated cell state and the output gate as inputs, the predicted state corresponding to the multiple hard disks is obtained through an activation function.
5. The hard disk failure prediction method according to claim 1, characterized in that, The convolutional long short-term memory (LSTM) network layer includes a first LTM network layer and a second LTM network layer. The step of obtaining the predicted state corresponding to the multiple hard drives through the LTM network layer in the lightweight prediction model further includes: Based on the multidimensional array, the first hidden state corresponding to the multiple hard disks is obtained through the first convolutional long short-term memory network layer; Based on the first hidden state, the predicted state corresponding to the multiple hard disks is obtained through the second convolutional long short-term memory network layer.
6. The hard disk failure prediction method according to claim 1, characterized in that, The step of processing the predicted states of multiple hard drives through the prediction layer in the lightweight prediction model and outputting the fault prediction results of the multiple hard drives includes: Based on the predicted status, historical data, and current data of the multiple hard drives, a health score is obtained for each of the multiple hard drives. Based on the health scores of the multiple hard drives, output the failure prediction results of the multiple hard drives.
7. The hard disk failure prediction method according to claim 6, characterized in that, The expression for obtaining the health score of the multiple hard drives based on their predicted status, historical data, and current data is as follows: ; in, Let be the absolute error between the predicted and actual values at time t, and α be the weighting coefficient. Given a sequence of predicted values for the next K time steps, The variance of the predicted sequence is given by β, where β is the weighting coefficient. These are the actual observations from the past k time points. The slope is obtained by linearly fitting the historical sequence, and γ is the weighting coefficient.
8. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the hard disk failure prediction method as described in any one of claims 1 to 7 when executing the computer program.
Citation Information
Patent Citations
Hard disk fault prediction method and system and computer readable storage medium
CN116820888A
Solid state disk health monitoring method and system
CN120066903A