Solid state disk health state monitoring method and system
By performing write feature analysis and heat transfer feature identification on the monitoring data of solid-state drives, combined with the health status evaluation model, the problem of poor monitoring of solid-state drive health status in the existing technology is solved, effective fault prediction and data migration is achieved, and the service life of the hard disk is extended and data security is ensured.
Patent Information
- Application Number
- CN202510675517.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-23
AI Technical Summary
The prior art is difficult to effectively monitor and predict the health status of solid-state drives, resulting in the failure to detect potential failures in time, which may lead to data loss or hard disk damage.
By collecting monitoring data from solid-state hard disks, writing features are extracted and fluctuation analysis is performed to identify overwrite load characteristics; abnormal temperature characteristics are identified based on the correlation between heat transfer characteristics and overwrite load; abnormal temperature characteristics are input into the pre-constructed health status evaluation model, predict the health status of the hard disk and formulate a data migration mechanism.
Effectively identify and optimize storage management strategies, extend the service life of hard disk, avoid data loss and hard disk damage, improve equipment performance, optimize cooling design, and identify potential failure risks in advance.
Smart Images

Figure CN120179189A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of solid state drives, and more specifically, to a method and system for monitoring the health status of a solid state drive. Background Art
[0002] A solid state drive (SSD) is a computer storage device based on flash memory (usually NAND flash memory) as a storage medium, used to replace traditional hard disk drives (HDDs). Compared with traditional hard disks, solid state drives have no moving parts, so they have higher read and write speeds, lower power consumption, stronger durability, and lower noise. Solid state drives store data through storage chips (NAND flash memory), and these chips are composed of multiple storage units, each of which stores binary information (0 or 1).
[0003] The status monitoring of solid state drives is crucial, especially monitoring indicators such as their health status, temperature, write volume, etc. Monitoring the health status of solid state drives can effectively predict hard disk failures and avoid data loss. By monitoring information such as temperature, read and write errors, and bad blocks, users can detect potential problems of the hard disk, such as overheating and wear, early, and take measures such as migrating data and reducing the load to extend the service life of the hard disk and ensure data security. If the status of the solid state drive is not monitored, it may lead to the inability to detect the hard disk in time when a failure occurs, increasing the risk of data loss or hard disk damage. For example, overheating of the hard disk may lead to performance degradation and even permanent damage; frequent write operations will cause wear of NAND flash memory cells, and the opportunity to replace the hard disk may be missed when not monitored.
[0004] In view of the problems in the related art, no effective solution has been proposed yet. Summary of the Invention
[0005] In view of the problems in the related art, the present invention proposes a method and system for monitoring the health status of a solid state drive to overcome the above technical problems existing in the existing related art.
[0006] To this end, the specific technical solutions adopted by the present invention are as follows: According to one aspect of the present invention, a method for monitoring the health status of a solid state drive is provided. The method: Collect monitoring data of the solid state drive, extract write characteristics from the monitoring data and perform fluctuation analysis, and identify the overwriting load characteristics of the solid state drive according to the fluctuation analysis results of the write characteristics; Analyze the heat transfer characteristics built in the solid state drive based on the monitoring data, and analyze the correlation between the heat transfer characteristics and the overwriting load to obtain the abnormal temperature characteristics during overwriting load; Take the abnormal temperature feature as the input of a pre - constructed solid - state drive health status evaluation model, predict the health status of the solid - state drive through the health status evaluation model, and formulate a data migration mechanism based on the health status.
[0007] Preferably, collect the monitoring data of the solid - state drive, extract the write features from the monitoring data and perform fluctuation analysis. According to the results of the fluctuation analysis of the write features, identify the over - write load features of the solid - state drive, including: Use hard - disk monitoring technology to collect the monitoring data of the solid - state drive. The monitoring data includes write volume, write frequency, erase - write cycle and temperature parameters; Perform time - series mining analysis on the write volume and write frequency of the fixed hard disk respectively to obtain the write features of the solid - state drive, and perform Fourier transform analysis on the write features to obtain the results of the fluctuation analysis of the write features; Input the results of the fluctuation analysis of the write features into a predefined variational auto - encoder to learn the normal mode of the write features, and compare the difference mode generated by the reconstruction error of the variational auto - encoder with the normal mode to identify the over - write load features of the solid - state drive from the write features.
[0008] Preferably, perform time - series mining analysis on the write volume and write frequency of the fixed hard disk respectively to obtain the write features of the solid - state drive, and perform Fourier transform analysis on the write features to obtain the results of the fluctuation analysis of the write features, including: Use the tree - shaped index technology to obtain the similar time - series sequences of the write volume and write frequency of the fixed hard disk within a preset period, and assign weights to the similar time - series sequences respectively; Sort the similar time - series sequences according to the weights, and select the combination of write volume and write frequency within the preset sequence range to generate a write feature sequence; Perform Fourier transform on the write feature sequence, and identify the frequency components of the write feature sequence based on the frequency fluctuation trend of the write feature sequence; Establish a time - series data curve according to the frequency components of the write feature sequence, assign weights to the write feature sequence based on the degree of interest of the time - series data curve, and calculate the time - series points of the write feature sequence to obtain the results of the fluctuation analysis of the write feature time - series.
[0009] Preferably, perform Fourier transform on the write feature sequence, and identify the frequency components of the write feature sequence based on the frequency fluctuation trend of the write feature sequence, including: Store the continuous write volume and write frequency of the write feature sequence into a predefined channel - stream cache; Take each time - series point of several write features as amplitude components, and insert phase components between every two adjacent amplitude components in turn to obtain the composite values of the channel - stream cache; Use a number of composite values in the time domain in the channel stream buffer as the input of decimation in time, and use the method of decimation in time to transform the time-domain signal to obtain a number of composite values in the frequency domain; Calculate the modulus length of the composite value in the frequency domain, convert the amplitude value to a decibel value to obtain the amplitude values of several discrete frequency lines, and use the frequency corresponding to the discrete frequency line as the input in the next time period; Iteratively execute the calculation process of the composite value in the time domain until the processed composite values are summarized to obtain the frequency components written into the feature sequence.
[0010] Preferably, the expression of the timing point written into the feature sequence is: ; In the formula, D represents the timing point written into the feature sequence; A represents the similar timing sequence; WA represents the weight of the similar timing sequence; Q represents the feature sequence written; WQ represents the weight of the feature sequence written.
[0011] Preferably, based on the monitoring data analysis, the heat transfer characteristics built in the solid-state drive are analyzed, and the correlation between the heat transfer characteristics and the overwriting load is analyzed. The abnormal temperature characteristics during overwriting load include: Divide the built-in area of the solid-state drive, and calculate the temperature rise rate of each area in the solid-state drive under each time window based on the temperature parameters; Based on the temperature rise rate, establish a temperature gradient transfer mechanism, and combine the physical structure of the solid-state drive and the thermal conductivity of the materials in each area to extract the heat transfer characteristics built in the solid-state drive; Use the Pearson correlation coefficient to calculate the correlation between the heat transfer characteristics and the overwriting load, and compare the correlation with the preset threshold; If the correlation is greater than or equal to the prediction threshold, it means that the temperature rise rate increases abnormally during overwriting load, then use the load balancing algorithm to regulate the temperature parameters of each area built in the solid-state drive. Otherwise, it means that the temperature rise rate is not affected during overwriting load.
[0012] Preferably, the calculation formula of the Pearson correlation coefficient is: ; In the formula, r represents the Pearson correlation coefficient; represents the mean value of the heat transfer characteristics; represents the overwriting load; represents the i th heat transfer characteristic; represents the iAn overwriting load; Represents a weight value; n Represents the total number of features.
[0013] Preferably, a temperature gradient transfer mechanism is established based on the temperature rise rate, and combined with the physical structure of the solid-state drive and the thermal conductivity of the materials in each region, the heat transfer characteristics of the heat in the solid-state drive are extracted, including: Use finite element analysis software to establish a finite element model of the solid-state drive, and the finite element model includes the layout of each region of the solid-state drive and the contact interface; Define the physical properties and boundary conditions of each region of the solid-state drive in the finite element model, use finite element analysis software to perform heat conduction simulation, obtain the temperature distribution of the solid-state drive under normal working conditions, and obtain the preliminary temperature field of the solid-state drive; Generate the heat transfer path of the heat in each region inside the solid-state drive based on the preliminary temperature field of the solid-state drive, use the genetic algorithm to estimate the fitness of each heat transfer path, and select the optimal heat transfer path; Obtain the temperature distribution of each region in the solid-state drive under the optimal heat transfer path, and based on the temperature distribution, use finite element analysis software to update the temperature field of the solid-state drive to obtain the heat transfer characteristics of the heat inside the solid-state drive.
[0014] Preferably, using the genetic algorithm to estimate the fitness of each heat transfer path and select the optimal heat transfer path includes: Encode the heat transfer path of the solid-state drive as an individual to obtain a number of initial individuals; Define a fitness function, and evaluate the initial individuals based on the fitness function. According to the fitness evaluation results and the predefined screening strategy, select the individuals with high fitness to enter the next generation; Perform crossover on the selected individuals to generate new individuals, and mutate the individuals after crossover to introduce small random changes to avoid falling into local optima; Evaluate the fitness of the new generation of individuals, update the individuals in the population, and iteratively perform crossover and mutation processing until the termination condition is reached and then stop to obtain the heat transfer path with the best fitness.
[0015] According to another aspect of the present invention, there is also provided a solid-state drive health status monitoring system, which includes: A write load feature recognition module, which is used to collect the monitoring data of the solid-state drive, extract the write features from the monitoring data and perform fluctuation analysis, and identify the overwriting load features of the solid-state drive according to the fluctuation analysis results of the write features; A temperature feature analysis module, which is used to analyze the heat transfer characteristics inside the solid-state drive based on the monitoring data, and analyze the correlation between the heat transfer characteristics and the overwriting load to obtain the abnormal temperature characteristics during overwriting load; A health status assessment module, which is used to input the abnormal temperature feature into a pre-built solid-state drive health status assessment model, predict the health status of the solid-state drive through the health status assessment model, and formulate a data migration mechanism based on the health status.
[0016] The beneficial effects of the present invention are as follows: 1. The data collected by the present invention includes write volume, write frequency, temperature, etc., which helps to identify the trends and abnormalities of write operations. By extracting write features, it is possible to clarify the write volume per unit time and its fluctuations, and further analyze the changes in write frequency. By identifying these overwriting load features, the storage management strategy can be effectively optimized, unnecessary write operations can be reduced, and the service life of the hard disk can be extended. Overwriting not only accelerates the wear of storage units in the SSD, reduces the hard disk life, but also may lead to performance degradation and data loss risks. Timely identification and handling help prevent this situation. In addition, identifying overwriting load features can also improve the performance of the device and avoid problems caused by performance bottlenecks and heat accumulation.
[0017] 2. By analyzing the heat transfer process inside the hard disk, the present invention identifies the heat transfer paths and patterns. By analyzing the heat transfer characteristics, it can help identify these high-temperature regions and the direction of heat flow, and reveal the heat accumulation phenomenon caused by overwriting load. Further analyzing the correlation between heat transfer characteristics and overwriting load, it can be found that when there is an overwriting load, some areas of the hard disk may have abnormal temperature fluctuations due to continuous high-frequency write operations. Therefore, by identifying the abnormal temperature features during overwriting load, it can provide an important basis for the temperature control and heat dissipation design of the hard disk, optimize the heat dissipation system of the hard disk, avoid failures caused by overheating, and effectively extend the service life of the hard disk at the same time.
[0018] 3. The present invention takes the abnormal temperature feature as an input, and predicts the health status of the hard disk through a pre-built solid-state drive health status assessment model, which can effectively identify potential failure risks of the hard disk in advance, so as to discover problems such as performance degradation or damage that may be caused by abnormal temperature, take timely measures, and based on the results of the health status assessment, a data migration mechanism can be formulated to migrate the data to another hard disk with a better health status, thus avoiding data loss or hard disk damage. Description of the Drawings
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0020] Figure 1 is a flowchart of a method for monitoring the health status of a solid - state drive according to an embodiment of the present invention; Figure 2 is a schematic block diagram of a system for monitoring the health status of a solid - state drive according to an embodiment of the present invention.
[0021] In the figure: 1. Write load characteristic recognition module; 2. Temperature characteristic analysis module; 3. Health status evaluation module. Detailed implementation manners
[0022] To further illustrate the embodiments, the present invention provides accompanying drawings, which are part of the disclosure of the present invention. They are mainly used to illustrate the embodiments and can be combined with the relevant descriptions in the specification to explain the operating principles of the embodiments. With reference to these contents, those of ordinary skill in the art should be able to understand other possible implementation manners and the advantages of the present invention.
[0023] According to an embodiment of the present invention, a method and a system for monitoring the health status of a solid - state drive are provided.
[0024] Now, the present invention will be further described in conjunction with the accompanying drawings and specific implementation manners. As Figure 1 shown, for the method for monitoring the health status of a solid - state drive according to an embodiment of the present invention, the method: S1. Collect the monitoring data of the solid - state drive, extract the write characteristics from the monitoring data and perform fluctuation analysis, and identify the over - write load characteristics of the solid - state drive according to the fluctuation analysis results of the write characteristics.
[0025] Among them, collecting the monitoring data of the solid - state drive, extracting the write characteristics from the monitoring data and performing fluctuation analysis, and identifying the over - write load characteristics of the solid - state drive according to the fluctuation analysis results of the write characteristics include: Using hard - disk monitoring technology to collect the monitoring data of the solid - state drive, and the monitoring data includes write volume, write frequency, erase - write cycle and temperature parameters.
[0026] It should be noted that using hard - disk monitoring technology to collect the monitoring data of the solid - state drive, and the monitoring data includes write volume, write frequency, erase - write cycle and temperature parameters includes: Step 1. Construction of the hard - disk monitoring system: Use hard - disk monitoring tools or APIs (such as smartctl, CrystalDiskInfo, HWMonitor, or monitor through custom scripts and operating - system interfaces). These tools can access the health status and performance metrics of the hard disk through S.M.A.R.T. (Self - Monitoring, Analysis and Reporting Technology) data.
[0027] Step 2. Data collection: Collection Write Volume: Obtain the amount of data written to the hard drive through "Total Host Writes" in the S.M.A.R.T. metrics or other appropriate items. This is usually expressed in bytes, megabytes (MB), or gigabytes (GB); collect the write volume that occurs on the hard drive within a specified time period (such as per hour, daily), and record the cumulative value of each write operation.
[0028] Collection Write Frequency: Write frequency usually refers to the number of write operations per unit time, which can be obtained by real-time monitoring of the number of I / O operations on the hard drive or the number of read / write operations per second (IOPS, Input / Output Operations Per Second).
[0029] Collection Erase / Write Cycle: The erase / write cycle refers to the number of times a block of the SSD is completely erased and rewritten. After the SSD reaches a certain erase / write cycle, its performance and durability will decrease. The erase / write cycle can be obtained through "Program / Erase Cycle" in S.M.A.R.T. Monitor the change of the erase / write cycle, record the number of erase / write operations for each block of the SSD, so as to understand the durability and health status of the hard drive.
[0030] Collection Temperature Parameter: SSDs usually have built-in temperature sensors that can report the operating temperature of the hard drive in real time. The temperature data of the hard drive can be directly read through a hard drive monitoring tool (such as smartctl). It is necessary to collect and record the operating temperature of the hard drive, including the maximum temperature and the current temperature, and track the temperature fluctuations to avoid high temperatures affecting the performance of the hard drive.
[0031] Step 3: Data Processing and Storage: Data Storage: Store the monitoring data (write volume, write frequency, erase / write cycle, temperature) in chronological order in a database or local storage. SQL databases (such as MySQL, PostgreSQL) or NoSQL databases (such as MongoDB) can be used to store these time series data.
[0032] Data Cleaning and Filtering: Clean the collected data to remove errors or outliers. Ensure the accuracy of the data, filter out invalid data points, set thresholds and alarm mechanisms to avoid interference from extreme values on data analysis.
[0033] Perform time series mining analysis on the write volume and write frequency of the fixed hard drive respectively to obtain the write characteristics of the solid-state hard drive, and perform Fourier transform analysis on the write characteristics to obtain the fluctuation analysis results of the write characteristics.
[0034] Among them, the write amount and write frequency of the fixed hard disk are respectively subjected to time series mining analysis to obtain the write characteristics of the solid-state hard disk, and the Fourier transform analysis is performed on the write characteristics to obtain the fluctuation analysis results of the write characteristics, including: The tree index technology is used to obtain the similar time series of the write amount and write frequency of the fixed hard disk within the preset period, and weight values are respectively assigned to the similar time series.
[0035] It should be noted that using the tree index technology to obtain the similar time series of the write amount and write frequency of the fixed hard disk within the preset period and respectively assigning weight values to the similar time series includes: Step 1: Establish a tree index structure: Tree indexes such as R-tree, KD-tree, and quadtree can be used for indexing high-dimensional data. Here, the R-tree or K-D tree can be selected. They are particularly suitable for handling the spatial index of time series data and can effectively organize and query similar time series data.
[0036] Construct the index: The time series data of the collected write amount and write frequency are used to construct a tree index according to the time stamp. Each time series data point (such as the write amount and frequency) is regarded as a node in the tree.
[0037] The tree structure can efficiently query and retrieve similar time series data, so that similar sequences to the current time series data within the time period can be quickly found.
[0038] Step 2: Retrieve similar time series: Methods such as dynamic time warping (DTW), Euclidean distance, and cosine similarity are used to define the similarity between time series data. Dynamic time warping (DTW) is a commonly used similarity measurement method in time series analysis, which can measure the non-linear matching degree of time series. For the time series data of the write amount and write frequency, after selecting an appropriate similarity measurement method, the similarity of data sequences in different time periods can be measured.
[0039] Retrieve similar time series: Using the constructed tree index, by calculating the similarity with the current time series (such as the write amount or frequency sequence in a certain time period), the time series data most similar to the current sequence is retrieved from the tree. These similar time series can help analyze the long-term trends and patterns of the data.
[0040] Step 3: Assign weight values to similar time series: Calculate the weight values: According to the results of the similarity calculation, a weight value is assigned to each similar time series. The weight value is usually proportional to the similarity: the higher the similarity of the sequence, the greater the weight value.
[0041] To control the computational complexity and improve the accuracy of the model, a minimum similarity threshold can be set for similar sequences, and time series sequences below this threshold will be ignored and not participate in subsequent analysis.
[0042] Sort the similar time series sequences according to the weights, and select the combination of the write amount and write frequency within the preset sequence range to generate a write feature sequence; Perform a Fourier transform on the write feature sequence, and identify the frequency components of the write feature sequence based on the frequency fluctuation trend of the write feature sequence.
[0043] Among them, performing a Fourier transform on the write feature sequence and identifying the frequency components of the write feature sequence based on the frequency fluctuation trend of the write feature sequence includes: Store the continuous write amount and write frequency of the write feature sequence into a predefined channel stream buffer.
[0044] It should be noted that the channel stream refers to the channel or stream for data transmission, especially in the processing of data streams. For example, in multi-channel audio processing, each channel can be regarded as an independent signal stream. The buffer refers to the area for temporarily storing data. The role of the buffer is to improve the processing speed of the system and avoid frequent calculation or storage operations. If these features are to be stored in the "channel stream buffer", these write features are stored as a sequence of time-domain signals, and it is planned to perform a Fourier transform (or similar signal analysis) on these data to identify the frequency features, and then use them to optimize the storage management or predict the health status of the hard disk.
[0045] Take each time point of several write features as the amplitude component, and insert the phase component between every two adjacent amplitude components in turn to obtain the composite value of the channel stream buffer; Take several composite values in the time domain in the channel stream buffer as the input of time decimation, and use the time decimation method to transform the time-domain signal to obtain several composite values in the frequency domain.
[0046] It should be noted that taking several composite values in the time domain in the channel stream buffer as the input of time decimation and using the time decimation method to transform the time-domain signal to obtain several composite values in the frequency domain includes: Step 1: Creation of the channel stream buffer: Obtain a set of time series data representing the hard disk write amount or write frequency (such as the write amount per second, frequency, etc.), and these data will be stored in the channel stream buffer. The data value of each time point can be represented as a composite value (in complex form), whose amplitude represents the intensity of the signal, and the phase component represents the phase information of the signal.
[0047] Data representation: The collected data can be composed of complex values in the time domain, where each data point can be a complex number. For example, the time series data points representing the hard disk write load x ( t ) are represented as complex numbers: , where A ( t ) represents the amplitude; represents the phase.
[0048] Create a channel stream buffer: Collect a set of time-domain data into the buffer. Assuming that each time series point stores a complex value, one or more such time series sequences are stored in the buffer.
[0049] Step 2. Decimation operation: Select some time series data (i.e., decimation) extracted from the time-domain signal. These data will be used as inputs for frequency-domain analysis. Decimation usually refers to selecting a part of the data from the overall time-domain signal for Fourier transform processing. The decimation methods include: 1. Select a time window of fixed length: Select consecutive time segments (windows) in the channel stream buffer. Assuming that the sampling frequency of the time-domain signal is f s , then the number of samples corresponding to a time window length T is N = T × f s .
[0050] For example, the hard disk write frequency data per second can be selected, or the data within a certain fixed period can be selected as a time window for processing.
[0051] 2. Sliding window: Use the sliding window technique to slide the window at a fixed step size and sequentially select different parts of the signal. Each part of the data selected each time will be used as an input for processing.
[0052] 3. Decimation interval: The decimation interval of the data can be fixed (such as per second, per minute), or based on certain rules (for example, decimation is performed when the write load reaches a certain threshold).
[0053] Step 3. Transform the time signal into a frequency-domain signal: After completing the decimation, convert the time-domain signal into a frequency-domain signal. Specifically, use the discrete Fourier transform (DFT), especially its efficient implementation, i.e., the fast Fourier transform (FFT). FFT can quickly calculate the frequency-domain representation of the time-domain signal and obtain its spectrum.
[0054] FFT process: Perform an FFT transformation on the time-domain data after time decimation to convert the time-domain signal into a frequency-domain signal.
[0055] The calculation result will obtain a series of complex values. The magnitude of each complex value represents the amplitude of the signal at a specific frequency, and the phase part of the complex number represents the phase information of that frequency component. The expression of the Fourier transform is: ; In the formula, f k represents the frequency points in the frequency domain, x ( t ) represents the time-domain data, N represents the length of the window, t represents the time.
[0056] Step Four: Obtain the complex values in the frequency domain: The frequency-domain data obtained by the FFT transformation can be expressed as several complex values. These complex values represent different frequency components of the signal, including their amplitude and phase information. Specifically: Frequency-domain amplitude and phase: The amplitude ∣ X ( f k )∣ and the phase arg( X ( f k )) can be extracted from each complex value in the frequency domain; the amplitude ∣ X ( f k )∣ describes the intensity of this frequency component and reflects the energy size of the signal at this frequency; the phase arg( X ( f k )) describes the phase information of this frequency component.
[0057] Spectrum: Plot the amplitude values of all frequency-domain data points as a spectrum diagram. The spectrum reflects the intensity distribution of the signal at each frequency.
[0058] Step Five: Analyze and use the frequency-domain data: Analyze the frequency distribution of the signal through the spectrum diagram to identify which frequency components are the most prominent in the write load of the hard disk, so as to be used to monitor the working mode of the hard disk and identify high-load periods or abnormal frequency patterns.
[0059] Frequency-domain analysis helps to extract certain specific frequency components in the operation of the hard disk, and further analyze whether these frequency components are related to hard disk failures, performance degradation, overload, etc.
[0060] The write volume data of a certain hard disk was collected continuously for 30 minutes, and the sampling frequency was 1 Hz, that is, the write volume was recorded once per second. Assume that within a certain period of time, the write volume data of the hard disk (in GB / s) is: [1.2, 1.3, 1.5, 1.4, 1.6].
[0061] Time decimation: Select a 10-second window for time decimation. Assume that the data from the 1st second to the 10th second is selected as a time series window for FFT transformation.
[0062] Perform FFT: Apply FFT to the selected 10-second data to obtain the composite values in the frequency domain. For example .
[0063] Frequency domain analysis: Calculate the amplitude and phase of the obtained composite values. For example, , ∣ X ( f k )∣ = 3.6, arg( X ( f k )) = tan -1 (3 / 2), indicating the intensity and phase information of this frequency component.
[0064] Calculate the modulus length of the composite values in the frequency domain, convert the amplitude values to decibel values to obtain the amplitude values of several discrete frequency lines, and use the frequencies corresponding to the discrete frequency lines as the input for the next time period; Iteratively execute the calculation process of the composite values in the time domain until the processed composite values are summarized to obtain the frequency components of the write feature sequence.
[0065] Based on the frequency components of the write feature sequence, establish a time series data curve, assign weights to the write feature sequence based on the degree of interest in the time series data curve, and calculate the time series points of the write feature sequence to obtain the fluctuation analysis result of the write feature timing.
[0066] Among them, the expression for the time series points of the write feature sequence is: ; In the formula, D represents the time series points of the write feature sequence; A represents the similar time series sequence; WA represents the weight of the similar time series sequence; Q represents the write feature sequence; WQ represents the weight of the write feature sequence.
[0067] Input the fluctuation analysis results of the write characteristics into a predefined variational autoencoder to learn the normal mode of the write characteristics, and compare the difference mode generated by the reconstruction error of the variational autoencoder with the normal mode to identify the over-write load characteristics of the solid-state drive from the write characteristics.
[0068] It should be noted that inputting the fluctuation analysis results of the write characteristics into a predefined variational autoencoder to learn the normal mode of the write characteristics, and comparing the difference mode generated by the reconstruction error of the variational autoencoder with the normal mode to identify the over-write load characteristics of the solid-state drive from the write characteristics includes: Step 1: Write fluctuation analysis: Calculate the fluctuation or change amplitude of the write volume within each time window. For example, the change amount of the write volume per second or per minute compared with the previous moment can be calculated to measure the volatility of the write load.
[0069] For example, assuming the write volume sequence: Write Volume = [1.2, 1.3, 1.5, 1.4, 1.6], calculate its fluctuation (difference) as: 。
[0070] Step 2: Feature extraction: Extract time-domain features (such as mean, standard deviation, maximum value, minimum value, etc.) from the original data, and extract additional features (such as the amplitude of fluctuations, periodicity, etc.) according to the fluctuation analysis, and vectorize these features for input into the variational autoencoder for learning.
[0071] Step 3: Train the variational autoencoder (VAE): The variational autoencoder (VAE) is a deep learning model that can learn the latent distribution of data and automatically generate a low-dimensional representation of features during training. In this scenario, the VAE is used to learn the normal mode of the hard disk write characteristics, which specifically includes: 1. Build and configure the variational autoencoder: Encoder: The encoder compresses the input write characteristic data into the latent space to learn the low-dimensional representation of the normal write characteristics.
[0072] Decoder: The decoder reconstructs the input data from the representation in the latent space.
[0073] Regularization: The variational autoencoder regularizes the latent space through the Kullback-Leibler Divergence to ensure that the learned latent representation has good distribution properties.
[0074] 2. Data input and training: Input the write features that have undergone fluctuation analysis and feature extraction into the VAE, and use a large amount of normal data (i.e., normal write features without overwriting) to train the VAE model. The goal is to enable the model to learn the latent distribution of normal write patterns.
[0075] Step 4: Reconstruction and error calculation: After the VAE training is completed, the model can reconstruct the input data through the decoder. In practical applications, new input data (including potential abnormal data) is input into the VAE, and the reconstruction error is calculated to identify whether there is an overwriting load.
[0076] 1. Input new write feature data: During the detection process, the real-time collected write feature data is input into the trained VAE. If the reconstruction error is large, it indicates that the current input data has a large difference from the normal mode, and there may be an abnormality of overwriting load.
[0077] 2. Comparison with the normal mode: Compare the reconstruction error with the error threshold of the normal write mode. If the error exceeds the preset threshold, it indicates that the input write data exhibits the characteristics of overwriting load, and an alarm needs to be triggered or further analysis is required.
[0078] Step 5: Identify the characteristics of overwriting load: Through the analysis of the reconstruction error, identify the overwriting load characteristics in the write features. If the reconstruction error accumulates beyond a certain threshold within a time window, it can be determined that there is an abnormality in the current write load of the hard disk, and then the possible overwriting load characteristics can be identified. Once an overwriting load is detected, the system can trigger an alarm to notify the administrator to take corresponding measures (such as reducing the load, adjusting the write strategy, etc.).
[0079] S2. Analyze the built-in heat transfer characteristics of the solid-state drive based on the monitoring data, and analyze the correlation between the heat transfer characteristics and the overwriting load to obtain the abnormal temperature characteristics under overwriting load.
[0080] Among them, analyzing the built-in heat transfer characteristics of the solid-state drive based on the monitoring data, and analyzing the correlation between the heat transfer characteristics and the overwriting load to obtain the abnormal temperature characteristics under overwriting load includes: Divide the built-in area of the solid-state drive, and calculate the temperature rise rate of each area in the solid-state drive under each time window based on the temperature parameters; Based on the temperature rise rate, establish a temperature gradient transfer mechanism, and combine the physical structure of the solid-state drive and the thermal conductivity of the materials in each area to extract the heat transfer characteristics of the heat in the solid-state drive.
[0081] Among them, a temperature gradient transfer mechanism is established based on the temperature rise rate, and combined with the physical structure of the solid-state drive and the thermal conductivity of the materials in each region, the heat transfer characteristics inside the solid-state drive are extracted, including: Use finite element analysis software to establish a finite element model of the solid-state drive, and the finite element model includes the layout of each region of the solid-state drive and the contact interface; Define the physical properties and boundary conditions of each region of the solid-state drive in the finite element model, use finite element analysis software to perform heat conduction simulation, obtain the temperature distribution of the solid-state drive under normal working conditions, and obtain the preliminary temperature field of the solid-state drive.
[0082] It should be noted that defining the physical properties and boundary conditions of each region of the solid-state drive in the finite element model, using finite element analysis software to perform heat conduction simulation, obtaining the temperature distribution of the solid-state drive under normal working conditions, and obtaining the preliminary temperature field of the solid-state drive includes: Step 1: Establish a geometric model of the solid-state drive, including its internal structure and each component. Different regions may have different material properties: 1. Determine the geometric shape of the hard disk: The shape of the solid-state drive: Usually includes components such as a housing, a PCB board, storage chips, circuits, etc. The geometric model of the hard disk should include the dimensions and relative positions of these different regions.
[0083] Divide each region: There may be different regions inside the hard disk (for example: storage chips, controllers, heat sinks, etc.), and the material properties and heat conduction properties of each region may be different. Therefore, the geometric model needs to be divided according to these regions.
[0084] 2. Create a 3D model: Use finite element software (such as ANSYS, COMSOL, Abaqus, etc.) to draw a 3D geometric model of the hard disk in the software. The model can be built according to the actual design drawings of the hard disk, or imported from a CAD model.
[0085] Divide different regions of the hard disk (such as storage chips, PCB, power modules, etc.) into different mesh elements to form a finite element mesh.
[0086] Step 2: Define material properties: The physical properties (such as thermal conductivity, specific heat capacity, density, etc.) of each region should be defined according to the material type. Different materials have different thermal conductivities and thermal responses, which have an important impact on the temperature distribution.
[0087] 1. Definition of material thermal physical properties: Thermal conductivity: Describes the ability of heat conduction. Different materials (such as metals, silicon, plastics, etc.) have different thermal conductivities.
[0088] Specific heat capacity: The amount of heat required for a unit mass of a substance to increase its temperature by 1°C.
[0089] Density: Affects heat capacity calculation and heat conduction.
[0090] Coefficient of thermal expansion: Considering the expansion effect of materials at high temperatures, it may need to be considered in some simulations with high precision requirements.
[0091] Define the physical properties of these regions according to the actual materials of each part of the hard disk. For example: Storage chip: Made of silicon material, with a thermal conductivity of about 148 W / m·K; PCB (printed circuit board): A composite material of fiberglass and epoxy resin, with a thermal conductivity of about 0.2 W / m·K; Housing: Made of metal material, with a relatively high thermal conductivity, usually aluminum or stainless steel.
[0092] 2. Material property assignment: In the model, correspond the thermal physical properties of the materials to the geometric parts of each region. For example: Assign the thermal conductivity of silicon to the storage chip region; Assign the thermal conductivity of fiberglass composite material to the PCB region; Assign the thermal conductivity of metal to the housing region.
[0093] Step 3. Define boundary conditions: In heat conduction simulations, boundary conditions are very important because they define the way heat is transferred and interacts with the environment.
[0094] 1. Set external boundary conditions: Convection boundary condition: Considering that solid-state drives are usually installed in a closed environment, the heat conduction between the housing and the outside air is often achieved through convection. It is necessary to define the convective heat transfer coefficient with the environment. Assuming little air flow, the standard convective heat transfer coefficient (usually 5 - 20 W / m²·K) can be used. If there are cooling devices such as fans, the coefficient may be higher.
[0095] Radiation boundary condition: If radiation effects are considered (especially in high-temperature environments), it is necessary to apply radiation heat conduction conditions on the external surface. The calculation of radiation heat conduction usually depends on the surface emissivity of the object.
[0096] 2. Set internal boundary conditions: Internal heat source: The storage chips, controllers, and other electronic components of the solid-state drive generate heat during operation. It is necessary to define the internal heat source according to the actual power consumption (such as the power consumption during writing and reading). For example, the storage chips may generate more heat under high load.
[0097] Estimate the internal heat source from the power consumption data of the hard disk (such as the performance indicators provided by the manufacturer).
[0098] Heat conduction model: Inside each region, it is necessary to define the heat conduction model of the solid-state drive. Usually, Fourier's law of heat conduction is used, which describes the heat conduction process based on the thermal conductivity of the material and the temperature gradient. Its expression is: , where q represents the heat flux density; k represents the thermal conductivity of the material; represents the temperature gradient.
[0099] Step 4: Mesh generation: Mesh generation is a key step in finite element analysis, directly affecting the accuracy and computational efficiency of the simulation results. It specifically includes: 1. Meshing: Divide the geometric model of the solid-state drive into a finite number of elements. The shape of each element may be a tetrahedron, hexahedron, etc., depending on the complexity of the model and the requirements for computational accuracy. For regions with higher accuracy requirements (such as the storage chips or circuit boards with intensive heat sources), finer meshes are needed.
[0100] 2. Mesh quality inspection: Ensure the quality of the mesh generation and avoid generating irregular meshes (such as long and thin elements, extremely small elements, etc.). These irregular meshes will lead to computational errors and instability.
[0101] Step 5: Heat conduction simulation and calculation: 1. Select the appropriate simulation type: Select steady-state or transient heat conduction analysis. Steady-state analysis is applicable to the temperature distribution under long-term stable operation, while transient analysis is applicable to the temperature changes of the hard disk under different operating conditions.
[0102] 2. Run the simulation: Run the heat conduction simulation to calculate the temperature distribution of the solid-state drive in different regions.
[0103] The software will calculate the temperature of each node (or element) based on the input physical properties, boundary conditions, and mesh generation.
[0104] 3. Output of the temperature field results: Output the simulation results, including the temperature distribution diagrams of each region of the hard disk.
[0105] The temperature values of different regions (such as storage chips, PCBs, enclosures) can be obtained to understand the thermal characteristics of the hard disk under normal operating conditions.
[0106] Generate the heat transfer paths of the internal heat of the solid-state drive in each area based on the preliminary temperature field of the solid-state drive, use the genetic algorithm to estimate the fitness of each heat transfer path, and screen out the optimal heat transfer path.
[0107] It should be noted that generating the heat transfer paths of the internal heat of the solid-state drive in each area based on the preliminary temperature field of the solid-state drive includes: The heat transfer path refers to the heat transfer path from the heat source (such as memory chips, controllers, power modules, etc.) to the outside of the hard disk or other areas with lower temperature. Heat is transferred between the inside and outside environment of the hard disk through three main ways: heat conduction, convection, and radiation.
[0108] Heat conduction: Heat is transferred through molecular collisions inside the material, usually occurring inside solids. The areas with higher thermal conductivity inside the solid-state drive (such as metal casings, storage chips, PCBs, etc.) are the main channels for heat transfer.
[0109] Convection: Heat is transferred through the flow of gas or liquid fluids, usually occurring between the solid-state drive and the external environment (air or cooling system).
[0110] Radiation: Heat is transferred in the form of electromagnetic waves (radiation). Although the influence of radiation inside the hard disk is small, at higher temperatures, radiation may become an important way of heat transfer.
[0111] Among them, using the genetic algorithm to estimate the fitness of each heat transfer path and screen out the optimal heat transfer path includes: Encode the heat transfer paths of the solid-state drive as individuals to obtain a number of initial individuals; Define the fitness function, and evaluate the initial individuals based on the fitness function. Select the individuals with high fitness to enter the next generation according to the fitness evaluation results and the predefined screening strategy; Perform crossover on the selected individuals to generate new individuals, and mutate the individuals after crossover to introduce small random changes to avoid falling into local optima; Evaluate the fitness of the new generation of individuals, update the individuals in the population, and iteratively perform the crossover and mutation process until the termination condition is reached and then stop to obtain the heat transfer path with the best fitness.
[0112] Obtain the temperature distribution of each area in the solid-state drive under the optimal heat transfer path, and based on the temperature distribution, use finite element analysis software to update the temperature field of the solid-state drive to obtain the heat transfer characteristics of the internal heat of the solid-state drive.
[0113] Calculate the correlation degree between the heat transfer characteristics and the overwriting load using the Pearson correlation coefficient, and compare the correlation degree with the preset threshold.
[0114] Among them, the calculation formula of the Pearson correlation coefficient is as follows: ; In the formula, r represents the Pearson correlation coefficient; represents the mean value of the heat transfer characteristics; represents the overwriting load; represents the i th heat transfer characteristic; represents the i th overwriting load; represents the weight value; n represents the total number of characteristics.
[0115] If the correlation degree is greater than or equal to the prediction threshold, it means that the temperature rise rate increases abnormally during overwriting load. Then, the load balancing algorithm is used to regulate the temperature parameters of each area inside the solid-state drive. Otherwise, it means that the temperature rise rate is not affected during overwriting load.
[0116] It should be noted that regulating the temperature parameters of each area inside the solid-state drive by using the load balancing algorithm includes: Step 1. Judge the overwriting load: If the correlation degree (such as the relationship between the load and the temperature rise rate) is greater than or equal to the set prediction threshold, it means that there is overwriting load and the load balancing mechanism needs to be enabled.
[0117] Temperature rise rate: Calculate the current temperature rise rate through the real-time temperature change. The temperature rise rate is used as the ratio of the temperature change to the time change. If the temperature rise rate increases abnormally, it indicates that the temperature of this area may rise too fast due to overwriting load.
[0118] Dynamic load adjustment: Calculate the load distribution of each area according to the temperature rise rate and the load data, and dynamically adjust the load distribution.
[0119] High-load area: If the temperature of some areas rises relatively fast, it means that the load of these areas is too large and the load needs to be reduced or dispersed.
[0120] Low-load area: Transfer the load from the high-temperature area to the low-temperature area to relieve the heat accumulation in the high-load area.
[0121] Step 2. Regulation mechanism: 1. Data migration: Migrate the data writing task from the area with too heavy load to the area with lower load. For example, when the temperature of the storage chip is too high, some data can be written to the area with lower load or the cache.
[0122] Write Scheduling: By optimizing the write scheduling algorithm, the write operations are made more uniform to avoid excessive concentrated writing in a certain area. For example, use a polling algorithm or a priority-based scheduling algorithm to allocate write tasks.
[0123] 2. Dynamic Thermal Control: Thermostatic Strategy: By monitoring the temperature and load conditions of each area, the thermal control measures of the area are dynamically adjusted. For example, increase the heat dissipation of the high-temperature area or limit the write operations at too high temperatures to reduce the load of the high-temperature area.
[0124] Internal Heat Dissipation Management: By adjusting the operation of the radiator inside the hard disk, the heat transfer from the high-temperature area to other areas is controlled. For the high-temperature area, active heat dissipation measures can be increased, such as starting the fan and adjusting the heat sink.
[0125] 3. Load Optimization Goals: Thermal Balance Optimization: Through load transfer and scheduling, the thermal load of the hard disk is evenly distributed to ensure that the temperature changes in each area are within a safe range. By adjusting the load distribution, the optimization of the temperature field is achieved.
[0126] Trade-off between Performance and Temperature: When optimizing the load, the balance between performance and temperature must be considered. For example, allocating too much load to the low-temperature area can reduce the temperature rise rate of the high-temperature area, but it may also affect the read and write speed of the hard disk.
[0127] It should be noted that the temperature gradient transfer mechanism combined with finite element analysis and genetic algorithm can provide significant advantages in solving complex heat transfer problems. Finite element analysis (FEA) is a numerical calculation method commonly used to simulate the heat transfer in complex structures, which can accurately calculate the distribution of the temperature field and help understand how heat flows in solid-state drives or other complex systems. Through finite element analysis, the geometric structure of the hard disk or device is discretized into small units, and then the temperature changes of each unit are calculated to obtain the global temperature field data.
[0128] However, finite element analysis itself cannot directly optimize the thermal design. Therefore, combining with genetic algorithm (GA) can further optimize on this basis. Genetic algorithm is an optimization method that simulates the process of natural selection and gene inheritance. By simulating operations such as selection, crossover, and mutation, the optimal solution can be found among a large number of possible solutions. After combining finite element analysis and genetic algorithm, multi-objective optimization can be carried out for heat transfer problems.
[0129] Through the optimization of the genetic algorithm, the optimal thermal management configuration in the temperature gradient transfer process can be explored, improving the thermal efficiency and reducing the energy consumption. At the same time, it can also extend the service life of the device and ensure its stability and reliability under high-load working conditions.
[0130] S3. Input the abnormal temperature characteristics into a pre - constructed solid - state drive (SSD) health status assessment model, predict the health status of the SSD through the health status assessment model, and formulate a data migration mechanism based on the health status.
[0131] It should be noted that inputting the abnormal temperature characteristics into a pre - constructed SSD health status assessment model, predicting the health status of the SSD through the health status assessment model, and formulating a data migration mechanism based on the health status include: Step 1. Construct a health status assessment model: The health status assessment model can predict the health of the hard disk based on machine learning, statistical analysis, or deep learning methods. The model includes: Supervised learning models: such as logistic regression, support vector machine (SVM), decision tree, etc.
[0132] Neural network models: such as deep neural network (DNN) or long short - term memory (LSTM) network, which are suitable for predicting time - series data.
[0133] Hybrid models: Combine multiple different machine learning models to enhance the accuracy of prediction.
[0134] Step 2. Dataset preparation and training: Feature selection: Select temperature - related features such as temperature rise rate, temperature peak, load fluctuation, and other hard disk health indicators (such as read / write error rate, power consumption, etc.).
[0135] Data labeling: Mark the health status according to historical data, usually marked as "normal", "warning", or "fault".
[0136] Normal state: The hard disk temperature is normal, and the temperature rise rate is within the expected range.
[0137] Warning state: The temperature is close to the upper limit, and the temperature rise rate is on the high side.
[0138] Fault state: The hard disk is overheated, and the temperature rise rate is abnormal, which may lead to hard disk failure.
[0139] Training process: Use the labeled dataset (including abnormal temperature characteristics and health status labels) to train the model.
[0140] Use cross - validation and optimization algorithms (such as grid search, random search) to adjust the model parameters and select the optimal model.
[0141] For example, use the support vector machine (SVM) model for training, select features such as temperature rise rate and peak temperature in the input data, and train a classification model that can judge the hard disk health status based on the input temperature characteristics.
[0142] Step 3: Health Status Prediction and Analysis: 1. Input Abnormal Temperature Features: In actual use, when the solid-state drive is running, the temperature data of each area can be obtained in real time, and its temperature rise rate and temperature peak can be calculated.
[0143] Suppose the current temperature of a certain hard disk is 72°C and the temperature rise rate is 5°C / min.
[0144] Input these features into the trained health status evaluation model, and the model will predict the health status of the hard disk according to the rules learned during the training process.
[0145] 2. Health Status Evaluation: Model Output: The model outputs the health status of the hard disk, such as "Normal", "Warning", or "Fault".
[0146] If the output is "Warning", it means that the hard disk temperature is close to or exceeds the safety threshold, and there is a risk of overheating.
[0147] If the output is "Fault", it indicates that the hard disk has already had a serious temperature anomaly and may malfunction.
[0148] For example, set a threshold. When the health status evaluation model outputs "Warning" or "Fault", immediately trigger the next load scheduling and data migration mechanism. Suppose the temperature rise rate of a certain hard disk is 4°C / min, and the model predicts that the health status of the hard disk is "Warning", then trigger the data migration mechanism.
[0149] It should be noted that the dynamic data migration mechanism based on the health status includes: Step 1: Trigger Conditions for Data Migration: High Temperature or Overloaded Areas: When the health status evaluation model detects that the hard disk is in the "Warning" or "Fault" state, the data migration mechanism can be started. At this time, it is necessary to migrate the data from the overloaded or high-temperature area to the area with lower load or lower temperature.
[0150] For example, among multiple areas of the SSD (such as storage chips, caches, controllers, etc.), the area with higher temperature may cause overheating, so it is necessary to migrate the data to the area with lower temperature.
[0151] Step 2: Data Migration Strategy: Data Migration Based on Priority: Set priorities according to the load and temperature conditions of each area of the hard disk. Give priority to migrating the data in the overloaded and high-temperature areas.
[0152] For example, migrate the data of the storage chip with higher write load and higher temperature to the controller or cache area with lower temperature.
[0153] Temperature - Load Joint Optimization: In addition to the load, the temperature distribution of the hard disk can also be considered. Adopt a temperature - load joint optimization strategy to ensure that after data migration, the load and temperature in each area are within a reasonable range.
[0154] Dynamic Migration and Cache Management: During the data migration process, combine the cache management mechanism built into the hard disk to reduce the performance impact caused by migration while ensuring data security.
[0155] Step Three: Data Migration Process: Execution of Migration Strategy: When the health status of the hard disk shows an anomaly, implement the data migration strategy. For example, migrate all the data in a certain storage area to another area with a lower temperature, or reduce the workload of the overheated area by distributing part of the write load to other areas.
[0156] Real - Time Monitoring and Adjustment: Monitor the status of the hard disk after migration in real - time and adjust the migration strategy to ensure the health of the hard disk.
[0157] Step Four: Evaluation of Migration Effect: Temperature Balance: After migration is completed, monitor the temperature changes in each area to ensure temperature equilibrium and avoid temperature concentration again.
[0158] Evaluation of Performance Impact: Evaluate the impact of data migration on the performance of the hard disk to ensure that the migration operation does not have too much impact on the read / write performance of the hard disk.
[0159] For example, conduct experiments on solid - state drives to obtain their temperature data under different loads and train a health status evaluation model. The following are some assumed data and results: Normal Working State: The temperature fluctuates between 50°C - 55°C, the temperature rise rate is 2°C / min, and the predicted health status is "normal".
[0160] Abnormal State: Under high - load conditions, the temperature exceeds 70°C, the temperature rise rate reaches 5°C / min, and the predicted health status is "warning".
[0161] Fault State: Continuous high - load, the temperature reaches 80°C, the temperature rise rate is 8°C / min, and the predicted health status is "fault".
[0162] When the health status of the hard disk is "warning", the storage area of the hard disk is migrated to an area with a lower load and a lower temperature. For example, part of the data of the storage chip originally in the high - temperature area is migrated to the controller area with better cooling effect, successfully reducing the temperature rise rate and avoiding the occurrence of faults.
[0163] According to another embodiment of the present invention, as Figure 2As shown, a solid-state drive health status monitoring system is also provided, and the system includes: A write load feature recognition module 1, configured to collect monitoring data of the solid-state drive, extract write features from the monitoring data and perform fluctuation analysis, and identify the over-write load features of the solid-state drive according to the fluctuation analysis results of the write features; A temperature feature analysis module 2, configured to analyze the heat transfer features built in the solid-state drive based on the monitoring data, and analyze the correlation between the heat transfer features and the over-write load to obtain abnormal temperature features during over-write load; A health status evaluation module 3, configured to input the abnormal temperature features into a pre-constructed solid-state drive health status evaluation model, predict the health status of the solid-state drive through the health status evaluation model, and formulate a data migration mechanism based on the health status.
[0164] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for monitoring the health status of a solid-state drive, characterized in that, The method: Collect the monitoring data of the solid-state drive, extract the write characteristics from the monitoring data and perform fluctuation analysis, and identify the over-write load characteristics of the solid-state drive according to the fluctuation analysis results of the write characteristics; Analyze the built-in heat transfer characteristics of the solid-state drive based on the monitoring data, and analyze the correlation between the heat transfer characteristics and the over-write load to obtain the abnormal temperature characteristics under the over-write load; Use the abnormal temperature characteristics as the input of the pre-constructed solid-state drive health status evaluation model, predict the health status of the solid-state drive through the health status evaluation model, and formulate a data migration mechanism based on the health status; The extracting the write characteristics from the monitoring data and performing fluctuation analysis includes: Perform time series mining analysis on the write volume and write frequency of the solid-state drive respectively to obtain the write characteristics of the solid-state drive, and perform Fourier transform analysis on the write characteristics to obtain the fluctuation analysis results of the write characteristics; The analyzing the built-in heat transfer characteristics of the solid-state drive based on the monitoring data includes: Divide the built-in area of the solid-state drive, and calculate the temperature rise rate of each area in the solid-state drive under each time window based on the temperature parameters; Establish a temperature gradient transfer mechanism based on the temperature rise rate, and combine the physical structure of the solid-state drive and the thermal conductivity of the materials in each area to extract the heat transfer characteristics built in the solid-state drive.
2. The method for monitoring the health status of a solid-state drive according to claim 1, characterized in that, The collecting the monitoring data of the solid-state drive includes: Use hard disk monitoring technology to collect the monitoring data of the solid-state drive, and the monitoring data includes write volume, write frequency, erase-write cycle and temperature parameters; The identifying the over-write load characteristics of the solid-state drive according to the fluctuation analysis results of the write characteristics includes: Input the fluctuation analysis results of the write characteristics into a predefined variational autoencoder to learn the normal mode of the write characteristics, and compare the difference mode generated by the reconstruction error of the variational autoencoder with the normal mode to identify the over-write load characteristics of the solid-state drive from the write characteristics.
3. The method for monitoring the health status of a solid-state drive according to claim 1, characterized in that, The performing time series mining analysis on the write volume and write frequency of the solid-state drive respectively to obtain the write characteristics of the solid-state drive, and performing Fourier transform analysis on the write characteristics to obtain the fluctuation analysis results of the write characteristics includes: Adopt a tree index technology to obtain the similar time series of the write volume and write frequency of the solid-state drive within a preset period, and assign weights to the similar time series respectively; Sort the similar time series according to the weights, and select the write volume and write frequency combinations within the preset sequence range to generate a write characteristic sequence; Perform Fourier transform on the write characteristic sequence, and identify the frequency components of the write characteristic sequence based on the frequency fluctuation trend of the write characteristic sequence; Establish a time series data curve based on the frequency components of the write characteristic sequence, assign weights to the write characteristic sequence based on the degree of interest of the time series data curve, and calculate the time series points of the write characteristic sequence to obtain the fluctuation analysis results of the write characteristic time series.
4. The method for monitoring the health status of a solid-state drive according to claim 3, characterized in that, The performing Fourier transform on the write characteristic sequence, and identifying the frequency components of the write characteristic sequence based on the frequency fluctuation trend of the write characteristic sequence includes: Store the continuous write volume and write frequency of the write characteristic sequence into a predefined channel stream cache; Take each time point of several writing features as amplitude components, and insert phase components between every two adjacent amplitude components in sequence to obtain the composite values of the channel stream buffer; Take several composite values in the time domain in the channel stream buffer as the input of time decimation, and use the time decimation method to transform the time domain signal to obtain several composite values in the frequency domain; Calculate the modulus length of the composite values in the frequency domain, convert the amplitude values to decibel values to obtain the amplitude values of several discrete frequency lines, and use the frequencies corresponding to the discrete frequency lines as the input in the next time period; Iteratively execute the calculation process of the composite values in the time domain until the processed composite values are summarized to obtain the frequency components of the writing feature sequence.
5. The method for monitoring the health status of a solid-state drive according to claim 3, characterized in that, The expression of the timing points for writing the characteristic sequence is as follows: ; Wherein, D represents the timing point for writing the feature sequence; A represents the similar timing sequence; WA Indicates the weights of similar time series; Q Indicates the write feature sequence; WQ Indicates the weight of the written feature sequence.
6. The method for monitoring the health status of a solid-state drive according to claim 1, characterized in that, The analysis of the correlation between the heat transfer characteristics and the overwriting load, and the abnormal temperature characteristics under the overwriting load include: Calculate the correlation between the heat transfer characteristics and the overwriting load using the Pearson correlation coefficient, and compare the correlation with a preset threshold; If the correlation is greater than or equal to the prediction threshold, it means that the temperature rise rate increases abnormally under the overwriting load, then use the load balancing algorithm to adjust the temperature parameters of each area inside the solid state drive. Otherwise, it means that the temperature rise rate is not affected under the overwriting load.
7. The method for monitoring the health status of a solid-state drive according to claim 6, wherein The calculation formula of the Pearson correlation coefficient is: ; In the formula, r represents the Pearson correlation coefficient; represents the mean value of the heat transfer characteristics; represents the overwriting load; represents the i th heat transfer characteristic; represents the i th overwriting load; represents the weight value; n represents the total number of characteristics.
8. The method for monitoring the health status of a solid-state drive according to claim 1, wherein The establishment of the temperature gradient transfer mechanism based on the temperature rise rate, and combined with the physical structure of the solid state drive and the thermal conductivity of the materials in each area, the heat transfer characteristics of the heat inside the solid state drive are extracted, including: Use finite element analysis software to establish a finite element model of the solid state drive, and the finite element model includes the layout of each area of the solid state drive and the contact interface; Define the physical properties and boundary conditions of each area of the solid state drive in the finite element model, and use the finite element analysis software to perform heat conduction simulation to obtain the temperature distribution of the solid state drive under normal working conditions, and obtain the preliminary temperature field of the solid state drive; Generate the heat transfer paths of the heat inside the solid state drive in each area based on the preliminary temperature field of the solid state drive, use the genetic algorithm to estimate the fitness of each heat transfer path, and select the optimal heat transfer path; Obtain the temperature distribution of each area in the solid state drive under the optimal heat transfer path, and based on the temperature distribution, use the finite element analysis software to update the temperature field of the solid state drive to obtain the heat transfer characteristics of the heat inside the solid state drive.
9. The method for monitoring the health status of a solid-state drive according to claim 8, wherein The use of the genetic algorithm to estimate the fitness of each heat transfer path and select the optimal heat transfer path includes: Encode the heat transfer paths of the solid state drive as individuals to obtain several initial individuals; Define the fitness function, and evaluate the initial individuals based on the fitness function. According to the fitness evaluation results and the predefined screening strategy, select the individuals with high fitness to enter the next generation; Perform crossover on the selected individuals to generate new individuals, and mutate the individuals after crossover to introduce small random changes to avoid falling into local optima; Evaluate the fitness of the new generation of individuals, update the individuals in the population, and iteratively execute the crossover and mutation process until the termination condition is reached and stop to obtain the heat transfer path with the best fitness.
10. A system for monitoring the health status of a solid-state drive, which is used to implement the method for monitoring the health status of a solid-state drive according to any one of claims 1-9, wherein The system includes: A write load feature recognition module, which is used to collect the monitoring data of the solid-state drive, extract write features from the monitoring data and perform fluctuation analysis, and identify the over-write load features of the solid-state drive according to the fluctuation analysis results of the write features; A temperature feature analysis module, which is used to analyze the heat transfer features built in the solid-state drive based on the monitoring data, and analyze the correlation between the heat transfer features and the over-write load, so as to obtain the abnormal temperature features during over-write load; A health status evaluation module, which is used to input the abnormal temperature features into a pre-constructed solid-state drive health status evaluation model, predict the health status of the solid-state drive through the health status evaluation model, and formulate a data migration mechanism based on the health status.
Citation Information
Patent Citations
Solid state drive health status monitoring method and device
CN106528377A
Building health monitoring and evaluation method and system based on physical neural network
CN119249073A
Cited By
High-speed large-capacity information storage system
CN120704609A
Time sequence data partition storage method for high-frequency acquisition system
CN121233596A
Memory particle life prediction-oriented multi-dimensional access load evaluation method and system, electronic equipment and storage medium
CN121523915A