A method and system for health monitoring of a solid-state drive
Through multi-dimensional data acquisition and precise synchronization methods, the health assessment and self-repair of solid-state drives are solved, and the reliability and accuracy of health monitoring in the existing technology is insufficient, and the accurate health assessment and self-repair of solid-state drives are achieved, thereby extending the life of the hard disk.
Patent Information
- Application Number
- CN202510545558.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-28
AI Technical Summary
The existing technology cannot fully reflect the actual health status of solid-state drives, resulting in low reliability and accuracy of health monitoring, and the inability to accurately synchronize multiple data sources, ignoring the interaction between different factors.
By collecting multi-dimensional monitoring data on solid state hard disks, including S.M.A.R.T attribute data, temperature sensor data, programmatic erase delay data and power supply voltage ripple data, time stamp synchronization and electronic tunneling current waveform analysis, charge well health index data is generated, and combined with multivariate degradation modeling and self-repair strategies, accurate health assessment and self-repair are achieved.
It improves the reliability and accuracy of solid-state drive health monitoring, can identify potential degradation trends in the early stage, extend the life of the hard disk, and reduces the cost of hard disk replacement and maintenance caused by misjudgment.
Smart Images

Figure CN120066903B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hard disk monitoring, and particularly to a method and system for health monitoring of a solid-state drive. Background Art
[0002] Early SSD health monitoring methods mainly relied on hardware monitoring interfaces, such as S.M.A.R.T. (Self-Monitoring, Analysis, and Reporting Technology), to monitor the basic health status of the device. These methods mainly detected indicators such as the temperature, power-on count, and number of bad blocks of the SSD. However, in actual applications, they often cannot fully reflect the actual health status of the SSD. With the progress of technology, SSD health monitoring has gradually developed into a multi-dimensional monitoring system. For example, by carefully analyzing real-time data in aspects such as write amplification effect, wear leveling algorithm, and flash memory life, the service life of the SSD can be predicted more accurately. At the same time, with the introduction of artificial intelligence and big data technologies, machine learning-based fault prediction and intelligent diagnosis methods have also become research hotspots. These technologies can continuously analyze the usage history and health data of storage devices to early warn of potential fault risks and improve the reliability and maintainability of SSDs. However, currently, existing technologies usually cannot accurately synchronize multiple data sources, resulting in the loss of data correlation in analysis. At the same time, they only rely on basic health indicators (such as the number of bad blocks) to predict hard disk failures, ignoring the interaction between different factors, thus leading to low reliability and accuracy in the health monitoring of solid-state drives. Summary of the Invention
[0003] Based on this, it is necessary to provide a method and system for health monitoring of a solid-state drive to solve at least one of the above technical problems.
[0004] To achieve the above object, a method for health monitoring of a solid-state drive includes the following steps:
[0005] Step S1: Collect multi-dimensional monitoring data of the storage unit of the solid-state drive to generate standard multi-dimensional monitoring data, where the standard multi-dimensional monitoring data includes S.M.A.R.T attribute data, temperature sensor data, programming / erasing delay data, and supply voltage ripple data;
[0006] Step S2: Use the supply voltage ripple data to synchronize the timestamps of the programming / erasing delay data, and perform electronic tunneling current waveform analysis on the programming / erasing delay data and the supply voltage ripple data according to the synchronized timestamps to generate charge trap health index data;
[0007] Step S3: Extract the number of bad blocks of S.M.A.R.T attribute data; extract the historical temperature distribution of temperature sensor data; perform multivariate degradation modeling based on the number of bad blocks, historical temperature distribution, and charge trap health indicator data to generate health score data;
[0008] Step S4: Based on a preset health threshold, perform abnormal health discrimination on the health score data to obtain an abnormal health discrimination result; make a self-repair strategy decision for the solid-state drive according to the abnormal health discrimination result to generate hard drive self-repair decision data;
[0009] Step S5: Perform visual encoding on the hard drive self-repair decision data to generate a user-readable hard drive health report.
[0010] The present invention realizes refined monitoring of storage units by collecting multi-dimensional data such as S.M.A.R.T attributes, temperature sensors, programming erase latency, and power supply voltage ripple, improves the comprehensiveness of hard disk health management, generates standard multi-dimensional monitoring data, ensures the uniformity of data formats from different sources, facilitates subsequent modeling and analysis, and improves the compatibility and usability of data. By using power supply voltage ripple data to monitor voltage fluctuations and evaluate the impact of power supply quality on the health status of NAND flash memory, it helps to discover potential hardware problems. Through timestamp synchronization, the time consistency of programming erase latency data and power supply voltage ripple data is improved, data deviation is reduced, and analysis accuracy is enhanced. Based on the analysis of the electronic tunneling current waveform, the state of the charge trap is studied in depth, the degradation of NAND flash memory is evaluated, and a more physically meaningful quantitative index is provided for health scoring. Through the change of the electronic tunneling current waveform characteristics, the potential degradation trend of the solid-state drive can be detected earlier, which is more forward-looking than the traditional bad block statistics method. Combining the number of S.M.A.R.T bad blocks, historical temperature distribution, and charge trap health indicators, a multi-variable degradation model is established to improve the predictive ability of health scoring and achieve accurate health assessment. As the data is updated, the health score can be dynamically adjusted to reflect the health status of the solid-state drive over time and avoid misjudgment caused by single-time-point assessment. The historical temperature distribution can be used to analyze the impact of high temperature on NAND degradation, help optimize the data storage environment, and extend the device life. Compared with the traditional experience-based health assessment, this method uses machine learning or statistical models to build a health scoring system based on real data, improving scientificity and objectivity. The health threshold can be optimized according to different application scenarios (such as data centers, high-performance computing, consumer storage) to make abnormal discrimination more in line with specific requirements. By performing anomaly detection on the health scoring data, hard disk failure warning can be achieved, and different levels of alarms can be issued based on the score level, improving maintenance efficiency. The abnormal discrimination results can be used to trigger different levels of self-repair strategies, such as data migration, block replacement, erase optimization, etc., to improve the availability of the hard disk. Based on accurate health assessment, the self-repair strategy can extend the hard disk life, reduce hard disk replacement caused by misjudgment, and reduce maintenance costs. Therefore, the present invention improves the reliability and accuracy of solid-state drive health monitoring through multi-dimensional data collection, precise synchronization, accurate degradation modeling, and intelligent self-repair strategies.
[0011] Preferably, step S1 includes the following steps:
[0012] Step S11: Access the S.M.A.R.T monitoring system of the solid-state drive through an interface to obtain the health status data of the hard disk; extract the S.M.A.R.T attributes of the health status data of the hard disk to obtain S.M.A.R.T attribute data;
[0013] Step S12: Monitor the temperature during the operation of the hard disk using the temperature sensor built in the hard disk to obtain temperature sensor data;
[0014] Step S13: Monitor the erase operation and programming process of the solid-state drive to obtain programming and erase latency data;
[0015] Step S14: Record voltage fluctuation data through the power management module, and perform voltage ripple analysis on the voltage fluctuation data to generate supply voltage ripple data;
[0016] Step S15: Integrate S.M.A.R.T attribute data, temperature sensor data, programming and erase latency data, and supply voltage ripple data into multi-dimensional monitoring data; perform data preprocessing on the multi-dimensional monitoring data to generate standard multi-dimensional monitoring data, where data preprocessing includes data cleaning, data denoising, missing value filling, and data standardization.
[0017] The present invention accesses the S.M.A.R.T monitoring system of a solid-state drive through an interface, can obtain the health status data of the hard drive, and extract S.M.A.R.T attributes. This monitoring method can comprehensively understand the health status of the hard drive, detect potential faults in a timely manner, and improve the reliability of the hard drive. The built-in temperature sensor of the hard drive is used to monitor the operating temperature of the hard drive in real time, which can effectively prevent overheating problems, avoid hardware damage or performance degradation caused by high temperature, and ensure the stability of the hard drive. By monitoring the erase operation and programming process of the hard drive, programming erase delay data is obtained. Programming and erase operations are important indicators of solid-state drives. Monitoring the delay helps identify performance bottlenecks in the read and write operations of the hard drive and detect potential performance degradation at an early stage. The power management module records voltage fluctuation data and performs voltage ripple analysis on the voltage fluctuation, which helps identify voltage instability situations and avoid the impact of power fluctuations on the operation of the hard drive, thereby improving the power adaptability and service life of the hard drive. Integrating the S.M.A.R.T attribute data, temperature sensor data, programming erase delay data, and power supply voltage ripple data forms multi-dimensional monitoring data. This data integration method helps comprehensively evaluate the health status and performance of the hard drive and facilitates the comprehensive analysis of the correlation between different monitoring dimensions. By performing data cleaning, denoising, missing value filling, and data standardization on the multi-dimensional monitoring data, the quality of the data is ensured. These preprocessing steps can eliminate noise and inconsistencies, fill in missing data, improve the accuracy and usability of the data, and provide high-quality input data for subsequent analysis. By comprehensively integrating different data sources (such as S.M.A.R.T attributes, temperature, delay, voltage fluctuations, etc.) and performing standardized preprocessing, the system can accurately evaluate the health status of the hard drive and implement more scientific health management and prediction based on this data. This helps identify potential faults in advance, perform timely maintenance or replacement, thereby extending the service life of the hard drive and reducing the occurrence of unexpected failures. The multi-dimensional monitoring data provides a comprehensive perspective on the performance of the hard drive, which can help developers or technicians better understand the operating status and performance bottlenecks of the hard drive. Analysis based on this data can provide guidance for the optimization, tuning, and problem diagnosis of solid-state drives, thereby improving the overall performance of the hard drive.
[0018] Preferably, the step of synchronizing the time stamps of the programming erase delay data using the power supply voltage ripple data in step S2 includes:
[0019] Performing voltage fluctuation trend analysis on the power supply voltage ripple data to obtain the voltage fluctuation peak value, valley value, and zero crossing point, and using them as the voltage reference time benchmark;
[0020] Extracting the key time points of the power supply voltage ripple data through the voltage reference time benchmark to obtain the power supply voltage ripple time stamp;
[0021] Extract the programming timestamp and the erasure timestamp from the programming erasure delay data; perform a delay time interval calibration on the programming timestamp and the erasure timestamp to obtain the programming erasure calibration timestamp;
[0022] Calculate the time difference between the power supply voltage ripple timestamp and the programming erasure calibration timestamp, and perform a sliding window matching on the power supply voltage ripple timestamp and the programming erasure calibration timestamp through the time difference to generate the optimal time alignment point;
[0023] Based on the optimal time alignment point, perform dynamic time warping on the power supply voltage ripple data and the programming erasure delay data to generate synchronized timestamps.
[0024] Through the voltage fluctuation trend analysis of the power supply voltage ripple data, the present invention can identify the peak value, valley value and zero crossing point of the voltage fluctuation, providing an accurate voltage reference time base for subsequent timestamp synchronization. This analysis method ensures the reliability and stability of the time base of the power supply voltage ripple data. Extract the key time points of the power supply voltage ripple data by using the voltage reference time base, so as to generate accurate power supply voltage ripple timestamps. In addition, the programming timestamp and the erasure timestamp in the programming erasure delay data are also extracted, providing basic data for further calibration. This timestamp extraction method effectively ensures the time consistency between different data sources. By performing a delay time interval calibration on the programming timestamp and the erasure timestamp, the time error in the data can be eliminated, thus ensuring the accuracy of the programming erasure calibration timestamp. This step improves the precision of time synchronization and avoids data inconsistency caused by time delay. By calculating the time difference between the power supply voltage ripple timestamp and the programming erasure calibration timestamp and adopting the sliding window matching technology, the power supply voltage ripple data and the programming erasure delay data can be accurately aligned. This process generates the optimal time alignment point by adaptively adjusting the time difference, ensuring the efficiency and precision of time synchronization. Based on the optimal time alignment point, perform dynamic time warping to effectively synchronize the power supply voltage ripple data and the programming erasure delay data. This step can eliminate the time misalignment caused by different sampling rates, delays or data missing, making the final synchronized timestamp more accurate and ensuring the accuracy of subsequent analysis and processing. By efficiently synchronizing different data sources, the power supply voltage ripple data and the programming erasure delay data can be compared and analyzed with the same time base, which can improve the accuracy of subsequent analysis and reduce the error caused by different time axes, thus improving the data quality. Through timestamp synchronization, the timing consistency between the power supply voltage ripple data and the programming erasure delay data can be ensured, which is crucial for in-depth analysis of the impact of voltage fluctuation on the hard disk performance and other related fault predictions. The precise alignment of the synchronized data will improve the accuracy of overall health monitoring and fault prediction.
[0025] Preferably, in step S2, the electronic tunneling current waveform analysis of the programming and erasing delay data and the supply voltage ripple data according to the synchronized timestamps includes:
[0026] Performing timing matching on the programming and erasing delay data and the supply voltage ripple data according to the synchronized timestamps to generate a timing matching data set, where the timing matching data set includes the programming and erasing delay data and the supply voltage ripple data at the same timestamp;
[0027] Using the supply voltage ripple data, calculating the tunneling current of the solid-state drive NAND cell according to the electronic tunneling equation, where the formula of the electronic tunneling equation is as follows:
[0028] ;
[0029] In the formula, is the tunneling current, is the supply voltage during programming or erasing, A is a constant related to the material characteristics of the NAND cell, and B is an exponential decay factor related to the material characteristics of the NAND cell;
[0030] Performing ultra-high-speed simulation of the current waveform on the tunneling current to obtain the transient tunneling current waveform; extracting the waveform rising edge slope, peak jitter, and charge integration of the transient tunneling current waveform as the original waveform feature data;
[0031] Performing waveform deviation analysis on the supply voltage ripple data according to the original waveform feature data to generate waveform deviation data;
[0032] Performing charge trap health analysis on the transient tunneling current waveform through the waveform deviation data to generate charge trap health index data.
[0033] By performing timing matching on the programmed erase delay data and the power supply voltage ripple data according to the synchronized timestamps, the present invention can ensure the synchronization between different signals, avoid analysis errors caused by timing deviations, and thus improve the accuracy and reliability of data processing. The tunneling current of the solid-state drive NAND cell is calculated using the power supply voltage ripple data and the electron tunneling equation, and the formula includes constants and exponential decay factors related to the material properties of the NAND cell, making the calculation of the tunneling current more accurate and able to truly reflect the current characteristics of the hard disk. By performing ultra-high-speed simulation of the current waveform on the tunneling current, the transient tunneling current waveform is obtained, thereby obtaining more detailed current waveform data, which provides a high-quality waveform basis for subsequent charge trap health analysis and ensures the accuracy of the analysis. The waveform rising edge slope, peak jitter, and charge integral of the transient tunneling current waveform are extracted as the original waveform feature data, which can extract key features from the current waveform, reveal the changes in the charge trap and the current tunneling process, and provide detailed data for further analyzing the health state of the charge trap. Through waveform deviation analysis, the influence of the power supply voltage ripple on the tunneling current waveform can be identified, further revealing the working state of the hard disk and the current waveform deviation. This not only helps to understand the performance of the hard disk in different working environments but also can detect the potential impact of voltage ripple on the hard disk performance. According to the waveform deviation data, charge trap health analysis is performed on the transient tunneling current waveform to generate charge trap health index data, providing a quantitative assessment of the health status of the hard disk. This helps to identify hard disk failures in advance, improve the maintenance efficiency of the hard disk, and reduce the failure rate. By generating the charge trap health index data, real-time monitoring and warning of the hard disk health state can be achieved. This method not only improves the accuracy of fault detection but also provides support for the optimization of solid-state drives, ensuring the stability and durability of the hard disk in high-load environments.
[0034] Preferably, the charge trap health analysis of the transient tunneling current waveform by using the waveform deviation data includes:
[0035] Performing high-frequency feature decomposition on the transient tunneling current waveform to extract well potential spectrum data; performing energy level transition analysis on the transient tunneling current waveform according to the well potential spectrum data to obtain transition feature data;
[0036] Calculating the trap density based on the transition feature data to generate trap density distribution data;
[0037] Performing spatio-temporal correlation analysis on the waveform deviation data and the trap density distribution data to obtain correlation feature data;
[0038] Performing multi-dimensional degradation mode recognition on the correlation feature data to generate degradation mode data, where the multi-dimensional degradation mode recognition includes linear degradation mode recognition and non-linear degradation mode recognition;
[0039] Quantify the charge trap health based on the degradation mode data for the transient tunneling current waveform to generate charge trap health index data.
[0040] In the present invention, by performing high-frequency feature decomposition on the transient tunneling current waveform and extracting the well potential energy spectrum data, the energy state of the charge trap can be accurately identified, providing an accurate data basis for subsequent health analysis, and enhancing the accuracy and reliability of the charge trap health assessment. Based on the well potential energy spectrum data for energy level transition analysis, it is possible to deeply understand the internal dynamic process of the charge trap, reveal the behavior of charges during tunneling, and provide rich dynamic characteristic information for analyzing the health status of the charge trap. By calculating the trap density and generating trap density distribution data, the distribution of traps in the charge trap can be evaluated, which helps to judge the influence of the charge trap on current tunneling and further reveals the health level of the charge trap under different working conditions. By performing spatio-temporal correlation analysis on the waveform deviation data and the trap density distribution data, a more accurate spatio-temporal model can be established to capture the variation law of the charge trap health status in time and space, improving the comprehensiveness and timeliness of the health assessment. Through linear and non-linear degradation mode recognition, the degradation process of the charge trap can be comprehensively identified. Whether it is a steady degradation mode or a complex non-linear degradation behavior, it can be accurately captured and identified to ensure a comprehensive grasp of the charge trap health status. Based on the degradation mode data for charge trap health quantification, the generated charge trap health index data has high precision and reliability, and can be used as the core index for hard disk health assessment and predictive maintenance to ensure the stability and long-term reliability of the hard disk during use.
[0041] Preferably, step S3 includes the following steps:
[0042] Step S31: Extract the number of bad blocks of the S.M.A.R.T attribute data;
[0043] Step S32: Extract the historical temperature distribution of the temperature sensor data;
[0044] Step S33: Perform Weibull degradation modeling on the number of bad blocks and the historical temperature distribution to generate degradation probability distribution data; perform Markov state transition analysis on the degradation probability distribution data to generate health state transition matrix data; perform steady-state probability calculation on the health state transition matrix data to generate dynamic health score data;
[0045] Step S34: Perform FFT spectrum analysis on the charge trap health index data to generate main frequency energy value data; use the multi-dimensional weighted health score fusion formula to perform weighted fusion on the dynamic health score data and the main frequency energy value data to generate health score data.
[0046] The present invention comprehensively evaluates the health status of a device through multiple data sources such as S.M.A.R.T bad block data, temperature history data, and charge trap health indicators, improving the comprehensiveness and accuracy of the evaluation. By using Weibull degradation modeling and Markov state transition analysis, the dynamic tracking of the device's health status is achieved, and an accurate health score is provided through steady-state probability calculation, which helps to early warn of potential failures. By extracting the main frequency energy characteristics of the charge trap through FFT spectrum analysis, the health score not only depends on time-domain data but also combines frequency-domain characteristics, improving the sensitivity to the device's health status. Using a multi-dimensional weighted health score fusion formula, multiple health indicators are fused into a comprehensive health score, ensuring that the evaluation result is more representative and providing a reliable basis for maintenance decisions. This method can adjust the health score according to real-time data, is applicable to long-term monitoring and predicting the degradation trend of the device, thereby supporting preventive maintenance and extending the service life of the device.
[0047] Preferably, the multi-dimensional weighted health score fusion formula in step S34 is specifically as follows:
[0048] ;
[0049] In the formula, is the final health score, is the dynamic health score at time t, is the main frequency energy value at time t, is the i-th additional influence factor at time t, is the weight of the dynamic health score, is the weight of the main frequency energy value, is the weight of the i-th additional influence factor, is the time decay factor, is the regularization coefficient, is the sensitivity coefficient, is the change rate of the health score at time t, T is the time upper limit, n is the number of additional influence factors, is the decay rate parameter.
[0050] The present invention analyzes and integrates a multi-dimensional weighted health score fusion formula, the core idea of which is dynamic weight fusion + time decay + regularization balance. Specifically, this formula integrates multiple health-related factors and ensures the rationality and stability of the health score through mechanisms such as time decay and regularization. The dynamic health score in the formula reflects the actual operating state of the hard disk at a certain moment, such as temperature, bad sectors, read / write latency, IOPS, etc. The immediate impact of these indicators on the health state of the hard disk is very important, so higher weights should be assigned. The operating frequency or frequency domain energy characteristics of the hard disk (such as vibration signals, head movement, etc.) can also reflect the health state of the hard disk. Abnormal frequencies indicate mechanical failures or other problems in the hard disk. Additional influencing factors include the usage time of the hard disk, workload, frequency of read / write requests, etc. Long-term operation or high-load work will affect the health of the hard disk. The time decay factor ensures that the score is mainly based on the recent state, and the influence of long-term data gradually weakens. The health state of the hard disk changes over time, so the decay factor can help reflect the current health condition. The change rate of the health score reflects the change range of the health state of the hard disk, and the regularization term is used to control the smoothness of the hard disk score and avoid drastic fluctuations in the score caused by instantaneous fluctuations or abnormal data. When using the conventional multi-dimensional weighted health score fusion formula in the art, the health score of the hard disk can be obtained. By applying the multi-dimensional weighted health score fusion formula provided by the present invention, the health score of the hard disk can be calculated more accurately. The formula synthesizes the real-time health state of the hard disk (such as temperature, number of bad sectors, read / write speed, etc.), frequency domain characteristics, and additional factors such as workload and usage time to provide a more comprehensive health score. Through the time decay mechanism, it is ensured that the score pays more attention to the current health state of the hard disk and ensures that the health assessment is real-time and effective. The regularization mechanism avoids the instability of the score due to instantaneous fluctuations (such as load changes within a short period of time) and provides a more reliable health prediction. The weights of each factor (such as , and ) can be flexibly adjusted according to the characteristics of the specific hard disk or the working environment to meet the health assessment needs of different types of hard disks.
[0051] Preferably, step S4 includes the following steps:
[0052] Step S41: Based on a preset health threshold, perform abnormal health discrimination on the health score data. When the health score data is greater than or equal to the preset health threshold, no processing is performed;
[0053] Step S42: When the health score data is less than the preset health threshold, mark the corresponding health score data as the abnormal health discrimination result;
[0054] Step S43: Calculate the error correction iteration times of the low-density parity-check code according to the abnormal health discrimination result and adjust the iteration times to generate iteration time adjustment data; Write voltage compensation parameters to the solid-state drive using the charge trap health index data to obtain voltage compensation write parameter data;
[0055] Step S44: Integrate the iteration time adjustment data and the voltage compensation write parameter data, and make a self-repair strategy decision for the solid-state drive to generate hard disk self-repair decision data.
[0056] Through classifying the health score data based on a preset health threshold, the present invention can timely identify devices with poor health status, ensuring a quick response when problems occur. For the case where the health score data is less than the threshold, the system can automatically mark it as abnormal, reducing the risk of human judgment errors. By using the error correction mechanism and iterative adjustment of the low-density parity-check code, the system can automatically correct errors occurring in the data transmission or storage process, improving the system's self-repair ability and reducing the risk of data corruption. By performing voltage compensation on the charge trap health index, the working state of the solid-state drive can be optimized, reducing the negative impact of voltage fluctuations on the hard disk performance, thereby enhancing the stability and long-term usage performance of the solid-state drive. After integrating the data of the iteration time adjustment and the voltage compensation write parameters, the system can make a self-repair strategy decision, automatically perform fault repair or optimization operations, reduce manual intervention, improve the device operation efficiency, and extend the service life. Through abnormal health discrimination and self-repair strategies, early warning and maintenance prevention of the device health status can be achieved, reducing the occurrence of sudden failures and improving the device reliability and working efficiency.
[0057] Preferably, the writing of voltage compensation parameters to the solid-state drive using the charge trap health index data in step S43 includes:
[0058] Extract the ripple characteristics of the waveform deviation data using the charge trap health index data to obtain reverse ripple reference data;
[0059] Perform amplitude compensation calculation on the reverse ripple reference data to generate amplitude compensation data;
[0060] Perform phase matching analysis according to the amplitude compensation data to obtain phase matching data;
[0061] Calculate the NAND word line drive parameters of the solid-state drive based on the phase matching data to generate drive parameter data;
[0062] Obtain injection control data by designing the injection timing of the solid-state drive using the drive parameter data;
[0063] Write voltage compensation parameters to the solid-state drive through the injection control data to obtain voltage compensation write parameter data.
[0064] By extracting the ripple characteristics of the waveform deviation data and obtaining the reverse ripple reference data, the present invention can accurately identify the ripple characteristics in the voltage signal, thereby effectively handling the impact of voltage fluctuations on the performance of the solid-state drive and ensuring power stability. Through the amplitude compensation calculation of the reverse ripple reference data, the system can adjust the amplitude of the voltage, reduce the potential damage of unstable voltage fluctuations to the solid-state drive components, and thus improve the reliability and service life of the hard disk. Through the phase matching analysis, the phase of the voltage waveform can be optimized to ensure that the voltage signal matches the control circuit in the hard disk, reduce signal interference, and improve the working efficiency and stability of the hard disk. According to the amplitude compensation data and the phase matching data, the NAND word line drive parameters are calculated, so that when the hard disk performs read and write operations, the voltage control is more accurate, effectively avoiding the performance degradation or failure risk caused by voltage mismatch. By injecting the timing design and injecting the control data, the timing of the hard disk can be optimized and adjusted to ensure that the hard disk maintains a good voltage supply under high-load or high-frequency read and write conditions, thereby improving the performance and durability of the hard disk. The entire voltage compensation process combines amplitude compensation, phase matching, and timing design to ensure that the solid-state drive can obtain stable voltage support in any working state, which not only improves the health status of the hard disk but also enhances the self-repair ability of the hard disk to automatically correct the problems caused by voltage fluctuations.
[0065] In this specification, a health monitoring system for a solid-state drive is provided for performing the above-mentioned health monitoring method for the solid-state drive. The health monitoring system for the solid-state drive includes:
[0066] A data acquisition module for collecting multi-dimensional monitoring data of the storage units of the solid-state drive to generate standard multi-dimensional monitoring data, where the standard multi-dimensional monitoring data includes S.M.A.R.T attribute data, temperature sensor data, programming and erasing delay data, and power supply voltage ripple data;
[0067] A health index analysis module for synchronizing the time stamps of the programming and erasing delay data using the power supply voltage ripple data, and performing an electronic tunneling current waveform analysis on the programming and erasing delay data and the power supply voltage ripple data according to the synchronized time stamps to generate charge trap health index data;
[0068] A health scoring module for extracting the number of bad blocks in the S.M.A.R.T attribute data; extracting the historical temperature distribution of the temperature sensor data; performing a multi-variable degradation modeling based on the number of bad blocks, the historical temperature distribution, and the charge trap health index data to generate health scoring data;
[0069] An adaptive repair module, which is used to perform abnormal health discrimination on health score data based on a preset health threshold to obtain an abnormal health discrimination result; make a self-repair strategy decision for the solid-state drive according to the abnormal health discrimination result, and generate hard disk self-repair decision data;
[0070] A visualization module, which is used to perform visual encoding on the hard disk self-repair decision data to generate a user-readable hard disk health report.
[0071] The beneficial effects of the present invention are as follows: The data acquisition module generates standard multi-dimensional monitoring data including S.M.A.R.T attribute data, temperature sensor data, programming erase latency data, and power supply voltage ripple data by performing multi-dimensional monitoring on the storage units of the solid-state drive. The acquisition of such multi-dimensional data can comprehensively and real-time monitor the operating status of the hard disk, providing sufficient data support for subsequent health analysis. The health index analysis module synchronizes the timestamp of the programming erase latency data by using the power supply voltage ripple data, and performs electronic tunneling current waveform analysis on both based on the synchronized timestamp to generate charge trap health index data. This step can accurately capture the relationship between the programming erase operation and voltage fluctuation, thereby providing a more accurate assessment of the charge trap health status. The health scoring module combines the number of bad blocks, historical temperature distribution, and charge trap health index data, and uses a multi-variable degradation modeling method to generate health score data. Through multi-dimensional degradation modeling, the system can more scientifically evaluate the health status of the hard disk, improve the accuracy of the health score, and effectively predict the failure risk of the hard disk. The adaptive repair module performs abnormal health discrimination on the health score data based on a preset health threshold, and formulates a self-repair strategy for the solid-state drive according to the discrimination result. This module can automatically adjust the strategy according to the hard disk health status, realize the self-repair of the hard disk, reduce the occurrence of hard disk failures, and lower the maintenance cost. The visualization module encodes the hard disk self-repair decision data and generates a user-readable hard disk health report. This report can intuitively display the health status and repair decision of the hard disk, facilitating users to understand the real-time status and maintenance requirements of the hard disk. Therefore, the present invention improves the reliability and accuracy of the health monitoring of the solid-state drive through multi-dimensional data acquisition, precise synchronization, accurate degradation modeling, and intelligent self-repair strategies. Description of the Drawings
[0072] Figure 1 It is a schematic diagram of the step flow of a health monitoring method for a solid-state drive;
[0073] Figure 2 It is Figure 1 A detailed implementation step flow diagram of step S3 in
[0074] Figure 3 It is Figure 1 A detailed implementation step flow diagram of step S4 in
[0075] The realization, functional features and advantages of the present invention will be further described in conjunction with the embodiments with reference to the accompanying drawings. Specific embodiments
[0076] The technical method of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative work fall within the scope of protection of the present invention.
[0077] In addition, the accompanying drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings represent the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.
[0078] It should be understood that although terms such as "first" and "second" may be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit may be referred to as the second unit, and similarly the second unit may be referred to as the first unit. The term "and / or" used herein includes any and all combinations of one or more of the listed related items.
[0079] To achieve the above object, please refer to Figures 1 to 3 , a method for health monitoring of a solid-state drive, the method comprising the following steps:
[0080] Step S1: Collect multi-dimensional monitoring data of the storage units of the solid-state drive to generate standard multi-dimensional monitoring data, where the standard multi-dimensional monitoring data includes S.M.A.R.T attribute data, temperature sensor data, program erase delay data, and supply voltage ripple data;
[0081] Step S2: Synchronize the time stamps of the program erase delay data using the supply voltage ripple data, and perform electron tunneling current waveform analysis on the program erase delay data and the supply voltage ripple data according to the synchronized time stamps to generate charge trap health index data;
[0082] Step S3: Extract the number of bad blocks of S.M.A.R.T attribute data; extract the historical temperature distribution of the temperature sensor data; perform multivariate degradation modeling based on the number of bad blocks, the historical temperature distribution, and the charge trap health indicator data to generate health score data;
[0083] Step S4: Based on a preset health threshold, perform abnormal health discrimination on the health score data to obtain an abnormal health discrimination result; make a self-repair strategy decision for the solid-state drive according to the abnormal health discrimination result to generate hard drive self-repair decision data;
[0084] Step S5: Perform visual encoding on the hard drive self-repair decision data to generate a user-readable hard drive health report.
[0085] The present invention realizes refined monitoring of storage units by collecting multi-dimensional data such as S.M.A.R.T attributes, temperature sensors, programming erase latency, and power supply voltage ripple, improves the comprehensiveness of hard disk health management, generates standard multi-dimensional monitoring data, ensures the uniformity of data formats from different sources, facilitates subsequent modeling analysis, and improves the compatibility and usability of data. Using the power supply voltage ripple data to monitor the voltage fluctuation situation and evaluate the impact of power supply quality on the health status of NAND flash memory helps to discover potential hardware problems. Through timestamp synchronization, the time consistency between the programming erase latency data and the power supply voltage ripple data is improved, data deviation is reduced, and the analysis accuracy is enhanced. Based on the analysis of the electronic tunneling current waveform, the state of the charge trap is studied in depth, the degradation of NAND flash memory is evaluated, and a more physically meaningful quantitative index is provided for the health score. Through the change of the electronic tunneling current waveform characteristics, the potential degradation trend of the solid-state drive can be detected earlier, which is more forward-looking than the traditional bad block statistics method. Combining the S.M.A.R.T bad block count, historical temperature distribution, and charge trap health index, a multi-variable degradation model is established to improve the prediction ability of the health score and achieve accurate health assessment. As the data is updated, the health score can be dynamically adjusted to reflect the health status of the solid-state drive over time and avoid misjudgment caused by the evaluation at a single time point. The historical temperature distribution can be used to analyze the impact of high temperature on NAND degradation, help optimize the data storage environment, and extend the device life. Compared with the traditional experience-based health assessment, this method uses machine learning or statistical models to construct a health score system based on real data, improving the scientificity and objectivity. The health threshold can be optimized according to different application scenarios (such as data centers, high-performance computing, consumer storage) to make the anomaly discrimination more in line with specific requirements. By performing anomaly detection on the health score data, hard disk failure early warning can be achieved, and different levels of alarms can be issued based on the score level to improve the maintenance efficiency. The anomaly discrimination result can be used to trigger different levels of self-repair strategies, such as data migration, block replacement, erase optimization, etc., to improve the availability of the hard disk. Based on accurate health assessment, the self-repair strategy can extend the hard disk life, reduce the hard disk replacement caused by misjudgment, and reduce the maintenance cost. Therefore, the present invention improves the reliability and accuracy of the health monitoring of solid-state drives through multi-dimensional data collection, precise synchronization, accurate degradation modeling, and intelligent self-repair strategies.
[0086] In an embodiment of the present invention, with reference to Figure 1 as shown, it is a schematic diagram of the step flow of a method for monitoring the health of a solid-state drive according to the present invention. In this example, the method for monitoring the health of a solid-state drive includes the following steps:
[0087] Step S1: Collect multi-dimensional monitoring data for the storage units of the solid-state drive to generate standard multi-dimensional monitoring data, where the standard multi-dimensional monitoring data includes S.M.A.R.T attribute data, temperature sensor data, programming / erasing latency data, and supply voltage ripple data;
[0088] In the embodiments of the present invention, the real-time acquisition of various types of monitoring data is ensured by using a solid-state drive (SSD) that supports such acquisition. A hard disk monitoring tool and a temperature sensor that support the S.M.A.R.T (Self-Monitoring, Analysis, and Reporting Technology) protocol are selected, and the hardware that supports the acquisition of programming erase latency data and power supply voltage ripple data at the solid-state drive controller level is also selected. A suitable hard disk working environment is chosen, and the working load of the hard disk (such as read / write frequency, data transfer rate, etc.) is set to ensure the reliability and representativeness of the monitoring data. Tools that support the S.M.A.R.T protocol (such as smartctl, CrystalDiskInfo, etc.) are used to regularly extract the S.M.A.R.T attribute data of the hard disk, which includes the health status of the hard disk, read errors, write errors, power-on duration, temperature information, bad track count, etc. According to the usage situation and acquisition objectives of the hard disk, an appropriate acquisition frequency is selected (for example, once per minute or once per hour). Modern solid-state drives usually come with built-in temperature sensors. The working temperature data of the hard disk is obtained in real time through the drivers or monitoring tools provided by the hard disk manufacturer (such as CrystalDiskInfo, HWMonitor, etc.). If the hard disk does not have a built-in temperature sensor, an external temperature sensor (such as the digital temperature sensor DS18B20, etc.) can be used to collect the temperature of the hard disk enclosure to provide ambient temperature data. According to the working state of the hard disk, temperature data is obtained regularly (for example, once per second, once per minute, etc.). The write and erase operations of the hard disk are monitored. Especially during the erase block and write operations, the operation latency data is obtained through the hard disk controller, which can be captured through specific SSD monitoring software or the API provided by the solid-state drive manufacturer. When the hard disk performs write and erase operations, the operation latency is monitored in real time, and the latency value is recorded. A digital oscilloscope, a power supply voltage monitoring tool, or the power management system of the hard disk driver is used to measure the fluctuation of the power supply voltage in real time. These tools can capture the ripple data of the power supply voltage, especially when the hard disk is under high load. According to the load situation of the hard disk power supply and the power frequency, an appropriate acquisition frequency is selected. Generally, it is recommended to collect data once per second, or collect data at specific moments through trigger conditions. The S.M.A.R.T attribute data, temperature sensor data, programming erase latency data, and power supply voltage ripple data are integrated into a standard multi-dimensional monitoring data format (such as CSV, JSON, etc.). Various types of data can be aligned through timestamps for further analysis. The collected data is cleaned to remove invalid values, error data, and abnormal data to ensure the quality and accuracy of the data.
[0089] Step S2: Synchronize the programming / erasing delay data with a timestamp using the supply voltage ripple data, and perform an analysis of the electronic tunneling current waveform on the programming / erasing delay data and the supply voltage ripple data based on the synchronized timestamp to generate charge trap health indicator data;
[0090] In the embodiments of the present invention, by ensuring that the timestamps of the power supply voltage ripple data and the programming / erasing delay data are accurate, these data should have timestamp information, which can usually be accurately marked through the clock of the acquisition hardware or an external synchronization module. If the data is out of sync, it is necessary to align the timestamps of the two acquired data sets (power supply voltage ripple data and programming / erasing delay data). Common timestamp synchronization methods include time window matching, interpolation method, etc. Match the timestamps of the power supply voltage ripple data and the programming / erasing delay data. In some cases, if the time intervals of the two data are not exactly aligned, linear interpolation or the nearest neighbor method can be used for synchronization alignment. If the data sampling frequencies are different, interpolation processing can be performed as needed to ensure that each data point corresponds to an accurate timestamp. In the flash memory storage unit of a solid-state drive, the electron tunneling current refers to the phenomenon of electrons tunneling through the gate of the storage unit under the condition of an applied electric field. Both voltage ripple and erase / write operations will affect this current, thereby affecting the state of the storage unit and further leading to a decline in the performance of the storage unit. There is a correlation between the change in voltage ripple and the programming / erasing delay. The power supply voltage fluctuation will affect the current, and thus affect the electron tunneling behavior. This process will form a specific current waveform for evaluating the health status of the storage unit. According to the synchronized power supply voltage ripple data and programming / erasing delay data, extract the current waveform characteristics during each programming and erasing operation. At this time, a detailed analysis of the time response of the waveform is required, including the current peak value, current rise time, fall time, etc. Based on the current characteristics during voltage ripple and programming / erasing, establish an electron tunneling current model. This model can be based on classical physical principles (such as tunneling effect theory) or be fitted through experimental data. Use the electron tunneling current waveform analysis method to perform time-domain analysis on the current waveform during each erase / write process, including the fluctuation range of the current, the time difference between the peak and valley of the wave, etc. Based on the waveform characteristics, perform spectral analysis (such as FFT, fast Fourier transform) to identify the influence of high-frequency noise and voltage ripple on the current waveform. A charge trap is a place in a flash memory cell where charges are stored. The process of charge capture and release determines the health status of the storage unit. By analyzing the electron tunneling current waveform, the state of the charge trap can be obtained. If the charge trap is frequently affected, problems such as data loss and programming / erasing failure will occur. In this step, the health status of the charge trap will be comprehensively analyzed through features such as the change in the current waveform, the stability of the waveform, and the change in the instantaneous current. According to the change in the electron tunneling current waveform, extract the features related to the health of the charge trap, such as the stability of the current waveform, the change in the peak current, the rise and fall slopes, etc. If the waveform shows instability or the current peak value is too high, it indicates that there is a problem with the charge trap. Use the charge trap characteristic data (such as the spectral analysis of the waveform, the average current value, the fluctuation range, etc.) to calculate the Charge Trap Health Index.Based on the charge trap health indicator obtained through calculation, the health of the charge traps in the solid-state drive is evaluated. The lower the evaluation value, the worse the health state of the charge traps and the lower the durability of the storage cells.
[0091] Step S3: Extract the number of bad blocks in the S.M.A.R.T attribute data; extract the historical temperature distribution of the temperature sensor data; perform multivariate degradation modeling based on the number of bad blocks, the historical temperature distribution, and the charge trap health indicator data to generate health score data;
[0092] In the embodiments of the present invention, information related to the number of bad blocks of storage units is extracted from the S.M.A.R.T. (Self-Monitoring, Analysis, and Reporting Technology) attribute data of a solid-state drive. The S.M.A.R.T. attributes provide key data on the health status of the hard disk, including the number of bad blocks (Reallocated Sector Count), etc. The S.M.A.R.T. attribute data is obtained through the health monitoring interface of the hard disk drive. This data can usually be obtained through hard disk diagnostic tools (such as CrystalDiskInfo, etc.) or the command-line interface of the operating system (such as the smartctl command in Linux). The attribute of "Reallocated Sector Count" is extracted from the S.M.A.R.T. data, which represents the number of bad blocks that have been reallocated during the use of the hard disk. An increase in this number usually indicates a gradual deterioration in the health of the hard disk. The temperature sensor data of the hard disk is obtained. This data is usually included in the S.M.A.R.T. attribute data or can be obtained through dedicated software provided by the hard disk manufacturer. Historical temperature data helps to evaluate the thermal management of the hard disk during operation, because high temperatures usually lead to a decline in hard disk performance and accelerate degradation. The historical temperature distribution of the hard disk is obtained from the temperature sensor. These data are usually stored in the form of a time series, representing the operating temperature of the hard disk at different time points. It usually includes the maximum temperature, minimum temperature, and average temperature. The historical temperature distribution data of the hard disk is extracted and statistically analyzed (such as maximum value, minimum value, average value, standard deviation, etc.) to understand whether the hard disk has overheated during long-term operation. The charge trap health indicator is a key indicator obtained through the analysis of the electron tunneling current waveform, indicating the health status of the charge traps (Charge Trap) in the flash memory cells. Usually, the decline of the charge traps will affect the performance and lifespan of the hard disk. The charge trap health indicator data is extracted according to the method in step S2, usually including the status evaluation of the charge traps (such as the charge trap health index), reflecting the degree of decline of the charge traps. The health status of the charge traps is evaluated to determine whether the charge traps have excessive damage or decline. This indicator is closely related to the durability and storage performance of the hard disk. Multivariate degradation modeling is used to evaluate the overall health status of the hard disk by combining multiple factors (the number of bad blocks, historical temperature distribution, and charge trap health indicator). The goal of degradation modeling is to establish a comprehensive model that can predict the future health status of the hard disk based on these input data. A method suitable for multivariate degradation modeling is selected, such as regression analysis, support vector machine (SVM), random forest, neural network, etc. The appropriate modeling method is selected according to the complexity and characteristics of the data. The regression model is suitable for linear relationships, while support vector machines and neural networks are suitable for complex non-linear relationships.Normalize the data of the number of bad blocks, temperature distribution, and charge trap health metrics to eliminate the influence between different dimensions and magnitudes. For example, the number of bad blocks, temperature range, and charge trap health index can be standardized to the interval from 0 to 1. Extract meaningful features such as the growth trend of the number of bad blocks, the fluctuation range of historical temperatures, and the decay rate of charge trap health metrics. These features will be used as input features for the model. Input the processed data into the selected model for training. Use historical data for model training and use cross-validation to ensure the generalization ability of the model. Use the test set to validate the prediction ability of the model and adjust the model parameters to optimize the performance. If a neural network is used, techniques such as early stopping can be used to prevent overfitting. Through multivariate degradation modeling, generate a health score data, which represents the current health state of the hard disk and predicts its future lifespan. The health score can be a comprehensive index, usually ranging from 0 to 100, and a higher value indicates a healthier hard disk.
[0093] Step S4: Based on a preset health threshold, perform abnormal health discrimination on the health score data to obtain an abnormal health discrimination result; make a self-repair strategy decision for the solid-state drive according to the abnormal health discrimination result, and generate hard disk self-repair decision data;
[0094] In the embodiments of the present invention, multiple threshold levels are set according to the health scoring model of the hard disk (for example, above 90 is healthy, 70 - 90 is a warning, and below 70 is dangerous). These thresholds can be set based on the manufacturer's recommendations, historical data, or the user's usage requirements. According to the health scoring data of the hard disk, it is compared with the preset health threshold. If the health score is lower than the set threshold, it is determined to be in an abnormal health state. According to the comparison result of the score and the threshold, an abnormal health discrimination result is generated. If the health score is in the "abnormal" or "warning" state, subsequent self - repair decisions are triggered. If the health score of the hard disk is normal (H≥90), no repair measures are required, but monitoring can be carried out. If the health score of the hard disk is in the warning range (70≤H<90), it is recommended to regularly check the hard disk, optimize performance, and storage allocation to prevent potential failures. If the health score of the hard disk is abnormal (H<70), self - repair strategies need to be taken, such as restoring the performance of the hard disk by repairing bad blocks, optimizing the data storage method, or performing firmware upgrades. For hard disks in an abnormal state, first, bad block marking and repair operations can be performed through S.M.A.R.T. tools or the repair tools built into the hard disk, and bad blocks are re - allocated. Adjust the load of the storage unit to reduce the read - write pressure in some high - load areas to prevent further damage to the hard disk. If the temperature history shows that the hard disk is operating at a high temperature, temperature control optimization can be carried out, and it is recommended to improve the heat dissipation capacity of the hard disk or reduce the operating load. If the power supply voltage ripple data shows abnormal fluctuations, repair measures can be taken, such as optimizing power management, increasing battery power, or using a voltage - stabilizing power supply. During the self - repair process, the firmware version of the hard disk can be checked to see if there are new optimization and repair patches available for upgrade. According to the evaluation result of the health state, the system will trigger the corresponding repair process and monitor the health changes of the hard disk during the repair process to ensure the effectiveness of the repair measures. After completing the health score discrimination and making self - repair decisions, self - repair decision data of the hard disk is generated, recording each step of the repair process, the repair status, and the health status after repair. The self - repair decision data can be used as subsequent decision - making support and maintenance reference.
[0095] Step S5: Visually encode the self - repair decision data of the hard disk to generate a user - readable hard disk health report.
[0096] In the embodiments of the present invention, by collecting the hard disk self-repair decision data generated in the previous steps, information including health scores, repair strategies, repair results, repair histories, and subsequent operation suggestions is ensured. Ensure that the data contains clear chronological records, such as the timestamps of each health score change, the repair measures taken, the health score after repair, etc. Design a user-friendly hard disk health report structure so that users can intuitively understand the health status of the hard disk and the repair measures taken. The report content should be concise and easy to read. The beginning of the report should include basic information about the hard disk, such as the hard disk model, serial number, total storage space, health score, etc. Display the trend of the hard disk health score, for example, present the change of the health score in a chart. If the health score is in a warning or abnormal state, it should be clearly marked and the corresponding explanation should be provided. Describe the repair history of the hard disk, the specific repair measures and their effects. The detailed steps of the repair can be shown using a timeline or a table. Based on the repair history, give subsequent operation suggestions for the hard disk, such as regular inspections, whether to continue using, or whether to replace the hard disk. A summary of the diagnosis of the hard disk health status, including the analysis of the current health state and future preventive measures.
[0097] Preferably, step S1 includes the following steps:
[0098] Step S11: Access the S.M.A.R.T monitoring system of the solid-state drive through an interface to obtain the health status data of the hard disk; extract the S.M.A.R.T attributes of the health status data of the hard disk to obtain S.M.A.R.T attribute data;
[0099] Step S12: Use the temperature sensor built in the hard disk to monitor the temperature during the operation of the hard disk to obtain temperature sensor data;
[0100] Step S13: Monitor the erase operation and programming process of the solid-state drive to obtain programming erase delay data;
[0101] Step S14: Record the voltage fluctuation data through the power management module, and perform voltage ripple analysis on the voltage fluctuation data to generate supply voltage ripple data;
[0102] Step S15: Integrate the S.M.A.R.T attribute data, temperature sensor data, programming erase delay data, and supply voltage ripple data into multi-dimensional monitoring data; perform data preprocessing on the multi-dimensional monitoring data to generate standard multi-dimensional monitoring data, where the data preprocessing includes data cleaning, data denoising, missing value filling, and data standardization.
[0103] In the embodiments of the present invention, it is connected to a solid-state drive through SATA, PCIe or NVMe interfaces to access the S.M.A.R.T (Self-Monitoring, Analysis, and Reporting Technology) monitoring system of the hard disk. The S.M.A.R.T data is accessed through standard protocols (such as S.M.A.R.T commands) using the hard disk management tool of the operating system or dedicated software (such as the smartctl tool). The health status information of the hard disk is obtained through commands, including S.M.A.R.T attribute data such as the operating condition of the hard disk, error logs, wear level, bad block count, operating temperature, read / write error rate, etc. The key S.M.A.R.T attributes of the hard disk health status are extracted, such as the number of bad blocks, read / write error rate, reallocated sector count, etc. The S.M.A.R.T attribute data related to the hard disk health status (such as the number of read / write errors, running time, temperature, number of bad blocks, etc.) is obtained and stored. The hard disk is usually equipped with a temperature sensor inside, which can monitor the real-time temperature of the hard disk during operation. The operating temperature data of the current hard disk is obtained through the sensor interface built into the hard disk. The temperature sensor data is collected regularly to obtain the temperature change of the hard disk under different operating conditions. The temperature data is read through the interface provided by the operating system or the hard disk driver to obtain the real-time temperature, maximum temperature and historical temperature data. The erasure and programming operations are monitored by the solid-state drive controller, and the relevant latency data is recorded. The latency time during the erasure operation and programming process is obtained using the hard disk driver or the tools built into the hard disk. The erasure and programming processes are continuously monitored, and the latency time of each erasure and programming cycle is recorded. The performance of the erasure and programming operations is analyzed based on the data to generate relevant latency data and evaluate the hard disk write performance. The voltage fluctuation of the power input is monitored in real time through the power management module of the hard disk. The power supply voltage fluctuation data is recorded, such as the peak value, valley value and fluctuation frequency of the voltage, etc. The spectrum analysis is performed on the voltage fluctuation data to identify the power noise, interference and frequency distribution of the voltage fluctuation, and the power supply voltage ripple data is generated to characterize the stability of the hard disk power supply. The S.M.A.R.T attribute data, temperature sensor data, programming and erasure latency data, and power supply voltage ripple data are integrated according to the time stamp to form a multi-dimensional monitoring data set. Each data record should include multi-dimensional information such as time stamp, health status of the hard disk, temperature, programming and erasure latency, power supply voltage ripple, etc. The invalid or incorrect data is removed, and the outliers are excluded, such as missing values, noise data, etc. The signal smoothing technology (such as moving average filtering, low-pass filtering, etc.) is used to denoise the collected sensor data. The missing part of the data caused by certain faults or sensor problems is filled using interpolation methods (such as linear interpolation, Lagrange interpolation, etc.). The data is standardized so that different monitoring data (such as temperature, erasure latency, voltage fluctuation, etc.) can be compared on the same scale, and the multi-dimensional monitoring data is obtained.
[0104] Preferably, the timestamp synchronization of the programming and erasing delay data using the supply voltage ripple data in step S2 includes:
[0105] Performing voltage fluctuation trend analysis on the supply voltage ripple data to obtain the voltage fluctuation peak value, valley value, and zero crossing point, and using them as the voltage reference time benchmark;
[0106] Extracting the key time points of the supply voltage ripple data through the voltage reference time benchmark to obtain the supply voltage ripple timestamp;
[0107] Extracting the programming timestamp and erasing timestamp in the programming and erasing delay data; calibrating the delay time interval of the programming timestamp and erasing timestamp to obtain the programmed and erased calibration timestamp;
[0108] Calculating the time difference between the supply voltage ripple timestamp and the programmed and erased calibration timestamp, and performing sliding window matching on the supply voltage ripple timestamp and the programmed and erased calibration timestamp through the time difference to generate the best time alignment point;
[0109] Based on the best time alignment point, performing dynamic time warping on the supply voltage ripple data and the programming and erasing delay data to generate synchronized timestamps.
[0110] In the embodiments of the present invention, waveform trend analysis is performed on the power supply voltage ripple data by using time series data analysis methods (such as smoothing filtering or high-frequency component filtering) to identify the rules of voltage fluctuations. By finding the maximum points in the voltage signal, the peak values of the voltage fluctuations are obtained. By finding the minimum points in the voltage signal, the valley values of the voltage fluctuations are obtained. By detecting the zero-crossing points in the voltage fluctuation signal, the periodic characteristics of the voltage waveform are identified. Taking these peak values, valley values, and zero-crossing points as the time reference, the main time points of the voltage fluctuations are determined, and these key time points will be used for subsequent timestamp synchronization. According to the interval between zero-crossing points or the peak-valley time, the periodic characteristics of the voltage ripple are deduced, providing a basis for subsequent timestamp alignment. Using the voltage reference time reference obtained in step S21, key time points are extracted from the power supply voltage ripple data, and these key time points usually include the peak values, valley values, and zero-crossing points of the voltage fluctuations. Each key time point is recorded as a timestamp to calibrate the time series of the power supply voltage ripple data. Each timestamp corresponds to a specific state change in the voltage fluctuation signal (such as from peak to valley or zero-crossing point). Timestamps of the programming process and the erasing process are extracted from the programming erase delay data. The programming timestamp corresponds to the start time of the write operation, while the erase timestamp corresponds to the start time of the erase operation. The time interval between the programming timestamp and the erase timestamp is calculated. Usually, this interval is the delay caused by the hard disk controller or the operating system. This time interval is used to calibrate the programming and erase timestamps to ensure the consistency of these two timestamps during synchronization. The calibration method can adopt time difference alignment technology to correct the time offset caused by clock deviation or system delay. The time difference between each power supply voltage ripple timestamp and the programming erase calibration timestamp is calculated. The time difference is obtained by comparing the absolute time values of each time point. The sliding window method is used for timestamp matching. The power supply voltage ripple timestamp and the programming erase calibration timestamp are compared one by one, and the difference between them is calculated within a time window. For each set of timestamps, the best matching point for time alignment is found by calculating the minimum value of the time difference. This sliding window matching method can adjust the time axes of the two data sets to make them better synchronized. The Dynamic Time Warping (DTW) algorithm is used to perform time warping on the power supply voltage ripple data and the programming erase delay data. DTW is a time series alignment algorithm that aligns two time series on the time axis by minimizing the difference between the time series. Based on the optimal time alignment point, the DTW algorithm is used to dynamically adjust the power supply voltage ripple data and the programming erase delay data. This process will generate synchronized timestamps between the two data sets to ensure that their time series are completely aligned on the time axis.Finally, synchronized timestamp data is generated for subsequent health assessment and analysis. These synchronized timestamp data can be used to analyze the correlation between the power supply voltage fluctuation and the programming / erasing delay, providing a more accurate time match for subsequent health assessment.
[0111] Preferably, in step S2, the electronic tunneling current waveform analysis of the programming / erasing delay data and the power supply voltage ripple data based on the synchronized timestamp includes:
[0112] Performing timing matching on the programming / erasing delay data and the power supply voltage ripple data according to the synchronized timestamp to generate a timing matching data set, where the timing matching data set includes the programming / erasing delay data and the power supply voltage ripple data at the same timestamp;
[0113] Using the power supply voltage ripple data, calculating the tunneling current of the NAND cells in the solid-state drive according to the electronic tunneling equation, where the formula of the electronic tunneling equation is as follows:
[0114] ;
[0115] In the formula, is the tunneling current, is the power supply voltage during programming or erasing, A is a constant related to the material characteristics of the NAND cells, and B is an exponential decay factor related to the material characteristics of the NAND cells;
[0116] Performing ultra-high-speed simulation on the tunneling current waveform to obtain the transient tunneling current waveform; extracting the waveform rising edge slope, peak jitter, and charge integral of the transient tunneling current waveform as the original waveform feature data;
[0117] Performing waveform deviation analysis on the power supply voltage ripple data according to the original waveform feature data to generate waveform deviation data;
[0118] Performing charge trap health analysis on the transient tunneling current waveform through the waveform deviation data to generate charge trap health index data.
[0119] In the embodiment of the present invention, by using the synchronized timestamp in step S2, it is ensured that the power supply voltage ripple data and the programming / erasing delay data are accurately aligned in time. Pairing the timestamps in the two data sets so that at the same timestamp, the programming / erasing delay data and the power supply voltage ripple data correspond to a data pair. Based on the synchronized timestamp, a new data set is created, which contains the corresponding programming / erasing delay data and power supply voltage ripple data at each timestamp, and these data will be used as the basis for subsequent analysis. Using the voltage information in the power supply voltage ripple data, calculating the tunneling current according to the electronic tunneling equation. The electronic tunneling equation is as follows: ; In the formula, is the tunneling current, V is the supply voltage during programming or erasing, A is a constant related to the material characteristics of the NAND cell, and B is an exponential decay factor related to the material characteristics of the NAND cell; extract the supply voltage value at each timestamp from the supply voltage ripple data , substitute it into the electron tunneling equation, and calculate the tunneling current at each time point . Due to the different material characteristics of NAND cells, the constant A and the exponential decay factor B need to be obtained through experiments or modeling to ensure accurate calculation of the tunneling current. Use high-precision numerical simulation methods (such as the finite element method or the fast Fourier transform) to perform ultra-high-speed simulation of the tunneling current in the time domain. The transient characteristics of the current waveform are considered during the simulation process, and a series of tunneling current waveform data at a series of time points are generated. Through ultra-high-speed simulation, the waveforms of the transient tunneling current are obtained, and these waveforms reflect the current change trends of the NAND cell during programming and erasing operations. Calculate the slope of the rising edge of the waveform, that is, the change rate between the lowest current value and the highest current value. The slope of the rising edge can reflect the response speed of the current waveform. Detect the stability of the waveform peak value, and obtain the peak jitter by calculating the change amount of the peak value. The peak jitter can reflect the stability of the current waveform. Calculate the charge amount under the transient tunneling current waveform, usually by integrating the current waveform to obtain the charge amount flowing through per unit time, which helps to evaluate the charge storage capacity of the NAND cell. Extract the rising edge slope, peak jitter, and charge integral amount as characteristic data, and use them as the input for subsequent analysis. Use statistical methods or signal processing techniques (such as root mean square error, correlation coefficient analysis, etc.) to analyze the relationship between the original waveform characteristic data and the supply voltage ripple data. Calculate the deviation degree of the waveform to evaluate the difference between the transient tunneling current waveform and the theoretical waveform, and reflect the influence of the supply voltage ripple on the tunneling current waveform. According to the deviation degree analysis result, generate waveform deviation degree data. This data can be used to evaluate the influence of the supply voltage ripple on the tunneling current waveform and further infer the health status of the charge trap. Use the waveform deviation degree data to judge the health status of the charge trap by analyzing the change of the current waveform. Identify whether the charge trap has declined or been damaged by calculating the deviation degree and characteristic data of the waveform. According to the charge trap health analysis result, generate charge trap health index data. This data can be classified by setting a health threshold, for example, to judge whether the charge trap is in a healthy state, or whether there are problems such as abnormal wear and failure
[0120] Preferably, the charge trap health analysis of the transient tunneling current waveform through the waveform deviation degree data includes:
[0121] Perform high-frequency feature decomposition on the transient tunneling current waveform to extract well potential spectrum data; perform energy level transition analysis on the transient tunneling current waveform according to the well potential spectrum data to obtain transition characteristic data;
[0122] Calculate the trap density based on the transition characteristic data to generate trap density distribution data;
[0123] Perform spatio-temporal correlation analysis on the waveform deviation data and the trap density distribution data to obtain correlation characteristic data;
[0124] Perform multi-dimensional degradation mode recognition on the correlation characteristic data to generate degradation mode data, where the multi-dimensional degradation mode recognition includes linear degradation mode recognition and non-linear degradation mode recognition;
[0125] Quantify the charge trap health based on the degradation mode data for the transient tunneling current waveform to generate charge trap health index data.
[0126] In the embodiments of the present invention, by using high-frequency signal processing techniques (such as wavelet transform or Fourier transform) to decompose the transient tunneling current waveform and extract high-frequency components, these high-frequency components can reveal the rapidly changing parts in the current waveform, which are usually related to the activities of charge traps. Perform spectral analysis on the transient tunneling current waveform to obtain well potential energy spectrum data, which reflect the energy level structure of charge traps in the NAND cell, especially the energy range related to the charge capture and release processes. By analyzing the energy distribution in different frequency bands, extract the energy spectrum features related to the charge trap sites, providing a basis for subsequent analysis. According to the extracted well potential energy spectrum data, use a physical model to analyze the transition process of charges between different energy levels. A quantum mechanics model (such as electron tunneling theory) can be used to describe the probability of charge transition from one energy level to another. Analyze the energy level transition conditions under different voltage conditions to determine the key features in the charge transition process, such as transition frequency, transition energy difference, etc. Extract transition feature data from the energy level transition analysis, mainly including transition frequency, transition energy difference, etc., which can help judge the state of the charge trap, reflect the capture and release efficiency of the charge trap, and the degradation trend. Based on the transition feature data, calculate the trap density by quantifying the transition frequency and energy difference. The trap density reflects the number and distribution of charge traps in the NAND cell. Use statistical methods or physical models to calculate the trap density at different energy levels, and analyze the change trend of the trap density according to the transition data at different time points to generate trap density distribution data, which contains the trap density at different energy levels and its change over time. These data can provide a quantitative basis for the health status of the charge trap. Conduct spatio-temporal correlation analysis on the waveform deviation data and the trap density distribution data. Specifically, time series correlation analysis (such as cross-correlation analysis) can be used to evaluate the relationship between the degree of waveform deviation and the trap density distribution. Conduct a joint analysis of the spatial distribution and time series of the waveform deviation data and the trap density data to obtain their correlation features at different time points and spatial positions. Through spatio-temporal correlation analysis, extract correlation feature data, including the correlation coefficient, trend change, etc. between the waveform deviation and the trap density. These data help to reveal the impact of waveform deviation on the degradation of charge traps. Use linear regression analysis or other linear pattern recognition algorithms to conduct linear degradation mode analysis on the correlation feature data. By observing the linear relationship of the correlation feature data, identify the change trend of the charge trap at different degradation stages. Based on the linear model, identify the healthy degradation mode of the charge trap and analyze its degradation speed and process. Conduct non-linear regression analysis on the correlation feature data, or use machine learning algorithms (such as support vector machines, neural networks, etc.) for non-linear pattern recognition. By analyzing the non-linear relationship, identify the non-linear behaviors (such as accelerated degradation) that occur during the charge trap degradation process and extract the non-linear degradation mode.Based on the degradation mode recognition results, generate degradation mode data, which can reveal the performance of the charge trap under different degradation modes and provide a reliable basis for health monitoring. Based on the degradation mode data, use a quantification method (such as a health index or score) to quantitatively evaluate the health status of the charge trap. Health quantification can be achieved by calculating the degradation degree of the charge trap or classifying it according to a preset health threshold. According to the charge trap health quantification results, generate charge trap health indicator data. This data assigns a health value or score to each charge trap to evaluate whether it is in a healthy state or whether maintenance or replacement is required.
[0127] As an example of the present invention, refer to Figure 2 shown, in this example, step S3 includes:
[0128] Step S31: Extract the number of bad blocks of S.M.A.R.T attribute data;
[0129] Step S32: Extract the historical temperature distribution of the temperature sensor data;
[0130] Step S33: Perform Weibull degradation modeling on the number of bad blocks and the historical temperature distribution to generate degradation probability distribution data; perform Markov state transition analysis on the degradation probability distribution data to generate health state transition matrix data; perform steady-state probability calculation on the health state transition matrix data to generate dynamic health score data;
[0131] Step S34: Perform FFT spectrum analysis on the charge trap health indicator data to generate main frequency energy value data; use a multi-dimensional weighted health score fusion formula to perform weighted fusion on the dynamic health score data and the main frequency energy value data to generate health score data.
[0132] In the embodiments of the present invention, S.M.A.R.T data of a solid-state drive is obtained by using a hard disk controller interface or a hard disk monitoring tool. S.M.A.R.T (Self-Monitoring, Analysis, and Reporting Technology) is a hard disk self-monitoring and reporting technology that can provide hard disk health status data. The number of bad blocks (Bad Block Count) is extracted from the S.M.A.R.T data. These are physical storage areas in the hard disk that cannot be read from or written to. The number of bad blocks is one of the important indicators of hard disk health and reflects the degree of hard disk failure. Historical data of the number of bad blocks is recorded for subsequent degradation modeling and health assessment. Temperature sensor data built into the hard disk is obtained. These sensors usually monitor the temperature change of the hard disk in real time and store the data in the internal storage of the hard disk or access it through a hard disk monitoring tool. Historical temperature records are extracted from the temperature sensor data to generate temperature distribution data. Statistical methods (such as mean, variance, etc.) can be used to analyze the fluctuation range and change trend of the temperature. The temperature distribution within different time periods is recorded to provide basic data for subsequent degradation modeling. Using the number of bad blocks and historical temperature distribution data, Weibull distribution is used for degradation modeling. Weibull distribution is a commonly used statistical distribution for describing the life of an object and can effectively represent the degradation process of a hard disk or other devices. The impact of the number of bad blocks and temperature on the hard disk life is calculated to obtain degradation probability distribution data, which can predict the failure probability of the hard disk in different degradation states. Markov state transition analysis is performed on the degradation probability distribution data, and the hard disk health state is divided into multiple discrete health states (for example, healthy, mildly degraded, severely degraded, failed, etc.), and the transition probabilities between each state are analyzed. According to the degradation probability distribution, health state transition matrix data is generated to reflect the evolution trend of the hard disk health state over time. Steady-state probability calculation is performed on the health state transition matrix data to determine the probability of the hard disk being in each health state after long-term operation. Through the steady-state probability, the current or future health status of the hard disk can be evaluated and used as a reference for health management. Based on the steady-state probability, dynamic health score data is generated. This data reflects the health status of the hard disk at different time points and can be used as a basis for evaluating the performance and reliability of the hard disk. The dynamic health score combines the number of bad blocks, temperature distribution, and degradation process, and can comprehensively reflect the health state of the hard disk. Fast Fourier Transform (FFT) is used to perform spectral analysis on the charge trap health indicator data. FFT is a conversion method from the time domain to the frequency domain, which can reveal the relationship between the charge trap health state and the frequency components in the current waveform. Through FFT analysis, the distribution of the charge trap health indicator data in the frequency domain is obtained, and the main frequency components and their corresponding energy values are extracted. The energy values of the main frequency components are extracted from the FFT results. These frequency components are usually related to the capture and release characteristics and degradation speed of the charge trap.The main frequency energy value data reflects the health status of the charge trap in different frequency ranges. Generating the main frequency energy value data provides frequency domain feature support for subsequent health scoring. Fusing the dynamic health scoring data with the main frequency energy value data, using the multi-dimensional weighted health scoring fusion formula, combining the dynamic health scoring data and the main frequency energy value data to generate the final health scoring data. The health scoring data can comprehensively consider the degradation status of the hard disk, the health condition of the charge trap and its frequency domain features, so as to provide more comprehensive decision-making support for the maintenance, warning and replacement of the hard disk.
[0133] Preferably, the multi-dimensional weighted health scoring fusion formula in step S34 is specifically as follows:
[0134] ;
[0135] In the formula, is the final health score, is the dynamic health score at time t, is the main frequency energy value at time t, is the i-th additional influence factor at time t, is the dynamic health score weight, is the main frequency energy value weight, is the weight of the i-th additional influence factor, is the time decay factor, is the regularization coefficient, is the sensitivity coefficient, is the health score change rate at time t, T is the time upper limit, n is the number of additional influence factors, is the decay rate parameter.
[0136] The present invention analyzes and integrates a multi-dimensional weighted health score fusion formula, the core idea of which is dynamic weight fusion + time decay + regularization balance. Specifically, this formula integrates multiple health-related factors and ensures the rationality and stability of the health score through mechanisms such as time decay and regularization. The dynamic health score in the formula reflects the actual operating state of the hard disk at a certain moment, such as temperature, bad sectors, read / write latency, IOPS, etc. The immediate impact of these indicators on the health state of the hard disk is very important, so higher weights should be assigned. The working frequency or frequency domain energy characteristics of the hard disk (such as vibration signals, head movement, etc.) can also reflect the health state of the hard disk. Abnormal frequencies indicate mechanical failures or other problems in the hard disk. Additional influencing factors include the usage time of the hard disk, workload, frequency of read / write requests, etc. Long-term operation or high-load work will affect the health of the hard disk. The time decay factor ensures that the score is mainly based on the recent state, and the influence of long-term data gradually weakens. The health state of the hard disk changes over time, so the decay factor can help reflect the current health status. The change rate of the health score reflects the change amplitude of the hard disk health state, and the regularization term is used to control the smoothness of the hard disk score, avoiding drastic fluctuations in the score caused by instantaneous fluctuations or abnormal data. When using the conventional multi-dimensional weighted health score fusion formula in the art, the health score of the hard disk can be obtained. By applying the multi-dimensional weighted health score fusion formula provided by the present invention, the health score of the hard disk can be calculated more accurately. The formula comprehensively considers the real-time health state of the hard disk (such as temperature, number of bad sectors, read / write speed, etc.), frequency domain characteristics, and additional factors such as workload and usage time to provide a more comprehensive health score. Through the time decay mechanism, it is ensured that the score pays more attention to the current health state of the hard disk, ensuring that the health assessment is real-time and effective. The regularization mechanism avoids the instability of the score due to instantaneous fluctuations (such as load changes within a short period of time) and provides a more reliable health prediction. The weights of each factor (such as , and ) can be flexibly adjusted according to the characteristics of the specific hard disk or the working environment to meet the health assessment needs of different types of hard disks.
[0137] As an example of the present invention, as shown in Figure 3 , in this example, step S4 includes:
[0138] Step S41: Based on a preset health threshold, perform abnormal health discrimination on the health score data. When the health score data is greater than or equal to the preset health threshold, no processing is performed;
[0139] Step S42: When the health score data is less than the preset health threshold, mark the corresponding health score data as the abnormal health discrimination result;
[0140] Step S43: Calculate the error correction iteration times of the low-density parity-check code according to the abnormal health discrimination result and perform iteration times adjustment to generate iteration times adjustment data; write voltage compensation parameters to the solid-state drive by using the charge trap health index data to obtain voltage compensation write parameter data;
[0141] Step S44: Integrate the iteration times adjustment data and the voltage compensation write parameter data, and make a self-repair strategy decision for the solid-state drive to generate hard drive self-repair decision data.
[0142] In the embodiments of the present invention, by obtaining the health score data (Hfinal) of the hard disk, which is the final health score after the fusion of dynamic health scoring and spectrum analysis. Determine the preset health threshold (Tthreshold), which is set according to the actual application situation of the hard disk and its health assessment criteria. Compare the health score data with the preset health threshold: If Hfinal≥Tthreshold, it is considered that the hard disk is in good health and no further processing is required. If Hfinal<Tthreshold, it is considered that the hard disk is in an abnormal health state and further processing is needed. For the case where the health score data is greater than or equal to the threshold, output a "healthy" signal and do not perform any processing. For the case where the health score data is less than the threshold, enter step S42 for abnormal health discrimination. If the health score data is less than the preset health threshold (Hfinal<Tthreshold), mark the health score of the hard disk as "abnormal". The marked content may include the abnormal health score data and relevant fault warning information, such as hard disk damage or degradation status. After marking the hard disk as abnormal, prepare for subsequent repair and adjustment measures and start the hard disk self-repair process. The abnormal discrimination result will affect subsequent repair measures such as the adjustment of the number of iterations and voltage compensation. Low-Density Parity-Check (LDPC) code is an error-correcting code widely used in hard disk data storage. Its main function is to correct errors in stored data and ensure data integrity. According to the abnormal health discrimination result of the hard disk, calculate the number of iterations required for error correction. The number of iterations usually depends on the health status of the hard disk and the severity of the error occurrence. When the hard disk health score is low, more iterations are required to repair data errors. Use a calculation formula to estimate the number of error-correction iterations: Iterations=f(Hfinal,Tthreshold,error rate); if the health score is low, increase the number of iterations to ensure better correction of errors in stored data. The adjusted number of iterations will affect the performance and repair effect of the hard disk, generate iteration number adjustment data, including the adjusted number of error-correction iterations, for subsequent repair processes. Based on the charge trap health indicator data, calculate and adjust the voltage compensation parameters of the hard disk. The charge trap health status directly affects the voltage requirements of the hard disk, especially during the degradation process, and voltage adjustment can help improve the performance of the hard disk. Calculate the voltage compensation parameters (such as voltage gain or voltage compensation coefficient) and write these parameters into the solid-state drive. After completing the calculation of the voltage compensation parameters, generate voltage compensation write parameter data for adjusting the voltage configuration of the hard disk. Integrate the iteration number adjustment data with the voltage compensation write parameter data, so as to ensure that during the hard disk repair process, there are both appropriate adjustments to the number of error-correction times and voltage compensation support. Based on the integrated data, formulate a self-repair strategy for the hard disk.The self - repair strategy includes selecting appropriate repair means (such as increasing the number of error - correction iterations, optimizing voltage configuration, etc.) to extend the hard - disk life and restore the hard - disk performance. The self - repair strategy can also be dynamically adjusted based on factors such as the current health score of the hard disk, the abnormal health discrimination result, the number of iterations, and voltage adjustment. After implementing the self - repair strategy, hard - disk self - repair decision data is generated, recording the key parameters, adjustment methods, and expected repair effects during the decision - making process.
[0143] Preferably, the voltage compensation parameter writing for the solid - state drive using the charge - trap health - indicator data in step S43 includes:
[0144] Extracting the ripple characteristics of the waveform deviation data using the charge - trap health - indicator data to obtain the reverse - ripple reference data;
[0145] Performing amplitude - compensation calculation on the reverse - ripple reference data to generate amplitude - compensation data;
[0146] Performing phase - matching analysis based on the amplitude - compensation data to obtain phase - matching data;
[0147] Calculating the NAND word - line drive parameters of the solid - state drive based on the phase - matching data to generate drive - parameter data;
[0148] Designing the injection timing for the solid - state drive using the drive - parameter data to obtain injection - control data;
[0149] Writing the voltage - compensation parameters to the solid - state drive using the injection - control data to obtain the voltage - compensation writing - parameter data.
[0150] In the embodiments of the present invention, waveform deviation data (ΔWwave) under the marked abnormal health state of the hard disk is obtained. The waveform deviation data reflects the change in the current waveform caused by the charge trap behavior during the health degradation of the hard disk. Ripple feature extraction is used to analyze the voltage fluctuation part in the waveform, especially the components of high-frequency ripples. Information such as the amplitude and frequency of the ripple signal is extracted using a filter or Fourier transform technology. The reverse ripple reference data is the comparison reference for the ripple feature and is used for subsequent compensation processing. Usually, it is the reverse analysis (reverse feature) of the original signal, that is, the reverse fluctuation trend is deduced from the perspective of the charge trap health of the hard disk. The reverse ripple reference data (RippleBaseline) is calculated from the ripple feature, which reflects the influence of the charge trap on the power supply voltage fluctuation during the degradation or abnormal health state of the hard disk. The voltage fluctuation of the hard disk is compensated using the reverse ripple reference data. The corresponding amplitude compensation coefficient (Kamplitude) is calculated according to the reverse ripple reference data, and this coefficient is used to adjust the amplitude of the voltage fluctuation. The amplitude compensation formula is: Vcomp = Vref × Kamplitude; where Vcomp is the compensated voltage value, Vref is the original voltage reference value, and Kamplitude is the calculated amplitude compensation coefficient. The calculated amplitude compensation coefficient and the compensated voltage value will generate amplitude compensation data and be used for subsequent phase matching analysis and voltage compensation applications. According to the amplitude compensation data, the phase information of the voltage fluctuation is analyzed. The purpose of this step is to ensure that the compensation of the voltage fluctuation is not only repaired in amplitude but also synchronized with the phase of the original signal to avoid causing excessive timing errors. A phase matching algorithm (such as least squares matching) is used for phase compensation. The matching process adjusts the phase of the voltage fluctuation to make it consistent with the timing requirements of the hard disk. This process generates a phase matching coefficient ( ) and the final phase matching data. After completing the phase matching, the generated phase matching data (including the matched phase value and timing adjustment parameters) will be used for subsequent calculation of the NAND word line drive parameters. Based on the phase matching data, the NAND word line drive of the hard disk is optimized and calculated. The word line drive parameters determine the voltage control accuracy of the NAND memory cell, so the influence of amplitude compensation and phase matching needs to be considered. The NAND word line drive parameters (such as word line voltage, pulse width, etc.) are generated through a calculation formula to ensure that the working state of the NAND memory cell is not damaged after voltage compensation and the read and write performance can be optimized. The specific formula is: ; Through calculation, the optimized word line drive parameter data is obtained, including voltage and timing adjustments for the word line. The generated drive parameter data is used for the timing design of the storage units of the solid-state drive. This step is used to adjust the read and write operations of the storage units through the hardware timing control signal. According to information such as the working frequency, voltage fluctuation, and phase matching of the hard disk, the operation timing of the storage units is designed, and it is ensured that the drive voltage and signals are accurately controlled in terms of timing. After completing the injection timing design, injection control data is generated, which includes the timing of the control signal, drive voltage, etc., for driving the operation of the NAND storage units. The calculated voltage compensation parameters are applied to the circuit system of the solid-state drive using the injection control data. This process optimizes the performance of the charge trap by adjusting the internal voltage supply of the solid-state drive to adapt to the health status and operation requirements of the hard disk. Through voltage compensation writing, voltage compensation writing parameter data is generated, and it is ensured that these parameters are written into the control system of the hard disk to support voltage compensation and stabilize the performance of the hard disk.
[0151] Therefore, from any perspective, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the application document are intended to be encompassed within the present invention.
[0152] The above description is only the specific implementation manners of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will conform to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for health monitoring of a solid-state drive, characterized in that, It includes the following steps: Step S1: Collect multi-dimensional monitoring data of the storage units of the solid-state drive to generate standard multi-dimensional monitoring data, where the standard multi-dimensional monitoring data includes S.M.A.R.T attribute data, temperature sensor data, program erase latency data, and supply voltage ripple data; Step S2: Use the supply voltage ripple data to synchronize the time stamps of the program erase latency data, and perform electron tunneling current waveform analysis on the program erase latency data and the supply voltage ripple data according to the synchronized time stamps to generate charge trap health index data; wherein, performing electron tunneling current waveform analysis on the program erase latency data and the supply voltage ripple data according to the synchronized time stamps includes: Perform time sequence matching on the program erase latency data and the supply voltage ripple data according to the synchronized time stamps to generate a time sequence matching data set, where the time sequence matching data set includes the program erase latency data and the supply voltage ripple data at the same time stamp; Use the supply voltage ripple data to calculate the tunneling current of the NAND cells of the solid-state drive according to the electron tunneling equation, where the formula of the electron tunneling equation is as follows: ; Wherein, is the tunneling current, is the power supply voltage during programming or erasing, A is a constant related to the material characteristics of the NAND cell, and B is an exponential decay factor related to the material characteristics of the NAND cell; Perform ultra-high-speed simulation of the current waveform of the tunneling current to obtain the transient tunneling current waveform; extract the waveform rising edge slope, peak jitter, and charge integral of the transient tunneling current waveform as the original waveform feature data; Perform waveform deviation analysis on the supply voltage ripple data according to the original waveform feature data to generate waveform deviation data; Perform charge trap health analysis on the transient tunneling current waveform through the waveform deviation data to generate charge trap health index data; Step S3: Extract the number of bad blocks in the S.M.A.R.T attribute data; extract the historical temperature distribution of the temperature sensor data; perform multi-variable degradation modeling based on the number of bad blocks, the historical temperature distribution, and the charge trap health index data to generate health score data; Step S4: Perform abnormal health discrimination on the health score data based on a preset health threshold to obtain an abnormal health discrimination result; make a self-repair strategy decision for the solid-state drive according to the abnormal health discrimination result to generate hard drive self-repair decision data; Step S5: Perform visual encoding on the hard drive self-repair decision data to generate a user-readable hard drive health report.
2. The health monitoring method of the solid state drive according to claim 1, wherein, Step S1 includes the following steps: Step S11: Access the S.M.A.R.T monitoring system of the solid-state drive through an interface to obtain the health status data of the hard drive; extract the S.M.A.R.T attributes of the health status data of the hard drive to obtain S.M.A.R.T attribute data; Step S12: Use the temperature sensor built in the hard drive to monitor the temperature during the operation of the hard drive to obtain temperature sensor data; Step S13: Monitor the erase operation and programming process of the solid-state drive to obtain program erase latency data; Step S14: Record the voltage fluctuation data through the power management module, and perform voltage ripple analysis on the voltage fluctuation data to generate supply voltage ripple data; Step S15: Integrate S.M.A.R.T attribute data, temperature sensor data, programming erase delay data, and power supply voltage ripple data into multi-dimensional monitoring data; perform data preprocessing on the multi-dimensional monitoring data to generate standard multi-dimensional monitoring data, where data preprocessing includes data cleaning, data denoising, missing value filling, and data standardization.
3. The health monitoring method of the solid-state drive according to claim 1, characterized in that, The time stamp synchronization of the programming erase delay data using the power supply voltage ripple data in Step S2 includes: Perform voltage fluctuation trend analysis on the power supply voltage ripple data to obtain the voltage fluctuation peak value, valley value, and zero crossing point, and use them as the voltage reference time benchmark; Extract the key time points of the power supply voltage ripple data through the voltage reference time benchmark to obtain the power supply voltage ripple time stamp; Extract the programming time stamp and erase time stamp from the programming erase delay data; calibrate the delay time interval of the programming time stamp and erase time stamp to obtain the programming erase calibration time stamp; Calculate the time difference between the power supply voltage ripple time stamp and the programming erase calibration time stamp, and perform a sliding window match on the power supply voltage ripple time stamp and the programming erase calibration time stamp through the time difference to generate the best time alignment point; Perform dynamic time warping on the power supply voltage ripple data and the programming erase delay data based on the best time alignment point to generate synchronized time stamps.
4. The method for health monitoring of a solid-state drive according to claim 1, wherein The charge trap health analysis of the transient tunneling current waveform through the waveform deviation data includes: Perform high-frequency feature decomposition on the transient tunneling current waveform to extract the well potential energy spectrum data; perform energy level transition analysis on the transient tunneling current waveform according to the well potential energy spectrum data to obtain the transition feature data; Calculate the trap density based on the transition feature data to generate trap density distribution data; Perform spatio-temporal correlation analysis on the waveform deviation data and the trap density distribution data to obtain correlation feature data; Perform multi-dimensional degradation mode recognition on the correlation feature data to generate degradation mode data, where multi-dimensional degradation mode recognition includes linear degradation mode recognition and non-linear degradation mode recognition; Quantify the charge trap health of the transient tunneling current waveform according to the degradation mode data to generate charge trap health index data.
5. The method for health monitoring of a solid-state drive according to claim 1, wherein Step S3 includes the following steps: Step S31: Extract the number of bad blocks in the S.M.A.R.T attribute data; Step S32: Extract the historical temperature distribution of the temperature sensor data; Step S33: Perform Weibull degradation modeling on the number of bad blocks and the historical temperature distribution to generate degradation probability distribution data; perform Markov state transition analysis on the degradation probability distribution data to generate health state transition matrix data; perform steady-state probability calculation on the health state transition matrix data to generate dynamic health score data; Step S34: Perform FFT spectrum analysis on the charge trap health index data to generate main frequency energy value data; use the multi-dimensional weighted health score fusion formula to perform weighted fusion on the dynamic health score data and the main frequency energy value data to generate health score data.
6. The method for health monitoring of a solid-state drive according to claim 5, wherein, The multi-dimensional weighted health score fusion formula in Step S34 is as follows: ; Wherein, is the final health score, is the dynamic health score at time t, is the main frequency energy value at time t, is the i-th additional influence factor at time t, is the weight of the dynamic health score, is the weight of the main frequency energy value, is the weight of the i-th additional influence factor, is the time decay factor, is the regularization coefficient, is the sensitivity coefficient, is the change rate of the health score at time t, T is the time upper limit, and n is the number of additional influence factors, is the decay rate parameter.
7. The method for health monitoring of a solid-state drive according to claim 1, wherein Step S4 includes the following steps: Step S41: Perform abnormal health discrimination on the health score data based on a preset health threshold. When the health score data is greater than or equal to the preset health threshold, no processing is performed; Step S42: When the health score data is less than the preset health threshold, mark the corresponding health score data as an abnormal health discrimination result; Step S43: Calculate the error correction iteration times of the low-density parity-check code according to the abnormal health discrimination result and adjust the iteration times to generate iteration time adjustment data; Write voltage compensation parameters to the solid-state drive using the charge trap health index data to obtain voltage compensation write parameter data; Step S44: Integrate the iteration time adjustment data and the voltage compensation write parameter data, and make a self-repair strategy decision for the solid-state drive to generate hard disk self-repair decision data.
8. The method for health monitoring of a solid-state drive according to claim 1, wherein The writing of voltage compensation parameters to the solid-state drive using the charge trap health index data in Step S43 includes: Extract the ripple characteristics of the waveform deviation data using the charge trap health index data to obtain reverse ripple reference data; Perform amplitude compensation calculation on the reverse ripple reference data to generate amplitude compensation data; Perform phase matching analysis according to the amplitude compensation data to obtain phase matching data; Calculate the NAND word line drive parameters of the solid-state drive based on the phase matching data to generate drive parameter data; Obtain injection control data by using the drive parameter data to design the injection timing of the solid-state drive; Write voltage compensation parameters to the solid-state drive through the injection control data to obtain voltage compensation write parameter data.
9. A health monitoring system for a solid state drive, characterized in that, For executing the health monitoring method of the solid-state drive as described in Claim 1, the health monitoring system of the solid-state drive includes: A data acquisition module, configured to collect multi-dimensional monitoring data of the storage unit of the solid-state drive to generate standard multi-dimensional monitoring data, where the standard multi-dimensional monitoring data includes S.M.A.R.T attribute data, temperature sensor data, program erase delay data, and power supply voltage ripple data; A health index analysis module, configured to synchronize the time stamps of the program erase delay data using the power supply voltage ripple data, and perform electron tunneling current waveform analysis on the program erase delay data and the power supply voltage ripple data according to the synchronized time stamps to generate charge trap health index data; A health scoring module, configured to extract the number of bad blocks in the S.M.A.R.T attribute data; extract the historical temperature distribution of the temperature sensor data; perform multi-variable degradation modeling based on the number of bad blocks, the historical temperature distribution, and the charge trap health index data to generate health score data; An adaptive repair module, configured to perform abnormal health discrimination on the health score data based on a preset health threshold to obtain an abnormal health discrimination result; make a self-repair strategy decision for the solid-state drive according to the abnormal health discrimination result to generate hard disk self-repair decision data; A visualization module, configured to perform visual encoding on the hard disk self-repair decision data to generate a user-readable hard disk health report.
Citation Information
Patent Citations
A method and system for predicting the life of a solid-state drive based on big data
CN119739351A
Scanning tunnel microscopy information processing system with noise detection to correct the tracking mechanism
US5371727A