A hard disk data reliability test method and system

By using predefined load and environmental data sequences, combined with I/O performance and SMART data, a dynamic failure characteristic model is used to identify early signs of hard drive failure and apply targeted enhanced stress. This solves the problem of inaccurate hard drive test results in existing technologies, enabling earlier and more accurate failure prediction and more efficient testing.

CN120849202BActive Publication Date: 2026-01-27GUIZHOU SHUSUAN INTERNET TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511353267.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2026-01-27
Estimated Expiration
2045-09-22

AI Technical Summary

Technical Problem

Existing hard drive data reliability testing methods fail to effectively simulate complex loads, resulting in test results that cannot accurately reflect performance and reliability in actual deployment environments, and cannot accurately identify the early signs of hard drive failure.

Method used

By predefined load test data and environmental data sequences, combined with I/O load performance indicators and SMART data, a dynamic failure feature model is constructed using a gated recurrent unit neural network. This model monitors the hard drive status in real time, identifies early signs of failure, and applies targeted enhanced stress for testing.

Benefits of technology

It enables accurate modeling and prediction of the individualized, evolutionary degradation process of hard drives, significantly improving the sensitivity and testing efficiency of identifying early signs of failure and reducing testing costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849202B_ABST
    Figure CN120849202B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computer data storage system, specifically to a hard disk data reliability test method and system. The specific implementation steps include: first, predefine various parameters to form a policy set for monitoring the state of the storage subsystem, and generate a test environment according to the policy set for testing; second, online analysis is performed on the data generated during the test process, and a dynamic failure characteristic model of the to-be-tested hard disk for predicting data loss risk is established; then, the failure precursor characteristics are identified through the model, and targeted enhanced stress is applied according to the identified failure characteristics to accelerate the test process; finally, a warning is issued to confirm whether to immediately perform data migration, fault-tolerant operation of device replacement. The present application can efficiently and accurately perform testing in a short time through dynamic identification and targeted enhancement, and provide a dynamic closed-loop test process, thereby saving test time and cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer data storage system technology, specifically to a hard disk data reliability testing method and system. Background Technology

[0002] In today's digital age, data has become a core asset for enterprises and institutions. As the primary carrier of data storage, hard drives directly impact business continuity, data security, and even a company's reputation. Enterprise data centers are increasingly adopting NAND flash-based solid-state drives (SSDs) to replace traditional mechanical hard drives, achieving higher performance and lower latency. However, hard drives face complex and demanding operating environments in real-world applications. Differences in internal controller firmware, NAND flash memory chips, and caching strategies among different brands, models, and batches of hard drives lead to vastly different reliability performance over long-term use. Therefore, before purchasing and deploying hard drives, enterprises and institutions cannot rely solely on the product specifications provided by suppliers; they must establish a scientific, comprehensive, and practical reliability testing method and system.

[0003] Existing testing methods largely rely on standardized benchmarking tools, which generate synthetic and idealized test loads, such as purely sequential read / write or completely random read / write. While this approach facilitates comparisons between different products, it deviates significantly from real-world application scenarios. Existing testing methods fail to effectively simulate such complex loads, resulting in test results that do not accurately reflect the true performance and reliability of hard drives in actual deployment environments, thus misleading device selection.

[0004] Therefore, a method and system for testing hard disk data reliability are proposed. Summary of the Invention

[0005] The purpose of this invention is to provide a method and system for testing the reliability of hard disk data, so as to solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution:

[0007] A method for testing hard disk data reliability includes:

[0008] Based on different testing requirements, various load test data and environmental data sequences including temperature, voltage and vibration are predefined to obtain a set of strategies for monitoring the status of the storage subsystem; test parameters are selected for the set of strategies to generate a benchmark test environment, and the hard drive under test is tested based on the benchmark test environment.

[0009] During normal system operation, multimodal data including I / O load performance indicators and SMART data values ​​are monitored in real time and collected periodically. Combining the multimodal data, the multi-dimensional change trend of the working status of the hard drive under test under different time, test environment and load conditions is analyzed online. Based on the multi-dimensional change trend, a dynamic failure feature model of the hard drive under test is constructed to predict the risk of data loss.

[0010] The dynamic failure feature model identifies various early failure features of the hard drive under test in real time. Based on the early failure features, the environmental data sequence is adjusted to obtain targeted enhanced pressure. The response data of the hard drive under test under the targeted enhanced pressure is recorded, and a warning is issued to confirm whether to immediately perform fault-tolerant operations such as data migration and device replacement.

[0011] Preferably, the specific implementation process for obtaining the strategy set for the monitoring storage subsystem status includes:

[0012] Predefined I / O workload types, including sequential read / write, random read / write, and mixed read / write I / O access modes, are generated. The duration of different workloads, including normal operation, high-intensity continuous operation, and intermittent operation, is also predefined to generate predefined load test data. Temperature variation ranges, including normal operating temperature and hard drive operating limit temperature, voltage fluctuation ranges including nominal voltage, overvoltage, and undervoltage, and vibration data including constant vibration frequency and composite vibration frequency are set. Temperature, voltage, and vibration sequences are constructed using linear and power functions, respectively, forming an environmental data sequence. Together with the predefined load test data, this constitutes a strategy set for monitoring the storage subsystem status.

[0013] Preferably, the specific implementation process for generating the benchmark testing environment includes:

[0014] Based on the initial performance indicators of the hard drive under test and its specific usage scenario, test parameters are selected from the strategy set and combined to simulate the actual working state of the hard drive and generate a benchmark test environment.

[0015] Preferably, the specific analysis process of the multi-dimensional change trend includes:

[0016] Time series analysis is performed on the multimodal data, and correlation analysis is conducted between the strategy set and the multimodal data to obtain multidimensional trends in the working status of the hard disk under test under different time, test environment and load conditions.

[0017] Preferably, the specific construction process of the dynamic failure feature model of the hard disk under test includes:

[0018] Based on a gated recurrent unit neural network, the multimodal data and historical failure data are standardized into input vectors. These vectors are then input into a GRU neural network for training. Stacked GRU layers with built-in update and reset gates capture the nonlinear long-term temporal dependencies in the data. After training, the model parameters, including the complete network topology, weights, and bias parameters in all stacked GRU layers and fully connected layers, are obtained. The optimal model state is obtained by serializing the model parameters. Based on the optimal model state, a dynamic failure feature model of the hard drive under test is constructed to predict the risk of data loss. During the testing process, newly collected data is continuously collected and standardized in the same way as during the training phase. This data is then input into the loaded baseline model to incrementally update the model.

[0019] Preferably, the specific identification process for each precursory failure feature of the hard drive under test includes:

[0020] The real-time collected multimodal data is cleaned to form feature vectors, which are then input into the dynamic failure feature model to analyze the evolution of the hard drive from health to failure, and the model output results are obtained. Key indicators of hard drive failure characteristics, including the decline in I / O load performance indicators and abnormal SMART data values, are monitored in the multimodal data. Dynamic thresholds are calculated based on the weighted values ​​of the key indicators. When the model output results and the key indicators of hard drive failure characteristics exceed the dynamic thresholds, the type of failure precursor features, including the decline in various performance characteristics of the hard drive, is identified based on the dynamic threshold anomaly detection algorithm.

[0021] Preferably, the specific implementation process of the targeted enhanced pressure includes:

[0022] Based on the identified types of pre-failure warning signs, a pre-built hard drive test history database is invoked to analyze the correlation between the pre-failure warning signs and environmental data. Based on the correlation, the test intensity of the corresponding environmental factors in the environmental data sequence is dynamically adjusted, and the application mode is changed to continue testing the hard drive under test. This forms an intelligent closed-loop test method that accelerates aging by analyzing the pre-failure warning signs in real time based on the changes in the internal state of the hard drive under test, thereby applying targeted enhancement pressure to the hard drive under test.

[0023] A hard disk data reliability testing system, comprising:

[0024] The test strategy definition module is used to predefine various load test data and environmental data sequences including temperature, voltage and vibration conditions based on different test requirements, and obtain a set of strategies for monitoring the status of the storage subsystem.

[0025] The test environment generation module is used to select test parameters from the test scheme parameter set and generate a benchmark test environment.

[0026] The test execution module is used to receive the generated benchmark test environment and test the hard drive.

[0027] The data acquisition module is used to monitor and periodically collect multimodal data, including I / O load performance indicators and SMART data values, in real time during normal system operation.

[0028] The analysis and modeling module is used to combine the multimodal data to analyze the multi-dimensional change trend of the working status of the hard drive under test under different time, test environment and load conditions, and to construct a dynamic failure feature model of the hard drive under test based on the multi-dimensional change trend.

[0029] The adjustment processing module is used to identify various failure precursor features of the hard drive under test in real time according to the dynamic failure feature model, adjust the environmental data sequence based on the failure precursor features to obtain targeted enhancement pressure, and record the response data of the hard drive under test under the targeted enhancement pressure.

[0030] The warning module is used to issue warnings and confirm whether to immediately perform fault-tolerant operations such as data migration and device replacement.

[0031] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0032] 1. Regarding the selection of the testing environment, by predefining a multimodal test parameter set that integrates I / O load performance indicators and environmental parameters, the sensitivity and completeness of the data foundation for identifying pre-failure indicators are fundamentally improved. Existing technologies, such as the SMART system, have a single monitoring dimension, limited to certain hardware attributes, and are insensitive to early failure indicators caused by performance degradation, resulting in a weak information foundation. This invention innovatively incorporates performance indicators that are more sensitive to faults, such as I / O latency and throughput, to construct a feature space that is richer in information and more comprehensive in dimensions. This multimodal data fusion provides high-quality, highly relevant input for the subsequent accurate modeling of dynamic failure feature models, enabling the earlier and more accurate capture of potential failure signs.

[0033] 2. Regarding testing effectiveness, this invention employs a dynamic failure characteristic model based on gated cyclic units, achieving accurate modeling and prediction of the individualized, evolutionary degradation process of hard drives. Existing SMART systems use static thresholds preset by the manufacturer; this mechanism cannot adapt to the unique, time-evolving dynamic failure process of each hard drive, resulting in low prediction accuracy. The dynamic model constructed in this invention can learn and track the nonlinear temporal dependencies of each hard drive under multimodal data streams, constructing a personalized health degradation trajectory for each hard drive. This data-driven, adaptive prognostic diagnostic method breaks free from the constraints of static rules, significantly improving the accuracy and reliability of remaining lifespan prediction.

[0034] 3. Regarding testing time and cost, the invention achieves a fundamental improvement in testing efficiency and relevance by establishing targeted enhanced stress. Existing technologies, such as high-acceleration life testing, employ pre-defined, indiscriminate, uniform testing methods. This process is often aimless, time-consuming, energy-intensive, and frequently induces failure modes unrelated to actual operating conditions, leading to distorted test results. This invention, however, utilizes the real-time prognostic diagnostic capabilities of a dynamic failure characteristic model to accurately identify the most likely failure mode and establishes an intelligent closed-loop feedback mechanism to dynamically adjust the test type and intensity, targeting and accelerating the evolution of the relevant failure mode in the most efficient way. This shift in testing methodology significantly shortens the testing cycle, ensures the effectiveness of failure analysis, and substantially reduces testing costs. Attached Figure Description

[0035] Figure 1 This is a flowchart of a hard disk data reliability testing method proposed in an embodiment of this invention application;

[0036] Figure 2 This is a schematic diagram illustrating the process of obtaining targeted enhanced pressure based on failure precursor characteristics as proposed in an embodiment of this invention application;

[0037] Figure 3 This is a flowchart of a hard disk data reliability testing system proposed in an embodiment of this invention. Detailed Implementation

[0038] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0039] Please see Figures 1-2 The present invention relates to a method for testing the reliability of hard disk data, the specific implementation steps of which are as follows:

[0040] Based on different testing requirements, various load test data and environmental data sequences including temperature, voltage and vibration are predefined to obtain a set of strategies for monitoring the status of the storage subsystem; test parameters are selected for the set of strategies to generate a benchmark test environment, and the hard drive under test is tested based on the benchmark test environment.

[0041] During normal system operation, multimodal data including I / O load performance indicators and SMART data values ​​are monitored in real time and collected periodically. Combining the multimodal data, the multi-dimensional change trend of the working status of the hard drive under test under different time, test environment and load conditions is analyzed online. Based on the multi-dimensional change trend, a dynamic failure feature model of the hard drive under test is constructed to predict the risk of data loss.

[0042] The dynamic failure feature model identifies various early failure features of the hard drive under test in real time. Based on the early failure features, the environmental data sequence is adjusted to obtain targeted enhanced pressure. The response data of the hard drive under test under the targeted enhanced pressure is recorded, and a warning is issued to confirm whether to immediately perform fault-tolerant operations such as data migration and device replacement.

[0043] The technical solution of the present invention will be further described in detail below with reference to specific embodiments. Example 1

[0044] This application discloses a hard disk data reliability testing method and system, and the optimization process for hard disk data reliability testing is described in the following reference. Figure 1 The specific implementation steps of the method proposed in this invention include: Step S1, predefining multiple test data sequences to obtain a strategy set for monitoring the status of the storage subsystem; Step S2, generating a benchmark test environment and testing the hard drive under test based on the benchmark test environment; Step S3, periodically collecting multimodal data including I / O load performance indicators and SMART data values; Step S4, analyzing the multidimensional change trend of the working status of the hard drive under test and constructing a dynamic failure feature model; Step S5, identifying various failure precursor features of the hard drive under test in real time based on the dynamic failure feature model; Step S6, obtaining targeted enhancement pressure based on the failure precursor features, continuing testing under targeted enhancement pressure and issuing warnings.

[0045] Furthermore, a set of strategies for monitoring the status of the storage subsystem is obtained by predefining various load test data and environmental data sequences including temperature, voltage, and vibration conditions; corresponding to step S1 above; the specific implementation process includes:

[0046] The process of predefining various load test data employs industry-recognized I / O load generation tools, such as fio, which is widely used for SSD benchmarking. This tool generates multiple I / O workload types, including sequential read / write, random read / write, and mixed read / write I / O access modes, and predefines the duration of different workloads, including normal operation, high-intensity continuous operation, and intermittent operation.

[0047] The predefined environment data sequence is based on the hard drive's operating status and includes:

[0048] The temperature series is defined based on JEDEC standards JESD218 and JESD22-A104. Typical operating temperatures for enterprise-class SSDs range from 0°C to 70°C, but the maximum allowable temperature for hard drive specifications is 85°C. The temperature series will cycle within this range and may include rapid temperature changes at a rate of 10-15°C per minute. The test also considers the impact of temperature on NAND data retention capabilities. JESD218 specifies that enterprise-class SSDs must retain data for at least three months under power-off conditions at 40°C after reaching their rated write lifespan. Therefore, the temperature series will include high-temperature operating and high-temperature storage phases to evaluate data retention performance.

[0049] The voltage series is defined based on ITIC curves to evaluate the stability of the SSD controller and power management integrated circuit under voltage fluctuations. The voltage series will include sustained overvoltage (+10%) and undervoltage (-10%) at nominal voltages (e.g., +12V, +3.3V), as well as transient voltage drops and surges simulating grid disturbances.

[0050] Vibration sequences are performed according to the IEC 60068-2-6 standard. Although SSDs have no moving parts and are not sensitive to vibration, vibration testing is still necessary to verify the solder strength of components such as capacitors and crystal oscillators on the circuit board and the reliability of connectors. Tests typically cover a frequency range from 10Hz to 500Hz, with acceleration levels reaching 5g or higher.

[0051] These predefined data sets are used to form a set of strategies for monitoring the status of the storage subsystem.

[0052] By predefining a rich database containing various dynamic sequences, this invention constructs a test environment capable of more accurately simulating the complex and fluctuating operating conditions of real-world data centers. This overcomes the insufficient fidelity problem caused by oversimplification of conditions in JEDEC / SNIA testing. This comprehensive set of strategies is a necessary foundation for subsequent steps to achieve intelligent and adaptive testing. It enables the system to build highly targeted, high-stress test scenarios on demand, scenarios that are impossible to achieve with the limited definitions of traditional testing methods.

[0053] Furthermore, test parameters are selected from the policy set for monitoring the status of the storage subsystem, a benchmark test environment is generated, and the hard drive under test is tested based on the benchmark test environment; corresponding to step S2 above, the specific implementation process includes:

[0054] The strategy is selected and combined from the aforementioned set of strategies for monitoring the status of the storage subsystem. For the operational requirements of the hard drive under test in a complex system, such as when conducting reliability testing on a batch of 14TB enterprise-grade SSDs used in a storage service server, the module can select the "Database OLTP Simulation" load and match it with a "24-hour temperature cycle sequence" and a "standard voltage fluctuation sequence" simulating high-density deployment to generate a benchmark test environment that highly simulates its real production environment. Failures in complex systems are often emergent failures caused by the interaction of multiple stress factors. This test scheme can proactively induce and discover multi-factor synchronously and dynamically coupled failure causes targeting emergent failures. The test execution module receives the generated benchmark test environment and applies it to the solid-state drive under test.

[0055] By combining highly simulated I / O patterns with dynamically changing temperature, voltage, and vibration environmental factors, a composite environment is created to simulate the most realistic hard drive operating conditions. The test conditions obtained in this environment are more likely to induce complex failure modes that only occur in the real world. This results in an unprecedented improvement in the realism and relevance of the test benchmark.

[0056] Furthermore, multimodal data, including I / O load performance metrics and SMART data values, are periodically collected; corresponding to step S3 above, the specific implementation process includes:

[0057] Throughout the testing process, the data acquisition module periodically collected multimodal data at a frequency of one minute. I / O load performance metrics were collected using tools such as fio. IOPS, average / maximum latency, 99.99% latency percentile, and throughput (MB / s) were recorded. For SSDs, even small latency fluctuations are a key precursor to performance degradation. The raw values ​​of the SSD-specific SMART attribute were also obtained. Table 1 shows the SMART attribute values ​​that require key monitoring:

[0058] Table 1: Key SSD SMART Attributes Monitored

[0059] SMART attribute value ID (Hex) Attribute Name The meaning of this attribute 5(0x05) ReallocatedSectors_Count Number of bad blocks in NAND flash memory. 173(0xAD) Wear_Leveling_Count Maximum number of erase / write cycles for a NAND block. 171 / 181(0xAB / B5) Program_Fail_Count Total number of programming failures. 172 / 182(0xAC / B6) Erase_Fail_Count Total number of erase failures. 187(0xBB) Reported_Uncorrectable_Errors The number of errors that ECC cannot correct. 194(0xC2) Temperature_Celsius SSD internal temperature. 233(0xE9) Media_Wearout_Indicator Media wear indicator. A comprehensive health score defined by the manufacturer, ranging from 100 to 0. 241(0xF1) Total_LBAs_Written The total number of LBAs written to the host. The write amplification factor can be calculated by combining this with the internal NAND write volume.

[0060] These attribute values ​​provide in-depth insights into NAND flash health, controller behavior, internal errors, data retention capabilities, and component lifespan, forming the basis for subsequent analysis.

[0061] By fusing high-frequency performance and health data, this invention creates an unprecedented, highly information-dense multimodal time-series dataset. This dataset enables the system to detect subtle, interconnected precursors to failure that are invisible to methods analyzing only a single data stream or analyzing two data streams at low frequencies. For example, the system can detect a correlation between a small increase in Reallocated_Sectors_Count and a 5% increase in 99.99% read latency, and this correlation only becomes significant when the temperature exceeds 75°C—a complex pattern that would almost certainly be overlooked in daily log analysis.

[0062] Furthermore, the multi-dimensional variation trend of the working status of the hard drive under test is analyzed online under different time, test environment, and load conditions, and a dynamic failure characteristic model of the hard drive under test is constructed based on the multi-dimensional variation trend to predict the risk of data loss; corresponding to step S4 above, the specific implementation process includes:

[0063] Time series analysis was performed on the collected multimodal data, and its correlation with test environment parameters such as temperature, voltage, and vibration type was analyzed. Based on the time series analysis results, a dynamic failure characteristic model for this type of hard drive was constructed. The model uses the Z-score normalization method to process the multimodal data that integrates I / O performance indicators and internal health status.

[0064] By acquiring data in real time and performing online analysis, more targeted guidance can be provided for subsequent modeling. This ensures that the established model is not a general analysis model, but a specialized model that can learn and adapt to the unique characteristics of the device under test for specialized analysis, greatly improving the model's relevance and predictive accuracy.

[0065] A deep neural network architecture based on gated recurrent units (GRUs) was chosen. As an advanced recurrent neural network, the GRU model's internal "update gate" and "reset gate" structures effectively capture complex, non-linear, long-term temporal dependencies in data, making it ideal for processing time-varying sequential data such as hard drive health status. The system employs a stacked GRU layer architecture, enabling the model to learn hierarchical failure characteristics from raw data, ranging from fluctuations in individual SMART values ​​to performance degradation occurring simultaneously with specific error codes.

[0066] The model is first pre-trained on a dataset containing a large amount of SSD field operation and failure data. Public datasets from various large cloud service providers can be used, containing daily SMART logs and failure records of hundreds of thousands of SSDs, providing the model with rich failure modes. The GRU network is then trained offline using historical data containing tens of thousands of SSDs of the same series or similar technology, including complete sensor data streams and final failure labels. At each time step t, the model input is a standardized feature vector composed of multimodal data. After training, a baseline model containing the complete network topology, weights, and bias parameters is obtained. This model functionalizes and vectorizes the health state, deriving a continuous and smooth hard drive health status degradation trajectory that evolves over time in a multidimensional feature space. By analyzing the degradation trajectory, not only the current state of the hard drive can be determined, but also the rate and trend of change in hard drive health status can be analyzed. This allows for more rational planning of spare parts inventory and maintenance plans, optimizing operation and maintenance costs, and achieving earlier and more accurate prognostic diagnosis.

[0067] Furthermore, during each test task, the system continuously inputs newly collected data into the loaded benchmark model, performing online, incremental updates to the model. This allows the model to gradually specialize from a general group model to the specific behavior and potential unique failure modes of the SSD under test, improving the personalization and accuracy of predictions and solving the problem of data distribution differences between different SSD models and firmware versions.

[0068] Through an online incremental update mechanism, the general prediction model can learn and adapt to the unique characteristics of the device under test (DUT), including its specific manufacturing process variations, firmware version behavior, and emerging unique failure modes. This transformation method greatly improves the specificity and accuracy of the prediction. Because the model is highly specialized to a single DUT, it can capture the unique, subtle failure precursors of that device, which might be ignored as noise in a general model trained on large-scale, heterogeneous population data. Therefore, compared to static models, the dynamic model proposed in this invention can significantly save testing costs and improve testing efficiency when predicting failures of specific DUTs.

[0069] Furthermore, based on the dynamic failure characteristic model, various early failure features of the hard drive under test are identified in real time; corresponding to step S5 above, the specific implementation process includes:

[0070] The system inputs real-time acquired multimodal data into the dynamic failure feature model. The normalized output values ​​between [0, 1] from the dynamic failure feature model are directly used as the anomaly score for the model part. Weighted penalty scores for key performance indicators are calculated based on the raw values ​​of the key SMART data in Table 1. A dynamic threshold is calculated by combining these two scores. This value can be optimized during the system validation phase through ROC curve analysis to quantify and visualize the model's predictive performance, find the optimal decision threshold, and provide guidance for model iteration, achieving the best balance between precision and recall. When the model's health score falls below this dynamic threshold, or when key indicators such as Uncorrectable_Error_Count become non-zero, the system determines that a failure precursor has been identified and records the type of the precursor, such as "early NAND wear," "controller response delay," or "thermal-related performance degradation." For example, it is assumed that after 500 hours of testing, the model output shows a sharp increase in the probability of "media-related errors" within the next 24 hours. Meanwhile, the system monitored that the weighted values ​​of two key indicators, uncorrectable_error_count and read_latency, exceeded the thresholds dynamically calculated based on recent data. Based on this, the system identified a clear precursor to failure: increased read instability caused by high temperature.

[0071] By employing multi-task learning and other methods, specific types of failure precursors can be identified, such as "increased read instability caused by high temperature" or "early wear of NAND". This diagnostic output, with clear physical or logical meaning, is a key prerequisite for the next step of "targeted" stress enhancement, enabling earlier and more sensitive failure warnings.

[0072] Furthermore, based on the characteristics of failure precursors, the environmental data sequence is adjusted to obtain targeted enhanced stress, and the response data of the hard drive under test under targeted enhanced stress is recorded. A warning is then issued to confirm whether to immediately perform data migration and device replacement fault-tolerant operations. Corresponding to step S6 above, the specific implementation process includes:

[0073] After identifying early signs of failure, a pre-built SSD test history database is queried. This database contains various I / O load performance metrics and key monitored SMART attribute values, linking the type of early failure to the most likely environmental or load stressors that triggered the failure. A relative anomaly assessment of the hard drive is then performed based on a dynamic peer baseline. Determining whether a hard drive is abnormal is not based solely on its own behavior or comparison with a static model, but rather on a horizontal comparison of its current real-time behavior with that of other hard drives of the same batch and type undergoing the exact same test environment and load. This peer group constitutes a dynamic, time- and environment-dependent, and highly comparable benchmark. When a device significantly deviates from the behavior of its peer group, it is considered abnormal. This method effectively filters out common-mode interference caused by changes in the test environment itself. Subsequently, the system dynamically adjusts test parameters and applies targeted enhancement stress. Faults that might otherwise require thousands of hours of conventional testing to manifest are successfully induced under tens of hours of targeted enhancement stress. The system meticulously records the complete response data of the hard drive under test under this stress, including the rate of error rate increase and performance degradation curves. This data reflects the dynamic changes of various indicators of the hard drive from health to failure. Using this response data, a high-resolution, multi-dimensional failure signature can be generated. This failure signature provides pattern-matching conclusions for future monitoring of hard drives of the same model in production environments. When a hard drive in the production environment exhibits early behavior similar to a certain inventory failure signature, the system can determine with high confidence what type of failure it is heading towards, greatly facilitating subsequent model updates and hard drive testing. For example, a clear precursor to failure is identified as increased read instability caused by high temperature. Historical cases stored in the database show that this precursor is highly correlated with the combined stress of high temperature and high-intensity random reads. Based on this correlation, the adjustment module no longer continues to execute the original benchmark test environment but dynamically adjusts the environmental data sequence. It commands the test execution module to rapidly raise the ambient temperature to 80°C and maintain it, while simultaneously switching the I / O load to 100% high-intensity random reads.

[0074] By implementing targeted stress enhancement, once a specific precursor to failure is identified, the system immediately interrupts the scheduled benchmark test and, based on an existing database, dynamically adjusts the combination of test parameters that best accelerates the specific failure mode. This precise approach avoids wasting time on irrelevant factors and can successfully induce faults that would normally require thousands of hours of routine testing in just tens of hours, achieving an exponential increase in testing efficiency.

[0075] Example 2

[0076] Company A used the system architecture based on this solution to conduct reliability tests on a batch of 14TB enterprise-class SSDs used in storage service servers. Please refer to [link / reference]. Figure 3 The present invention relates to a hard disk data reliability testing system, comprising:

[0077] The test strategy definition module is used to predefine various load test data and environmental data sequences including temperature, voltage and vibration conditions based on different test requirements, and obtain a set of strategies for monitoring the status of the storage subsystem.

[0078] The test environment generation module is used to select test parameters for the strategy set and generate a benchmark test environment.

[0079] The test execution module is used to receive the generated benchmark test environment and test the hard drive.

[0080] The data acquisition module is used to monitor and periodically collect multimodal data, including I / O load performance indicators and SMART data values, in real time during normal system operation.

[0081] The analysis and modeling module is used to combine the multimodal data to analyze the multi-dimensional change trend of the working status of the hard drive under test under different time, test environment and load conditions, and to construct a dynamic failure feature model of the hard drive under test based on the multi-dimensional change trend.

[0082] The adjustment processing module is used to identify various failure precursor features of the hard drive under test in real time according to the dynamic failure feature model, adjust the environmental data sequence based on the failure precursor features to obtain targeted enhancement pressure, and record the response data of the hard drive under test under the targeted enhancement pressure.

[0083] The warning module is used to issue warnings and confirm whether to immediately perform fault-tolerant operations such as data migration and device replacement.

[0084] The devices under test consist of a batch of new 14TB NVMe interface enterprise-grade solid-state drives (SSDs). These SSDs are planned for use in the company's next-generation high-density storage servers, serving as a "warm data" and "hot data" layer in a hybrid storage pool, primarily supporting workloads from virtualization environments, containerized applications, and high-performance databases. Prior to large-scale deployment, the dynamic closed-loop testing method described in this invention is used to comprehensively and efficiently evaluate the long-term reliability, performance consistency, and tolerance to complex stress conditions of this batch of 14TB SSDs in a simulated real data center environment, identifying and reproducing potential firmware defects or hardware weaknesses.

[0085] First, a set of strategies for monitoring the status of the storage subsystem is predefined. In the test scheme definition module, various I / O workload types in the load test data are predefined.

[0086] Type 1: Database OLTP simulation, characterized by 80% random reads, 20% random writes, and a data block size of 8KB.

[0087] Type 2: Simulated video streaming write, characterized by 100% sequential write and a data block size of 1MB.

[0088] Type 3: General virtual machine storage simulation, characterized by 50% random reads and 50% random writes, with a mix of data block sizes ranging from 4KB to 64KB.

[0089] At the same time, define the duration pattern of the workload, such as "8 hours of high-intensity continuous operation" to simulate peak business periods and "4 hours of intermittent operation" to simulate nighttime data processing.

[0090] Next, the environmental parameter sequence is defined as follows:

[0091] Temperature sequence: The temperature range is set from the normal operating temperature of 40°C to the hard drive specification limit of 85°C. A linear function is used to construct a temperature sequence that gradually rises from 45°C to 70°C over 24 hours, then rapidly cycles to 85°C and falls back down, to simulate the temperature fluctuations and thermal shocks of the server under high load.

[0092] Voltage Sequence: The voltage range is set to ±10% of the nominal voltage of 12V. A voltage sequence including stable, slow-decreasing, and transient surges is constructed using a power function to simulate the instability of the power supply network.

[0093] Vibration sequence: Set a composite random vibration sequence that includes constant vibration at the server fan resonant frequency of 120Hz and simulated rack movement.

[0094] The two predefined parts are combined to form a set of strategies for monitoring the status of the storage subsystem.

[0095] The test environment generation module selects and combines strategies from the aforementioned set. For different specific needs under different working conditions, the module selects the "Database OLTP Simulation" load and matches it with a "24-hour temperature cycle sequence" and a "standard voltage fluctuation sequence" simulating high-density deployment, generating a benchmark test environment that highly simulates its real-world working environment. The test execution module receives the generated benchmark test environment and applies it to the solid-state drive under test.

[0096] Throughout the test, the data acquisition module periodically collected multimodal data once per minute. IOPS, average / maximum latency, 99.99% latency percentile, and throughput were recorded using the fio tool, and raw values ​​of all SSD-specific SMART attributes were obtained using specialized diagnostic tools provided by the hard drive manufacturer.

[0097] The analysis and modeling module performs time-series analysis on the collected multimodal data and correlates it with test environment parameters such as temperature, voltage, and vibration type. Based on the time-series analysis results, a dynamic failure characteristic model for this type of hard drive is constructed. The model processes the multimodal data, which integrates I / O performance indicators and internal health status, using the Z-score normalization method. For each feature dimension, its mean μ and standard deviation σ are calculated and saved from a large-scale historical failure dataset during the initial training phase of the model, and then used to form the time-series input vector. The GRU model input layer feature vector V_t is constructed, and during testing, each newly acquired feature vector V_t is subjected to the same normalization processing using the pre-stored μ and σ to ensure the consistency of data distribution.

[0098] The constructed GRU network input layer can receive input sequence tensors of shape (30, 12).

[0099] The first Bi-GRU layer contains 128 GRU units. As a bidirectional layer, it actually consists of a 64-unit forward GRU and a 64-unit backward GRU, whose outputs are concatenated along the feature dimension. This layer returns a complete output sequence for processing by the next layer.

[0100] The second Bi-GRU layer contains 64 GRU units: 32 forward units and 32 backward units. This layer only returns the hidden state output of the last time step of the sequence, which incorporates information from the entire input sequence.

[0101] The fully connected layer takes a 64-dimensional vector output from the second Bi-GRU layer and feeds it into a fully connected layer containing 32 neurons, using ReLU as the activation function.

[0102] The output layer is a fully connected layer containing a single neuron, using the sigmoid activation function. Its output is a continuous value between 0 and 1.

[0103] The activation functions within a GRU unit follow the standard GRU implementation, with update and reset gates using the Sigmoid activation function and candidate hidden states using the tanh activation function.

[0104] The Hardware_ECC_Recovered count of the tested solid-state drive showed a significant increase when the temperature exceeded 75°C, revealing its sensitivity to high temperatures. The GRU model identified a temperature-related periodic anomaly pattern.

[0105] The data obtained from the analysis and modeling module was input into the adjustment and processing module. It was found that whenever the ambient temperature cycled above 65°C, the SSD's 99.99% tail latency exhibited several spikes exceeding 500ms, far exceeding the normal value of <10ms. These latency spikes coincided closely with the increase in the Hardware_ECC_Recovered count. The system determined this to be a precursor to failure, characterized by "unstable NAND cell read performance and decreased data retention at high temperatures." Upon identifying this thermal sensitivity precursor, the adjustment and processing module immediately initiated targeted stress enhancement.

[0106] The test chamber was adjusted to raise the temperature of the faulty hard drive to its specification limit of 70°C and maintain it for 2 hours, while applying a full-disk random read load. Subsequently, the DUT-11 was powered off and left to stand at 40°C as specified by the JEDEC standard for 12 hours.

[0107] After power-on, a full disk data verification was performed on the hard drive. Multiple logical blocks were found to contain Reported_Uncorrectable_Errors, indicating that irreversible data bit flipping occurred after high-temperature read / write operations and power-off storage. This thermally related failure mode was successfully confirmed. The test then ended, and the warning module issued a warning, requesting confirmation of whether to immediately perform data migration and device replacement fault-tolerant operations.

[0108] Example 3

[0109] The dynamic failure characteristic model identifies various early failure features of the hard drive under test in real time, and adjusts the environmental data sequence based on the identified early failure features to obtain targeted stress enhancement. The specific implementation method is as follows:

[0110] The 14TB enterprise-class SSD, serial number DUT-08, has been running stably for 96 hours in a benchmark environment. This environment applies a simulated virtualized mixed read / write I / O load. The GRU neural network model in the analysis and modeling module, when processing the multimodal data time series of the most recent hour, has begun to show a statistically significant, slow but continuous downward trend in its output.

[0111] The system automatically performed correlation analysis and found that the decline in the health score was highly correlated with specific change patterns in two key indicators. SMART attribute analysis revealed that the raw value of Total_LBAs_Written grew at a normal rate, indicating stable host write volume. However, another internal indicator, Total_NAND_Writes, provided by vendor-specific SMART attributes or internal logs, grew at a rate far exceeding the host write rate. Simultaneously, the raw value of Program_Fail_Count began to show sporadic, non-periodic increases. During this period, the average write latency of the DUT-08 did not change significantly, but its WAF write amplification factor was calculated to be an abnormally high value of 8.x, far exceeding the average WAF of 3.x to 4.x for other SSDs in the same batch under the same load.

[0112] Based on the above multimodal data analysis, the system has identified a failure precursor characteristic of "write amplification runaway caused by inefficient firmware garbage collection algorithm". This precursor indicates that the SSD's flash translation layer is inefficient in its internal data organization and space reclamation when handling small block random writes. As a result, more than 8 NAND programming operations need to be performed internally to write 1 unit of host data, which will drastically consume the programming / erase life of the NAND flash memory.

[0113] Upon receiving the aforementioned warning, the adjustment module queries its internal hard drive test history database. The database records indicate that this type of write amplification-related failure precursor can be most effectively accelerated by applying an extremely high-intensity, small-block, 100% random-write I / O load. At this point, the adjustment module immediately generates a new, targeted I / O load configuration and instructs the test execution module to apply this stress only to the DUT-08.

[0114] Under this targeted stress test, the health of the DUT-08 deteriorated rapidly. The data acquisition module recorded that the raw values ​​of Program_Fail_Count and Erase_Fail_Count increased rapidly and continuously over several hours, while the normalized value of Media_Wearout_Indicator dropped sharply, indicating that the NAND lifespan was being rapidly depleted. After approximately four hours of targeted stress testing, the DUT-08's firmware triggered its internal protection mechanism, placing the hard drive in read-only mode and reporting a fatal hardware error to the host.

[0115] This embodiment uses intelligent closed-loop feedback to automatically apply precise stress load after identifying subtle signs of performance degradation. Within hours, it reproduces and confirms a serious firmware defect that might have taken weeks or even months to surface under standard testing.

[0116] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for testing the reliability of hard disk data, characterized in that, include: Based on different testing requirements, various load test data and environmental data sequences including temperature, voltage, and vibration are predefined to obtain a strategy set for monitoring the status of the storage subsystem; test parameters are selected for the strategy set to generate a benchmark test environment, and the hard drive under test is tested based on the benchmark test environment; when conducting reliability testing for the hard drive under test in a complex system, a database OLTP simulation load is selected, and a 24-hour temperature cycle sequence and a standard voltage fluctuation sequence simulating high-density deployment are matched to generate a benchmark test environment that highly simulates the real production environment; During normal system operation, multimodal data including I / O load performance indicators and SMART data values ​​are monitored in real time and collected periodically. Combining the multimodal data, the multi-dimensional change trend of the working status of the hard drive under test under different time, test environment and load conditions is analyzed online. Based on the multi-dimensional change trend, a dynamic failure feature model of the hard drive under test is constructed to predict the risk of data loss. The optimal model state is obtained by serializing the model parameters. Based on the optimal model state, a dynamic failure feature model of the hard disk under test is constructed to predict the risk of data loss. During the test, newly collected data is continuously collected and subjected to the same standardization process as in the training phase. The data is then input into the loaded baseline model to incrementally update the model. The dynamic failure characteristic model identifies various pre-failure features of the hard drive under test in real time. Based on these pre-failure features, the environmental data sequence is adjusted to obtain targeted enhancement pressure. The response data of the hard drive under test under the targeted enhancement pressure is recorded, and a warning is issued to confirm whether to immediately perform fault-tolerant operations such as data migration and device replacement. According to the type of the identified pre-failure features, a pre-built hard drive test history database is invoked to analyze the correlation between the pre-failure features and the environmental data. Based on the correlation, the test intensity of the corresponding environmental factors in the environmental data sequence is dynamically adjusted, and the application mode is changed to continue testing the hard drive under test. This forms an intelligent closed-loop test method that accelerates aging by analyzing the pre-failure features in real time based on the changes in the internal state of the hard drive under test, thereby applying targeted enhancement pressure to the hard drive under test.

2. The hard disk data reliability testing method according to claim 1, characterized in that, The specific implementation process of the predefined multiple load test data includes: predefining multiple I / O workload types for generating sequential read / write, random read / write, and mixed read / write I / O access modes; predefining the duration of different workloads including normal operation, high-intensity continuous operation, and intermittent operation; generating predefined load test data; setting temperature variation ranges including normal operation temperature and hard disk operating limit temperature, voltage fluctuation ranges including nominal voltage, overvoltage, and undervoltage, and vibration data including constant vibration frequency and composite vibration frequency; constructing temperature sequence, voltage sequence, and vibration sequence through linear functions and power functions respectively, forming an environmental data sequence, which together with the predefined load test data constitutes a strategy set for monitoring the status of the storage subsystem.

3. The hard disk data reliability testing method according to claim 1, characterized in that, The specific implementation process of generating the benchmark test environment includes: based on the initial performance indicators of the hard drive under test and the specific usage scenario of the hard drive under test, selecting test parameters from the strategy set and combining them to simulate the actual working state of the hard drive and generate the benchmark test environment.

4. The hard disk data reliability testing method according to claim 1, characterized in that, The specific analysis process of the multi-dimensional change trend includes: performing time series analysis on the multimodal data and performing correlation analysis between the strategy set and the multimodal data to obtain the multi-dimensional change trend of the working status of the hard disk under test under different time, test environment and load conditions.

5. The hard disk data reliability testing method according to claim 1, characterized in that, The specific identification process for each failure precursor feature of the hard drive under test includes: cleaning the real-time collected multimodal data to form a feature vector, which is then input into the dynamic failure feature model to analyze the evolution process of the hard drive from health to failure, and obtaining the model output results; monitoring key indicators of hard drive failure features in the multimodal data, including the decline in I / O load performance indicators and abnormal SMART data values; calculating a dynamic threshold based on the weighted calculation of the key indicators; when the model output results and the key indicators of hard drive failure features exceed the dynamic threshold, identifying the types of failure precursor features, including the decline in various performance aspects of the hard drive, based on the dynamic threshold anomaly detection algorithm.

6. A hard disk data reliability testing system, characterized in that, include: The test strategy definition module is used to predefine various load test data and environmental data sequences including temperature, voltage and vibration conditions based on different test requirements, and obtain a set of strategies for monitoring the status of the storage subsystem. The test environment generation module is used to select test parameters for the strategy set and generate a benchmark test environment. The test execution module is used to receive the generated benchmark test environment to test the hard drive. When performing reliability testing on the hard drive under test in a complex system, it selects the database OLTP to simulate the load and matches a 24-hour temperature cycle sequence and a standard voltage fluctuation sequence that simulates high-density deployment to generate a benchmark test environment that highly simulates the real production environment. The data acquisition module is used to monitor and periodically collect multimodal data, including I / O load performance indicators and SMART data values, in real time during normal system operation. The analysis and modeling module is used to combine the multimodal data to analyze the multi-dimensional change trend of the working status of the hard drive under test under different time, test environment and load conditions, and to construct a dynamic failure feature model of the hard drive under test based on the multi-dimensional change trend. The optimal model state is obtained by serializing the model parameters. Based on the optimal model state, a dynamic failure feature model of the hard disk under test is constructed to predict the risk of data loss. During the test, newly collected data is continuously collected and subjected to the same standardization process as in the training phase. The data is then input into the loaded baseline model to incrementally update the model. The adjustment processing module is used to identify various failure precursor features of the hard drive under test in real time according to the dynamic failure feature model, adjust the environmental data sequence based on the failure precursor features to obtain targeted enhancement pressure, and record the response data of the hard drive under test under the targeted enhancement pressure. The warning module is used to issue warnings and confirm whether to immediately perform fault-tolerant operations such as data migration and device replacement. Based on the type of identified failure precursor characteristics, it calls a pre-built hard drive test history database to analyze the correlation between the failure precursor characteristics and environmental data. Based on the correlation, it dynamically adjusts the test intensity of corresponding environmental factors in the environmental data sequence, changes the application mode, and continues to test the hard drive under test. This forms an intelligent closed-loop test method that accelerates aging by analyzing the failure precursor characteristics in real time based on the internal state changes of the hard drive under test, thereby applying targeted enhancement pressure to the hard drive under test.

Citation Information

Patent Citations

  • Method and device for testing pressure of hard disk of server

    CN116775396A

  • Solid state disk reliability analysis method, device and equipment and medium

    CN118899026A

  • Hard disk fault prediction method and device, electronic equipment and storage medium

    CN119806923A