A method for testing total write data volume of a solid state disk

By dynamically adjusting the test load and conducting multi-dimensional health assessments, the efficiency and accuracy issues of testing the total amount of data written to solid-state drives (SSDs) in existing technologies have been resolved. This enables predictive assessment of SSD health status and timely detection of failure risks, generating detailed test reports.

CN121260223BActive Publication Date: 2026-04-07SHENZHEN JINGCUN TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-02
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing solid-state drive (SSD) total write data volume testing schemes cannot effectively simulate the actual wear and tear process, lack multi-dimensional health assessment, and cannot detect potential performance degradation or failure risks in a timely manner.

Method used

A method for testing the total amount of data written to a solid-state drive is adopted, including initializing the test environment and configuring adaptive parameters, creating parallel data acquisition threads, establishing a health assessment model, dynamically adjusting the test load based on a dynamic acceleration stress model, and generating a test report.

Benefits of technology

By dynamically adjusting the test load, the test cycle is significantly shortened, the test efficiency and accuracy are improved, and predictive assessment of the health status of solid-state drives is achieved. Potential failure risks can be detected in a timely manner, and detailed test reports are provided, offering more valuable data support for SSD manufacturers and consumers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121260223B_ABST
    Figure CN121260223B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data storage, and discloses a solid state disk total write data volume test method, which comprises the following steps: initializing a test environment and adaptive parameter configuration of a solid state disk; creating a plurality of parallel data acquisition threads, collecting test data through the data acquisition threads, and storing the collected test data into a database; establishing a preset health degree evaluation model, analyzing the collected test data in real time according to a preset health degree index system, dynamically adjusting a test load according to a real-time state of the solid state disk based on a dynamic acceleration stress model, executing corresponding test operations according to the adjustment result, and recording test results; and analyzing the test results and automatically generating a test report according to the analysis result. Through the introduction of the dynamic acceleration stress model, the test load is dynamically adjusted according to the real-time state of the SSD, and the efficiency and accuracy of the test are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data storage technology, and in particular to a method for testing the total amount of data written to a solid-state drive. Background Technology

[0002] With the rapid development of emerging technologies such as cloud computing and big data, the widespread adoption of internet services, and the acceleration of enterprise digital transformation, solid-state drives (SSDs) are increasingly widely used in data storage. As a type of external computer storage device that uses flash memory chips as the non-volatile storage medium, SSDs offer advantages such as high speed, small size, low power consumption, and no noise. However, the lifespan of SSDs remains a primary concern for both enterprise users and consumers. The lifespan of an SSD is typically measured by its tolerable total bytes written (TBW), but actual lifespan can vary due to factors such as controller algorithms, operating temperature, and write amplification.

[0003] Currently, common TBW testing schemes mainly include installing the SSD under test as a data disk, using simple data filling tools or scripts to continuously write to the SSD, and periodically manually checking the "Total Writes" attribute in the SMART information. This method has the problem of poor test load realism and cannot effectively simulate the actual wear process [4]. In addition, existing technologies lack multi-dimensional health assessment of SSDs and cannot detect potential performance degradation or failure risks in a timely manner. Summary of the Invention

[0004] This invention provides a method for testing the total amount of data written to a solid-state drive (SSD), which can solve the technical problem of lacking multi-dimensional health assessment of SSDs and failing to detect potential performance degradation or failure risks in a timely manner.

[0005] To solve the above-mentioned technical problems, one technical solution adopted by the present invention is: to provide a method for testing the total amount of data written to a solid-state drive, the method comprising:

[0006] Initialize the test environment and adaptive parameter configuration for the solid-state drive;

[0007] Create multiple parallel data acquisition threads, collect test data through the data acquisition threads, and store the collected test data into the database;

[0008] Establish a preset health assessment model, analyze the collected test data in real time according to the preset health index system, and dynamically adjust the test load according to the real-time status of the solid-state drive based on the dynamic accelerated stress model. Based on the adjustment results, perform corresponding test operations and record the test results.

[0009] Analyze the test results and automatically generate a test report based on the analysis.

[0010] The beneficial effects of this invention are as follows: By introducing a dynamic accelerated stress model and abandoning the traditional fixed load mode, the test load (such as read / write ratio, queue depth, and block size) is dynamically adjusted according to the real-time status of the SSD (temperature, wear, performance degradation), effectively simulating extreme real-world usage scenarios, accelerating the aging process, significantly shortening the test cycle, and improving test efficiency and accuracy; a multi-dimensional health indicator system is constructed, going beyond single "write volume" and "error rate," introducing soft criteria such as performance consistency, latency stability, and SMART parameter trend analysis, achieving predictive assessment of the SSD's "health status," and timely identifying potential performance degradation or failure risks; a fully automated closed-loop testing ecosystem is realized, integrating hardware control through a central control script. The test execution, data collection, intelligent analysis, and decision-making processes are all automated, achieving a fully unattended workflow from test initiation to final report generation. This significantly improves the automation level of testing and reduces the need for manual intervention. A data integrity verification step has been added, embedding data validation functionality into read / write tests to ensure that data can not only be written but also read correctly at the end of the SSD's lifespan, detecting silent data errors and providing crucial assurance for assessing data reliability. Through multi-dimensional data collection and intelligent analysis, a detailed test report is generated, including configuration parameters, collected data, trend charts, termination criteria, and the final TBW value. This provides SSD manufacturers with richer and more valuable data support for quality assessment and lifespan prediction, as well as for consumers in product selection, compared to traditional TBW testing. Attached Figure Description

[0011] Figure 1 This is a flowchart illustrating the method for testing the total amount of data written to a solid-state drive according to the first embodiment of the present invention.

[0012] Figure 2 yes Figure 1 A flowchart illustrating step 1.

[0013] Figure 3 yes Figure 1 A flowchart illustrating step 2.

[0014] Figure 4 yes Figure 1 A flowchart illustrating step 3.

[0015] Figure 5 yes Figure 1 A flowchart illustrating step 4. Detailed Implementation

[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0017] The terms "comprising" and "having," and any variations thereof, used in this invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the steps or units listed, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such process, method, product, or apparatus.

[0018] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0019] Figure 1 This is a flowchart illustrating the method for testing the total amount of data written to a solid-state drive according to the first embodiment of the present invention. Figure 1 As shown, the system includes hardware and software components:

[0020] Step 1: Initialize the test environment and adaptive parameter configuration of the solid-state drive;

[0021] Step 2: Create multiple parallel data acquisition threads, collect test data through the data acquisition threads, and store the collected test data into the database;

[0022] Step 3: Establish a preset health assessment model, analyze the collected test data in real time according to the preset health index system, and dynamically adjust the test load according to the real-time status of the solid-state drive based on the dynamic accelerated stress model. Based on the adjustment results, perform corresponding test operations and record the test results.

[0023] Step 4: Analyze the test results and automatically generate a test report based on the analysis results.

[0024] To meet the presentation requirements of "no tables, no code," I will use a purely text-based narrative, retaining technical details while optimizing the expression logic. I will re-explain the four-step implementation plan for SSD testing from the dimensions of practical procedures, core logic, and tool selection:

[0025] Step 1: Initialize the SSD test environment and adaptive parameters. First, standardize the hardware environment: unify the SSD connection interface protocol, such as NVMe 1.4 or SATA III, and use an adapter card to ensure compatibility with multiple interface models and prevent performance deviations caused by interface differences; fix the test host configuration, such as using an Intel Xeon E5 series CPU and 64GB DDR4 memory, and disable CPU power saving mode and memory prefetch function to ensure stable output of computing resources; at the same time, connect high-precision temperature and voltage sensors to monitor the SSD casing temperature (accuracy controlled within ±1℃) and power supply voltage (±0.01V) in real time, providing environmental data support for subsequent load adjustments.

[0026] Next, perform SSD preprocessing: Use manufacturer-specific tools (such as Samsung Magician or Intel SSD Toolbox) to perform a Secure Erase operation, clear all residual data, and restore the SSD to its factory state; simultaneously record basic SSD information, including model, capacity, firmware version, flash memory type (TLC or QLC), and nominal PE cycle count. This information will serve as the core basis for subsequent parameter configuration.

[0027] The adaptive parameter configuration is dynamically generated based on SSD attributes: for example, for the initial test load, the initial IOPS is set to 10k and the read-write ratio to 3:1 when the capacity is ≤1TB, and the initial IOPS is increased to 20k and the read-write ratio to 1:1 when the capacity is >1TB, simulating real-world usage scenarios for SSDs of different capacities; the data sampling interval is adjusted according to the flash memory type, with QLC flash memory having a shorter lifespan and the sampling interval set to 10 seconds, while TLC flash memory is set to 30 seconds; the safety thresholds refer to the manufacturer's nominal values, such as the upper temperature limit being 5℃ lower than the manufacturer's nominal value, and the voltage fluctuation threshold being controlled within ±5%; the test duration baseline is set according to the capacity, with 24 hours for 1TB SSDs and 48 hours for 2TB SSDs, which can be manually extended according to actual needs.

[0028] Step 2: Parallel data acquisition thread setup and database storage. Multi-threaded parallel acquisition of full-dimensional data ensures data real-time performance and traceability, while database is used for standardized storage.

[0029] Threads are divided into three data categories: "Performance," "Health," and "Environment," to avoid resource contention. The performance acquisition thread primarily collects metrics such as random read / write speed, IOPS, average latency, and queue depth, obtaining data through the real-time output function of the fio tool. This thread has a high priority to ensure uninterrupted performance data acquisition. The health status thread collects remaining lifetime percentage, used PE cycle count, number of newly added bad blocks, and SMART attributes (such as 0x05 bad block count and 0xC2 temperature attribute), using the smartctl command to read internal SSD health data. This thread has a medium priority. The environment monitoring thread collects casing temperature, controller temperature (obtained via the manufacturer's API), and power supply voltage, using the sensor SDK and ipmitool tools for data acquisition. This thread also has a medium priority.

[0030] The thread implementation employs an independent thread execution mode. For example, three threads are created using Python's `threading` module, each focusing on collecting one type of data. Data is buffered using a queue to prevent data blocking during the collection process. The collected data is categorized and sent to a database. A combination of the time-series database InfluxDB and the relational database MySQL is used. InfluxDB stores high-frequency time-series data, such as read / write speeds and IOPS in performance data. Each data entry includes a timestamp, SSD model identifier, and specific metric values. MySQL stores static configuration information and test metadata, such as SSD basic parameters, test task information, and load adjustment logs, ensuring clear data categorization and quick retrieval for subsequent analysis.

[0031] Step 3: Health assessment model construction and dynamic load adjustment. The health status of the SSD is assessed by real-time data, and the test load is dynamically optimized based on the status to balance test efficiency and accuracy.

[0032] The health assessment adopts a hybrid model of "rules + machine learning". First, a multi-level indicator system is established and weights are assigned: performance degradation indicators account for 40% of the weight, including read and write speed degradation rate (weight 20%), with the health threshold set as degradation rate <10% for healthy, 10%-20% for warning, and >20% for failure; lifespan loss indicators account for 30% of the weight, with PE cycle usage rate as the core (weight 30%), usage rate <50% for healthy, 50%-80% for warning, and >80% for failure; reliability indicators account for 30% of the weight, including the number of new bad blocks (weight 20%) and temperature stability (weight 10%), with 0 new bad blocks and temperature fluctuation within ±2℃ for healthy, >3 new bad blocks or temperature fluctuation within ±5℃ for warning, and >10 new bad blocks or temperature fluctuation within ±8℃ for failure. The health score is calculated by summing the standardized values ​​of each secondary indicator × their weights. The standardized values ​​map the actual values ​​of the indicators to a score of 0-100. For example, a PE utilization rate of 30% corresponds to 70 points, and 80% corresponds to 20 points. The final score is ≥80 points for healthy, 60-80 points for warning, and <60 points for failure.

[0033] Dynamic load adjustment is based on health scores and real-time status (temperature, performance fluctuations) to formulate strategies: When the health score is ≥80 and the temperature is <70℃, the test load is increased, IOPS increases by 20%, and the queue depth is increased by 1 (maximum not exceeding 32) to speed up the test progress; when the health score is 60-80 or the temperature is 70-85℃, the current IOPS is maintained, and the read-write ratio is adjusted to 1:1 to increase write pressure to further verify SSD stability; when the health score is <60 or the temperature is ≥85℃, the load is reduced, IOPS decreases by 50%, and if the temperature continues to rise, the test is paused and the cooling mechanism is triggered; if performance fluctuations >15% are found in 3 consecutive samples, the test mode is switched from random read / write to sequential read / write to investigate whether performance fluctuations are caused by data fragmentation.

[0034] During test execution, the fio tool is called to load the dynamically adjusted parameter configuration file, synchronously recording the triggering conditions for each adjustment, the parameters before and after the adjustment (such as old IOPS, new IOPS), and the performance data before and after the adjustment. These records are stored in the MySQL test adjustment log table, which facilitates subsequent tracing of the impact of load adjustment on the SSD status.

[0035] Step 4: Test Result Analysis and Automatic Report Generation. By analyzing test data from multiple dimensions, intuitive and quantifiable reports are output to provide a basis for SSD quality assessment.

[0036] The results analysis unfolds across four core dimensions: health trend analysis, which uses time-series curves (with the help of matplotlib) to observe the rate of change in health over the testing period, such as "health decreases by 2% every 24 hours"; performance limit analysis, which records the maximum IOPS and minimum latency within the safe load range, compares them with the manufacturer's nominal values ​​to calculate the deviation rate, and determines whether the actual performance meets the standards; lifespan prediction analysis, which estimates the remaining lifespan (remaining PE cycles ÷ daily consumption) based on the PE cycle consumption rate during the testing period (e.g., "consuming 5 PE cycles per day") and the total nominal PE cycle count of the SSD; and abnormal event analysis, which statistically analyzes the number of load adjustments, overheating triggers, and performance drops during the testing process to identify the root causes of anomalies, such as "when the temperature is >80℃, the SSD read / write speed decreases by 15%, requiring optimization of the heat dissipation design."

[0037] The automatic report generation is implemented using a template engine. First, an HTML report template is built using Jinja2. The template contains four parts: test summary (basic SSD information, test duration, and environment parameters), core indicator data (health score, maximum performance value, and estimated remaining lifetime), trend charts (health change curve, performance fluctuation chart, and temperature change chart), and anomaly analysis and suggestions (optimization directions proposed for abnormal events). Then, the HTML report is converted into PDF format using the weasyprint tool to ensure a consistent report format, making it easy to share and archive.

[0038] In addition, the entire testing process needs to be equipped with an exception handling mechanism: if the data acquisition thread crashes, the thread will be automatically restarted and the logs will be recorded through the thread status monitoring mechanism; if the database connection is interrupted, the acquired data will be temporarily written to a local CSV file and then re-uploaded to the database after the connection is restored; if an unexpected power outage occurs, the test progress will be restored through the SSD's SMART logs and local backup data to avoid data loss.

[0039] If you need more details on specific tool operation methods for a particular step (such as fio parameter configuration methods, smartctl command usage tips) or additional scenarios for exception handling, please feel free to let me know.

[0040] Figure 2 yes Figure 1 The flowchart for step 1 is as follows: Figure 2As shown, step 1 includes: step 101, installing the solid-state drive under test as the test data disk; step 102, setting test environment parameters, which include at least temperature control and power quality monitoring; step 103, configuring the test software environment, which includes at least a central control script, a data filling tool, and a health management tool; and step 104, initializing the test database and establishing a data recording format.

[0041] Step 101: Install the SSD under test as the test data drive. Before installation, confirm the interface compatibility of the test host (e.g., NVMe interfaces must match the PCIe slot version, SATA interfaces must correspond to the SATA III bus). Wear anti-static gloves throughout the process to avoid electrostatic damage to the SSD. During installation, prioritize using a dedicated interface slot to avoid sharing the PCIe lane with other high-power devices (such as graphics cards), preventing performance deviations caused by bandwidth contention. If it is a 2.5-inch SATA SSD, it must be secured to the host's hard drive bracket with screws to prevent vibration from affecting data acquisition during testing. When connecting the power cable, use the host's native power connector instead of an expansion splitter to reduce the risk of power instability. After installation, confirm through the host BIOS that the SSD has been recognized as a "data drive" (not a system drive) to prevent system processes from consuming SSD resources and ensure that test data comes only from dedicated test operations.

[0042] Step 102: In setting the test environment parameters (temperature control, power quality monitoring), temperature control uses a constant temperature test chamber to build a closed test space. The test host and SSD are placed inside the chamber, and the target temperature is set (e.g., 25℃±1℃ for room temperature testing, 50℃±1℃ for high temperature stress testing). Two temperature sensors are placed inside the chamber, one close to the surface of the SSD casing and the other near the main control chip, to provide real-time feedback on local temperature and avoid temperature monitoring deviations caused by a single sensor. For power quality monitoring, a power analyzer needs to be connected to the host power input to monitor input voltage fluctuations (set threshold of ±2%, e.g., allowable fluctuation range of 215.6V-224.4V for 220V AC mains), current ripple coefficient (threshold ≤50mV), and instantaneous power peak. At the same time, a miniature ammeter is connected in series at the SSD power interface to monitor the real-time power consumption of the SSD separately (e.g., standby power consumption ≤0.5W, full load power consumption ≤8W). If the power parameters exceed the threshold, an audible and visual alarm is immediately triggered and the test is paused to prevent abnormal power supply from damaging the SSD.

[0043] Step 103: Configuring the test software environment (central control script, data population tool, health management tool). The central control script is written in Python, and its core functions include test process scheduling (such as starting / pausing acquisition threads, triggering load adjustments), logging (real-time recording of test events and exception information, with log levels of "info / warn / error"), and status monitoring (displaying the running status of each thread and the real-time health of the SSD). The script needs to integrate a configuration file reading module to support quick adjustment of test parameters (such as sampling interval and load baseline) by modifying the JSON file. The data population tool adopts a "full capacity population" strategy, using the dd command (such as dd if= / dev / zero of= / dev / nvme0n1 bs=1G) or a vendor-specific population tool (such as Intel Data Center Tool) to write full-capacity random data to the SSD, simulating the data distribution state in actual use and avoiding test data distortion in an empty disk state. The health management tool is based on smartctl and integrates SDKs provided by manufacturers (such as Samsung SSD Toolbox SDK and Kingston SSD Manager API). In addition to obtaining regular SMART attributes (such as bad block count and PE cycle count), it can also read deep health data (such as NAND flash memory erase and write distribution and cache hit rate), providing more comprehensive data support for health assessment.

[0044] Step 104: Initialize the test database and establish the data record format. First, complete the database deployment: Install InfluxDB (a time-series database used to store high-frequency collected data) and MySQL (a relational database used to store static data) on the test server, configure independent storage directories for both (it is recommended to use a high-speed SSD as the database storage disk to improve data writing speed), and create a dedicated test account and grant data read and write permissions (do not use the root account directly for testing). Data recording formats should be designed according to "data type - repository": In InfluxDB, create three measurements (performance data / health data / environment data). Each measurement should include a timestamp (accurate to milliseconds), a unique SSD identifier (e.g., a combination of "Model_Serial", such as "Samsung_990Pro_S1AXNS0KB00001"), and core metric fields (e.g., performance data includes read bandwidth / write bandwidth / IOPS, health data includes remaining lifetime / number of newly added bad blocks). In MySQL, create two basic tables (SSD information table / test task table). The SSD information table records static parameters such as model, capacity, and firmware version. The test task table records the test number, test type (e.g., room temperature test / high temperature test), start time / end time. All fields should have data type constraints (e.g., capacity field set to INT type, unit GB; time field set to DATETIME type) to ensure a consistent data format for easy subsequent analysis and retrieval. After the database initialization is complete, a data write test needs to be performed to simulate 10 minutes of data collection and writing to verify the data storage speed (InfluxDB write rate ≥ 100 records / second) and integrity (no data loss or field misalignment).

[0045] Based on the environmental parameters in step 102 (such as the set temperature of the constant temperature chamber) and the database configuration in step 104, the adaptive parameters are further optimized: if the current test is at a high temperature (50℃), the temperature safety threshold is lowered to the manufacturer's nominal value - 8℃ (e.g., the original threshold of 90℃ is adjusted to 87℃); combined with the SSD status after data filling in step 103 (full capacity data has been written), the read-write ratio of the initial test load is adjusted to 2:1 (closer to the actual read-write scenario after data filling), and other parameter configurations (such as sampling interval and test duration baseline) are still dynamically generated according to the SSD attributes. After the configuration is completed, the parameters are synchronously written to the configuration file of the central control script and the MySQL test task table.

[0046] Before the collection thread in step 2 starts, the health management tool in step 103 needs to be called through the central control script to read the initial health data of the SSD and write it to the database; the load adjustment command in step 3 needs to be issued by the central control script to the test tool (such as fio) in step 103; when the report is generated in step 4, complete test data needs to be retrieved from the database in step 104 to ensure that the analysis is based on the full and accurate collection results.

[0047] If the SSD is not recognized after installation in step 101, check the interface tightness and BIOS interface enable status; replace the interface / power cable if necessary. If temperature control fails in step 102 (e.g., temperature fluctuation in the constant temperature chamber exceeds ±2℃), calibrate the temperature sensor and clean the heat dissipation channels. If software environment configuration errors occur in step 103 (e.g., the script cannot call the SDK), verify SDK version compatibility and reconfigure environment variables. If database write fails in step 104, check the database service status and remaining storage directory space to ensure that all pre-test steps are normal before starting the test.

[0048] Figure 3 yes Figure 1 The flowchart for step 2 is as follows: Figure 3 As shown, step 2 includes: step 201, starting the test tasks within the test cycle according to the central control script and creating a data acquisition thread; step 202, performing performance data acquisition, health data acquisition and physical environment data acquisition through the data acquisition thread; step 203, ensuring the synchronization of timestamps of all acquired data through the NTP protocol, and writing all timestamped data into the InfluxDB time series database in real time.

[0049] Step 201: Before starting the test task and creating the data acquisition thread based on the central control script, the central control script (the Python script configured in step 103) needs to complete three pre-checks: "parameter verification - task initialization - thread creation". First, it verifies whether the optimized adaptive parameters (such as sampling interval and load baseline) in step 1 have been synchronized to the script configuration file, and at the same time confirms that the InfluxDB and MySQL services in step 104 are running. If the verification passes, the script automatically generates a unique test task ID (in the format "TEST_YYYYMMDD_001") and writes the task information (test type, target load, planned duration) into the "test task table" of MySQL.

[0050] Thread creation follows the principle of "one independent thread for one type of data." The script calls the `threading` module to generate three acquisition threads: a performance data acquisition thread, a health data acquisition thread, and a physical environment data acquisition thread. Each thread initializes with its own configuration: the performance thread reads dynamic parameters from the fio tool (such as IOPS and read / write ratio), the health thread loads the call paths between smartctl and the vendor's SDK, and the physical environment thread obtains the communication port between the sensor and the power analyzer (such as " / dev / ttyUSB0" for USB-to-serial conversion). To avoid thread resource contention, the script assigns different priorities to the three threads (highest for the performance thread, followed by the health and environment threads) and creates a temporary data buffer queue (1000 records) to ensure that acquired data is temporarily stored and not lost.

[0051] After startup, the script interface displays the status of each thread ("Running / Paused / Abnormal") and the progress of the test task (e.g., "Running for 1 hour / 23 hours remaining") in real time, and generates an initial log (level "info"), recording information such as "task start time, number of threads, and initial parameters". The log file is named according to the test task ID (e.g., "TEST_20251114_001.log") and stored in the specified directory.

[0052] Step 202: In the multi-dimensional data acquisition execution (performance / health / physical environment), three acquisition threads execute acquisition operations in parallel at a preset frequency. The acquisition logic and tool calls strictly rely on the software configuration in step 103.

[0053] Performance data collection: The performance thread periodically calls the fio tool (call interval = sampling interval set in step 1, such as 30 seconds for TLC SSD) to perform preset load tests (such as random read / write, IOPS=20k). fio outputs metrics such as read bandwidth (MB / s), write bandwidth (MB / s), IOPS, and average latency (ms) in real time. The thread extracts data by parsing fio's standard output stream (stdout), removes outliers (such as data with instantaneous bandwidth fluctuations exceeding 50%), and temporarily stores it in the cache queue.

[0054] Health data collection: The health thread collects data at a slightly lower frequency (e.g., once per minute to avoid frequent reads affecting SSD lifespan). First, it reads the basic SMART attributes of the SSD (e.g., 0x05 bad block count, 0xC2 temperature, 0xE9 PE cycle count) using the smartctl command. Then, it calls the manufacturer's SDK (e.g., Samsung SSD Toolbox SDK) to obtain deeper data (e.g., NAND flash memory block erase / write count distribution, SLC cache hit rate). After collection, the two types of data are integrated into a "health data package", supplemented with the SSD's unique identifier ("Model_Serial" set in step 104), and passed to the cache queue.

[0055] Physical environment data acquisition: The environment thread acquires data in real time (1 second / time to ensure capture of instantaneous fluctuations). The temperature of the SSD casing and the main controller is read through the sensor SDK (accuracy ±1℃). The host input voltage and current ripple coefficient are obtained through the power analyzer. The real-time power consumption of the SSD (in W) is read through the micro ammeter. If a parameter exceeds the threshold set in step 102 (such as temperature > 87℃, voltage fluctuation > ±2%), the thread immediately generates a "warn" level log and triggers the script's warning mechanism (interface pop-up prompt) synchronously, but does not interrupt the acquisition (the test is only paused after the parameter continues to exceed the threshold for 10 seconds).

[0056] During the data acquisition process, if a thread encounters an exception (such as a failed fio call or a sensor disconnection), the thread will automatically attempt to restart (up to 3 times). If the restart fails, it will be marked as "exception" and an "error" level log will be triggered to record the reason for the exception (such as "fio process crash" or "sensor port not responding"). Other threads will continue to run to avoid the overall test being interrupted due to a single thread failure.

[0057] Step 203: NTP time synchronization and real-time writing of collected data to InfluxDB. To ensure that all data timestamps are consistent (avoiding data misalignment caused by time deviation during subsequent analysis), NTP time synchronization configuration must be completed before the test starts: Configure both the test host (running the collection thread) and the database server (deploying InfluxDB) to synchronize to the same NTP server (preferably a high-precision NTP server within the enterprise; if none is available, use a public server such as "ntp.aliyun.com"). Set the synchronization frequency to 1 minute / time, ensuring that the time deviation between the two is ≤1 millisecond. During the test, the central control script checks the time synchronization status every 5 minutes. If the deviation exceeds 5 milliseconds, it automatically triggers forced synchronization. After synchronization is completed, a "time calibration log" is recorded.

[0058] Data writing employs a "cache queue batch write" strategy: The central control script starts an independent "data write thread" to extract data in batches from the cache queue (10 data entries per batch to reduce the number of InfluxDB connections). The data structure is organized according to the measurement format set in step 104 (different measurements correspond to three types of data: performance, health, and environment), and millisecond-level timestamps (taken from the system time synchronized with the test host) and test task IDs are added. A connection is established through the InfluxDB Python client library (such as "influxdb-client") to write the batch data to the corresponding measurement. After successful writing, the written data is deleted from the cache queue. If writing fails, the data is saved back to the end of the queue, waiting for the next batch to be retried (maximum of 5 retries; if it still fails, it is written to the local backup file "backup_YYYYMMDD.csv").

[0059] During the writing process, the script displays the amount of data written in real time (e.g., "1200 performance data entries, 120 health data entries, and 36000 environment data entries have been written") and synchronizes the writing status ("normal / retry / backup") to the log file to ensure the integrity of subsequent traceable data writing.

[0060] After completing the real-time data writing in step 203, subsequent processes (original step 3: health assessment and dynamic load adjustment, original step 4: result analysis and report generation) can directly retrieve data from InfluxDB: the health assessment model in step 3 calculates the health score by querying the "health data" and "performance data" in InfluxDB, and the dynamic load adjustment command is issued by the central control script based on real-time data; the result analysis in step 4 exports the full-cycle collected data from InfluxDB, combines it with the test task information in MySQL, and generates a complete test report, realizing the end-to-end data connection of "collection-storage-analysis-report".

[0061] If NTP time synchronization fails (e.g., the NTP server crashes), the test host automatically switches to the local hardware clock (RTC) and marks the time synchronization status as "abnormal". At the same time, it reduces the data write batch size (from 10 records / batch to 5 records / batch) to reduce the impact of time deviation on data association. Once the NTP server recovers, it immediately resynchronizes.

[0062] If the InfluxDB service is temporarily interrupted, the "data write thread" will write all data in the cache queue to the local backup file. It will attempt to reconnect to InfluxDB every 30 seconds. After the connection is restored, it will automatically rewrite the data in the backup file into InfluxDB in timestamp order. After the rewriting is completed, the backup file will be deleted to ensure that no data is lost.

[0063] If the data acquisition thread experiences a "data acquisition timeout" (e.g., no response after calling smartctl for more than 10 seconds), the thread will automatically skip the current acquisition cycle and enter the next cycle, while recording a "timeout log" to avoid thread blocking, and resume normal acquisition in subsequent cycles.

[0064] Figure 4 yes Figure 1 The flowchart for step 3 is as follows: Figure 4 As shown, step 3 includes: step 301, analyzing the collected data in real time according to the preset health index evaluation model; step 302, dynamically adjusting the test load by switching between multiple modes such as random read / write mode, sequential write mode, and mixed read / write mode according to the dynamic accelerated stress model, based on temperature, wear degree, and performance degradation; step 303, performing corresponding test operations according to the dynamic adjustment results and recording the test results.

[0065] Step 301: Real-time analysis of collected data based on a preset health index assessment model. The core of real-time analysis is to extract full-dimensional data from InfluxDB and calculate the health score according to the preset model, providing a basis for subsequent load adjustments. First, the central control script starts an independent "health analysis thread," with the analysis frequency synchronized with the performance data collection interval (e.g., 30 seconds / time for TLC SSDs to ensure data timeliness). Before each analysis, the integrity of InfluxDB data is verified (if more than 3 records of a certain type of data are missing within the past minute, the average of the previous 3 times is used to fill in the missing data, and a "warn" level log is generated to mark the data gap).

[0066] The analysis process strictly follows the logic of "multi-level indicator weighted calculation": First, extract health data (remaining lifetime percentage, PE cycle rate, number of new bad blocks) and performance data (read / write speed degradation rate, IOPS fluctuation range) from InfluxDB. Then, combine this with physical environment data (SSD casing temperature, temperature fluctuation range), and calculate the standardized value of each secondary indicator according to preset weights (performance degradation 40%, lifetime loss 30%, reliability 30%). For example, when the PE cycle rate reaches 60%, the standardized value is 40 points (100 - (60 / 80) × 100, 80% is the warning threshold); when the read / write speed degradation rate is 15%, the standardized value is 85 points (100 - 15). Then, sum up the standardized values ​​of all secondary indicators according to their weights to obtain the total health score (out of 100), and determine the current status: a score ≥ 80 is "healthy", 60-79 is "warning", and < 60 is "faulty".

[0067] After the analysis is completed, the script synchronously writes the health score, status judgment result and raw data of each indicator into the "Health Analysis Table" of MySQL (including test task ID, analysis timestamp, total score and scores of each secondary indicator). At the same time, the script interface updates the health status in real time (e.g., green to mark "healthy" and yellow to mark "warning"). If it is judged as "fault", it immediately triggers an audible and visual alarm (sharing the alarm module with the environmental warning in step 202) to remind the operation and maintenance personnel to pay attention.

[0068] Step 302: In the dynamic adjustment of the test load based on the dynamic accelerated stress model, the load adjustment is based on "health status as the core, temperature, wear, and performance degradation as auxiliary factors", combined with multiple read / write mode switching, to balance test intensity and SSD security. The adjustment logic is executed by the "load scheduling module" of the central control script. The specific strategy is as follows:

[0069] The SSD is in a "healthy" state (score ≥ 80 points): At this time, the SSD performance and lifespan are sufficiently redundant, and the load intensity is increased first to accelerate the exposure of potential problems. If the current temperature is < 70℃ (below the high temperature warning threshold of 85℃ set in step 102) and the PE cycle usage rate is < 50% (low wear), then the IOPS of the fio tool will be increased by 20% (e.g., from 20k to 24k), and the queue depth will be increased by 1 (not exceeding 32). At the same time, the "mode rotation mechanism" will be started, switching the read and write modes every 1 hour - first performing 30 minutes of random read and write (read-write ratio 3:1, simulating a server scenario), and then performing 30 minutes of mixed read and write (read-write ratio 1:1, simulating a personal terminal scenario), covering the load pressure of multiple scenarios through mode switching.

[0070] A health status of "Warning" (60-79 points) indicates that the SSD has experienced performance degradation or minor wear. It's necessary to maintain load intensity and focus on critical stress points. If the temperature is between 70-85℃ (close to the warning threshold) or the read / write speed decreases by 15%-20%, maintain the current IOPS and fix the read / write mode to "Sequential Write Mode" (continuous write pressure more easily exposes NAND flash stability issues). Simultaneously, shorten the mode switching interval to 30 minutes (switch between sequential write and mixed read / write every 15 minutes to avoid test bias caused by a single mode). If the PE cycle usage reaches 80%-90% (high wear), appropriately reduce IOPS (down 10%) to prevent excessive wear from affecting the accuracy of lifespan assessment.

[0071] A health status of "Fault" (score < 60 points) indicates that the SSD has exhibited significant abnormalities (e.g., more than 10 new bad blocks or temperature > 85°C). The load must be immediately reduced to protect the device and verify the reversibility of the anomaly. First, reduce IOPS by 50% (e.g., from 20k to 10k), and switch the read / write mode to "Low-Pressure Random Read / Write" (read / write ratio 4:1, reducing write pressure). If the temperature continues to exceed 85°C or performance degradation exceeds 20%, pause the load adjustment and trigger "Cooling Buffer" (reduce the temperature of the constant temperature chamber by 5°C for 5 minutes). After the temperature drops below 80°C, restart the load according to the "Fault" status strategy. If the health status does not recover after cooling, generate an "error" level log and prompt for manual intervention.

[0072] All adjustment strategies must be associated with the preceding parameters: the temperature threshold should be set according to step 102 (e.g., the warning temperature should be lowered to 80℃ during high temperature testing), and the wear level should be determined based on the nominal PE cycle count of the SSD recorded in step 101 (e.g., a QLC SSD is nominally rated for 1500 PE cycles, and high wear is determined when the usage reaches 1200 cycles), to ensure that the adjustment logic conforms to the properties of the SSD itself.

[0073] Step 303: Execute test operations and record test results based on the dynamic adjustment results. The test execution is centered on the "central control script scheduling tool and real-time recording of process data" to ensure that the adjusted load is accurately implemented and the results are traceable. First, after the load adjustment strategy is determined, the central control script automatically generates a new fio configuration file (including the adjusted IOPS, queue depth, read / write mode, and switching interval). The configuration parameters are sent to the fio tool through the "tool call interface" (if fio is in the middle of the previous test, the script waits for the current test cycle to end (e.g., waits for 10 seconds to complete) before loading the new configuration to avoid data anomalies caused by forced interruption).

[0074] During test execution, the script captures fio's output data in real time (such as read bandwidth and write latency in the first round of tests after adjustment), and records "adjustment trigger conditions" (such as "health score 72, temperature 78℃, triggering sequential write mode") and "parameter comparison before and after adjustment" (such as old IOPS=20k / new IOPS=20k, old mode=random read / write / new mode=sequential write). This information is integrated with the test results (average IOPS and maximum latency of 3 samples after adjustment) into "load adjustment record", and synchronously written to the "test adjustment log table" in MySQL (including test task ID, adjustment timestamp, trigger conditions, new and old parameters, and test results). At the same time, key results (such as the performance fluctuation range after adjustment) are added to the "performance data" measurement in InfluxDB, providing a complete data chain for the result analysis in the subsequent step 4.

[0075] If a tool malfunctions during execution (e.g., fio fails to load a new configuration), the script automatically rolls back to the previous load configuration (e.g., restores the IOPS and mode before adjustment), re-executes the test, and records the "adjustment failure log" (including the reason for failure, such as "fio configuration file syntax error"), and attempts to regenerate the configuration file (up to 3 retries; if it still fails, a manual intervention prompt is triggered). If the test results show that the performance drops by more than 30% after adjustment (e.g., IOPS drops sharply from 20k to below 14k), the current load is immediately paused, and the load is switched to "basic load" (50% of the initial IOPS). After the performance stabilizes (3 consecutive sampling fluctuations <5%), the health is reassessed and the load is adjusted to avoid irreversible damage to the SSD caused by abnormal load.

[0076] The analysis data in step 301 comes entirely from the performance, health, and environmental data written to InfluxDB in step 203. The temperature and wear thresholds in the adjustment strategy are referenced from steps 102 (environmental parameters) and 101 (SSD nominal attributes), ensuring data and parameter consistency. Connecting with subsequent steps: The "load adjustment log" recorded in step 303 and the test results data in InfluxDB will serve as the core basis for step 4 (result analysis and report generation). The analysis phase needs to extract performance change trends under different load modes and the decay pattern of health with load. The report should specifically present the "impact of dynamic load adjustment on SSD status" (e.g., "In sequential write mode, the health rate decreases 15% faster than random read / write"), achieving end-to-end data connectivity from "collection-analysis-adjustment-reporting".

[0077] If the health analysis thread makes a calculation error (such as a corrupted weight configuration file), the script will automatically load the default weight template (40% performance degradation, 30% lifespan loss, and 30% reliability), continue analysis, and generate a "warn" level log. After the configuration file is manually repaired during the test interval, the correct weights will be used in the next round of analysis.

[0078] If the SSD power consumption suddenly increases after dynamic load adjustment (exceeding the power consumption threshold set in step 102 by 120%, such as from 8W to more than 10W), the script will immediately pause the test and trigger the constant temperature chamber to provide powerful cooling (the cooling rate will be increased to 5℃ / minute). After the power consumption drops back to within the threshold, the load will be readjusted according to the "warning" state strategy to avoid power overload.

[0079] If the MySQL "Test Adjustment Log Table" fails to write (e.g., the database connection is interrupted), the script will temporarily write the adjustment records to a local JSON file (named "adjust_log_TEST_YYYYMMDD_001.json"). It will attempt to reconnect to the database every minute. After the connection is restored, the records will be rewritten in timestamp order. After the rewriting is completed, the local file will be deleted to ensure that no records are lost during the adjustment process.

[0080] Based on the matched rules, the script dynamically generates new fio job files and notifies the running fio instance to reload the job configuration via the fio remote control interface (--remote command) or a signal (SIGUSR2). This allows for load adjustments without interrupting the test process, achieving true dynamism. In the example above, the load will immediately shift from high-pressure sequential writes to a more moderate, validated mixed load.

[0081] Based on the original load balancing strategy, a new "dynamic fio job file generation module" has been added to ensure that adjustment parameters can be accurately converted into executable fio configurations and support uninterrupted loading. The job file generation rules are as follows: the central control script automatically generates new fio job files (in .fio format) based on the load balancing strategy (e.g., switching to "mixed load with verification" when a health "warning" occurs). Core parameters must include: basic configuration: [global] section specifies the test device (e.g., filename= / dev / nvme0n1), block size (e.g., bs=4k), IO engine (ioengine=libaio), queue depth (iodepth=16, set according to the adjustment strategy); load parameters: set rw (read / write mode, e.g., mixed load set to rw=randrw, rwmixread=50, sequential write set to rw=write) and IOPS (e.g., from 20k...). Reduce to 15k), verify (enable verification under mild load, such as verify=crc32 to prevent data corruption); dynamic identifier: add jobid=TEST_20251114_001_V2 (including test task ID and version number) to the [job1] section to facilitate subsequent tracking of configuration versions.

[0082] After the job file is generated, the script calls fio --validate<new job file>.fio to perform syntax validation. If there are parameter errors (such as iops settings exceeding the SSD hardware limit), it immediately rolls back to the previous version of the job file and records a "configuration generation failure log" (including error parameters) to avoid invalid configurations being distributed.

[0083] The original "waiting for the test cycle to end" execution logic has been optimized to achieve uninterrupted loading via the fio remote control interface or the SIGUSR2 signal, ensuring that the test process is not interrupted. Two uninterrupted loading methods are available:

[0084] Method 1: Remote control interface (--remote command)

[0085] Prerequisites: Remote control must have been enabled using `--remote<host IP>:<port>` when starting fio (e.g., `fio --remote 192.168.1.100:8765 test.fio`). During adjustments, the script sends a "reload configuration" command to the running fio instance using `fio --remote192.168.1.100:8765 --command=reload,<new job file>.fio`. fio will immediately stop the current job loop, load the new configuration, and continue execution. The entire process is uninterrupted (the data acquisition thread runs continuously, only load parameters change). Applicable scenario: Cross-host testing (fio runs on the test machine, the control script runs on the server), requiring a stable network connection. Method 2: Triggered by SIGUSR2 signal.

[0086] Prerequisite: The script records the process ID (PID) of the current fio instance (obtained and stored via `pgrep fio` when fio starts). During adjustments, the script uses `kill -SIGUSR2`.<fio PID> Sending a signal simultaneously writes the path to the new job file into fio's signal listener configuration (this needs to be preset beforehand at fio startup using `--signal=SIGUSR2:<new job file path>`). Upon receiving the signal, fio automatically reads the new file and loads the configuration, achieving millisecond-level switching. Applicable scenarios: Single-machine testing, no network dependency, faster response time.

[0087] After the configuration is loaded, the script uses fio --remote <ip>:<port> --command=status (remote mode) or read the fio output log (signal mode) to verify whether the new configuration has taken effect (e.g., confirm that IOPS has changed from 20k to 15k and rw mode has changed to randrw): If the verification is successful, write "Configuration loaded successfully" and the new parameters to the MySQL "Test Adjustment Log Table", and synchronously update the "Load Version" label of the InfluxDB performance data (e.g., load_version=V2); If the verification fails (e.g., fio does not respond to the command, and there is no status update for 3 seconds), immediately retry loading (maximum 2 times). If it still fails, trigger the "Test Execution Abnormal" fault judgment in step 304 and record the reason for "Configuration Loading Failure" (e.g., "fio process is unresponsive").

[0088] Add a "Configuration loading exception" type to the original step 304 to broaden the scope of fault coverage:

[0089] New trigger scenarios: Configuration loading fails after 2 retries (fio does not respond to remote commands or SIGUSR2 signal); fio verification fails after loading new configuration (e.g., verify=crc32 triggers data verification error, log shows "verify error at offset XXXX"); performance data shows "cliff-like anomaly" after loading (e.g., IOPS drops sharply from 15k to 0, and does not recover for 5 seconds).

[0090] Targeted Judgment and Recording: If "Loading unresponsive", extract the fio process status (e.g., check if it's frozen using ps -ef |grep fio) and system resource usage (CPU / memory usage) to determine if it's "fio process frozen" or "system resources exhausted". Recordings must include the fio PID and resource usage values. If "Verification failed", combine InfluxDB data before the failure (e.g., whether it was in a high-temperature environment or had recently undergone a large number of writes) to determine if it's "NAND flash data error" or "excessive load adjustment causing data instability". Recordings must include the offset address of the verification error and the load parameters before the failure. Post-failure Interlocking: If "Process frozen", the script automatically restarts fio and loads the basic load (50% of the initial IOPS) to avoid complete test interruption. If "Data verification error", immediately pause the test and read the SSD's internal error logs using the vendor's SDK to check for irreversible hardware damage.

[0091] During the configuration loading process, the data acquisition threads (performance / health / environment) run continuously without interruption, ensuring that performance fluctuations before and after loading are fully recorded (e.g., capturing the smooth transition of IOPS from 20k to 15k), providing continuous data for subsequent analysis of "the impact of load adjustment on performance"; the "load version" label (V1 / V2) in InfluxDB can be used to statistically analyze performance data by version, such as comparing the health decay rate under "high-pressure sequential write (V1)" and "mild mixed load (V2)", providing data support for "load optimization suggestions" in the report.

[0092] If the SIGUSR2 signal transmission fails (e.g., the PID does not exist), the script automatically re-acquires the fio PID (via pgrep fio). If it still cannot be obtained, it is determined that "the fio process has exited," triggering step 304 failure. At the same time, fio is restarted and the previous version configuration is loaded. If the remote control interface connection times out (e.g., network interruption), it automatically switches to SIGUSR2 signal mode (single-machine scenario only) and records "network interruption, switch to signal loading" to ensure redundancy in loading methods. After loading a load with verification, if verification errors are frequently triggered (more than 3 times within 1 minute), the verification strength is automatically reduced (e.g., from crc32 to md5) or verification is paused to avoid test interruption due to excessive verification. At the same time, "verification strategy adjustment" is marked and the reason is recorded.

[0093] After step 303, the method further includes:

[0094] Step 304: When a fault is detected, immediately determine the fault and record the cause of the fault.

[0095] Step 304 needs to be initiated at the "fault trigger moment" as a closed-loop link following Step 301 (fault status determination), Step 302 (load adjustment anomaly), and Step 303 (test execution anomaly). The core is to clarify the root cause of the fault through cross-validation of multi-dimensional data and ensure that the recorded information is traceable and analyzable.

[0096] Step 304: In the judgment and cause recording after fault triggering, the central control script needs to monitor three types of fault trigger signals in real time. Once captured, Step 304 is started immediately: Health trigger: Step 301 calculates the health score <60 points (judged as "fault"), and the trigger indicators include key anomalies (such as the number of newly added bad blocks >10, PE cycle utilization rate >90%, read / write speed decrease rate >20%); Real-time parameter trigger: Step 202 collects physical environment data exceeding the emergency threshold (such as SSD case temperature >90℃, power supply voltage fluctuation >±8%), Step 303 shows a sudden performance drop of more than 40% during test execution (such as IOPS dropping from 20k to below 12k) or the fio tool continuously crashes (still unable to start after 3 retries); Hardware signal trigger: reads internal hardware errors of the SSD through the manufacturer's SDK (such as controller overheat protection trigger, NAND flash memory verification error, cache data loss). These signals have the highest priority and fault judgment is started directly without waiting for health calculation.

[0097] The judgment process is executed by the script's "Fault Analysis Module," and is completed in three steps: "Scenario Classification - Data Extraction - Cross-validation," ensuring that no cause is missed or misjudged: Step 1: Scenario Classification: Determine the fault category based on the trigger signal—health-related triggers correspond to "performance / lifespan faults," real-time parameter triggers correspond to "environment / load faults," and hardware signal triggers correspond to "hardware faults." Step 2: Data Extraction: Extract full-dimensional data (such as temperature change curves, performance fluctuation trends, and health indicator changes) from InfluxDB for the 5 minutes prior to the fault, extract the most recent load adjustment parameters (such as adjusted IOPS and read / write modes) from the MySQL "Test Adjustment Log Table," and obtain the SSD's internal error codes (such as "0x0B" representing bad block errors and "0x12" representing controller errors for Samsung SSDs) from the vendor's SDK. Step 3: Cross-validation: Combine preceding parameters and data trends to determine the root cause—for example, a health "fault" with more than 10 newly added bad blocks, and the system was in "sequential write mode" for the 30 minutes prior to the fault (step 302). If the temperature remains between 85-90℃, it is determined that "continuous writing under high temperature environment leads to a surge in bad blocks in NAND flash memory"; if the hardware signal shows "cache data loss" and the power supply voltage fluctuates by more than ±10%, it is determined that "abnormal power quality causes cache power failure".

[0098] After the determination is completed, the "fault type" (such as "NAND flash memory fault", "power supply abnormality", "performance degradation exceeds limit") and "root cause description" (must include specific indicator values, such as "temperature of 88℃ for 15 minutes resulted in 5 new bad blocks and a 25% decrease in read and write speed") should be output to avoid vague descriptions.

[0099] The record needs to consider both "database storage" and "local backup" to prevent data loss. Specific content and storage methods are as follows: Record content: Includes a "unique fault ID" (format "FAULT_YYYYMMDD_001"), test task ID (associated with the corresponding test), fault trigger time (accurate to milliseconds, synchronized with NTP time), trigger scenario (e.g., "health score 58 points", "voltage fluctuation 12%"), judgment result (fault type + root cause), snapshot of key data at the time of the fault (e.g., temperature 88℃, number of bad blocks 12, IOPS 11k), and processing status ("pending / processed", initially "pending"); Storage location: First, write to the newly added "fault record table" in MySQL (fields correspond one-to-one with the above content), and simultaneously generate a local JSON format backup file (named "fault_record_FAULT_YYYYMMDD_001.json"), with the storage path consistent with the adjustment log backup in step 303; If the MySQL connection is interrupted, the script will update every 30 seconds. The system attempts to reconnect within seconds until a successful write operation is completed, after which the local backup is deleted. Log synchronization: The fault determination result is synchronously written to the test task log (e.g., "2025-11-14 15:30:02: Fault triggered, type = NAND flash memory fault, cause = high temperature + continuous writing leading to a surge in bad blocks"), with the log level set to "error" to ensure differentiation from other abnormal logs.

[0100] After recording, the script needs to trigger a linkage operation to connect with subsequent operation and maintenance processes: Test control: Immediately pause the current test task (stop the fio tool and all acquisition threads) to prevent the fault from escalating (such as continuing to test under high temperature causing damage to the main controller); Alarm escalation: In addition to the original audible and visual alarms, add "fault SMS notification" (sent to the mobile phones of operation and maintenance personnel through the SMS gateway interface). The notification content includes the fault ID, test task name, and core reason (such as "TEST_20251114_001 Task triggered fault: SSD temperature 92℃, 15 bad blocks added"); Status marking: Highlight "fault paused" in red in the script interface, and update the task status to "fault interrupted" in the MySQL "Test Task Table" to facilitate subsequent statistics on the number and type of fault tasks.

[0101] If InfluxDB data is missing during fault diagnosis (e.g., no data has been written in the last 5 minutes), the script automatically extracts data from the local backup file (backup files from steps 203 and 303), marks "Data Source = Local Backup," and generates a "warn" level log to prevent the diagnosis from failing due to missing data. If the vendor SDK cannot read the hardware error code (e.g., SDK connection timeout), the diagnosis logic is downgraded to "judgment based on surface data" (e.g., inferring the cause only from temperature and performance data), and the fault record is noted as "Hardware error code not obtained, cause inferred from surface data," ensuring that basic diagnosis can still be completed even if the SDK is abnormal. If the local backup file fails to write (e.g., insufficient disk space), the script immediately triggers a disk space check, cleans up old backup files from 3 days ago to free up space, and re-attempts to write, while simultaneously alerting the user with an audible and visual alarm: "Local backup of fault record failed, disk space needs to be checked," to prevent the loss of fault information.

[0102] Following step 304, the method further includes: step 305, statistically analyzing the types and number of faults during the testing process, and analyzing the causes and impact of various faults based on fault determination and causes; step 306, comparing the fault analysis results with historical data, and identifying potential failure modes based on machine learning algorithms; step 307, identifying the time series of fault occurrences under the failure modes, combining the time series of fault occurrences with stress change curves, identifying the fault concentration intervals and corresponding loads, evaluating the impact weight of different stress levels on the reliability of the solid-state drive, and optimizing the dynamic adjustment strategy of the dynamic acceleration stress model based on the impact weight to improve the accuracy and efficiency of the test.

[0103] Step 305: In the analysis of fault type statistics, causes, and impact, multi-dimensional statistics and in-depth analysis are used to clarify the fault distribution characteristics and their actual impact on testing, laying the foundation for subsequent failure mode identification.

[0104] Fault Statistics Dimensions and Methods: The central control script automatically extracts all fault data from the MySQL "Fault Record Table" for the current test period (or a specified time range, such as testing the same model of SSD within the last 30 days), and statistically analyzes them according to three dimensions: "Type - Trigger Scenario - Time": Type Statistics: Faults are divided into three categories (performance: such as sudden drop in read / write speed, abnormal IOPS; hardware: such as a surge in NAND bad blocks, controller errors; configuration: such as fio configuration loading failure, verification errors), and the occurrence frequency and percentage of each type of fault are counted (e.g., "Hardware faults 12 times, accounting for 40%; performance faults 8 times, accounting for 26.7%)"); Trigger Scenario Statistics: The number of faults in each scenario is counted according to the trigger scenarios in step 304 (health trigger, real-time parameter trigger, hardware signal trigger), such as "real-time parameter trigger (temperature exceeding 90℃) caused 7 faults". "58.3% of hardware failures"; Time statistics: Failures occurred during the testing phase (initialization phase, low load phase, high load phase), such as "15 failures during the high load phase (IOPS>25k), accounting for 50% of the total failures".

[0105] Cause and impact analysis: Combining the fault determination results (root cause) from step 304 with InfluxDB test data (performance and environmental changes before and after the fault), conduct in-depth analysis: After the analysis is completed, generate the "Fault Statistics and Impact Assessment Report", which includes a pie chart of fault type distribution, a list of high-frequency fault causes, and a table of the proportion of impact at each level, and simultaneously store it in the MySQL "Fault Analysis Report Table" for subsequent steps.

[0106] Cause tracing: For high-frequency fault types, delve into the related factors. For example, for "NAND bad block surge" faults, it is necessary to analyze the temperature curve before the fault (whether it is consistently >85℃), load mode (whether it is sequential write), PE cycle utilization (whether it exceeds 80%), and identify the core cause (such as "80% of bad block faults occur in scenarios where the temperature is >85℃ and the sequential write load lasts for more than 30 minutes").

[0107] Impact Classification: Based on the degree of interference of the fault on the test, it is divided into three levels: mild (such as a brief configuration loading failure, which is recovered after retry, with no data loss and an impact time of less than 5 minutes), moderate (such as a sudden drop in performance of more than 30%, requiring a restart of fio, affecting the continuity of test data, and requiring 1 hour of additional testing), and severe (such as irreversible hardware damage, the SSD cannot continue testing, the test task is terminated, and the device needs to be replaced). The percentage of impact and average loss time of each level are also calculated.

[0108] Step 306: Historical Data Comparison and Machine Learning for Potential Failure Mode Identification. Based on the statistical results of Step 305 and combined with historical test data, machine learning algorithms are used to uncover implicit correlations between failures and the environment and load, and to identify reusable failure modes.

[0109] Historical data extraction and standardization: Two types of historical data are extracted from the database: one is "similar data" (historical test records of the same model and flash type as the current SSD being tested, including fault records, performance data, and environmental parameters); the other is "similar data" (historical records of SSDs of different models but in the same application scenario, such as those used in high-write scenarios in servers).

[0110] The extracted data is standardized by mapping indicators such as temperature, IOPS, and PE cycle rate to a unified range (0-100) and converting fault types into labels (e.g., "NAND bad block" is labeled as 1 and "configuration loading failure" is labeled as 2) to ensure consistency with the current test data format and avoid the comparison results being affected by differences in units or ranges.

[0111] Historical data comparison logic: The current data is compared with historical data using the "feature matching" method. If the current high-frequency failure (such as "high temperature + high write → bad block") occurs ≥8 times in the historical data and the overlap of core features (temperature threshold, load intensity, failure interval) is ≥70%, it is determined as a "known failure mode". If the overlap is <50% or the current failure occurs <3 times in the historical data, it is marked as a "potential unknown failure mode" and further analysis is required through machine learning.

[0112] Machine learning algorithms identify failure modes:

[0113] For "potential unknown failure modes", two types of algorithms are selected for analysis: After identification, a "Failure Mode List" is generated, which is labeled with mode type (known / newly identified), core features and probability of occurrence, and stored in the MySQL "Failure Mode Table" to provide the basis for analysis in step 307. Clustering Algorithm (K-means): Using "temperature, IOPS, read / write mode, PE utilization, and fault type" as feature variables, clusters current and historical fault data. If more than 80% of the samples in a certain cluster contain the feature "low temperature (<60℃) + low IOPS (<10k) → performance fluctuation", it is identified as a new failure mode (such as "cache scheduling anomaly mode under low load"). Classification Algorithm (Random Forest): Using "environmental + load parameters" from historical data as independent variables and "whether a fault occurs" as the dependent variable, a fault prediction model is trained. The model accuracy is verified using current test fault data (≥85%). If the model's prediction probability for a certain type of fault is >90%, the feature combination corresponding to that fault (such as "voltage fluctuation > ±5% + queue depth > 20 → cache data loss") is confirmed as a stable failure mode.

[0114] Step 307: In time series analysis, stress influence weight assessment, and model optimization, the failure modes from step 306 are used as the core. Combining the failure time series and stress change curves, the impact of stress on reliability is quantified, and the dynamic accelerated stress model is ultimately optimized to improve test accuracy.

[0115] Correlation analysis between fault time series and stress curves: Two types of data were extracted from InfluxDB: first, the time series of fault occurrences (e.g., "NAND bad block faults occurred 3 times between 10:00 and 10:30 on November 14, 2025"); second, the stress change curves during the same period (temperature curve: from 82℃ to 88℃ between 10:00 and 10:30; load curve: IOPS increased from 22k to 26k). A "sliding window method" (window size 10 minutes) was used to analyze the correlation between the two: if the number of fault occurrences within a window is ≥2, and the corresponding stress curve shows a "sudden increase / decrease" (e.g., temperature increases by 6℃ within 10 minutes, IOPS increases by 4k), then the window is marked as a "fault concentration interval," and the average stress value within the interval is recorded (e.g., average temperature 85℃, average IOPS 24k), clarifying the temporal correspondence between faults and stress.

[0116] Stress influence weight assessment:

[0117] Using "failure probability" as the dependent variable and stress indicators such as "temperature, IOPS, read / write mode, and voltage fluctuation" as independent variables, "multiple linear regression analysis" was used to quantify the influence weight of each stress: Data preparation: 1000+ sets of "stress indicator - failure occurrence" samples were extracted from historical test data (e.g., "temperature 85℃, IOPS 24k, no failure → Sample 1; temperature 88℃, IOPS 26k, failure → Sample 2"); Weight calculation: The coefficients of each indicator were output through regression analysis (e.g., temperature coefficient 0.6, IOPS coefficient 0.3, voltage fluctuation coefficient 0.1). The larger the absolute value of the coefficient, the higher the influence weight on the stress (e.g., temperature has an influence weight of 0.6 on failure, making it a core stress indicator); Weight verification: The failure data from the current test were substituted into the regression model. If the deviation between the model's predicted failure probability and the actual occurrence was <10%, the weights were confirmed as effective; if the deviation exceeded 20%, additional samples were added and the weights were recalculated.

[0118] Dynamic accelerated stress model optimization strategy: Based on the stress influence weight, adjust the dynamic load adjustment logic in step 302, focusing the optimization direction on "early warning and precise load control":

[0119] Core stress priority adjustment: If the temperature weight is 0.6 (highest), the original threshold of "adjust load when temperature > 85℃" will be lowered to 82℃, and the adjustment range will be increased (original IOPS reduction of 20% will be changed to reduction of 30%) to avoid core stress accumulation leading to failure.

[0120] Multi-stress coordinated adjustment: If the temperature is 83℃ (close to the threshold) and IOPS is 23k (weight 0.3) in a certain scenario, then "coordinated load control" (IOPS decrease by 25% + constant temperature chamber temperature decrease by 2℃) is triggered, instead of adjusting the load alone, thus reducing the risk of multiple stress superposition;

[0121] Failure mode adaptation adjustment: For the "low load cache abnormal mode" (weight 0.2) identified in step 306, a new strategy of "switching the read and write mode once every 30 minutes when IOPS < 8k" is added to avoid cache scheduling rigidity caused by a single low load;

[0122] Optimization result feedback: The adjusted strategy parameters (such as temperature threshold of 82℃ and multi-stress collaborative rules) are written into the configuration file of the central control script and the MySQL "Dynamic Model Parameter Table" is updated synchronously. Subsequent tests directly call the new parameters to achieve iterative optimization of the model.

[0123] Step 305 uses the "Fault Record Table" from Step 304 as the data foundation; Step 306 uses the "Fault Statistics Results" from Step 305 plus historical data from the MySQL database as input; Step 307 uses the "Failure Mode List" from Step 306 plus "Stress Curve Data" from InfluxDB as the basis to ensure the integrity of the data chain. The optimized dynamic accelerated stress model parameters in Step 307 are directly fed back to Step 302 (Dynamic Load Adjustment) to make the load adjustment of subsequent tests more accurate (such as anticipating high-weight stresses). At the same time, they are fed back to Step 102 (Environmental Parameter Settings), such as adjusting the high temperature warning threshold of the constant temperature chamber from 85℃ to 82℃, forming a closed loop of "Test-Analysis-Optimization-Retest".

[0124] If there is insufficient historical data for the same type of SSD in step 306 (<50 records), automatically supplement historical data for similar flash memory types (such as TLC replacing QLC), and note in the report "Data source includes similar models, weight calculation needs further verification"; Algorithm identification bias: If the accuracy of the random forest model in step 306 is <80%, add "decision tree pruning" optimization, or supplement with failure samples from the past 3 months for retraining. If it still cannot be improved, temporarily use "manual labeling of failure modes" to replace algorithm identification; Abnormal weight calculation: If the regression analysis in step 307 shows "negative coefficients" (such as IOPS coefficient -0.1, which contradicts the actual logic), check whether there is abnormal data of "low IOPS and high failure" in the samples, remove them and recalculate to ensure that the weights conform to physical laws (such as higher temperature and greater load, higher probability of failure).

[0125] After step 307, the method further includes: step 308, associating fault data with response stress and creating a fault stress model through a continuous iterative feedback mechanism; step 309, calculating the failure boundary of the solid-state drive under different usage scenarios based on the fault stress model.

[0126] Step 308: Creating a Fault Stress Model (based on a continuous iterative feedback mechanism) uses the "stress influence weight" and "failure mode list" from Step 307 as core inputs. Through "data association - model building - iterative update", discrete fault data and stress indicators are transformed into a quantifiable and reusable fault stress model. The core is to achieve a precise mapping of "stress combination → fault probability".

[0127] Deep integration of fault and stress data: First, extract all related data from the database to form the basic dataset for model training. After data integration, preprocessing is required: remove outliers (such as stress indicators exceeding the physical reasonable range, such as temperature = -5℃), fill missing values ​​with the mean (such as a data point missing voltage fluctuation value, which is filled with the mean of 10 similar data points in the same scenario), and normalize the features (map to the 0-1 interval to eliminate dimensional differences, such as IOPS 0-50k→0-1, temperature 0-100℃→0-1).

[0128] Input feature set: Select the top 5 stress indicators with the highest weight in step 307 (such as temperature, IOPS, write ratio, voltage fluctuation, PE cycle usage rate), and supplement the failure mode labels identified in step 306 (such as "low load cache anomaly" labeled as feature variable "mode_cache=1") to ensure that the features cover "basic stress + failure mode".

[0129] Output label set: Use "whether a fault occurred" as a binary label (fault = 1, no fault = 0). If further refinement is needed, use "fault severity" (mild = 1, moderate = 2, severe = 3) as a multi-category label.

[0130] Data timing correlation: The average stress value in the 5 minutes before the fault occurred is used as the characteristic value of the corresponding fault (e.g., if a NAND bad block fault occurs at 10:00, the average temperature of 88℃ and the average IOPS of 28k from 9:55 to 10:00 are taken as the stress characteristics of the fault) to avoid correlation deviation caused by timing misalignment.

[0131] Fault stress model construction (supervised learning + rule constraints): A hybrid architecture of "main model + rule correction" is adopted to balance prediction accuracy and physical rationality. Main model selection: Gradient boosting tree (XGBoost) or lightweight neural network (such as MLP with 1 hidden layer) are preferred because these two types of models can capture the nonlinear correlation between stress features (such as the synergistic failure effect of "high temperature + high write"). Using normalized stress features as input and fault labels as output, a "stress → failure probability" prediction model is trained. The model evaluation metrics must meet the following requirements: test set accuracy ≥ 90%, fault recall ≥ 85% (to avoid missing faults). Rule constraint correction: Combining the stress influence weights in step 307, physical rules are added to restrict the model output (e.g., when the temperature is < 50℃ and IOPS < 10k, the model's predicted failure probability must be ≤ 1% because the theoretical failure risk is extremely low in this scenario; if the model output exceeds 1%, it is forcibly corrected to 1%) to prevent the model from making predictions that violate physical logic due to sample bias.

[0132] Continuous Iterative Feedback Mechanism (Ensuring Model Timeliness): Establishing an iterative closed loop of "New Data Supplementation → Incremental Model Training → Effect Validation → Parameter Update": Data Supplementation: After each new testing cycle (e.g., weekly testing of the same model of SSD), newly generated fault data and corresponding stress data are automatically added to the model training set to ensure that the data covers the latest failure scenarios (e.g., new failure modes caused by new firmware); Incremental Training: Adopting the method of "freezing the underlying parameters + fine-tuning the top-level weights" (e.g., XGBoost only updates the weights of the leaf nodes of the tree) to avoid the waste of resources caused by full retraining, and recalculating the model evaluation metrics after each incremental training; Feedback Update: If the model accuracy improves by ≥3% after incremental training, the new model parameters (e.g., the tree structure and feature weights of XGBoost) are stored in the MySQL "Fault Stress Model Table", and the model call interface of the central control script is updated synchronously; If the accuracy decreases or the recall rate is lower than 85%, the model is rolled back to the previous version, and the cause of the data anomaly is analyzed (e.g., whether it is a new failure mode that has not been identified).

[0133] Step 309: Calculate the failure boundary under different scenarios based on the fault stress model. Using the fault stress model in step 308 as a tool, quantify the upper limit of stress (i.e., failure boundary) for the "acceptable failure probability" in typical SSD usage scenarios, and provide a basis for load control in practical applications.

[0134] Typical Usage Scenarios Classification and Stress Characteristics Definition: First, clarify the core application scenarios of SSDs, classifying them into scenario types based on "load intensity, environmental requirements, and read / write characteristics." Each scenario corresponds to a unique combination of stress characteristics: For each scenario, determine the "priority of key stress indicators" (e.g., server scenarios prioritize IOPS and write ratio, industrial scenarios prioritize temperature range) to provide a focus for subsequent boundary calculations. Server scenario (high load continuous write): characteristics are "IOPS 20k-40k, write ratio 70%-90%, temperature 50-80℃, continuous operation throughout the year (fast PE cycle consumption)"; Consumer scenario (low load intermittent read / write): characteristics are "IOPS 1k-5k, write ratio 30%-50%, temperature 25-50℃, daily use 4-8 hours (slow PE cycle consumption)"; Industrial scenario (wide temperature range + stable load): characteristics are "IOPS 5k-15k, write ratio 40%-60%, temperature -20-70℃, uninterrupted operation (high environmental adaptability requirements)".

[0135] Failure boundary calculation method (failure probability threshold method): Using "failure probability ≤ 5%" as an acceptable threshold (adjustable according to the scenario, such as ≤ 2% for industrial scenarios), the failure boundary is calculated through "single indicator traversal + multi-indicator collaboration".

[0136] After the calculation is completed, a "Failure Boundary List" is generated for each scenario, which clarifies the threshold of a single indicator, the coordinated range of multiple indicators, and marks the failure probability corresponding to the boundary (e.g., "IOPS 32k (failure probability 5%)").

[0137] Single-indicator failure boundary (with other variables fixed): For a key stress indicator in a specific scenario, with other indicators fixed as the scenario average, the indicator value at which the failure probability reaches 5% is calculated using a fault stress model. This value is the single-dimensional failure boundary for that indicator. Example (server scenario): With fixed write ratio of 80%, temperature of 65℃, and PE utilization of 50%, IOPS increases from 20k to 40k. The model predicts that the failure probability reaches 5% when IOPS = 32k → The single-dimensional failure boundary for IOPS in the server scenario is 32k.

[0138] Multi-indicator collaborative failure boundary (variable interaction): Considering the synergistic effect between stress indices, a "grid search method" is used to traverse key index combinations to find all combinations with a failure probability ≤ 5%, forming the boundary interval. Example (consumer-grade scenario): Traversing temperatures of 25-50℃ and IOPS of 1k-5k, the model predicts a failure probability ≤ 5% when "temperature ≤ 45℃ and IOPS ≤ 4.2k" → the collaborative failure boundary for this scenario is "temperature ≤ 45℃ ∩ IOPS ≤ 4.2k".

[0139] Scenario-based verification and correction of failure boundaries: Verify the accuracy of the boundaries through actual testing to avoid deviations between model predictions and actual conditions: Verification method: In a certain scenario, control the SSD stress near the failure boundary (e.g., server scenario IOPS=31k, 32k, 33k), run each for 24 hours, and count the actual number of failures. If there are no failures within the boundary (31k), the number of failures at the boundary point (32k) is ≤1 (within 24 hours), and the number of failures outside the boundary (33k) is ≥2, then the verification is passed; Boundary correction: If the verification finds that failures still occur within the boundary (e.g., 1 failure at 31k), then the boundary of this index is lowered (e.g., 32k→30k), and the failure probability is recalculated by substituting it into the failure stress model to ensure that the failure probability within the boundary is ≤3% after correction; If there are no failures outside the boundary, then high-stress samples for this scenario (e.g., test data of IOPS 35k) need to be added, the model is retrained, and the boundary is recalculated.

[0140] The fault stress model in step 308 relies on the stress weights in step 307 (to determine core features), the failure modes in step 306 (to supplement mode labels), and the fault and stress data (basic training set) in steps 304-305 to ensure the comprehensiveness of the model input. Output value: The failure boundary in step 309 can feed back into two main stages—① feeding back into the dynamic load adjustment in step 302 (e.g., in server scenario testing, the load limit is set to 80% of the failure boundary, i.e., 32k × 80% = 25.6k, reserving safety redundancy); ② feeding back into SSD product design (e.g., recommending "temperature ≤ 40℃, IOPS ≤ 3.5k" for consumer-grade SSDs to reduce the risk of failure in user scenarios), realizing the value realization of "testing → model → application".

[0141] If there are fewer than 30 historical test data points for a certain scenario (such as an industrial wide-temperature scenario), which cannot support boundary calculation, then "parameter migration" is first performed based on the model parameters of a similar scenario (such as a server scenario) and combined with the environmental characteristics of this scenario (such as temperature range expansion) to calculate the temporary failure boundary. After supplementing 50+ scenario data points, the model is retrained and the boundary is corrected.

[0142] Model prediction bias: If the model prediction of the failure probability is found to be more than 10% different from the actual probability during the verification in step 309 (e.g., the model predicts a probability of 5% when it predicts 32k, but the actual probability reaches 12%), then the model training process in step 308 is traced back to check whether any key features are missing (e.g., the "humidity" indicator is not included in the industrial scenario), and the model is retrained after the features are added.

[0143] Boundary conflict: If there is a conflict between the failure boundaries of different scenarios (e.g., a stress index is 32k at the boundary of the server scenario and 28k at the boundary of the industrial scenario), it is necessary to reconfirm the definition of the scenario characteristics (e.g., whether the stress tolerance of the industrial scenario is reduced due to the wide temperature range). The conflict should be resolved by training a scenario-specific model (rather than a general model) to ensure that the boundary conforms to the actual needs of the scenario.

[0144] Figure 5 yes Figure 1 The flowchart for step 4 is as follows: Figure 5 As shown, step 4 includes: step 401, integrating all test data, fault statistics and analysis results to generate a structured test report, the report content including key performance indicator trend charts, health change curves and fault timelines; step 402, automatically pushing the test report to the designated platform via email or API interface, and triggering the next stage of the test process or ending the test task.

[0145] Step 401: Integrate full data to generate a structured test report. Using full-process database data (InfluxDB time-series data, MySQL business data) as input, a standardized report containing charts and conclusions is generated through "data extraction - integration and association - structured presentation." The core goal is to make test results intuitively interpretable, traceable, and reusable. Full data extraction and integration: The central control script extracts the following core data in batches through the database interface to form the report's data source. During data integration, relationships need to be established: use "test task ID" to link all data, use "timestamp" to associate performance data with load adjustment nodes, and use "fault ID" to associate fault records with corresponding stress data (such as temperature and IOPS when a fault occurs), ensuring logical data consistency.

[0146] Basic test information: from the MySQL "Test Task Table", including test task ID, SSD model / capacity / firmware version, test period (start / end time), test scenario (e.g., server scenario / consumer-grade scenario), and initial parameters (e.g., sampling interval, initial IOPS).

[0147] Core performance data: from InfluxDB "Performance Data" measurement, extracting time-series data of read / write bandwidth (MB / s), IOPS, and average latency (ms) within the test period, summarizing by "hourly average" (to avoid excessive data volume), and marking load adjustment nodes (e.g., "November 14, 10:00, IOPS adjusted from 20k to 15k").

[0148] Health data: Sourced from the MySQL "Health Analysis Table", extracting the time-series data of health scores (1 data point every 30 minutes), and the final values ​​and changes of key health indicators (remaining lifespan percentage, PE cycle rate, number of new bad blocks);

[0149] Fault Statistics and Analysis: Extracted from the MySQL "Fault Record Table" "Fault Statistics and Impact Assessment Report", the following data were extracted: fault type distribution (e.g., hardware-related faults 12 times, performance-related faults 8 times), high-frequency fault causes (e.g., "80% of bad block faults originate from high temperature + sequential write"), and fault impact classification statistics (mild 6 times, moderate 10 times, severe 4 times).

[0150] Failure boundary results: from the "Failure Boundary List" in step 309, including single indicator boundaries of the current test scenario (e.g., server scenario IOPS=32k), multi-indicator collaborative range (e.g., "temperature ≤45℃ ∩ IOPS≤4.2k") and boundary verification results (e.g., "31k no faults, 32k only 1 fault").

[0151] The report adopts a structured design (sectioned presentation): It follows a four-section structure of "Summary - Data Charts - Analysis Conclusions - Recommendations," balancing conciseness and completeness. The specific content is as follows:

[0152] Chapter 1: Test Overview (Pages 1-2)

[0153] Summarize the basic test information, briefly describe the core test objectives (such as "verify the reliability and failure boundary of a certain SSD model under high server load scenarios"), test environment (such as "constant temperature chamber temperature 50-80℃, power supply voltage fluctuation within ±2%), and attach an "SSD basic information table" (including key parameters such as model, capacity, and firmware version).

[0154] Chapter Two: Core Data Charts (Core Chapter, accounting for 60%)

[0155] Visual charts are presented in the categories of "Performance - Health - Failure". All charts are labeled with the time range and unit of the data source.

[0156] Key performance indicator trend charts: There are three charts: "Read / Write Bandwidth Trend", "IOPS Trend", and "Latency Trend". The horizontal axis represents the test time, and the vertical axis represents the indicator value. The load adjustment nodes are marked with red dashed lines (such as "IOPS reduced at 10:00") to intuitively show the impact of load changes on performance.

[0157] Health status change curve: The horizontal axis represents the test time, and the vertical axis represents the health status score (0-100). The "healthy (≥80)", "warning (60-79)" and "fault (<60)" intervals are marked with green / yellow / red areas respectively. Key nodes are marked next to the curve (e.g., "November 14, 15:00, health status dropped from 82 to 75 due to temperature rising to 85℃").

[0158] Fault timeline: The horizontal axis represents the test time, and the vertical axis represents the fault type (hardware / performance / configuration). Different colored dots mark the time of the fault occurrence. Next to the dots are the fault ID and the core cause (e.g., "FAULT_001, NAND bad block surge, temperature 88℃"). A pie chart of fault type distribution is attached below.

[0159] Failure Boundary Verification Chart: For the current scenario, a scatter plot is used to show the relationship between "stress index (such as IOPS) - failure probability". The red solid line marks the failure boundary (such as IOPS=32k), the blue dots are the data inside the boundary (no failure), and the orange dots are the data outside the boundary (failure). This visually verifies the accuracy of the boundary.

[0160] Chapter 3: Fault and Failure Boundary Analysis (Pages 2-3)

[0161] Text interpretation of failure patterns and failure boundaries: Analyze the causes of high-frequency failures (e.g., "Under high-temperature environments, sequential write workloads will accelerate the generation of NAND bad blocks, and the temperature needs to be controlled below 82°C"), the scenario value of failure boundaries (e.g., "The IOPS boundary in the server scenario is 32k, which can be used as a reference for the upper limit of the workload in actual applications. After reserving a 20% safety margin, it is recommended to be ≤25.6k"), and at the same time explain the limitations of the model (e.g., "The current boundary does not cover the impact of humidity, and subsequent supplementary tests are needed");

[0162] Chapter 4: Test Conclusions and Recommendations (1 page):

[0163] Refine 3 - 5 core conclusions (e.g., "In the server scenario, the health of this model SSD decreases linearly as IOPS increases. For every 5k increase in IOPS, the daily health degradation rate increases by 1.2%"), and give targeted recommendations: recommendations for test optimization (e.g., "Subsequent tests can add humidity variables to improve the failure boundary"), and recommendations for product application (e.g., "When used in the consumer scenario, it is recommended to control the temperature within 45°C to extend the service life").

[0164] Report format and generation tool: Generate in "HTML + PDF" dual formats: First, render the HTML report using the Jinja2 template engine (supporting online viewing, with charts drawn using ECharts and being interactive and zoomable), and then convert the HTML to PDF format through the WeasyPrint tool (for easy archiving and email sending). The report naming format is "SSD_TEST_<task ID><SSD model><date>.pdf" (e.g., "SSD_TEST_20251114_001_Samsung990Pro_20251114.pdf"), and after generation, it is automatically stored in the specified directory (e.g., " / test_reports / 202511 / "), and at the same time, the report storage path is written into the MySQL "test task table".

[0165] Step 402: Automatically push the report and trigger subsequent processes. Push the report through multiple channels and automatically trigger the next-stage operations based on the report conclusions to achieve the automated end and connection of the test process.

[0166] Multiple ways to push the report (email + API interface): The script automatically selects the push method according to the preset configuration (stored in the MySQL "push configuration table") to ensure that the report is accurately delivered:

[0167] Email push (general method): Push targets: Test leader, operations and maintenance team, product liaison (the recipient list can be set in the configuration table, supporting differentiation by scenario, e.g., in the server scenario, copy to the server R & D team);

[0168] Email content: The subject is "

SSD Test Completed

[0169] API Interface Push (Connecting to the Enterprise Platform): If it is necessary to connect to the enterprise test management platform (such as Jira, TestRail), the script calls the platform API through an HTTP POST request. The pushed content includes: report name, storage path (the platform can download the PDF through the path), and the core conclusion JSON string (such as {"task_id":"20251114_001","ssd_model":"Samsung990Pro","fault_count":20,"failure_boundary":{"iops":32000}}); Push verification: Receive the "Push Success" response code returned by the platform (such as 200 OK). If an error code is returned (such as 404 Interface Not Found, 500 Server Error), it will automatically record the error information (such as "API returned 500, platform server exception") and switch to email push as a fallback.

[0170] Trigger subsequent processes based on the report conclusions: The script automatically determines the next operation according to the "Fault Severity" and "Failure Boundary Verification Results" in the report, achieving a closed-loop process:

[0171] Trigger the next stage of testing (when there are no serious faults): If the report shows that "Number of Severe Faults = 0" and "Failure Boundary Verification Passed" (such as no faults within the boundary, ≤1 fault at the boundary point), it will automatically trigger the "Scenario Expansion Test" for SSDs of the same model (such as if the current is the server scenario, the next stage of testing will be the industrial scenario): The script creates a new task in the MySQL "Test Task Table", inherits the basic information of the current SSD, updates the test scenario parameters (such as changing the temperature range to -20 - 70°C), and starts the initialization process of the new task (corresponding to steps 101 - 104);

[0172] Triggering the Fault Review Process (when a serious fault exists): If the report shows "Number of serious faults ≥ 1" (such as irreversible damage to SSD hardware), the "Fault Review Task" will be automatically triggered: Create a review work order on the enterprise platform (such as Jira). The work order content includes the fault ID, fault cause, and stress data at the time of the fault. Assign it to the hardware R&D team, and suspend subsequent testing of the same model of SSD. After the review is completed (the work order status changes to "Resolved"), manual confirmation is required to restart the test.

[0173] End the test task (when the preset goal is achieved): If the current test is the last scenario of "full scenario verification of a certain SSD model" (such as server, consumer grade and industrial scenarios have been completed), and all scenario reports have no serious failures, then the test task will be automatically ended: the status of the MySQL "test task table" will be updated to "completed", the full scenario report will be archived (creating a dedicated folder according to the SSD model), and a "full scenario test summary" (summarizing the failure boundaries and core conclusions of each scenario) will be generated and pushed to the product decision-making team.

[0174] The report data source in step 401 covers the entire process from step 101 (SSD installation) to step 309 (failure boundary calculation), ensuring that the report is a result of the entire cycle, rather than data from a single stage; the process triggering in step 402 realizes the automation of "test-report-next action": if there is no serious failure, the test scenario is expanded (deepening verification); if there is a serious failure, a review is initiated (solving the problem); if the goal is achieved, the results are archived and concluded (preserving the results), avoiding "discontinuity" or "repetition" in the test process.

[0175] If a critical data point is found to be missing during data extraction (e.g., InfluxDB performance data is missing for 1 hour), the script automatically marks the corresponding chart in the report as "partially missing data" and explains the missing period in the "Test Conclusion" section (e.g., "Performance data was missing from 9:00 to 10:00 on November 14th, but this does not affect the overall trend judgment"). Simultaneously, it triggers a data recovery process (re-uploading the missing data from a local backup file). Once the data is successfully re-uploaded, the report can be regenerated. If both email and API push fail (e.g., email server downtime, platform API maintenance), the script stores the report on a shared file server (e.g., NAS) and sends an SMS reminder to the test manager (via an SMS gateway interface). The SMS message includes the report sharing path and the reason for the push failure, ensuring the report can be manually retrieved. If the report simultaneously meets the criteria of "no serious faults" and "completion of preset goals" (e.g., no serious faults in the last scenario test), it is executed according to the priority: "Completion of goals > Extended testing." This means the test task is terminated first, and then a manual decision is made regarding whether to initiate new test requirements (e.g., testing other SSD models) to avoid process confusion.

[0176] Based on the first embodiment, a second embodiment is proposed. In the second embodiment, a dynamic accelerated stress model is used to dynamically adjust the test load according to the real-time status of the solid-state drive in order to accelerate the aging process and achieve multi-dimensional health assessment.

[0177] 1. Test environment initialization and parameter configuration: First, the test environment is set up and initialized.

[0178] Hardware configuration: The solid-state drive under test (model: Samsung 870 EVO 1TB) was installed on the test platform. The platform was placed in a high and low temperature test chamber with a temperature control range of -10℃ to 85℃. The test platform was connected to the host computer via a PCIe adapter card. The host computer was configured with an Intel i7-12700K processor and 32GB of memory. A programmable power supply was used to power the SSD, and a power quality analyzer (model: Keysight N6705C) was connected to monitor voltage fluctuations and current ripple.

[0179] Software configuration: Ubuntu 22.04 LTS operating system is installed on the host machine. Central control scripts are written using Python 3.9. fio is used as the primary load testing and data population tool due to its support for flexible, dynamically adjustable I / O load. SMART health data is collected using the smartctl command-line tool.

[0180] Parameter initialization: The central control script reads a JSON-formatted configuration file and initializes key parameters, including:

[0181] Initial load profile: fio's initial job file, set to sequential write, block size 128KB, queue depth 64, no cache mode.

[0182] Dynamic Adjustment Rule Table: The core of this embodiment, some of its rules are shown in the table below. This table defines the mapping relationship from "status parameters" to "load adjustment actions".

[0183]

[0184] 2. In fully automated, multi-source data synchronous acquisition, after the central control script starts, multiple parallel data acquisition threads are created, and the timestamps of all acquired data are synchronized through the NTP protocol.

[0185] Performance data acquisition: Perform an fio status query every 5 minutes, and obtain real-time IOPS, throughput (MB / s) and average latency (us) by parsing its output.

[0186] Health data collection: Every 15 minutes, complete SMART data is obtained using the command `smartctl -a / dev / nvme0n1`, and key attributes such as Data_Units_Written, Host_Reads / Write, Temperature, and Media_Wearout_Indicator are parsed.

[0187] Physical environment data acquisition: Every minute, the current ambient temperature and SSD power supply voltage are read from the high and low temperature chamber and power quality analyzer via the Modbus protocol.

[0188] Data storage and preprocessing: All timestamped data is written to the InfluxDB time-series database in real time. Control scripts perform preliminary data cleaning (such as removing outliers caused by momentary interruptions in data collection).

[0189] 3. Intelligent multi-dimensional fault diagnosis and dynamic load adjustment (core step), which is a closed-loop control process.

[0190] Real-time status assessment: The control module periodically (e.g., every minute) queries the database for the latest SSD status data.

[0191] Rule matching and decision-making: The acquired status data is compared with the "trigger conditions" in the "dynamic adjustment rule table". For example, when the script detects that the SSD temperature reaches 71°C for the first time, "condition 1" is matched.

[0192] Dynamic load adjustment: Based on the matched rules, the script dynamically generates new fio job files and notifies the running fio instance to reload the job configuration via the fio remote control interface (--remote command) or a signal (SIGUSR2). This allows load adjustments to be made without interrupting the test process, achieving true dynamism. In the example above, the load will immediately shift from high-pressure sequential writes to a more moderate, validated mixed load.

[0193] Multi-dimensional fault diagnosis: Fault diagnosis not only relies on traditional "cannot write" (hard fault) or SMART error, but also includes the following "soft criteria":

[0194] Data integrity failure: If, during the data verification process, the read data does not match the written checksum, this type of failure is identified.

[0195] Performance degradation failure: If the average write latency continues to exceed 500% of the initial value for more than 10 minutes, the SSD is considered to have failed functionally, even if it does not report an error.

[0196] Predictive health failure: A machine learning model based on the trend of SMART parameters (such as the rate of decline of Available_Spare) can predict that the SSD will fail within the next 24 hours, and the test can be terminated in advance and the reason for the prediction can be recorded.

[0197] 4. Test Termination and Report Generation The test terminates when any predefined termination criterion is met (e.g., reaching the SSD's nominal TBW, a hard failure, a data integrity failure, or a performance degradation failure).

[0198] Automatic report generation: The control script calls Python's Jinja2 template engine to populate the entire process data from the database into an HTML report template, automatically generating a comprehensive test report.

[0199] Report content:

[0200] Test summary: SSD information, total test duration, total written data volume (final TBW value), and reason for termination.

[0201] Multidimensional data trend charts: Time series charts generated using Matplotlib show the coordinated changing trends of write volume, temperature, IOPS, latency, and key SMART parameters.

[0202] Fault Analysis: List all recorded errors and warnings in detail.

[0203] Health assessment: Scoring and commenting on the performance consistency and latency stability of the SSD throughout its entire lifespan.

[0204] This embodiment introduces a dynamic accelerated stress model, abandoning the traditional fixed load mode. It dynamically adjusts the test load (such as read / write ratio, queue depth, and block size) based on the SSD's real-time status (temperature, wear, performance degradation), effectively simulating extreme real-world usage scenarios, accelerating the aging process, significantly shortening the test cycle, and improving test efficiency and accuracy. It constructs a multi-dimensional health indicator system, going beyond simple "write volume" and "error rate," introducing soft criteria such as performance consistency, latency stability, and SMART parameter trend analysis to achieve predictive assessment of the SSD's "health status," promptly identifying potential performance degradation or failure risks. It realizes a fully automated closed-loop testing ecosystem, integrating hardware control and test execution through a central control script. The system integrates data collection, intelligent analysis, and decision-making, achieving a fully unattended process from test initiation to final report generation. This significantly improves the automation level of testing and reduces the need for manual intervention. A data integrity verification step has been added, embedding data validation functionality into read / write tests to ensure that data can not only be written but also read correctly at the end of the SSD's lifespan, detecting silent data errors and providing crucial assurance for assessing data reliability. Through multi-dimensional data collection and intelligent analysis, a detailed test report is generated, including configuration parameters, collected data, trend charts, termination criteria, and the final TBW value. This provides SSD manufacturers with richer and more valuable data support for quality assessment and lifespan prediction, as well as for consumers in product selection, compared to traditional TBW testing.

[0205] In the embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.

[0206] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0207] The above are merely embodiments of the present invention and do not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.< / ip>

Claims

1. A method for testing the total amount of data written to a solid-state drive, characterized in that, The method includes: Initialize the test environment and adaptive parameter configuration for the solid-state drive; Create multiple parallel data acquisition threads, collect test data through the data acquisition threads, and store the collected test data into the database; Establish a preset health assessment model, analyze the collected test data in real time according to the preset health index system, and dynamically adjust the test load according to the real-time status of the solid-state drive based on the dynamic accelerated stress model. Based on the adjustment results, perform corresponding test operations and record the test results. Analyze the test results and automatically generate a test report based on the analysis results; The steps of establishing a preset health assessment model, analyzing the collected test data in real time according to a preset health index system, dynamically adjusting the test load based on the real-time status of the solid-state drive according to a dynamic accelerated stress model, executing corresponding test operations based on the adjustment results, and recording the test results include: The collected data is analyzed in real time based on a preset health index assessment model. Based on the dynamic accelerated stress model, the test load is dynamically adjusted by switching between multiple modes such as random read / write mode, sequential write mode, and mixed read / write mode according to temperature, wear degree, and performance degradation. Based on the dynamic adjustment results, perform the corresponding test operations and record the test results; After the steps of performing corresponding test operations based on the dynamic adjustment results and recording the test results, the method further includes: When a fault is detected, the fault determination should be carried out immediately and the cause of the fault should be recorded. After the step of immediately determining the fault and recording the cause of the fault upon detection, the method further includes: Statistically analyze the types and number of faults during the testing process, and analyze the causes and impact of each type of fault based on fault identification and causes. Compare the failure analysis results with historical data, and identify potential failure modes based on machine learning algorithms; The system identifies the time series of failures under failure modes, combines the time series of failures with stress change curves to identify the concentrated failure range and corresponding load, evaluates the impact weight of different stress levels on the reliability of solid-state drives, and optimizes the dynamic adjustment strategy of the dynamic acceleration stress model based on the impact weight to improve the accuracy and efficiency of testing.

2. The method for testing the total amount of data written to a solid-state drive according to claim 1, characterized in that, The steps for initializing the solid-state drive's test environment and adaptive parameter configuration include: Install the solid-state drive under test as the test data disk; Set the test environment parameters, which include at least temperature control and power quality monitoring; Configure a test software environment, which includes at least a central control script, a data population tool, and a health management tool. Initialize the test database and establish the data record format.

3. The method for testing the total amount of data written to a solid-state drive according to claim 1, characterized in that, The steps of creating multiple parallel data acquisition threads, acquiring test data through these threads, and storing the acquired test data in a database include: The test tasks within the test cycle are initiated according to the central control script, and a data acquisition thread is created; The data acquisition thread performs performance data acquisition, health data acquisition, and physical environment data acquisition. The NTP protocol ensures that the timestamps of all collected data are synchronized, and all timestamped data is written to the InfluxDB time series database in real time.

4. The method for testing the total amount of data written to a solid-state drive according to claim 1, characterized in that, After the steps of identifying the failure occurrence time series under the failure mode, combining the failure occurrence time series with the stress change curve, identifying the failure concentration range and corresponding load, evaluating the impact weight of different stress levels on the reliability of the solid-state drive, and optimizing the dynamic adjustment strategy of the dynamic acceleration stress model based on the impact weight to improve the accuracy and efficiency of the test, the method further includes: By using a continuous iterative feedback mechanism, fault data is correlated with the stress of the response, and a fault stress model is created. The failure boundary of solid-state drives (SSDs) under different usage scenarios is calculated based on the fault stress model.

5. The method for testing the total amount of data written to a solid-state drive according to claim 4, characterized in that, The step of analyzing the test results and automatically generating a test report based on the analysis results includes: All test data, fault statistics and analysis results are integrated to generate a structured test report. The report includes trend charts of key performance indicators, health change curves and fault timelines. Test reports can be automatically pushed to a designated platform via email or API, triggering the next stage of the testing process or ending the testing task.

Citation Information

Patent Citations

  • SSD (Solid State Disk) testing method

    CN117271247A