A solid state disk compatibility test method

By automating the collection and analysis of SSD performance information using Python scripts, and combining power-on/off cycles and automatic identification and verification mechanisms after insertion, a cross-platform compatibility testing method is constructed. This method overcomes the limitations of existing testing methods, achieves efficient and accurate SSD testing, and is suitable for rapid iteration and high-quality development of enterprise-grade SSDs.

CN121301113BActive Publication Date: 2026-04-07SHENZHEN JINGCUN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-01
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing solid-state drive (SSD) testing methods suffer from fixed testing steps, limited configuration options, and a lack of flexibility, leading to undertesting, missed tests, and incorrect tests. Furthermore, they rely on imported hardware and software, resulting in high costs and making it difficult to meet the demands of rapid iteration and high-quality development. They are particularly inadequate in adapting to domestic hardware and software and domestic business scenarios.

Method used

Python scripts are used to automatically collect basic performance information of solid-state drives and generate standardized log files. Multi-user or multi-task concurrent tests are performed through automated scripts. Combined with power-on loop tests and automatic identification and verification mechanisms after insertion, a power-on loop test mechanism is formed. Test data is centrally stored and analyzed using a central server to build a cross-platform compatibility test method.

Benefits of technology

It enables cross-platform compatibility testing, improves the universality and forward-looking nature of testing, reduces labor costs, improves testing efficiency and accuracy, provides a more comprehensive and reliable quality assessment, and offers an intelligent solution for the R&D and maintenance of enterprise-level solid-state drives.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121301113B_ABST
    Figure CN121301113B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data storage, and discloses a solid state disk compatibility test method, which comprises the following steps: after the solid state disk is started, the basic performance information of the solid state disk is automatically collected based on a Python script and a standardized log file is generated; the performance operation of the solid state disk is run within a preset time period, and a plurality of user or multitask concurrent test operations are simultaneously executed based on the automatic script; a system startup command or a system shutdown command is called by using the script, and the next startup process is automatically triggered after shutdown, thereby forming a startup and shutdown cycle test mechanism, wherein the startup time, the initialization duration and the abnormal interruption record of the solid state disk are synchronously collected in each startup and shutdown process; and the solid state disk is safely removed and reinserted by using the Python script when the system is running. The application solves the problem that the existing test method platform has a narrow coverage range.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data storage, in particular to a solid state disk compatibility test method. BACKGROUND

[0002] With the rapid development of emerging technologies such as cloud computing and big data, the popularity of Internet services and the acceleration of enterprise digital transformation, data is in a state of explosion, and the market demand and scale of enterprise-level solid state disks (SSDs) are constantly rising. As a high-performance data storage device, solid state disks have the advantages of fast read and write speed, short access delay, excellent random access performance, etc., and can meet the needs of modern enterprises for data processing speed and system performance. However, despite the significant progress made in SSD technology, there are still some problems to be solved in actual application.

[0003] Traditional SSD testing methods mainly rely on manual operation and simple script testing, and have problems such as fixed testing steps, limited configuration options, lack of flexibility, etc. These methods are difficult to meet the needs of rapid iteration and high-quality development, and are prone to under-testing, missed testing and wrong testing, affecting product delivery and maintenance. In addition, existing automated testing solutions often rely on imported hardware and software and commercial testing tools, which are costly and lack of pertinence. Foreign technology may be insufficient in adapting to domestic hardware and software and domestic business scenarios. SUMMARY

[0004] The present application provides a solid state disk compatibility test method, which can solve the technical problem of meeting the needs of rapid iteration and high-quality development, and is prone to under-testing, missed testing and wrong testing.

[0005] To solve the above technical problems, one technical solution adopted by the present application is to provide a solid state disk compatibility test method, the method comprising:

[0006] After the solid state disk is started, the basic performance information of the solid state disk is automatically collected based on a Python script and a standardized log file is generated;

[0007] Run the performance operation of the solid state disk within a preset time period, and simultaneously perform multi-user or multi-task concurrent test operation based on an automated script;

[0008] Use the script to call the system startup command or shutdown command, and automatically trigger the next startup process after shutdown to form a startup and shutdown cycle test mechanism, wherein the startup time, initialization duration and abnormal interruption record of the solid state disk are synchronously collected during each startup and shutdown process;

[0009] In the system running, the solid state disk is removed and reinserted using a Python script, and after reinsertion, the device identification delay, drive letter normal rate, file system mounting success rate and data consistency check result are automatically recorded by the Python script, forming an automatic recognition and verification mechanism after insertion.

[0010] Based on the concurrent test operation, the on-off cycle test mechanism and the automatic recognition and verification mechanism, the target test is determined, different test cases are created for different solid state disk models, and the test cases are distributed to different test platforms connected in the local area network by one key, wherein the test nodes in all test processes automatically upload logs to the central server.

[0011] The beneficial effects of the present application are: by extending the test platform from the traditional Intel / AMD x86 architecture to the national security platform containing domestic Kunpeng, Feiteng and other ARM architectures, and covering the Windows and domestic OS (such as UOS, Kylin OS), the present application completely solves the problem of narrow coverage of the existing test method platform. This makes the compatibility test result of SSD can truly reflect its adaptability in the current complex and multiple computing ecology, providing reliable quality credentials for the product to enter a wider market, especially the security market, greatly improving the universality and forwardness of the test scheme. By using scripts and automation framework, the repetitive manual test is automated, the standardization and consistency of the test process are realized, the uncertainty caused by human operation is avoided, the test efficiency is significantly improved, and the labor cost is reduced. By simulating real users and enterprise-level loads (such as database operations, compression and decompression, multi-user concurrency, etc.), deep stability, data consistency and performance tests are carried out, breaking through the limitation of traditional test methods that only focus on single function, and comprehensively evaluating the performance of SSD in actual application scenarios, improving the practicality and accuracy of the test. The central test platform is constructed, the test data is automatically collected and analyzed, and the intelligent report is generated, realizing the automatic processing and standardized presentation of test results, avoiding the errors caused by manual intervention, improving the timeliness and accuracy of the test, and providing a more comprehensive and reliable reference for the application of enterprise-level SSD. By introducing enterprise-level application scenario simulation, deep reliability testing and intelligent data analysis platform, a set of forward-looking SSD quality evaluation standard is formed, which provides a more systematic, comprehensive and intelligent solution for the research, testing and operation of enterprise-level solid state disks. BRIEF DESCRIPTION OF DRAWINGS

[0012] Figure 1 is a flowchart of the solid state disk compatibility test method of the first embodiment of the present application.

[0013] Figure 2 is Figure 1Flowchart of step 1 in the middle.

[0014] Figure 3 is Figure 1 Flowchart of step 2 in the middle.

[0015] Figure 4 is Figure 1 Flowchart of step 3 in the middle.

[0016] Figure 5 is Figure 1 Flowchart of step 4 in the middle.

[0017] Figure 6 is Figure 1 Flowchart of step 5 in the middle. DETAILED DESCRIPTION

[0018] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0019] The terms "comprising" and "having" and any variations thereof in the present application are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to the process, method, product or device.

[0020] In this document, reference to "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the application. The appearances of the phrase in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily mutually exclusive of one another. It is explicitly contemplated that embodiments described herein can be combined with each other.

[0021] Figure 1 Flowchart of the solid state disk compatibility test method of the first embodiment of the present application. As shown in Figure 1 The system includes a hardware part and a software part:

[0022] Step 1, after the solid state disk is started, the basic performance information of the solid state disk is automatically collected based on the Python script and a standardized log file is generated;

[0023] Step 2: Run performance operations on the solid state disk within a preset time period, simultaneously execute multi-user or multi-task concurrent test operations based on automated scripts;

[0024] Step 3: Use scripts to call system startup or shutdown commands, automatically trigger the next startup process after shutdown, and form a startup and shutdown cycle test mechanism, where the startup time, initialization duration, and abnormal interruption records of the solid state disk are synchronously collected during each startup and shutdown process;

[0025] Step 4: When the system is running, use Python scripts to safely remove and reinsert the solid state disk. After reinsertion, automatically check the device recognition delay, drive letter normal rate, file system mounting success rate, and data consistency verification results through Python scripts, forming an automatic recognition and verification mechanism after insertion;

[0026] Step 5: Based on the concurrent test operation, startup and shutdown cycle test mechanism, and automatic recognition and verification mechanism, determine the target test, create different test cases for different solid state disk models, and distribute the test cases to different test platforms connected in the local area network with one key. During all test processes, the test nodes automatically upload logs to the central server.

[0027] In step 1, after the system where the solid state disk is located is started, the pre-deployed Python script automatically starts the information collection process. This script will call system interfaces or hardware monitoring tools to collect basic performance parameters of the solid state disk, such as read and write speed benchmark values, cache size, used / remaining storage space, power-on times, health status (such as bad block number, wear level), etc. After collection, the script will organize these information into standardized log files in a preset format (such as fixed field, structured text, or JSON format), ensuring the standardization and traceability of the information, and providing a reference for the subsequent test based on the basic data.

[0028] In step 2, within the set test duration (such as several hours or days), the actual use scenario is simulated through automated scripts to perform continuous performance operations on the solid state disk. The script will trigger multi-user access or multi-task parallel processing, such as multiple processes simultaneously reading and writing files, large file transfer, random data writing, frequent small file operations, etc., to test the performance stability of the solid state disk in high load and high concurrency scenarios (such as read and write speed fluctuations, response delay, whether to appear stuttering or speed drop). The entire process is automatically controlled by the script to start, execute and monitor the status of the task without human intervention.

[0029] In Step 3, the system's power-on command is invoked by the script to build a loop test mechanism: after the system is shut down, the next boot is automatically triggered by hardware wake-up (such as BIOS timing start) or external control tools, forming a cycle of "boot-up - running - shutdown - reboot". During each cycle of power-on and power-off, the script synchronously records key data:

[0030] Boot-up time: the time taken from booting to the system recognizing the solid state disk; initialization duration: the time taken for the solid state disk to complete self-checking and be ready to receive read and write instructions; abnormal interruption record: if there are abnormalities such as the solid state disk not being recognized, system freezing, power failure, etc. during the power-on and power-off process, the script will record the interruption time, error prompt, etc. for analysis of hardware stability.

[0031] In Step 4, when the system is running normally, the Python script calls the system device management interface to safely remove the solid state disk (ensuring that data writing is completed to avoid damage), and then controls the physical plug-in (which may need to be combined with a hardware control module to achieve automation). After reinsertion, the script automatically triggers the verification process: records the device recognition delay from the system detecting the solid state disk to fully recognizing it; calculates the normal rate of whether the drive letter is allocated as expected; checks the mounting success rate of the file system (such as NTFS, EXT4); verifies data consistency (ensures that the data is not lost or damaged during the plug-in process) by comparing the hash values of key files before and after plug-in. These results are recorded in real time to form the basis for evaluating the compatibility and stability of the solid state disk hot plug.

[0032] In Step 5, based on the test content of the previous four steps (concurrent performance, power-on and power-off cycle, plug-in verification), different test cases are designed for different brands and models of solid state disks (such as adjusting the concurrent task intensity and cycle number according to their nominal performance parameters). Through the management tool within the local area network, these test cases are distributed to multiple test platforms (such as computers or servers with different configurations) connected in one key. During all test execution processes, each test node (i.e. each test platform) automatically collects test logs, including operation records in text form, system screenshots during abnormal times, real-time collected performance data (such as read and write speed curves), etc., and automatically uploads them to the central server through the network, realizing centralized storage, unified analysis and visual display of test data, facilitating horizontal comparison of the performance and stability of different solid state disks.

[0033] Figure 2 is Figure 1 The flowchart of Step 1 in the middle is as follows: Figure 2As shown, the step S1 includes: step 101, after the solid state disk is started, the basic performance information of the solid state disk is automatically collected based on the Python script and a standardized log file is generated, wherein the basic performance information includes: step 102, all SMART information output according to the smartctl command, verification of the initial value of specific key SMART information, disk drive detailed information in the device manager, and disk capacity, serial number, firmware version recognized by the operating system.

[0034] When the system (such as Windows, Linux) where the solid state disk is located is started (manifested as the system entering the desktop / command line interface, the disk has no "initializing" prompt, and the file operation can be normally executed), the Python script pre-deployed in the system is automatically started - no manual intervention is required, and the script calls the hardware interface and tool through system permission to ensure the "non-interference" of the collection process (avoiding the influence of manual operation on the initial state).

[0035] The script organizes the collected various types of information according to the "preset standardized format", and the log needs to include key metadata: collection timestamp (accurate to seconds, used to record the initial state time point); classification label (such as "SMART information", "device manager details", "system identification attribute", convenient for subsequent retrieval); original information and structured description (such as "disk capacity: 1TB (system identification value)", instead of only retaining the original data). The generated log file (such as TXT, JSON format) needs to have "readability" and "traceability", which is convenient for manual viewing of the initial state, and also supports subsequent scripts to automatically read the baseline data for comparison.

[0036] Based on the full-amount SMART information collection of the smartctl command, SMART (Self-Monitoring, Analysis and Reporting Technology) is a hardware-level state monitoring mechanism built into the SSD, and smartctl is a cross-platform SMART information reading tool. Through this command, the underlying running data of the SSD can be obtained, which is the core basis for judging the hardware health status.

[0037] The collected content includes but is not limited to: health status identifier (such as "healthy", "warning", "fault"), cumulative power-on time, total erase count (TBW, a core indicator of SSD life), remaining life percentage, bad block number (used / backup sector number), read / write error count (such as CRC error, seek error), temperature fluctuation record, cache state, etc.; The collection logic is realized by Python script through calling system command line to execute smartctl (tool path and permission need to be configured in advance), capturing all fields of command output, avoiding missing key underlying data, and ensuring that the initial hardware state of the SSD can be traced later.

[0038] In the full SMART information, filter the key indicators that have a significant impact on life and stability, verify whether the initial value meets the "normal range", and avoid using SSDs with potential hardware problems for subsequent testing.

[0039] Key indicator filtering usually includes - remaining life percentage (normal initial value ≥ 95%, some new disks are 100%), number of bad blocks (initial value should be 0, no spare sector enable record), total erase count (new disk should be 0 or very low value, meet "unused / light usage" state), error count (all error values should be 0, no historical failure record); The script compares the key indicator values collected with the "preset normal threshold", if all indicators are within the normal range, mark "key SMART initial value verification passed" in the log; If there is an exception (such as initial bad block number > 0), mark "initial state exception" and pause subsequent testing, and first investigate hardware problems.

[0040] Device Manager is the entry of "hardware identification and driver management" at system level, collecting information here can verify the hardware compatibility and driver matching of SSD and system, to avoid deviation in subsequent testing caused by identification exception. The collection content includes but is not limited to - SSD device model (such as "Samsung 990 PRO NVMe SSD 2TB", confirm hardware model), interface type (such as "PCIe 4.0 x4" "SATA III", match hardware specifications), driver version (such as "10.0.19041.1", confirm whether the driver is the latest stable version), hardware ID (system's unique identification code for SSD, used to distinguish different devices), device status (such as "the device is running normally", exclude driver failure);

[0041] Collection logic: In Windows system, the script reads the target SSD information in the "Disk Drive" category of Device Manager by calling WMI interface (Windows Management Instrumentation); In Linux system, it is obtained by reading the device attribute file under / sys / class / block / directory, to ensure that the information is consistent with the system identification result.

[0042] Disk capacity, serial number, and firmware version collected by the operating system identification are used to confirm hardware specifications, distinguish different test devices, and trace production batches, while the firmware version directly affects the functional stability and performance of the SSD.

[0043] Data collected includes: Disk capacity: Total capacity identified by the system (e.g., "1.81TB"; note the difference from the nominal capacity, as reasonable deviations due to different calculation methods (base 1000 / base 1024) are normal), and available capacity (initially close to the total capacity, with no redundant files occupying it); Serial number: Unique hardware identifier for the SSD (e.g., "SN1234567890", used to locate the target disk in multi-device testing to avoid confusion); Firmware version: SSD underlying control program version (e.g., "BKV7"; different versions may fix performance vulnerabilities or optimize compatibility, the initial version should be recorded to avoid "firmware differences" affecting the comparison results in subsequent tests).

[0044] In Windows systems, obtain the information through "This PC - Disk Properties" or the WMI interface; in Linux systems, obtain the information through the lsblk "hdparm" command, ensuring that the information matches the actual identification result of the system and that there are no identification errors (such as abnormal capacity display or missing serial number).

[0045] During testing, a central server aggregates SMART data, power consumption fluctuations, and bit error rate changes from each node in real time. A multi-dimensional stress-coupled model is constructed by combining vibration frequency and temperature change rate to dynamically evaluate the reliability performance of solid-state drives (SSDs) under complex operating conditions. After testing, a comprehensive analysis report is automatically generated, including trend charts of key indicators, timestamps of abnormal events, and consistency verification results. This supports horizontal comparisons by model, batch, and testing cycle, providing data support for product optimization. This method, by integrating multi-stress coupling of temperature gradient changes, mechanical vibration, and I / O load, accurately simulates the real operating conditions of SSDs in harsh environments such as automotive and industrial control systems, significantly improving fault prediction capabilities. Test data shows that under conditions of 200 power-on / off cycles and a temperature change range of ±40℃ to 85℃ combined with 5-500Hz random vibration, it can effectively capture potential defects such as hidden bad blocks and controller response delays, improving the defect detection rate by 3.2 times compared to traditional temperature cycling tests. This provides a standardized technical path for verifying high-reliability storage products.

[0046] Figure 3 yes Figure 1 The flowchart for step 2 is as follows: Figure 3 As shown, step S2 includes: Step 201, installing a lightweight database and performing CRUD operations and transaction processing to verify the data consistency when the processed solid-state drive is used as the system disk and data disk; Step 202, repeatedly compressing and decompressing a large set of files using compression / decompression tools to test the mixed read / write capability of the solid-state drive; Step 203, calculating the hash value of the same algorithm before and after copying the file. If the hash values ​​output by the two steps are completely consistent, the file copy is correct; if they are inconsistent, it indicates that the file is corrupted or the copying process is incorrect, and it needs to be copied again.

[0047] Lightweight database operations and dual-bay data consistency verification: When the solid-state drive is used as both the system disk (storing database programs / system files) and the data disk (storing database data files), it ensures the integrity of data during high-frequency CRUD (Create, Read, Update, Delete) and transaction processing, avoiding data loss or corruption due to disk performance fluctuations.

[0048] Lightweight databases with low resource consumption and easy operation (such as SQLite and MySQLCommunity Lightweight Edition) are selected and deployed in two disk configurations: Configuration 1: The solid-state drive is used only as the "system disk", where the database program, system files and database data files are stored; Configuration 2: The solid-state drive is divided into two roles - it is used as the "system disk" to store the database program and system files, and at the same time as the "data disk" to store the database data files separately (achieved through partitioning or independent mounting).

[0049] Automated CRUD and transaction processing execute a series of standardized database operations through preset scripts (not at the code level, but focusing on operational logic): CRUD operations: batch insertion of structured data (such as 100,000 simulated business data records with timestamps and unique identifiers), high-frequency queries (filtering data by conditions), batch updates (modifying specified field values), and targeted deletion (deleting part of the data), covering daily high-frequency database operation scenarios; Transaction processing: executing transactions with "atomicity requirements" (such as "transfer" logic - deducting the amount from account A while increasing the amount in account B, and triggering a transaction rollback if an error occurs in the intermediate steps), ensuring that the operation is either fully executed or completely rolled back. After the operation is completed, data integrity is confirmed through dual verification: Internal verification: Query the database "key data indicators" (such as the total number of data records, the calculation result of the summation field, and the logical consistency of transaction-related data) and compare them with the preset "standard result set" to confirm that there is no missing or tampered data after CRUD; Dual disk comparison: Repeat the above operation under "system disk single disk configuration" and "system disk + data disk dual disk configuration" respectively, and compare the "data result set" and "transaction execution log" of the two operations to ensure that the solid-state drive can stably support database operations under different roles (system disk / data disk) without data inconsistency caused by disk position differences.

[0050] The specific logic for compression / decompression loop testing and mixed read / write capability verification is as follows: The test file set is selected as a "large mixed file set" instead of a single file to simulate real-world usage scenarios. The file set contains files of different types (documents, images, video clips) and different sizes (from KB-level small files to GB-level large files), with the total capacity controlled at 30%-50% of the available capacity of the solid-state drive (to avoid disk full load affecting the fairness of the test). The cyclic compression / decompression operation uses standardized compression tools (such as 7-Zip, WinRAR Lite) and performs cyclic operations according to fixed rules: Compression logic: Compress the file set into a single compressed package according to a preset format (such as ZIP, 7Z), and set the compression level to "medium" (balancing speed and compression ratio, simulating daily use scenarios); Decompression logic: Delete the source file set immediately after compression, and then decompress the compressed package to the original path to restore the file set; Number of cycles: Set the number of cycles according to test requirements (such as 10 times) to ensure that the disk is in a continuous mixed read and write state (read source files → write compressed package during compression, read compressed package → write decompressed file during decompression, and alternately trigger disk read and write operations).

[0051] Record the time taken for each compression / decompression operation and observe whether the time is within a reasonable fluctuation range (e.g., the time deviation for a single operation does not exceed 10%) to avoid "sudden increase in the time taken for a certain operation" (indicating disk read / write lag); after each decompression is completed, confirm that there are no file corruption or missing due to disk read / write errors by checking the "file quantity" (the number of decompressed files is consistent with the original file set) and "file size" (the size of a single file after decompression is consistent with the original file).

[0052] The specific logic of file copy hash value comparison and data integrity verification is as follows: Test premise and file selection: Select representative files (such as the large mixed file set used in step 202, or a single GB-level large file) to ensure that the file content is not redundant and the structure is stable (to avoid the test results being affected by the file itself being damaged).

[0053] Hash value calculation and copy operation: Before copying: Calculate the hash value of the "source file" (stored on the source path of the solid-state drive) using the specified algorithm (such as SHA256, which is more secure than MD5), and record the hash value as the "baseline value" (the calculation time and file path need to be noted); Copy operation: Perform standardized copying (such as copying within a local partition or copying across partitions to simulate different usage scenarios) to ensure that the copying process is uninterrupted (avoid test deviations caused by human intervention); After copying: Calculate the hash value of the "target file" (the copied file, stored on the target path of the solid-state drive) using the same algorithm as the source file, and record it as the "verification value".

[0054] If the "baseline value" and the "verification value" are completely consistent (no difference in character order or case), it means that there was no byte-level damage during the copying process (source file reading → data transfer → target file writing), and the solid-state drive's write / read function is normal. For exception handling: If the two are inconsistent, after ruling out human factors such as "incorrect copy path" or "target file being tampered with," it can be determined as "disk read / write anomaly" (such as sector errors when reading the source file or data packet loss when writing to the target file). The copying process needs to be repeated, and the disk hardware status needs to be checked (e.g., using the SMART information from step 102 to determine if bad blocks exist).

[0055] After step S203, the method further includes: step 204, starting the load through an automated script, running multi-user or multi-task concurrent test operations in the background and continuously monitoring system indicators and solid-state drive IO latency, and automatically terminating all test processes after a preset time, wherein the multi-user or multi-task concurrent test operations include at least background large file copying, foreground office software running, and simultaneous web browsing and video playback.

[0056] The automated load start-up, multi-task concurrent testing, and full-process monitoring termination mainly simulate the user's daily "multi-task mixed use" scenario to verify the performance stability of the solid-state drive under high IO load. By running background data transmission, foreground office work, web browsing, video playback, and other operations simultaneously, the IO latency (reflecting read and write response speed) and system resource usage of the SSD are monitored to determine whether it can meet the smoothness requirements of real-world use. At the same time, the standardization and repeatability of the test process are ensured through automated control.

[0057] The script needs to complete the "environment initialization" and "task scheduling configuration" in advance to ensure the accuracy and concurrency effectiveness of load startup. The specific logic is as follows: In the data preparation stage, multiple sets of large files need to be generated in advance, with the size of a single file controlled between 2-10GB. The file types include mixed formats such as documents, video clips, and compressed packages (simulating real user data). The storage path and the target copy path are both set within the test SSD (for example, the source path is set to the "test data area" of the SSD, and the target path is set to the "temporary storage area" of the SSD) to avoid cross-disk transfer interfering with the test results.

[0058] Regarding software and resource configuration, the executable file paths for office software (such as Office, WPS), browsers (such as Chrome, Edge), and video players (such as PotPlayer, VLC) need to be confirmed. Test resources should also be prepared, including an Excel document with formulas, a multi-image PowerPoint presentation, three different types of webpage URLs (such as a text news page, a multi-image e-commerce page, and a live video streaming page), and one 1080P local video file. For permission configuration, the script needs to obtain system administrator privileges (ensuring it can start background processes, monitor system resources, and terminate processes), and functions that may interfere with testing, such as automatic system hibernation and screensavers, should be disabled.

[0059] The script triggers all test tasks simultaneously through the "process scheduling module," employing a "background / foreground collaboration" mode to simulate real-world usage and avoid task blocking. The background tasks operate without a user interface, consuming only system resources. Specifically, the script calls system commands to initiate multi-threaded large file copying—using the `robocopy` command on Windows and the `cp -r` command on Linux—setting 3-5 concurrent copy threads to copy different large file sets simultaneously, ensuring the SSD remains under continuous read / write load.

[0060] The foreground tasks simulate user operations. Although there is an interface, they are executed automatically. Specifically, these include: office software operations, automatically opening Excel documents (performing scrolling, formula calculations, and saving operations) and PPT documents (automatically playing slides), performing simple interactions every 30 seconds (such as Excel cell editing and PPT page switching) to simulate user usage; web browsing operations, automatically opening a browser and loading three different types of tabs, refreshing the tabs every minute to simulate resource re-requests; and video playback operations, starting a player and looping a local 1080P video in the background (set to silent mode to avoid noise interference), while setting the playback quality to "HD" to generate stable I / O requests.

[0061] The design of each concurrent task needs to match the IO characteristics of users' daily use to ensure that the load can truly reflect the actual stress capacity of the SSD. The background large file copy task involves 3-5 threads copying a mixed set of 2-10GB files (both the source and destination paths are on the test SSD) simultaneously, and running in the background without a headline. This type of operation requires the SSD to continuously perform high-throughput read and write operations (belonging to the large file continuous IO scenario), which mainly tests the continuous transmission stability of the SSD.

[0062] Foreground office software operations, including Excel's automatic formula calculations (requiring data file reading), PowerPoint's automatic playback (reading image / animation resources), and scheduled document saving, involve frequent small file read / write operations (such as reading configuration files, generating temporary caches, and saving / writing documents), primarily testing the SSD's random I / O response speed. Web browsing tasks involve loading multiple tabs (reading browser cache files and web resource files) and scheduled refreshes (re-requesting and reading resources), representing a mixed I / O scenario (both random I / O for reading small cached files and continuous I / O for loading large web resources), simulating SSD performance under fragmented I / O scenarios. Video playback tasks, such as playing local 1080P videos in the background in high definition (reading video file segments), require stable continuous read I / O (continuous reading of the video stream), primarily testing the SSD's continuous read performance and cache scheduling capabilities. By combining these tasks, a mixed scenario of "continuous I / O + random I / O" and "high background load + foreground interaction" can be achieved, covering the main sources of I / O pressure on SSDs in daily user experience.

[0063] The script needs to synchronously start the "monitoring module" to collect system resource metrics and SSD core metrics in real time, and write them to the test log in the format of "timestamp + metric value" to ensure that the correlation between load changes and metric fluctuations can be traced later. System resource metrics include CPU utilization (including total utilization and utilization of each core), memory utilization (including used capacity and free capacity), and SSD disk utilization. In Windows systems, these metrics are collected through the WMI interface and Task Manager API; in Linux systems, the output data is captured periodically through the top, vmstat, and df -h commands, with a collection frequency of 5 seconds / time. SSD IO latency metrics include average read latency (unit: ms), average write latency (unit: ms), and IO queue length. In Windows systems, the disk IO module of Resource Monitor is used in conjunction with the diskpart tool for auxiliary collection; in Linux systems, the iostat -x command and iotop tool are used for collection, with a collection frequency of 1 second / time (due to the rapid changes in IO latency, a higher frequency is required to ensure data accuracy).

[0064] Log entries must include "task start time - current timestamp - metric values ​​- task status", for example, "2025-11-11 14:30:00 | 10 seconds after task start | CPU utilization 65% | Memory utilization 70% | SSD read latency 3.2ms | Write latency 4.5ms | Large file copy progress 30%". If any abnormal metrics occur (such as IO latency suddenly exceeding 50ms, CPU or memory utilization remaining ≥90% for 10 seconds), the script must mark the "abnormal time point" in the log and record it in bold for easy analysis of performance bottlenecks later.

[0065] The script uses a "time control module" to automatically terminate the test process, ensuring that all test tasks are smoothly terminated after a preset time (e.g., 30 minutes, 1 hour, configurable as needed). This prevents residual processes from consuming resources or corrupting test data. The main trigger condition is time-based. When the script starts, it records the "start time" and calculates the "run time" in real time. When the run time reaches the preset time (e.g., the preset duration is set to 30 minutes), the termination process is automatically triggered. In addition, optional abnormal trigger conditions can be set. If SSD IO latency is detected to be ≥100ms for 30 seconds (or the system is unresponsive), the script can trigger termination in advance to avoid hardware overload and mark "abnormal termination" in the log.

[0066] Termination Steps: First, the script records the "last monitoring data before termination" to ensure the integrity of the test segment log. Second, terminate processes in the order of "foreground tasks → background tasks"—when terminating foreground tasks, use the command `taskkill / PID processID / F` on Windows and `kill -9 processID` on Linux to close office software, browsers, and video players (the PIDs of each process at startup need to be recorded beforehand). When terminating background tasks, terminate the `robocopy` process on Windows and the `cp` process on Linux, while deleting temporary files generated during copying (to avoid occupying SSD space). Third, after termination, check for "residual processes" by using the `tasklist` command on Windows or the `ps` command on Linux to confirm that all test-related processes have been closed. If any remainders are found, perform a second termination. Fourth, output a "termination summary," recording the total test duration, the completion status of each task (e.g., "2 large file copies completed, 1 remaining unfinished"), and peak monitoring metrics (e.g., "maximum IO latency 12ms") in the log to provide summary information for subsequent performance analysis.

[0067] Simulating a real-world multitasking scenario where a user is copying files, writing documents, and watching videos simultaneously, the study focuses on verifying three key aspects: first, the latency stability of the SSD under mixed I / O loads to determine if multitasking causes a sudden increase in latency, thus affecting operational smoothness; second, the coordination capability between the SSD and system resources (CPU, memory) to determine if system lag is caused by disk I / O bottlenecks; and third, the hardware reliability of the SSD under prolonged high loads to determine if I / O errors, process crashes, or other anomalies occur.

[0068] Figure 4 yes Figure 1 The flowchart for step 3 is as follows: Figure 4 As shown, step S3 includes: Step 301: Using a script to call the system power-on or power-off command, sending power-on / off commands to the programmable power socket via network protocols (such as SNMP, HTTP, or custom TCP commands); Step 302: Cutting off / turning on the programmable power supply according to the acquired power-on / off commands, simulating unplugging / not unplugging the power, automatically recording the current state after power-off, and triggering a self-test process by the script after power-on to verify the solid-state drive mounting status and restore the running state before power-off, repeating the power-on / off cycle test mechanism a preset number of times; Step 303: Synchronously collecting the solid-state drive's startup time, initialization time, and abnormal interruption records during each power-on / off process, statistically analyzing file system errors, disk recognition failures, and log error information after abnormal restarts, and evaluating the stability and data retention capabilities of the solid-state drive under extreme power supply environments.

[0069] The power control command sending in collaboration with scripts and network protocols requires pre-written automated scripts (such as Python and PowerShell scripts) to have dual capabilities: "calling system power on / off commands" and "sending network commands." On the one hand, the script can directly call native system commands (Windows' shutdown command, Linux's shutdown -h now command) to achieve "normal shutdown," and then achieve "normal power on" through power control. On the other hand, for "unexpected power outage" scenarios, the script does not need to call the system shutdown command, but directly sends a "power outage command" to the power outlet through the network protocol to simulate a sudden power interruption.

[0070] The script needs to complete permission configuration in advance (such as obtaining system administrator privileges to call the shutdown command and obtain power socket control permissions), and preset key parameters: the IP address of the programmable power socket, communication port, authentication information (such as username / password, key, to avoid unauthorized control), to ensure a stable communication link between the script and the power socket.

[0071] The script sends power on / off commands to the programmable power socket through three common network protocols. The specific selection needs to match the power socket's support capabilities. The application logic of each protocol is as follows: SNMP protocol (Simple Network Management Protocol): Applicable to standardized power sockets that support SNMP. The script calls the power control node through the SNMP "Management Information Base (MIB)" and sends "port power off" or "port power on" commands. For example, the snmp-set command specifies a specific port of the power socket (the port connected to the test host) and sets its power supply status to "off" or "on". At the same time, it can synchronously obtain the current power supply status of the power socket (such as voltage and current) to help determine whether the command was executed successfully.

[0072] The HTTP / HTTPS protocol is suitable for power outlets that provide a web management interface. The script calls the power control API by sending HTTP requests (such as POST or GET)—for example, sending JSON format commands ({"status":"off"} or {"status":"on"}) to the power outlet's web interface (such as http: / / [power IP] / api / power / port1) to control power on / off. This method does not require complex protocol configuration, only interface authentication (such as token or cookie), and is compatible with most lightweight programmable power supplies. Custom TCP commands are suitable for power outlets that do not support standard protocols and only provide private communication interfaces. The script establishes a TCP connection (specifying the power outlet's IP and port) and sends pre-agreed binary or text commands (such as "POWER_OFF_1" representing port 1 is off, and "POWER_ON_1" representing port 1 is on). The command format needs to be confirmed with the power outlet manufacturer in advance to ensure that the commands can be correctly parsed and executed. Regardless of the protocol used, after the script sends a command, it must wait for "command confirmation feedback" (such as the power socket returning an "OK" status code, or SNMP query showing a change in port status) to avoid control failure due to communication packet loss.

[0073] The power-on / off scenario simulation and system self-test recovery process simulates two core power supply scenarios through the power-on / off control of a programmable power supply. Upon power restoration, an automated self-test is triggered to ensure the SSD can resume normal operation. The specific logic is as follows: Simulation execution of the two power supply scenarios: Scenario 1: Simulating normal power on / off without unplugging (corresponding to shutting down and powering on in daily use): The script first calls the normal system shutdown command (such as Windows shutdown / s / t0, Linux shutdown). -rnow), wait for the system to completely shut down (confirmed by monitoring the host indicator lights and network connection status), and then send a "power off → delay 5 seconds → power on" command through the power socket (a brief power outage simulates a "reboot," but without completely cutting off power before restoring), achieving a "normal power on / off" cycle; Scenario 2: Simulate an unexpected power outage (corresponding to sudden power outages or power failures): When the system is running normally (e.g., executing the multi-task test in step 204), the script does not call any system shutdown commands, but directly sends an "immediate power off" command through the power socket to cut off the host power supply; after the power outage lasts for a preset duration (e.g., 10 seconds, simulating the duration of the power outage), it sends a "power on" command to simulate a "power restoration after an unexpected power outage" scenario. Both scenarios need to be repeated a preset number of times (e.g., 50 times, 100 times, the number of times can be configured according to test requirements). Before each loop, the script needs to record the "current loop number" and "scenario type" to ensure test traceability.

[0074] After powering on, the system performs a self-test and restores its state. Once the host is powered on and boots up, the script automatically triggers a "self-test process," which primarily verifies the "availability" and "state continuity" of the SSD. The specific process is as follows: Step 1: Verify SSD hardware recognition—The script checks whether the target SSD is successfully recognized by the system by reading the system's Device Manager (Windows) or using the `lsblk / fdisk` command (Linux). If not recognized, it marks "Disk recognition failed" and records it in the log. Step 2: Verify file system integrity—The script calls file system checking tools (Windows' `chkdsk` command, Linux's `fsck` command) to scan the SSD partition's file system, checking for bad sectors, file index errors, data block corruption, etc. If errors are found, it records the "file system error type" and "repair result." Step 3: Restore the operating state before power failure— If the self-test shows no abnormalities, the script automatically restarts the test process before the power outage (such as the multi-task concurrent test in step 204), restoring the test progress to the state before the power outage (e.g., restarting the large file copy thread, office software automation operations), ensuring the continuity of the loop test; if the self-test reveals unrecoverable errors (such as file system corruption, SSD inability to mount), the test is paused and marked as "critical anomaly," awaiting manual investigation. In another embodiment, the motherboard BIOS is configured to support "power-on auto-boot," or the boot process is triggered through other out-of-band management methods (such as IPMI).

[0075] Key data collection and extreme environment stability assessment provide a "data-driven summary" of the power-on / off cycle test results. By collecting core metrics and statistically analyzing anomalies, the performance of the SSD under extreme power supply conditions is evaluated. Real-time data collection dimensions and methods include: SSD boot time: The time from when the host is powered back on to when the system recognizes the SSD via the Device Manager / lsblk command is recorded as the "boot time" (unit: seconds). For example, if the system recognizes the SSD 15 seconds after power-on, the boot time is 15 seconds. This metric reflects the speed of SSD hardware initialization. SSD initialization time: The time from when the system recognizes the SSD until the SSD completes partition mounting and executable file read / write operations is recorded as "initialization time" (unit: seconds). For example, if the SSD partition is successfully mounted and test files can be opened 8 seconds after recognition, the initialization time is 8 seconds. This metric reflects the readiness efficiency of the SSD software layer (driver, file system). Abnormal interruption recording: If situations such as "unable to power on after power failure," "system blue screen after power-on," "SSD recognition failure," or "file system error" occur in each loop, the script needs to record the "time of occurrence of the abnormality," "loop scenario type," and "description of the abnormal phenomenon" (e.g., "23rd unexpected power failure loop, SSD not recognized after power-on"). Simultaneously, it should capture SSD-related error information (such as "disk I / O error" or "partition table corruption") from the system logs (Windows Event Viewer, Linux / var / log / syslog).

[0076] After all loop tests are completed, the script performs statistical analysis on the collected data. The core evaluation dimensions are as follows: Anomaly incidence rate: The script calculates the proportion of "file system error count", "SSD recognition failure count", and "serious anomaly pause count" to the total number of loops. For example, if there are 2 minor file system errors in 100 loops, the anomaly incidence rate is 2%. This proportion is usually required to be below 5% (adjustable as needed); Startup and initialization time fluctuation: The script calculates the average and standard deviation of "startup time" and "initialization time" in all normal loops. If the standard deviation is too large (e.g., the average startup time is 15 seconds, and the standard deviation exceeds 5 seconds), it indicates that the SSD has poor initialization stability under power fluctuations; Data retention capability verification: For each recovery scenario after an abnormal interruption, the script verifies the data integrity by comparing the "hash value of the key test file before power failure" with the "hash value of the file after power failure". If the hash values ​​are consistent, it indicates good data retention capability; if they are inconsistent, it indicates that the power failure caused data corruption, and it is necessary to evaluate whether the SSD's cache power loss protection function is effective (e.g., whether there is a spare capacitor to prevent data loss).

[0077] After step S303, the method further includes: step 304, before entering hibernation / sleep, creating a temporary file with a timestamp and a unique check value through a script; when the system wakes up, the script automatically runs and reads the temporary file, checks whether the temporary file exists and whether the check value is correct, and records the wake-up time; step 305, continuously executing the power-on / off cycle and wake-up cycle within a preset number of times, calculating the wake-up success rate and drawing a wake-up time distribution chart, and checking for any abnormal delays.

[0078] The process of creating temporary files before hibernation and verifying them after wake-up uses the logic of "pre-hibernation data embedding - post-wake-up data verification" to ensure that the SSD does not lose or corrupt data during hibernation, while quantifying wake-up efficiency. The specific operation logic is as follows:

[0079] Temporary file creation and information recording before hibernation: Before the system triggers the "hibernate / sleep" command (the script can automatically start 10 seconds before the hibernation command is executed by calling the system API or task scheduler), the script will perform a "temporary file generation" operation. The core is to ensure the "uniqueness" and "verifiability" of the file: Step 1: Generate a temporary file with a precise timestamp - The file naming format includes a millisecond-level timestamp (such as "ssd_test_temp_20251111153022888.txt", where "20251111153022888" represents 15:30:22.888 on November 11, 2025), ensuring that the file generated each time hibernation is not duplicated; the file content is written with fixed test data (such as a 100KB random character sequence), and metadata such as "generation timestamp" and "target SSD serial number" are appended to the end of the file for easy association and confirmation after wake-up. Step 2: Calculate the file's unique checksum – The script calls a hash algorithm (such as SHA256, consistent with step 203 to ensure consistent verification logic) to calculate the hash value of the generated temporary file. This "unique checksum," along with the file path and generation timestamp, is recorded in the "Hibernation / Wake-up Log" (independent of the power-on / off log for easy traceability). Simultaneously, the checksum is backed up in the system registry (Windows) or the / var / log directory (Linux) to prevent data loss in case of temporary file loss. Step 3: Confirm file writing completion – The script reads the actual size of the temporary file and compares the file hash values ​​before and after writing to ensure the file has been completely written to the target SSD (not just cached). This prevents file corruption after hibernation due to cache synchronization issues. After confirmation, the script triggers the system hibernation / sleep command (Windows' `shutdown / h` command, Linux's `systemctl suspend` command).

[0080] File verification and wake-up time recording after wake-up: When the system is woken up by user operation (such as pressing the power button) or remote script command, the script will run automatically through the "system startup item preset" or "wake-up task trigger" mechanism, and perform the following verification and recording operations: Step 1: Check the existence of temporary files - The script searches for temporary files on the target SSD based on the file path recorded before hibernation. If the file does not exist, it immediately marks "file lost after hibernation" in the log and triggers subsequent file system checks (such as calling chkdsk / fsck to check for SSD partition errors); if the file exists, proceed to the next verification step. Step 2: Verify file integrity - The script recalculates the SHA256 hash value of the temporary file and compares it with the "unique checksum" recorded before hibernation: If the two are completely consistent, it means that the data on the SSD was not tampered with or damaged during hibernation, and it is marked "data verification passed"; if they are inconsistent, it is marked "data corrupted", and the "difference in checksum before and after hibernation" and "file corruption location (located by file segment verification)" are recorded to analyze whether data loss was caused by the lack of cache protection during SSD hibernation. Step 3: Record the wake-up time - Start timing from when the system begins to respond to the wake-up command (such as the power button light turning on or the fan starting) until the system fully enters the operating system desktop and the script successfully reads the temporary file and completes the verification. Record this as the "wake-up time" (unit: seconds). For example, if it takes 7 seconds from pressing the power button to the verification being completed, the wake-up time is 7 seconds. This metric reflects the SSD's response efficiency from hibernation to a usable state (including the entire process of SSD hardware wake-up, driver reloading, and file system readiness).

[0081] Multi-loop execution and wake-up stability statistical analysis transforms the single-time verification result of step 304 into a "batch stability conclusion" through "continuous loop testing + data statistical visualization," quantifying the SSD's performance under frequent hibernation and wake-up. The specific logic is as follows: The script continuously executes the "power-on loop" and "wake-up loop" according to preset rules to ensure coverage of alternating scenarios of "complete power failure" and "low-power hibernation," avoiding the limitations of single-scenario testing. Loop count configuration: Preset total number of loops (e.g., 100 times, which can be adjusted as needed), where the "power-on loop" and "wake-up loop" are allocated proportionally (e.g., 50 power-on loops + 50 wake-up loops, or one wake-up loop after every two power-on cycles, simulating the user's usage habit of "frequent wake-ups on weekdays - long-term shutdown on weekends"). Loop Execution Logic: Before each loop, the script clears the temporary files from the previous iteration (to avoid file accumulation affecting the test), and then executes according to the scenario: If it is a "wake-up loop", the "sleep-wake-verification" process in step 304 is executed; if it is a "power-on / off loop", the "programmable power supply on / off" process in step 302 is used. The execution results of both scenarios are uniformly summarized in the "Comprehensive Test Log", with the scenario type and loop number marked (e.g., "Loop 35 - Wake-up Scenario" "Loop 36 - Power-on / off Scenario"). Exception Handling Mechanism: If a serious exception such as "wake-up failure (system cannot wake up)" or "SSD cannot be recognized after power-on / off" occurs in a certain loop, the script will not immediately terminate all tests, but will mark the loop as an "abnormal loop", skip the subsequent process and directly enter the next loop, while recording the "abnormal trigger scenario" and "abnormal phenomenon". After all loops are completed, the number of exceptions will be counted centrally to avoid a single exception causing the overall test to be interrupted.

[0082] After all loop tests are completed, the script performs statistical analysis on the collected "wake-up related data". The core is to intuitively present the SSD's wake-up stability through "quantitative indicators + visual charts": The first dimension: wake-up success rate calculation - to calculate the proportion of "successful wake-up times" (file verification passed after wake-up, no hardware recognition errors) to "total wake-up loop times". For example, if 48 out of 50 wake-up loops are successful, the wake-up success rate is 96%. Usually, this proportion is required to be no less than 95% (which can be adjusted according to the application scenario, such as ≥99.9% for server scenarios). If the success rate is lower than the threshold, it is necessary to combine log analysis to see if there is "specific loop number clustered anomalies" (such as failure every 10th wake-up) to locate whether it is related to hardware fatigue caused by the accumulation of SSD hibernation times. The second dimension: Wake-up time distribution analysis – Group all successful wake-up times by interval (e.g., 0-5 seconds, 5-10 seconds, 10-15 seconds, over 15 seconds), and calculate the percentage of wake-ups in each interval to generate a "wake-up time distribution chart" (e.g., bar chart or frequency distribution histogram). Under normal circumstances, wake-up times should be concentrated in the 0-10 second interval and evenly distributed. If "the percentage of abnormal delays exceeding 15 seconds exceeds 10%", or if the wake-up time shows an "increasing trend" with the number of cycles (e.g., an average of 5 seconds for the first 20 cycles and an average of 12 seconds for the next 20 cycles), it indicates that the SSD may have problems such as "slower driver response", "file system fragmentation accumulation", or "decreased hardware initialization efficiency" after frequent wake-ups. The third dimension: Anomaly correlation analysis – Cross-compare "wake-up anomalies" (file loss, data corruption, wake-up failure) with "power-on / off anomalies" (SSD recognition failure, file system error) to analyze whether there is a correlation pattern of "wake-up anomalies immediately followed by power-on / off anomalies". This helps determine whether the anomalies are caused by underlying hardware vulnerabilities of the SSD (e.g., unstable cache chips) rather than occasional problems in a single scenario.

[0083] Before step S4, the method further includes: step 401, using the FIO tool to preprocess the solid-state drive according to the Global Network Storage Industry Association standard, and testing the solid-state drive's random read / write IOPS, sequential read / write bandwidth, and latency at different queue depths under steady state; step 402, using the FIO tool to continuously perform random writes to the solid-state drive, monitoring the growth of the "Media Wear Indicator" or "Host Writes" in the SMART information, and making a preliminary assessment of the solid-state drive's lifespan performance.

[0084] Step 401 Preprocessing based on the SNIA standard: The core goal of preprocessing is to "exhaust the SSD's temporary cache" and "trigger the wear leveling mechanism," allowing the SSD to enter a stable "steady-state phase." The SNIA standard explicitly requires that preprocessing write "2-5 times the nominal capacity of the SSD" and use "random write mode" (simulating fragmented write scenarios in real applications). The specific operation logic is as follows: FIO preprocessing script configuration: Use the FIO tool to write a preprocessing task, set the block size to 4KB (SNIA's recommended "typical random IO block size," close to the IO characteristics of databases, operating systems, and other applications), the IO mode to "random write (randwrite)," and the queue depth (QD) to 32 (covering medium to high load scenarios). At the same time, disable the temporary acceleration function of "write cache" (by using the FIO direct=1 parameter, ensure that data is written directly to the SSD flash memory chip, rather than the system or SSD cache); Preprocessing execution and steady-state judgment: Start the FIO preprocessing task and continuously write data (e.g., 2-5TB of data needs to be written for a 1TB SSD), recording the write performance (throughput) every 30 minutes during the process. When the throughput fluctuation range recorded for three consecutive times is ≤5%, the SSD is determined to have entered a steady state (indicating that the cache has been exhausted, the wear leveling mechanism is running stably, and the performance no longer fluctuates due to the initial state). At this point, preprocessing is stopped, and the performance testing phase begins.

[0085] Core performance metrics testing under steady state: After entering steady state, the SSD's "random read / write IOPS", "sequential read / write bandwidth", and "latency at different queue depths" are tested using the FIO tool in different scenarios. Each test item must be run for a sufficient duration (≥30 minutes) according to the SNIA standard to avoid short-term fluctuations affecting the results. The specific test logic is as follows: Random read / write IOPS test: IOPS (input / output operations per second) reflects the SSD's ability to process random small files (such as office software, database queries). The FIO configuration is "4KB block size + random read / write (randrw, read / write ratio set to 7:3, simulating a read-heavy, write-light scenario in real applications)". The IOPS at queue depths (QD) of 1, 8, 32, and 64 are tested respectively. Values—QD1 corresponds to light desktop user load (such as opening files), QD32 / 64 corresponds to high-concurrency server scenarios (such as multiple users accessing a database). Record the "random read IOPS" and "random write IOPS" under each QD, and take the average value over 30 minutes as the steady-state result; Sequential read / write bandwidth test: Bandwidth (MB / s) reflects the SSD's ability to process consecutive large files (such as video copying, large file transfer). FIO is configured as "128KB block size + sequential read / write," with random mode disabled. Bandwidth values ​​are also tested under different QDs (1, 8, 16)—the sequential read test simulates a "reading local video" scenario, and the sequential write test simulates a "backing up large files" scenario. Record the "sequential read bandwidth" and "sequential write bandwidth" under each QD to ensure the results meet SNIA's testing requirements for "continuous IO performance"; Different queue depth latency test: Latency (milliseconds / microseconds) reflects the SSD's speed in responding to IO requests (directly affecting user operation smoothness). FIO is configured as "4KB..." "Block Size + Random Read / Write" tests were conducted for "average read latency" and "average write latency" at QD1, 4, 16, and 32 respectively. The focus was on latency at QD1 (the "instant response speed" most commonly perceived by desktop users, such as the waiting time for double-clicking to open a file). The average latency under steady-state conditions was required to be ≤10 milliseconds (excellent SSDs can achieve microsecond levels). The "99th percentile" of the latency was also recorded (reflecting latency performance under extreme conditions, avoiding the possibility of individual high-latency instances slowing down the overall experience). All test results were recorded in the SNIA standard format, including "test time, SSD model, preprocessed data volume, steady-state values ​​of each indicator, and test parameter configuration" to ensure cross-comparison of test results from different brands of SSDs. FIO Continuous Random Write and SMART Wear Indicator Monitoring applied a "continuous write load" through FIO to simulate the wear process of an SSD during long-term use, while simultaneously monitoring key wear indicators in the SMART information in real time to preliminarily estimate the SSD's lifespan. The core principle was to establish a correlation between "write volume" and "wear level," providing a quantitative basis for lifespan assessment.

[0086] FIO Continuous Random Write Load Configuration: To simulate "high-frequency write scenarios" in real-world use (such as server log writing, database updates, and video surveillance storage), FIO needs to be configured with "high-pressure, long-duration" random write tasks. The specific logic is as follows: Load parameter settings: Block size is set to 4KB (fragmented writing, which causes more significant wear on the SSD), IO mode is "pure random write (randwrite)", queue depth (QD) is set to 16 (medium to high load, to avoid extreme load causing test distortion), and the cache is bypassed by using the direct=1 parameter to ensure that each write directly affects the SSD flash memory chip; at the same time, "infinite loop write" (FIO's runtime=0 parameter) is set, or a long duration is preset (such as 24 hours, 72 hours, adjusted according to the test cycle) to ensure that a significant change in the SMART wear index can be observed during the test. Load stability control: During the test, the write throughput is recorded every hour. If the throughput drops sharply (e.g., the drop is greater than 20%), it is necessary to check whether the SSD has entered "write protection mode" (triggered by excessive wear). If it has not been triggered, the test continues. If it has been triggered, the test is stopped and the "total write volume at the time of triggering" is recorded as a key data point for subsequent lifespan assessment.

[0087] In real-time monitoring and lifespan assessment of SMART wear indicators, the focus is on monitoring two core indicators in SMART information that are directly related to the "degree of wear". Through regular collection and analysis, the lifespan performance of the SSD can be preliminarily judged. The specific logic is as follows: Key SMART indicator identification: Media wear indicator (SMART attribute ID: 177 [Samsung], 233 [Intel / Micron], etc., the ID may be different for different brands): This indicator is usually presented in the form of "percentage". The initial value is 100%, which represents "no wear". As the write volume increases, the percentage gradually decreases - when it drops to the threshold (usually 10%), it means that the SSD is close to the end of its lifespan and needs to be replaced; Host write volume (TBW, SMART attribute ID: 241 [most brands]): Records the "total amount of data written" (unit: TB) from the time the SSD leaves the factory to the present. It directly corresponds to the "TBW lifespan" claimed by the SSD manufacturer (e.g., a 1TB SSD is claimed to have 600TBW, which means that it can theoretically write 600TB of data). It is the core quantitative indicator for lifespan assessment.

[0088] Indicator monitoring and data recording are performed using the smartctl tool (consistent with step 102 to ensure tool uniformity). SMART information is collected every hour, and the "media wear indicator percentage" and "host write volume" are extracted and recorded in the "life test log" to form a corresponding table of "time-write volume-wear percentage" (e.g., "12 hours of test, host write volume 1.2TB, media wear indicator 98%").

[0089] Preliminary lifespan assessment logic: Based on TBW: If the host write volume is 2.4TB after 24 hours of testing, the "average daily write volume = 2.4TB" can be calculated. If the SSD's nominal TBW is 600TB, then the theoretical lifespan ≈ 600TB ÷ 2.4TB / day = 250 days (Note: This is a "theoretical value under extreme write scenarios"; in actual use, the write volume is lower, and the lifespan will be longer). Based on wear indicator: If the media wear indicator drops from 100% to 97% after 72 hours of testing, a 3% decrease in wear corresponds to 7.2TB of write volume. Therefore, it can be calculated that "every 1% decrease in wear requires 2.4TB of write volume." When the wear drops to 10%, the total write volume ≈ (100% - 10%) × 2.4TB = 216TB. If the nominal TBW is 600TB, this indicates that the wear rate in this scenario is slower than the theoretical value, resulting in a better lifespan. This assessment is a "preliminary prediction." In actual use, factors such as write intensity, data type (random / sequential), and ambient temperature will affect the actual lifespan. The final lifespan needs to be judged comprehensively based on long-term actual use data. However, this step can quickly screen out unqualified SSDs with "abnormally fast wear rate."

[0090] Figure 5 yes Figure 1 The flowchart for step 4 is as follows: Figure 5 As shown, step S4 includes: Step 403, during system operation, using a Python script to safely remove and reinsert the SSD; Step 404, using a Python script to automatically check the system's automatically recorded device recognition latency, drive letter normality rate, file system mount success rate, and data consistency verification results, forming an automatic recognition and verification mechanism after insertion; Step 405, using a script to automatically and repeatedly execute the safe removal and reinsertation operation, recording the latency and mount status of each SSD re-recognition; Step 406, calculating the average recognition latency and mount success rate after multiple operations, and combining the logs recorded in the system to analyze the cause of abnormal interruptions, verifying the stability and data integrity guarantee capability of the SSD in frequent hot-swapping scenarios.

[0091] Python script-driven safe removal and re-insertion of SSDs serves as the "operation entry point" for hot-swapping tests. The script achieves "software-level safety preprocessing + hardware-level precise insertion and removal," preventing data loss or hardware damage due to forceful insertion and removal. The software preprocessing for safe removal ensures the SSD is in a "safely removable" state before execution, preventing power loss during incomplete read / write operations. The specific process is as follows: Step 1: Check the current read / write status of the SSD—by calling system interfaces (Windows' WMI interface, Linux's iotop / fuser commands) to query whether there are active read / write processes (such as file copying or database operations) on the target SSD partition. If so, wait for the process to end or forcibly terminate it (a "list of non-critical processes allowed to terminate" must be configured beforehand to avoid affecting core system services); Step 2: Trigger a safe eject—In Windows, the script calls the "eject" method of the "disk drive" via WMI (or simulates the "safely remove hardware" operation in Device Manager) to ensure the system unloads the SSD driver and clears cached data; In Linux, it executes `eject / dev / sdX` (where X is the SSD device identifier) ​​or `umount`. The ` / mnt / ssd` command (unmount point) completes the safe removal at the software level. After successful removal, the script will receive a "Device safely removed" confirmation signal from the system (such as Windows event logs or Linux dmesg information). The third step is hardware-level re-insertion—For physical machine testing, automated physical insertion and removal must be achieved using a "programmable hardware module" (such as a relay with USB control or a robotic arm plug-and-play device): the script sends an "insertion command" to the hardware module via serial port / network, and the control module reconnects the SSD to the USB / SATA interface. For virtual machine testing, the "disconnect-reconnect device" operation can be simulated through virtualization platform APIs (such as VMware PowerCLI or VirtualBox SDK), without requiring physical hardware intervention.

[0092] The script needs to be designed with adaptation schemes for the differences between Windows and Linux systems: Windows should prioritize the use of the WMI interface (strong compatibility, no additional tools required), and if the interface call fails, it should fall back to using the devcon command (Microsoft's official device management tool); Linux should rely on udev rules (to monitor device plug-in and unplug events) and the mount / umount command to ensure that software operation is synchronized with hardware status and avoid the false state of "software is shown to be removed but hardware is not disconnected".

[0093] The core logic of the automatic identification and verification mechanism after insertion is the "quality verification unit" of the hot-plug test. It automatically checks four core dimensions through a script to verify the "availability" and "data security" of the SSD after re-insertion. The specific verification logic is as follows: Device recognition latency statistics: The time from when the script sends the "re-insertion command" (or the hardware module responds with an "inserted" signal) to when the system successfully recognizes the SSD device (displaying the device model and capacity) through Device Manager (Windows) or the lsblk / fdisk command (Linux) is recorded as "device recognition latency" (unit: seconds). For example, if the system displays the SSD in the lsblk results 3 seconds after insertion, the recognition latency is 3 seconds. This metric reflects the system's hardware enumeration speed for the SSD. If the latency exceeds 10 seconds (the threshold for typical scenarios), it is marked as "slow recognition".

[0094] Drive letter accuracy verification: The script records the original drive letter (e.g., "D:" in Windows, " / dev / sdb1" in Linux) before the SSD is removed. After reinsertion, it checks if the system-assigned drive letter matches the original. If they match, the drive letter is considered "normal." If they don't match (e.g., "E:" in Windows, " / dev / sdc1" in Linux), further investigation is needed to determine if it's due to the original drive letter being occupied (e.g., other devices being connected). If the change is without reason, it's marked as "abnormal." The drive letter accuracy rate is subsequently calculated as "number of normal drives / total number of drives" to reflect the stability of system drive letter allocation.

[0095] File system mount success rate verification: The script verifies the mount status in the following ways: In Windows, check if the SSD partition icon is displayed in "This PC", or check if the partition is in an "available" state using the wmic logicaldisk get name,description command; In Linux, execute mount|grep / mnt / ssd (assuming the mount point is / mnt / ssd). If the output contains mount information and there is no "read-only mount" marker (ro), then the mount is considered "successful"; If a "mount failure" occurs (e.g., the mount command returns an error code) or "read-only mount" occurs (the file system detects an error and automatically protects itself), then a "mount error" is marked, and the dmesg / event log is called to check the specific error cause (e.g., partition table corruption, file system error).

[0096] Data consistency verification: The script follows the "hash value comparison" logic (consistent with steps 203 and 304 to ensure a unified verification standard): A "verification baseline file" (such as a 1GB mixed format file containing text, images, and compressed files) is created in advance on the SSD, and its SHA256 hash value is calculated and stored; after re-insertion, the script automatically reads the file and recalculates the hash value, comparing it with the baseline value: if they match, "data integrity" is determined; if they do not match or the file is lost, "data corruption / loss" is marked, and file system errors are checked using chkdsk (Windows) / fsck (Linux) to locate whether data block corruption was caused by unplugging and plugging.

[0097] Automated loop execution and data logging constitute the "batch data acquisition unit" of hot-plug testing. Scripts control the number of loops and intervals to ensure test coverage of "high-frequency plugging / unplugging" scenarios, while recording key data for each loop. The specific logic is as follows: Loop parameter configuration: The script presets core loop parameters: total number of loops (e.g., 100 times, adjustable according to the test cycle; ≥500 times is recommended for server scenarios), single loop interval (e.g., 30 seconds, to avoid excessive system resource consumption due to frequent plugging / unplugging in a short period), and number of retry attempts for exceptions (if a plugging / unplugging fails, it automatically retryes twice; if it still fails, it is marked as "loop exception" and proceeds to the next iteration). Parameters can be modified through the script configuration file without modifying the code, improving flexibility.

[0098] Each loop executes in the order of "safe removal → re-insertion → four-dimensional verification," and the script writes data to the "hot-plug test log" in real time. Each log entry includes: loop number (e.g., "loop 23"), execution timestamp, identification latency, drive letter status (normal / abnormal), mount status (success / failure), data consistency result (complete / corrupted), and anomaly description (e.g., "loop 23: mount failed, dmesg displays 'ext4 file system superblock error'"). If a serious anomaly such as "data corruption" occurs in a loop, the script automatically backs up the system logs (Windows event log, Linux / var / log / syslog) for later traceability.

[0099] To avoid frequent plugging and unplugging damaging the SSD interface, the script needs to monitor the "interface temperature" and "connection stability" data fed back by the hardware module (if the hardware supports it): if the interface temperature exceeds 60℃ (the normal safety threshold) or "hardware connection timeout" occurs 3 times in a row, the script will automatically pause the loop for 10 minutes (to cool down / recover) and continue after the status returns to normal to ensure hardware safety.

[0100] After executing a preset number of hot-plug cycles, the data consistency verification results are compared with the initial baseline value each time to detect whether there is data offset or corruption. Combining the system dmesg log and SSD firmware log, the latency fluctuations in the driver loading, device enumeration, and file system response stages during abnormal timing are analyzed. The identified latency data is fitted according to a normal distribution, and outliers exceeding three times the standard deviation are marked, and associated with power management policies and PCIe link status changes. Finally, a hot-plug stability assessment report is generated, covering the mount success rate, average identification latency, data integrity pass rate, and frequency of key error codes, completing the comprehensive reliability verification of the solid-state drive in dynamic plugging and unplugging scenarios.

[0101] Statistical analysis and hot-swappable stability assessment, through quantitative statistics and log analysis, establishes conclusions regarding SSD stability in frequent hot-swappable scenarios. Key quantitative indicators include: Average recognition latency: Calculating the average recognition latency across all "successful recognition" loops. If the average is ≤5 seconds (excellent threshold), the system's recognition efficiency is high; if the average is >8 seconds, it's necessary to analyze whether it's due to an outdated driver version or insufficient interface bandwidth. Mount success rate: Calculated as "number of successful mounts / total number of loops × 100%". For example, 98 successful mounts out of 100 loops represent a success rate of 98%. For typical scenarios, a success rate ≥95% is required; for server / industrial scenarios, ≥99.9% is needed. If the success rate is below the threshold, file system compatibility and driver stability issues should be thoroughly investigated.

[0102] Analysis of abnormal interruption causes: The script categorizes and summarizes the "abnormal records" in the logs, and combines system logs and hardware feedback to locate common abnormal causes: Hardware level: poor interface contact (hardware module reports "weak connection signal"), unstable SSD power supply (programmable power supply records voltage fluctuations); Software level: driver conflict (Windows event log shows "driver loading failed"), file system error (Linuxdmesg shows "superblock corrupted"), insufficient system resources (CPU / memory usage remains ≥90% during loop, causing identification timeout); Operation level: not completely and safely removed (script records "still read / write processes" but forced plugging and unplugging causes data cache not to be synchronized).

[0103] If the average recognition latency is ≤5 seconds, the mount success rate is ≥95%, and there are no data corruption records, then the SSD is considered to have "good stability" in frequent hot-swapping scenarios. If "data corruption" or "mount success rate <90%" occurs, further investigation is required: for data corruption, first check the SSD's cache power-loss protection function (whether there is a spare capacitor). If the mount success rate is low, test the compatibility of different interfaces (USB3.0 / USB4 / SATA) with the system version, and finally form a "scenario adaptation recommendation" (such as "this SSD is suitable for desktop hot-swapping, and is not recommended for industrial high-frequency hot-swapping scenarios").

[0104] Figure 6 yes Figure 1 The flowchart for step 5 is as follows: Figure 6 As shown, step S5 includes: Step 501, determining the target test based on concurrent test operations, power-on / off loop test mechanisms, and automatic identification and verification mechanisms, creating different test cases based on different solid-state drive models, and distributing the test cases to different test platforms connected to the local area network with one click; Step 502, all test platforms receive and automatically execute the test cases, and collect the running status, performance indicators, and abnormal logs of each node in real time; Step 503, during the test, each terminal periodically uploads encrypted text logs, screenshot evidence, and performance data to a designated directory on the central server.

[0105] Step 501 primarily integrates the target scope based on the previous testing mechanism, creates test cases adapted to multiple SSD models, and achieves automated distribution via the local area network to ensure the "targeting" and "efficiency" of the tests. The specific logic is as follows:

[0106] Based on the preliminary mechanism, the target test scope is determined as follows: The target test needs to integrate the core test scenarios of the first three stages to form a complete "test matrix" to avoid missing key verification points: Integrate "concurrency performance testing" (steps 201-204): Include scenarios such as lightweight database CRUD operations, file compression and decompression, and multi-task parallel load in the target test, and specify the test duration (e.g., 30 minutes per scenario) and load intensity (e.g., starting 5 sets of parallel data copy tasks simultaneously); Integrate "power on / off and hibernation loop testing" (steps 301-305): Include power on / off loops controlled by programmable power (e.g., executed 100 times) and hibernation and wake-up loops (e.g., executed 50 times), and specify the loop interval (e.g., 5 minutes for power on / off and 2 minutes for hibernation and wake-up); Integrate "automatic hot-plug identification verification" (steps 403-406): Include SSD safe removal and re-insertion loops (e.g., executed 200 times), and specify the insertion / removal interval (e.g., 1 minute interval each time) and verification dimensions (device identification time, whether the drive letter allocation is normal, and whether the file system is successfully mounted, etc.). The ultimate goal of testing is to create a "quantifiable test list" that marks each test item with its "mandatory / optional" attribute (e.g., high concurrency testing is mandatory for server scenarios, while hot-swap testing is optional for desktop scenarios) to ensure that it meets the needs of different application scenarios.

[0107] Create differentiated test cases based on SSD models: For SSD models of different brands, interface types (such as SATA, PCIe), and capacities (such as 2TB, 4TB), test case parameters need to be adjusted to avoid a "one-size-fits-all" approach that could lead to distorted test results: Interface adaptation: The IO performance limit of SATA interface SSDs is lower than that of PCIe interfaces, so the "task queue depth" in concurrent tests needs to be set differently (SATA interface is set to 16, PCIe interface is set to 64) to ensure that the load intensity matches the hardware capabilities; Capacity adaptation: The "preprocessing data volume" (step 401) for 2TB capacity SSDs is set to 4-10TB, and for 4TB capacity it is set to 8-20TB, both maintaining the industry standard of "2-5 times the nominal capacity" to avoid insufficient preprocessing due to capacity differences; Functional adaptation: Enterprise-grade SSDs that support "cache power loss protection" need to add "sudden power failure data protection verification" (such as checking whether cache data is completely written to the storage chip after a power failure) in hot-swapping or power-on / off tests, while ordinary consumer-grade SSDs focus on basic stability verification. All test cases are stored in the form of "standardized script template + parameter configuration document". The parameter configuration document clearly specifies "SSD model, interface type, test item parameters, and expected pass threshold" to facilitate quick modification and distribution later.

[0108] One-click distribution of test cases within a local area network (LAN) is achieved using a "LAN test management platform" (or a lightweight tool suite) to automate test case distribution, ensuring that multiple test platforms receive and prepare for execution simultaneously. The specific logic is as follows: Distribution platform setup: Deploy management tools on a central server within the LAN, pre-enter the "network information" (IP address, remote login authentication information, test file storage path) of all test platforms, and verify the network connectivity of each platform (through network connectivity testing and remote login testing); One-click distribution execution: The administrator selects the "target test cases" and "test platforms to be distributed" on the management platform (batch selection is possible, such as selecting 10 test machines equipped with different SSDs simultaneously), clicks "Distribute," and the test cases are distributed to the designated test platforms. The platform pushes test cases in the following ways: Lightweight scenarios: Test scripts and parameter configuration documents are directly copied to the "preset test directory" of each platform (Intel, AMD, and domestic platforms) via remote file transfer. After copying, each platform returns a "successful receipt" confirmation message to the management platform. Complex scenarios: Through the "task script" of the automated operation and maintenance tool or the test pipeline tool, test cases are not only distributed, but the "test environment readiness status" of each platform is also automatically checked (such as whether the performance testing tool is installed, whether the script running dependencies are complete, and whether the programmable power supply is online). If the environment is missing, the installation and configuration process is automatically executed (such as installing the performance testing tool through the package management tool of the Linux system and configuring the script dependencies with the help of the package manager of Windows) to ensure that the test can be started after distribution. Distribution status monitoring: The management platform displays the distribution progress of each test platform in real time ("Pending distribution - Distributing in progress - Completed - Failed"). If the distribution of a platform fails (such as network interruption), it will automatically retry twice. If it still fails, it will mark it as an "abnormal platform" and prompt the administrator to investigate (such as checking whether the IP address has changed or whether the firewall is blocking the transmission).

[0109] Step 502 ensures that each test platform can complete the test without manual intervention and provides real-time feedback on its running status and key indicators. The specific logic is as follows: Automated execution triggering of test cases: After receiving the test cases, each test platform automatically triggers execution through a "preset startup mechanism" to avoid delays caused by manual operation; Timed triggering: After distribution, the management platform sends a "unified startup command" to each test platform to ensure that all platforms start testing at the same time, facilitating subsequent comparison of "performance differences within the same time period"; Event triggering: If a platform receives test cases late due to environment repair, it can be triggered separately through the "supplementary test command" of the management platform. Before execution, the script will automatically clear residual files from the previous test (such as temporary logs and verification files) to avoid affecting the current results; Abnormal continuation: If a platform's test is paused due to "temporary network interruption" or "SSD not being recognized for a short time" during the test, the script will automatically record the "pause node" (such as "35th power-on / off cycle pause"). After the abnormality is recovered (such as network reconnection or SSD being recognized again), execution will continue from the pause node without starting from the beginning, improving testing efficiency.

[0110] Real-time data acquisition dimensions and methods: The "data acquisition scripts" on each test platform run synchronously with the test cases, collecting three types of core data according to the "high-frequency real-time" principle to ensure the traceability of the testing process and the localization of problems. All collected data is first stored in the "local temporary directory" on each test platform, with folders named according to "test case name-timestamp" to avoid data chaos. Running status data: Collected every 10 seconds, including "current test stage" (e.g., "preprocessing - concurrent testing - hot-plugging loop"), "test progress" (e.g., "20th loop / total 100 loops"), and "system resource usage" (CPU usage, memory usage, SSD disk usage). Data is obtained by calling system interfaces or tools to ensure timely detection of "test lag caused by system resource exhaustion"; Performance indicator data: Collected differently according to the test scenario, such as recording "random read operation response speed and average write latency" every 30 seconds in concurrent testing, recording "startup time and initialization duration" for each loop in power-on / off testing, and recording "plug-in time and initialization duration" for each hot-plugging test. The system records "identification latency and mount status," and all metrics are bound to "timestamps and cycle numbers" to ensure that the data can be mapped to specific test stages. Abnormal log data: Once an abnormality is detected (such as an SSD not being recognized, a file system error, or a test process interruption), an "abnormal snapshot" is immediately triggered—recording "abnormal occurrence time, current test step, and system error information" (such as disk error records in Windows event logs and file system abnormal information in Linux systems). Simultaneously, it automatically captures "test interface screenshots" (such as the interface where the SSD is not displayed in Device Manager or the interface where performance testing tools report errors), providing intuitive evidence for subsequent abnormality analysis.

[0111] Step 503: By periodically uploading encrypted data, the test data scattered across various platforms is aggregated to a central server, laying the foundation for subsequent analysis and report generation. The specific logic is as follows: Data encryption mechanism design:

[0112] To ensure data security during local area network transmission and storage (preventing test data from being tampered with or leaked, especially concerning performance parameters of enterprise-grade SSDs), a dual protection system of "transmission encryption + file encryption" is adopted: Transmission encryption: When each test platform uploads data to the central server, a secure channel is established through an encrypted transmission protocol. The server is configured with a "trusted certificate," and the client verifies the certificate's legitimacy before transmission, preventing "man-in-the-middle attacks" that could lead to data interception; File encryption (optional, for sensitive data): If the test data contains sensitive information such as "unpublished SSD firmware versions or custom test parameters," the local data file can be encrypted before uploading—using an industry-standard high-strength encryption algorithm. The key is uniformly generated by the central server and distributed to each test platform (transmitted offline or via encrypted email to avoid key transmission over the network). After uploading, the server decrypts and stores the data using the key, ensuring that even if the file is stolen, its content cannot be deciphered.

[0113] A "timed incremental upload" strategy is adopted to balance "real-time performance" and "network bandwidth consumption," avoiding local area network congestion caused by concentrated uploads. Upload interval configuration: Different intervals are set according to the test scenario, such as uploading every 5 minutes for performance testing (large data volume, frequent updates), and every 10 minutes for power-on / hot-swap testing (small data volume, cyclical updates), or uploading according to "test phases" (e.g., uploading once after preprocessing, uploading once after concurrent testing). Incremental upload implementation: Each test platform records the "last upload timestamp." Each upload only packages and uploads "data added after the timestamp" (such as newly added performance indicator logs, anomaly screenshots), instead of transmitting the entire dataset. For example, if there are 10 log files in the local temporary directory, and the first 3 were uploaded last time, only the last 7 will be uploaded this time to reduce duplicate transmissions. Upload status confirmation: After each upload is completed, the central server returns a "data received successfully" confirmation message (including the list of received files and data verification results) to the test platform. The test platform compares the local files with the verification results returned by the server. If they match, it marks "uploaded" and deletes the local temporary files (releasing storage space). If they do not match, it re-uploads to ensure that no data is lost.

[0114] The central server stores all uploaded data in a structured directory for easy retrieval, filtering, and analysis. The directory hierarchy is as follows: the root directory is categorized by test platform (e.g., "Test Platform 01", "Test Platform 02"); each platform directory is further divided by SSD model (e.g., "Samsung 870 EVO", "Western Digital SN850X"); within each model directory, subdirectories are created based on test case type (e.g., "Concurrency Test", "Power-On / Off Loop Test") and test time (e.g., "20251112_1000", representing startup at 10:00 AM on November 12, 2025). These subdirectories store performance indicator logs, system status data, and exception screenshot folders, respectively. Simultaneously, a "Data Indexing Service" is deployed on the server to automatically index uploaded data (including test platform, SSD model, test cases, timestamps, and key metrics), supporting quick data location via "Platform Filtering", "Model Comparison", and "Time Range Query" (e.g., "Query the hot-swap mounting success rate of Samsung 870 EVO across all platforms").

[0115] Step S5 further includes: Step 504, automatically parsing text logs using hashing and archiving them by test number and timestamp; Step 505, using Python scripts to aggregate and analyze multi-source logs, extracting key event timelines, error codes, and response delays, and determining the pass / fail status of each test based on preset thresholds; Step 506, combining historical comparisons with hardware configurations and firmware versions recorded in the database to identify performance anomalies or degradation, and automatically generating a unified test report from the test results, wherein the test report includes at least a pass rate summary, performance comparison charts, error details, and a comprehensive PDF report of the test environment.

[0116] Step 504: Hash Mapping for Text Log Parsing and Structured Archiving. Automatic parsing of text logs is achieved through "hash mapping," followed by archiving according to the "test number + timestamp" rule. This ensures that log data transforms from "unordered text" into "ordered and searchable structured data," laying the foundation for subsequent analysis. The specific logic is as follows: Automatic Log Parsing Driven by Hash Mapping: Here, "hash parsing" does not calculate the file's hash value. Instead, it establishes a correspondence between "log features and parsing rules" through a "pre-set hash mapping table," enabling automatic identification and field extraction of different types of text logs, avoiding the inefficiency of manual line-by-line parsing. Hash Mapping Table Design: A "feature-rule" mapping is pre-established in the central server's database for each type of test log (such as concurrent performance logs, power-on / off cycle logs, and hot-plug verification logs). For example, the feature keywords of "concurrent performance logs" (such as "random read IOPS" and "average write latency") are used as hash keys, and the corresponding values ​​are "field extraction rules" (such as "extract the numeric portion from 'random read IOPS: XXX' as IOPS"). The value is "extracted from 'Write Delay: Xms' as the delay value in milliseconds"; for system error logs, the error code (such as "0x80070002" or "ext4 superblock error") is used as the hash key, and the corresponding value is "error type classification" (such as "file path does not exist" or "file system corruption"). Automated parsing execution: The central server periodically scans the text log directory centrally stored in the 503 step, and for each log file:

[0117] The first step is to read the "test identifier field" at the beginning of the log (such as "test number: TC-20251112-001" and "test type: hot-plug loop"), match it with the "log type characteristics" in the hash mapping table, and determine the parsing rule corresponding to the log. The second step is to extract key structured fields according to the parsing rules. For example, extract "test number, loop number, identification latency, mount status, and verification result" from the hot-plug log, and extract "timestamp, IOPS value, latency value, and system resource usage" from the performance log. Redundant information in the log (such as duplicate system prompts and blank lines) is automatically filtered out. The third step is to mark "to be manually supplemented mapping" if unmatched log characteristics are encountered during the parsing process (such as new error codes or unknown log formats) and temporarily store the log in the "abnormal log directory" to avoid data loss due to parsing failure.

[0118] The parsed structured data must be archived according to the rules of "traceability and easy retrieval," echoing the centralized storage directory structure of 503 mentioned earlier. Specific strategies include: Archive directory inheritance: Under the "Platform-Model-Test Case-Test Time" directory level of 503, a new subdirectory "Parsed Logs" is added to store the parsed structured data (e.g., stored in Excel or CSV format for easy reading by subsequent analysis tools); Unified naming rules: The naming format for structured data files is "Test Number_Log Type_Timestamp.Format," for example, "TC-20251112-001_Hot-plug Log_202511121530.csv," where the "Test Number" matches the test case number in 501, and the "Timestamp" is accurate to the minute, ensuring that different logs for the same test item can be associated by number and time; Synchronous index updates: Each time a log is archived, the "Log Index Library" on the central server is automatically updated. The index information includes "Test Number, Platform IP, SSD..." The system allows users to select the model, log type, archive path, and parsing completion time. It also supports quick retrieval operations such as "searching all related logs by test number" and "filtering all test logs of the same SSD model" in the future.

[0119] By integrating parsed logs from different sources and of different types, key events and performance indicators are extracted. These are then combined with preset thresholds to determine test pass / fail status, transforming "data" into "clear conclusions." The specific logic is as follows: Multi-source log aggregation and association logic: Multi-source logs include the "performance indicator logs, system status logs, and exception logs" collected in step 502, as well as the structured data parsed in step 504. The core of aggregation is "establishing associations based on timestamps and test numbers," forming a "full-process data view for a single test item": Timestamp association: For logs from the same test platform and the same test case, "millisecond-level timestamps" are used as the link to associate data at the same time point in different types of logs. For example, at the time "2025-11-12 15:30:22.123," the concurrent performance log's "IOPS=3000" and the system status log's "CPU utilization = The "65%" and "no error" log entries can be aggregated into comprehensive "performance-resource-status" data at that moment. Test number association: All logs (preprocessing logs, concurrent test logs, hot-plug logs) under the same "test number" (e.g., TC-20251112-001) are aggregated to form the "full lifecycle data chain" of the test item. For example, a complete data sequence from "test start → preprocessing → concurrent load → power on / off cycle → hot-plug → test end" is generated, which makes it easy to trace the related events before and after a certain anomaly (e.g., "the IO latency suddenly increased to 20ms within 30 seconds before the 50th hot-plug failure").

[0120] Key Information Extraction and Anomaly Classification: From the aggregated multi-source data, three types of core information are extracted to provide a basis for result judgment: Key Event Timing: "Key node times" in the entire testing process are extracted, such as "test start time, preprocessing completion time, first anomaly occurrence time, and test end time," as well as the "time consumption" of each testing stage (e.g., preprocessing takes 2 hours, concurrent testing takes 1 hour), forming an "event timing axis" to intuitively present the test progress and anomaly occurrence nodes; Error Code and Problem Classification: Error codes and phenomenon descriptions from all anomaly logs are summarized, and combined with the hash mapping table of the 504 steps, errors are classified into "hardware" categories (e.g., SSD recognition failure, interface power supply failure). The test categorizes errors into three types: "Abnormal," "Software-related" (e.g., driver loading failure, file system errors), and "Environment-related" (e.g., network interruption, power fluctuations). It also counts the frequency and percentage of each type of error (e.g., "Hardware errors occurred 3 times, accounting for 20% of the total errors"). For response latency and performance fluctuations, the test extracts the "core performance indicators" for each test stage, such as the "average random read / write IOPS and 99th percentile latency" for concurrent tests, the "average startup time and initialization duration fluctuation range" for power-on / off tests, and the "average recognition latency and mount success rate" for hot-plug tests. It then marks any fluctuations that exceed the normal range (e.g., "The average recognition latency of a certain batch of SSDs is 10 seconds, higher than the preset threshold of 5 seconds").

[0121] The test pass / fail judgment rules based on preset thresholds are consistent with the "expected pass thresholds for test cases" in step 501 to ensure standard uniformity. Specifically, the following procedures are followed: Single test item judgment: For each test (such as concurrent performance testing or hot-plug testing), the "multi-dimensional threshold comprehensive judgment" principle applies. For example, the hot-plug test must meet the following criteria: "average recognition latency ≤ 5 seconds + mount success rate ≥ 95% + data verification pass rate 100%". If all three criteria are met, the test item is judged "passed". If any criterion is not met, it is marked "failed", and the dimension of non-compliance is noted (e.g., "mount success rate 92%, below the 95% threshold"). Overall test judgment: The overall test results for a specific SSD model. The overall pass rate is determined by "all mandatory test items pass + optional test item pass rate ≥ 80%". If any mandatory test item (such as high concurrency test in server scenarios) fails, the overall pass rate is directly determined as "unqualified" and the problems of the mandatory test items are marked first (e.g., "concurrent IOPS is less than 30% of the nominal value, mandatory test item fails"). Anomaly marking and priority: For test items judged as "unqualified", priority is marked according to the "impact degree" (high / medium / low): "data corruption" and "unable to recognize" that affect the core function are marked as "high priority", and "slightly high recognition latency but does not affect use" are marked as "low priority", which facilitates the priority ranking of subsequent problem investigation.

[0122] Step 506: By comparing historical data, performance anomalies or degradation are identified, ultimately generating a PDF report containing key information to provide direct evidence for SSD quality assessment and decision-making. The specific logic is as follows: Historical comparison of hardware configuration and firmware version: The central server needs to establish a database linking "SSD hardware configuration - firmware version - historical test data" in advance, storing test records of different models, firmware versions, and batches of SSDs. The comparative analysis focuses on two core scenarios: Comparison of different firmware versions of the same model: For example, comparing test data of a certain SSD model with "firmware version V1.0" and "V2.0". If V2.0 shows "random write IOPS increased by 20%+ and error rate decreased by 50%", then "firmware upgrade optimization is effective"; if V2.0 shows "mount success rate decreased from 98% to 90%", then "firmware version V2.0 has stability issues"; Comparison of different batches with the same firmware version: Comparing different batches of the same model and firmware version. For SSDs produced in batches, if a batch's "average write latency is 15% higher than the historical average and the media wear indicator drops twice as fast," it is determined that "this batch may have performance degradation or hardware quality fluctuations," requiring further investigation into differences in manufacturing processes or components. Cross-model horizontal comparison: For SSDs of the same type from different brands (such as 2TB consumer-grade SSDs with SATA interfaces), compare "core performance indicators" (such as IOPS, lifespan expectation) and "stability performance" (such as abnormal error rate, power-on / off cycle pass rate) to form a "model comparison matrix," providing data support for model selection (e.g., "Brand A's IOPS is 10% higher than Brand B's, but its abnormal error rate is twice that of Brand B").

[0123] Identification of Performance Anomalies and Degradation: Anomaly identification uses a dual approach of "data fluctuation threshold" and "trend analysis" to ensure accuracy: Single Test Anomaly Identification: If the test data of a certain SSD deviates from the "historical average data of the same model" by more than a preset fluctuation range (e.g., ±10%), it is marked as "single performance anomaly". For example, "the random read IOPS of a certain SSD is 2800, while the historical average of the same model is 3500, a deviation of 20%, which is considered anomaly"; Long-Term Performance Degradation Identification: For multiple test records of the same SSD (or SSDs of the same batch) (e.g., two tests with a 3-month interval), if the core indicators show a "continuous decline", it is judged as "performance degradation". For example, "the TBW lifetime of a certain SSD was 600TB in the first test, and 550TB in the retest after 3 months, a degradation of 8.3%, and the media wear indicator dropped from 98% to 92%", it is necessary to assess whether the degradation rate is within the normal range (e.g., annual degradation ≤5% is normal); Preliminary Root Cause Locator: Based on historical comparison results, the root cause of the anomaly is preliminarily determined, such as "all SSDs of the same batch show IOPS..." "The data loss is due to the same firmware version, suggesting a common issue with that firmware version"; "A single disk experienced data corruption, while others in the same batch were normal, suggesting an individual hardware failure of that disk".

[0124] Unified PDF test report generation containing core information: The report must adhere to the principles of "clear structure, comprehensive information, and highlighting key points," and must cover at least the following content, outputting in PDF format (supporting watermarks and signatures to ensure authority): Basic report information: Cover (test project name, report generation time, testing organization), table of contents, test overview (test purpose, test scope, test cycle), test environment description (test platform configuration, network environment, test tools and standards used), ensuring readers can quickly understand the test background; Pass rate summary: Divided into "single test item pass rate" (e.g., "concurrency performance test pass rate 100%, hot-plug test pass rate 92%) and "overall pass rate" (e.g., "This test included 10 SSD models, 8 of which passed overall, with an overall pass rate of 80%)," and visually displayed using "pie charts / bar charts," highlighting the unqualified models and core issues;

[0125] Performance comparison charts: include "historical comparison charts of the same model" (such as a line chart of IOPS changes in different batches of a certain model), "horizontal comparison charts across models" (such as a bar chart of average identification latency of SSDs from different brands), and "key indicator fluctuation charts" (such as a box plot of latency fluctuation range of a certain test item). The charts should be labeled with "normal range threshold lines" to facilitate intuitive identification of anomalies.

[0126] Error Details Summary: All anomalies are listed by "error type" (such as hardware and software), including "error description, number of occurrences, affected models / batch, priority, and preliminary root cause speculation". For example, "5 hardware errors occurred, involving 2 batches, mainly SSD recognition failures, speculated to be related to poor interface contact, high priority";

[0127] Test conclusions and recommendations: Summarize the core findings of this test (such as "Significant optimization effect of firmware V2.0" and "Performance degradation risk exists in a certain batch"), provide recommendations for different scenarios (such as "Recommended for server scenarios to use brand A V2.0 firmware SSD" and "A certain batch of SSDs is not recommended for high-frequency hot-swappable scenarios"), and mark "Issues requiring further verification" (such as "The cause of data corruption of a certain model needs to be confirmed by hardware disassembly").

[0128] A second embodiment is proposed based on the first embodiment. In the second embodiment,

[0129] Phase 1: Environmental Preparation and Basic Validation

[0130] 1. Deployment of multiple operating systems

[0131] Install Windows 10 on the Intel 12th Gen platform, Windows 11 on the Intel 13th Gen platform, and the Ubuntu LTS operating system on the AMD Ryzen 5000 series platform; install Tongxin UOS on the Kunpeng 920 platform and Qiqi OS on the Feiteng D2000 platform. Create a test task on the central platform and select the 'Basic Verification', 'Power Cycle' (200 times), and 'Performance Benchmark' suites.

[0132] 2. Automatic basic information acquisition

[0133] The platform distributes the task to the corresponding test nodes. Write a Python script that runs automatically after the system starts, collects the following information, and generates a standardized log file:

[0134] smartctl -a / dev / nvme0n1 | grep -E 'SMART Health Status|Temperature|Power On Hours'>smart_info.log

[0135] lsblk -f>disk_info.log

[0136] cat / proc / diskstats>io_stats.log

[0137] wmic diskdrive get Model, SerialNumber, FirmwareVersion>os_info.log

[0138] Phase 2: Basic and advanced function verification

[0139] 1. Data integrity stress test

[0140] Install the MySQL 8.0 database and execute the following SQL operations:

[0141] insert into t1 (id, name, age) values (1, 'Zhang San', 18);

[0142] update t1 set age = 25 where id = 1;

[0143] delete from t1 where id = 1;

[0144] select from t1 where name = 'Zhang San';

[0145] commit;

[0146] 7z a test_data.7z -tzip -mmt -mx9 / home / test_data /

[0147] md5sum test_data.7z>hash_value.log

[0148] 2. Multi-user / Multi-task Concurrency Testing

[0149] Execute the following tasks simultaneously through a Python script:

[0150] python -m pip install pywin32

[0151] start / b notepad++.exe

[0152] start / b chrome.exe

[0153] start / b vlc.exe

[0154] iostat -x | find "nvme0n1">io_delay.log

[0155] The third stage: In-depth testing of power and status management, the Agent program on the node is executed in sequence

[0156] 1. Automated Power Cycle Testing

[0157] Use Python + pyautogui library to achieve full automation:

[0158] pyautogui.hotkey('ctrl', 'alt', 'del')

[0159] pyautogui.press('y')

[0160] pyautogui.write('password')

[0161] pyautogui.press('enter')

[0162] pyautogui.sleep(60)

[0163] pyautogui.hotkey('ctrl', 'hift', 'q')

[0164] pyautogui.press('enter')

[0165] pyautogui.write('password')

[0166] pyautogui.press('enter')

[0167] Run the program 500 times in a loop, recording the startup time and success rate each time.

[0168] 2. S3 / S4 State Depth Testing

[0169] Before entering hibernation, the Python script creates a file named "hibernation flag.txt";

[0170] Upon activation, the script runs automatically, checking if the file exists and if its contents are correct.

[0171] if [ -f "hibernation flag.txt" ]&&[ $(cat "hibernation flag.txt" | grep -E 'hibernation timestamp|hibernation checksum' ) == 'hibernation timestamp 2022-06-01 14:30:00|hibernation checksum 1234567890' ];then

[0172] echo "Wake-up successful"

[0173] else

[0174] echo "Wake-up failed"

[0175] fi

[0176] Record wake-up time and success rate.

[0177] Phase 4: Performance and Reliability Benchmark Testing: After all tests are completed, the platform automatically generates a report showing a 100% power cycle success rate, an average startup time of 15 seconds, and 98% of the nominal IOPS.

[0178] 1. Professional benchmarking kit

[0179] Use the FIO tool to preprocess according to the SNIA standard, and then test:

[0180] fio --filename=ssd_test_file --bs=4k --iodepth=64 --runtime=60 --group_reporting --name="random read / write IOPS"

[0181] fio --filename=ssd_test_file --bs=4k --iodepth=64 --runtime=60 --group_reporting --name="Sequential read / write bandwidth"

[0182] fio --filename=ssd_test_file --bs=4k --iodepth=64 --runtime=60 --group_reporting --name="random read latency"

[0183] 2. Hot-swap test

[0184] While the system is running, use a Python script to safely remove and reinsert an SSD that is being used as a data disk:

[0185] import shutil

[0186] shutil.move(" / mnt / data", " / mnt / temp")

[0187] shutil.move(" / mnt / bak", " / mnt / data")

[0188] os.system("echo 1> / sys / class / scsi_disk / 0_0:0_1 / delete")

[0189] os.system("echo 1> / sys / class / scsi_disk / 0_0:0_1 / eject")

[0190] os.system("echo 0> / sys / class / scsi_disk / 0_0:0_1 / delete")

[0191] A third embodiment is proposed based on the first embodiment. In the third embodiment,

[0192] This invention's testing system applies to a complete quality assessment process for an enterprise-grade NVMe SSD (hereinafter referred to as "the SSD under test") on a typical x86 platform. This embodiment will follow the five stages described above, unfolding step by step.

[0193] 1. Test environment configuration

[0194] Hardware configuration:

[0195] Test platform: A standard x86 test platform equipped with an Intel® Core™ i7-13700K processor and a Z790 chipset motherboard.

[0196] Device under test: SSD (1TB capacity) under test, installed in the M.2 slot of the motherboard, used as the system drive.

[0197] Auxiliary equipment: A programmable smart power strip that supports remote network control (such as a PDU based on SNMP or HTTPAPI), to which the power cord of the test host is connected.

[0198] Network environment: All devices are located on the same local area network to ensure reliable communication.

[0199] Software configuration:

[0200] Operating System: Install a clean Microsoft Windows 11 Professional operating system on the SSD under test.

[0201] Test toolset: All software required for testing is pre-installed, including:

[0202] System information tool: smartmontools (used to obtain SMART data).

[0203] Database simulation tool: MySQL Community Server 8.0.

[0204] Compression and verification tools: 7-Zip command-line version, CertUtil (built into Windows).

[0205] Performance benchmarking tool: FIO (Flexible I / O Tester).

[0206] Application simulation tools: Microsoft Office suite, Google Chrome browser.

[0207] Automated control scripts: A complete set of test control and verification scripts written in Python.

[0208] Central control platform: Jenkins is used as the core of continuous integration / continuous deployment (CI / CD), and corresponding test tasks and pipelines are configured.

[0209] 2. Detailed description of the test execution process: Phase 1: Multi-platform environment preparation and automated basic verification Task trigger: The test engineer starts a task named "SSD-X_Win11_Comprehensive Test" on the Jenkins central control platform.

[0210] Environment self-check: After receiving the task, the Jenkins Agent on the test platform first executes the environment self-check script. This script checks whether all necessary testing tools are ready, whether the network connection is normal, and whether the SSD under test is correctly identified.

[0211] Automated acquisition and recording of basic information:

[0212] The script automatically executes a series of commands to collect complete information about the SSD under test. For example:

[0213] Use the command `smartctl -x / dev / nvme0n1` to obtain all NVMe log information, including SMART health status, temperature, power-on time, power-on count, total read / write volume, media wear indicator, etc.

[0214] Use WMI commands to obtain the disk model, serial number, and firmware version that are recognized at the operating system level.

[0215] All raw data, along with timestamps, is saved as a text-formatted log file.

[0216] Data validation: The script parses SMART data and verifies that key attributes (such as whether the initial value of the "Media Wear Indicator" is 100) are within the expected range, serving as a benchmark for subsequent durability testing. All log files are automatically uploaded to the Jenkins server for archiving.

[0217] Phase Two: In-depth functional and stability testing based on real-world scenarios

[0218] Data integrity stress testing:

[0219] Database transaction load testing:

[0220] Steps: The script automatically installs and starts the MySQL service. A test database and table structure are created. Then, an automated, 30-minute mixed database load test is performed, simulating a typical OLTP (Online Transaction Processing) scenario, including high-frequency INSERT, UPDATE, DELETE, and SELECT operations, while ensuring the ACID properties of transactions.

[0221] Verification: Throughout the process, the script monitors the MySQL error log. After the load is complete, the script performs a final data consistency check, such as verifying the total number of records in the table and the checksum of specific fields to ensure that no data corruption or transaction failures have occurred.

[0222] Mixed read / write and data verification test:

[0223] Steps: The script uses the 7-Zip command-line tool to repeatedly perform extreme compression (-mx=9) and decompression operations on a mixed set of files (containing documents, images, small files, etc.) of about 10GB, and executes the operation 10 times in a loop.

[0224] Verification: Before and after each compression and decompression, the script uses CertUtil to calculate and compare the MD5 hash values ​​of the source and destination file directories. The hash value comparison results for all loops must be completely consistent.

[0225] Multi-task concurrency stability test:

[0226] Steps: The automated script will simultaneously start the following four task threads and run them continuously for 30 minutes:

[0227] Thread 1: Continuously copy and delete large files (such as an ISO image) in the background.

[0228] Thread 2: Control Microsoft Word to automatically open, edit, and save documents.

[0229] Thread 3: Control the Chrome browser to automatically and repeatedly access a set of predetermined web pages.

[0230] Thread 4: Run system performance monitoring in real time, using the typeperf command at 1-second intervals to capture performance counters related to the SSD under test (such as "\LogicalDisk(C:)\Avg. Disk sec / Read" average disk read latency).

[0231] Verification: The script analyzes the collected performance counter data to calculate the average latency and the peak latency (99th percentile). The criteria for passing the test are: no blue screens, crashes, or application unresponsiveness occur, and the average I / O latency remains below an acceptable threshold (e.g., less than 50 milliseconds).

[0232] Phase 3: Reliability Enhancement Testing of Power Supply and State Management

[0233] Automated power supply cycle test (500 cycles):

[0234] Single loop process:

[0235] Step (soft shutdown): The Python script calls the os.system('shutdown / s / t 0') command to shut down the operating system normally.

[0236] Steps (Power Off and Power On): The script sends a command to the programmable smart socket's API interface via an HTTP POST request to cut off the power to the test host. After waiting 15 seconds, it sends the command again to restore power.

[0237] Steps (Startup and Heartbeat Detection): Utilizing the motherboard's "Power-On Auto-Start" function, the host was tested to power on automatically. During system startup, a Python script was configured as a startup service. This script runs automatically after the system enters the desktop and sends a "heartbeat" signal to the Jenkins server, reporting a successful startup.

[0238] Monitoring and Judgment: The Jenkins platform records the timestamp of each loop, whether the startup was successful, and the time from power-on to receiving the heartbeat signal (i.e., startup time). After the test, the platform automatically generates a statistical report, requiring a 100% success rate for 500 loops, and analyzes the stability and trend of startup time.

[0239] Hibernation depth test (100 cycles):

[0240] Single loop process:

[0241] Steps (Pre-hibernation state settings): Before the system enters hibernation, the script creates a temporary file named hibernate_marker.txt in the root directory of drive C and writes the current system timestamp and a randomly generated unique verification code (UUID) into it.

[0242] Steps (Entering and Waking from Hibernation): The script calls the `shutdown / h` command to put the system into hibernation (S4) mode. After waiting 30 seconds, a brief "power outage and power on" operation is simulated via the smart plug, triggering the system to resume from hibernation.

[0243] Steps (Post-Wake-up Status Verification): After the system wakes up, another self-starting script runs automatically. This script first checks whether the system is logged in normally, then locates and reads the hibernate_marker.txt file.

[0244] Verification: The script verifies the existence of the file and checks whether the timestamp and checksum in the file are completely consistent with the values ​​written before entering hibernation. Simultaneously, it records the total time (i.e., wake-up time) required from power-on to the system being fully ready to execute the verification script.

[0245] Final determination: After 100 cycles, the wake-up success rate and data verification accuracy rate must both be 100%. Simultaneously, the distribution of wake-up times will be analyzed to check for any abnormal delays.

[0246] Phase 4: Enterprise-level performance and specialized reliability benchmark testing

[0247] Steady-state performance benchmark test:

[0248] Steps: Strictly follow the SNIA (Storage Networking Industry Association) solid-state drive performance testing specifications and use the FIO tool.

[0249] Preprocessing: Perform continuous and sufficient random writes on the SSD under test until its performance reaches a stable state.

[0250] Formal testing: Under steady-state conditions, random read / write IOPS were tested with block sizes of 4 KiB and queue depths ranging from 1 to 256, as well as sequential read / write bandwidth with block sizes of 128 KiB. Each test run lasted at least 5 minutes, and the average value was taken.

[0251] Record: Record all raw performance data and compare and analyze them with the manufacturer's advertised values.

[0252] Hot-swap durability test (for data disk role):

[0253] Steps: In this test, the SSD under test is used as a slave disk (data disk). The script executes the following process, looping 20 times:

[0254] In the operating system's disk management, a script command is used to take the SSD under test "offline".

[0255] Physically and safely remove and reinsert the SSD.

[0256] The script detects new hardware, brings the disk "online," and assigns a drive letter.

[0257] The script verifies whether a previously written checksum file to the disk is readable and contains correct information.

[0258] Verification: Count the number of times the system can correctly identify the device, assign drive letters, and ensure data integrity in 20 loops.

[0259] Phase 5: Integrated Intelligent Analysis Platform and Report Generation

[0260] Data aggregation: Logs, performance data, screenshots, and judgment results generated in all the aforementioned stages are automatically collected by JenkinsAgent and uploaded to the central server, where they are categorized and stored according to test cases and timestamps.

[0261] Intelligent Analysis: Automated Judgment: The Jenkins pipeline integrates analysis scripts that automatically determine the results of each test item. For example, it compares the power cycle startup time with a threshold (e.g., 30 seconds) and the performance test IOPS with the nominal value (e.g., minimum 95%).

[0262] Trend Analysis: The platform compares the FIO performance data from this test with historical test data of the same model SSD, and uses charts to visually display the consistency or deviation in performance.

[0263] One-click report generation: After all tests are completed, Jenkins invokes the report generation plugin to automatically create a detailed test report. The report includes: a test overview, environment configuration, pass / fail status for each test phase, key performance data charts, detailed error information for failed test cases, and overall conclusions. The report is provided in PDF and HTML formats for engineers to view and analyze immediately.

[0264] Summary of Implementation Results: Through the full-process testing described in this embodiment, a comprehensive and objective evaluation of the SSD under test was conducted, ranging from basic compatibility to in-depth reliability. The entire testing process is highly automated, greatly reducing manual intervention and ensuring the consistency and repeatability of the tests. The generated intelligent report not only provides a "pass / fail" conclusion, but more importantly, it provides rich quantitative data, offering solid data support for the selection, acceptance, and quality traceability of enterprise-level SSDs. This perfectly demonstrates the comprehensiveness, depth, and efficiency advantages of this invention compared to traditional testing methods.

[0265] This embodiment expands the testing platform from the traditional Intel / AMD x86 architecture to include domestic ARM architecture platforms such as Kunpeng and Phytium, and covers operating systems such as Windows and domestic OSes (e.g., UnionTech UOS, Kylin OS). This invention completely solves the problem of narrow platform coverage in existing testing methods. This allows SSD compatibility test results to truly reflect their adaptability in the current complex and diverse computing ecosystem, providing reliable quality credentials for products entering a broader market, especially the domestic IT innovation market, and greatly improving the versatility and forward-looking nature of the testing solution. By utilizing scripts and automation frameworks, repetitive manual tests are automated, achieving standardization and consistency in the testing process, avoiding the uncertainty caused by human operation, significantly improving testing efficiency, and reducing labor costs. By simulating real user and enterprise-level loads (such as database operations, compression / decompression, multi-user concurrency, etc.), in-depth stability, data consistency, and performance testing is conducted, breaking through the limitations of traditional testing methods that only focus on a single function. This comprehensively evaluates the performance of SSDs in real-world application scenarios, improving the practicality and accuracy of the testing. A central testing platform was built to automatically collect and analyze test data and generate intelligent reports. This automated processing and standardized presentation of test results avoided errors caused by manual intervention, improved the timeliness and accuracy of testing, and provided a more comprehensive and reliable reference for enterprise-level SSD applications. By introducing enterprise-level application scenario simulation, in-depth reliability testing, and an intelligent data analysis platform, a set of forward-looking SSD quality assessment standards was formed, providing a more systematic, comprehensive, and intelligent solution for the R&D, testing, and maintenance of enterprise-level solid-state drives.

[0266] In the embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.

[0267] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0268] The above are merely embodiments of the present invention and do not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A solid-state drive (SSD) compatibility testing method, characterized in that, The method includes: After the solid-state drive (SSD) boots up, it automatically collects basic performance information about the SSD and generates standardized log files based on a Python script. The system performs solid-state drive performance tests within a preset time period, using automated scripts to execute concurrent tests involving multiple users or tasks. The script calls the system's boot or shutdown commands, and automatically triggers the next boot process after shutdown, forming a boot-up loop test mechanism. During each boot-up process, the startup time, initialization time, and abnormal interruption records of the solid-state drive are collected synchronously. During system operation, a Python script is used to safely remove and reinsert the solid-state drive. After reinsertion, the Python script automatically checks the system's recorded device recognition latency, drive letter normality rate, file system mount success rate, and data consistency verification results, forming an automatic recognition and verification mechanism after insertion. The steps involved in safely removing and re-inserting the solid-state drive (SSD) using a Python script during system operation, followed by an automatic post-insertion identification and verification mechanism that uses the Python script to check the system's recorded device recognition latency, drive letter correctness rate, file system mount success rate, and data consistency verification results: While the system is running, use a Python script to safely remove and reinsert the SSD; The system automatically checks and records device recognition latency, drive letter normality rate, file system mount success rate, and data consistency verification results through Python scripts, forming an automatic recognition and verification mechanism after insertion. The script automates the safe removal and re-insertion operation in a loop, recording the latency and mount status of each solid-state drive re-identification. The average recognition latency and mounting success rate after multiple operations were statistically analyzed, and the causes of abnormal interruptions were analyzed in combination with the logs recorded in the system to verify the stability and data integrity guarantee capabilities of the solid-state drive in frequent hot-swapping scenarios. The target test is determined based on concurrent test operations, power-on / off loop test mechanisms, and automatic identification and verification mechanisms. Different test cases are created for different solid-state drive models, and the test cases are distributed to different test platforms connected in the local area network with one click. During the test process, all test nodes automatically upload logs to the central server.

2. The solid-state drive compatibility testing method according to claim 1, characterized in that, After the solid-state drive (SSD) boots up, the steps for automatically collecting basic performance information of the SSD and generating standardized log files based on a Python script include: After the solid-state drive (SSD) boots up, it automatically collects basic performance information about the SSD and generates standardized log files based on a Python script. Based on all SMART information output by the smartctl command, verify the initial values ​​of specific key SMART information, disk drive details in Device Manager, and disk capacity, serial number, and firmware version recognized by the operating system.

3. The solid-state drive compatibility testing method according to claim 1, characterized in that, The steps of performing solid-state drive performance operations within a preset time period, and simultaneously executing multi-user or multi-task concurrent test operations based on automated scripts, include: Install a lightweight database and perform CRUD operations and transaction processing to verify the data consistency when the processed solid-state drive is used as the system disk and data disk. To test the mixed read and write capabilities of the solid-state drive, a large set of files was repeatedly compressed and decompressed using compression / decompression tools. Before and after copying the file, calculate the hash value using the same algorithm. If the two output hash values ​​are exactly the same, the file copy is correct; if they are not the same, it means the file is corrupted or the copying process is faulty, and the file needs to be copied again.

4. The solid-state drive compatibility testing method according to claim 3, characterized in that, The method involves calculating the hash value using the same algorithm before and after copying the file. If the two output hash values ​​are completely identical, then the file copy is error-free. If there is a discrepancy, it indicates that the file is corrupted or the copying process has failed, requiring a recopy. Following this step, the method further includes: The logic involves starting a load through an automated script, running multi-user or multi-task concurrent test operations in the background, continuously monitoring system metrics and SSD IO latency, and automatically terminating all test processes after a preset time. The multi-user or multi-task concurrent test operations include at least background large file copying, foreground office software running, and simultaneous web browsing and video playback.

5. The solid-state drive compatibility testing method according to claim 1, characterized in that, The method of using a script to call the system boot or shutdown command, and automatically triggering the next boot process after shutdown, forms a boot-up / shutdown loop test mechanism. The steps of synchronously collecting the SSD's boot time, initialization duration, and abnormal interruption records during each boot-up / shutdown process include: Use scripts to call system power-on or power-off commands, and send power-on / off instructions to programmable power sockets via network protocols; The programmable power supply is cut off / on according to the acquired power-on / off command, simulating power disconnection / non-power disconnection. After power failure, the system automatically records the current state, and after power is restored, the script triggers a self-test process to verify the solid-state drive mounting status and restore the running state before power failure. The power-on / off cycle test mechanism is repeated a preset number of times. During each power-on and power-off process, the startup time, initialization duration, and abnormal interruption records of the solid-state drive are collected synchronously. File system errors, disk recognition failures, and log error information after abnormal restarts are statistically analyzed to evaluate the stability and data retention capabilities of the solid-state drive under extreme power supply environments.

6. The solid-state drive compatibility testing method according to claim 5, characterized in that, After the steps of synchronously collecting the SSD's boot time, initialization time, and abnormal interruption records during each power-on and power-off process, statistically analyzing file system errors, disk recognition failures, and log error information after abnormal restarts, and evaluating the SSD's stability and data retention capabilities under extreme power supply environments, the method further includes: Before entering hibernation / sleep mode, a temporary file with a timestamp and a unique checksum is created via a script. When the system wakes up, the script automatically runs and reads the temporary file, checks whether the temporary file exists and whether the checksum is correct, and records the time required to wake up. Continuously execute the power-on / off cycle and wake-up cycle within a preset number of times, calculate the wake-up success rate and draw a wake-up time distribution chart, and check for any abnormal delays.

7. The solid-state drive compatibility testing method according to claim 1, characterized in that, Before the step of using a Python script to safely remove and reinsert the solid-state drive during system operation, and then automatically checking the system's recorded device recognition latency, drive letter normality rate, file system mount success rate, and data consistency verification results after reinsertion via the Python script to form an automatic recognition and verification mechanism after insertion, the method further includes: Using the FIO tool, the solid-state drive was preprocessed according to the standards of the Global Network Storage Industry Association, and the random read / write IOPS, sequential read / write bandwidth and latency at different queue depths under steady state were tested. Use the FIO tool to perform continuous random writes to the SSD, monitor the media wear indicator or host write volume growth in the SMART information, and make a preliminary assessment of the SSD's lifespan performance.

8. The solid-state drive compatibility testing method according to claim 1, characterized in that, The method involves determining the target test based on concurrent test operations, power-on / off loop test mechanisms, and automatic identification and verification mechanisms. Different test cases are created for different SSD models, and these test cases are distributed to different test platforms connected within the local area network with a single click. The step of automatically uploading logs from all test nodes during the test process to the central server includes: Based on concurrent test operations, power-on / off loop test mechanisms, and automatic identification and verification mechanisms, target tests are determined. Different test cases are created for different solid-state drive models, and the test cases are distributed to different test platforms connected to the local area network with one click. All testing platforms receive and automatically execute test cases, and collect the running status, performance indicators and exception logs of each node in real time; During the test, each terminal periodically uploaded encrypted text logs, screenshot evidence, and performance data to a designated directory on the central server.

9. The solid-state drive compatibility testing method according to claim 8, characterized in that, The process involves determining the target test based on concurrent test operations, power-on / off loop test mechanisms, and automatic identification and verification mechanisms. Different test cases are created for different SSD models, and these test cases are distributed to different test platforms connected within the local area network with a single click. The step of automatically uploading logs from all test nodes during the test process to the central server also includes: Text logs are automatically parsed using hashing and archived by test number and timestamp; The system uses Python scripts to aggregate and analyze multi-source logs, extracting key event timings, error codes, and response delays. It determines the pass / fail status of each test based on preset thresholds. By combining historical comparisons with hardware configurations and firmware versions recorded in the database, it identifies performance anomalies or degradations and automatically generates a unified test report. The test report includes at least a comprehensive PDF report containing a pass / fail summary, performance comparison charts, error details, and a description of the test environment.

Citation Information

Patent Citations

  • Method and device for testing webpage performances

    CN106681926A

  • Hard disk test method and system based on PCIE link configuration, terminal and storage medium

    CN114116337A