A solid state disk test index optimization method, device, equipment and storage medium
Patent Information
- Application Number
- CN202511263376.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2045-09-05
AI Technical Summary
此类方法虽在一定程度上缩短测试周期,但在测试可靠性层面存在显著局限,具体而言:敏感参数识别易受瞬时电压波动等偶发干扰产生“伪敏感参数”,导致测试资源错配与可靠性验证偏差;浴盆曲线应用局限于初始失效期,无法覆盖SSD全生命周期的可靠性验证;测试策略仅针对常温环境设计,未考虑工业/车载场景下高低温、湿度波动对参数敏感性的影响,导致实验室测试结果与实际应用环境脱节
[0015]本申请的有益效果是:首先将待测试产品划分为样本组与全量组,基于预设的第一测试参数集和多环境参数矩阵启动测试;
Smart Images

Figure CN121281612B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of solid-state drive (SSD) testing, and in particular to a method, apparatus, device, and storage medium for optimizing SSD testing metrics. Background Technology
[0002] Solid State Drives (SSDs) are critical data storage devices, and their reliability is paramount in enterprise, industrial, and automotive applications. To assess and ensure the long-term reliability of SSDs, the industry commonly uses Reliability Demonstration Tests (RDTs). These tests accelerate product aging by applying high-intensity workloads, thereby simulating long-term usage conditions, identifying potential failure modes, and inferring product lifespan.
[0003] Previously, Chinese patent CN117637004A provided a method for optimizing test indicators based on test result data, specifically proposing a scheme of "RDT testing to find sensitive parameters + bathtub curve to determine aging time," which accelerates anomaly identification by increasing the probability of sensitive parameter selection. While such methods shorten the test cycle to some extent, they have significant limitations in terms of test reliability. Specifically: sensitive parameter identification is easily affected by occasional interference such as instantaneous voltage fluctuations, resulting in "pseudo-sensitive parameters," leading to misallocation of test resources and deviations in reliability verification; the application of bathtub curves is limited to the initial failure period and cannot cover the reliability verification of the entire SSD lifecycle; the test strategy is only designed for normal temperature environments and does not consider the impact of high and low temperatures and humidity fluctuations on parameter sensitivity in industrial / automotive scenarios, resulting in a disconnect between laboratory test results and actual application environments. Summary of the Invention
[0004] To overcome the shortcomings of existing technologies, this application provides a method, apparatus, device, and storage medium for optimizing solid-state drive (SSD) test metrics, thereby achieving more precise sensitive parameters, full-cycle testing strategies, and diversified scenario coverage for SSD test metric optimization, significantly improving the reliability and accuracy of test results.
[0005] The technical solution adopted by this application to solve its technical problem is: Firstly, this application provides a method for optimizing solid-state drive (SSD) test metrics, the method comprising: Identify the products to be tested, set a portion of the products to be tested as a sample group, and obtain a preset first test parameter set and a multi-environment parameter matrix; Based on the first test parameter set and the multi-environment parameter matrix, the sample group is subjected to initial screening test and multiple rounds of verification test, sample test data is collected, sensitive parameters in the first test parameter set are identified and their sensitivity is graded, and the selection probability of the corresponding parameters is updated according to the grading results to obtain the second test parameter set; The sample test data is fitted using a bathtub curve model to obtain a bathtub curve, and the stage features of the initial failure period, accidental failure period and wear-out failure period of the bathtub curve are extracted as a stage feature set. An environment adaptation test strategy is generated based on the stage feature set and the environment parameter matrix. The second test parameter set and the environment adaptation test strategy are then combined to perform tests on all products to be tested, and full test results are obtained.
[0006] Optionally, the multi-round verification test includes RDT verification test and HALT verification test; The steps of performing initial screening tests and multiple rounds of verification tests on the sample group based on the first test parameter set and the multi-environment parameter matrix, collecting sample test data, identifying sensitive parameters in the first test parameter set, and performing sensitivity classification include: The initial screening test is performed on the sample group according to the first test parameter set to identify abnormal SMART data and obtain suspected sensitive parameters through initial screening. After changing the RDT testing strategy, the RDT verification test is performed on the sample group. If the suspected sensitive parameter still triggers the SMART anomaly, the corresponding suspected sensitive parameter is updated to a candidate sensitive parameter. The HALT validation test is performed on the sample group based on the multi-environment parameter matrix, and the sensitivity of the candidate sensitive parameters is graded according to the HALT validation results.
[0007] Optionally, the step of identifying sensitive parameters in the first test parameter set and classifying their sensitivity, and updating the selection probability of the corresponding parameters based on the classification results, includes: Based on the HALT verification results, the candidate sensitive parameters are divided into core sensitive parameters and secondary sensitive parameters; The selection probability of the core sensitive parameter is adjusted to a first selection probability, and the selection probability of the secondary sensitive parameter is adjusted to a second selection probability, wherein the second selection probability is less than the first selection probability; The adjusted selection probabilities are updated in the first test parameter set to form the second test parameter set.
[0008] Optionally, after the step of obtaining the full test results, the method further includes: The core sensitive parameters are verified based on the full test results. The false positive rate is calculated based on the number of misjudged parameters to obtain the false positive rate of the sensitive parameters. The deviation is calculated by comparing the full test results with the predicted value of the bathtub curve, thus obtaining the bathtub curve prediction deviation. The sensitive parameter verification criteria are updated based on the false positive rate of the sensitive parameters, and the model parameters of the bathtub curve model are updated based on the prediction deviation of the bathtub curve.
[0009] Optionally, the step of extracting the stage features of the initial failure period, accidental failure period, and wear-out failure period of the bathtub curve as a stage feature set includes: The bathtub curve is analyzed to extract the duration of the initial failure period, the mean time between failures (MTBF) of the random failure period, and the parameter degradation rate of the wear-out failure period. All extracted stage features are integrated into the stage feature set.
[0010] Optionally, the step of generating an environment adaptation test strategy based on the stage feature set and the environment parameter matrix includes: The multi-environment aging time is determined based on the duration of the initial failure period, the RDT test interval frequency is determined based on the mean time between failures of the accidental failure period, and the key test parameters are determined based on the parameter degradation rate of the wear-out failure period, and integrated to form a basic test strategy. Based on the multi-environment parameter matrix, the basic test strategy is adjusted for environmental adaptation to generate the environment-adapted test strategy.
[0011] Optionally, the step of performing tests on all the products under test according to the second test parameter set and the environment adaptation test strategy to obtain full test results includes: Based on the multi-environment aging time in the environment adaptation test strategy, perform multi-environment aging tests on all the products under test. After completing the multi-environment aging test, the product under test is sampled and tested according to the RDT test interval frequency in the environmental adaptation test strategy and the selection probability in the second test parameter set. Based on the key test parameters in the environmental adaptation test strategy, and the test intensity is determined based on the sensitivity level of the key test parameters, targeted lifetime tests are performed on the corresponding parameters and the failure threshold is recorded. All test data are integrated to form the full test results.
[0012] Secondly, this application provides a device for optimizing solid-state drive test metrics, comprising: The sample group and parameter determination module is used to determine the products to be tested, set a portion of the products to be tested as a sample group, and obtain a preset first test parameter set and a multi-environment parameter matrix. The parameter sensitivity grading module is used to perform initial screening tests and multiple rounds of verification tests on the sample group based on the first test parameter set and the multi-environment parameter matrix, collect sample test data, identify sensitive parameters in the first test parameter set and perform sensitivity grading, update the selection probability of the corresponding parameters according to the grading results, and obtain the second test parameter set. The stage feature extraction module is used to fit the sample test data to obtain the bathtub curve by using the bathtub curve model, and extract the stage features of the initial failure period, accidental failure period and wear-out failure period of the bathtub curve as the stage feature set. The environment adaptation full-scale testing module is used to generate an environment adaptation testing strategy based on the stage feature set and the environment parameter matrix, and to perform tests on all products to be tested in combination with the second test parameter set and the environment adaptation testing strategy to obtain full-scale test results.
[0013] Thirdly, this application provides an electronic device, comprising: One or more processors; One or more memory units; And one or more computer programs, wherein the one or more computer programs are stored in the one or more memories, and the one or more computer programs include instructions that, when executed by the one or more processors, cause the electronic device to perform the methods described above.
[0014] Fourthly, this application provides a computer-readable storage medium storing a program or instructions that, when executed, implement the above-described method.
[0015] The beneficial effects of this application are: firstly, the products to be tested are divided into sample groups and full groups, and the test is started based on the preset first test parameter set and multi-environment parameter matrix; In the sensitive parameter identification stage, parameters are graded through initial screening tests and multiple rounds of verification tests. The selection probability is then updated based on the sensitivity grade to form the second set of test parameters. In this step, cross-validation is used to eliminate interference from testing strategies and to remove "pseudo-sensitive parameters" caused by occasional factors such as instantaneous voltage fluctuations, thereby avoiding resource mismatch and reliability verification deviations. Subsequently, the sample test data was fitted using a bathtub curve model to fully extract the characteristics of the initial failure period, the accidental failure period, and the wear-out failure period, thus obtaining a stage feature set. Then, an environment-adaptive testing strategy was generated based on the stage feature set and the environmental parameter matrix. For the initial failure period, differentiated aging times were determined by combining multiple environmental parameter matrices to ensure full exposure of early defects. For the accidental failure period, the RDT test interval was optimized based on the mean time between failures (MTBF) to reduce resource consumption during stable operation. For the wear-out failure period, targeted lifetime test parameters were set according to the degradation rate of core parameters. This full-cycle modeling allows the testing strategy to cover the entire SSD lifecycle, improving RDT testing resource consumption and lifetime prediction accuracy.
[0016] Finally, testing was performed on all products using an environmental adaptation testing strategy and a second set of test parameters. Specifically, in multi-environment aging tests, core sensitive parameters were prioritized to trigger extreme condition tests in corresponding scenarios; sampling tests achieved tiered resource allocation based on differences in selection probabilities, ensuring high-frequency detection of core parameters; and directional lifetime tests dynamically adjusted intensity based on parameter sensitivity, thereby improving the fit between test results and actual applications in complex scenarios such as industrial and automotive environments, and solving the problem of traditional room temperature testing being out of sync with real-world environments. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating the method for optimizing solid-state drive test metrics provided in the first embodiment of this application; Figure 2 This is a flowchart illustrating the method for optimizing solid-state drive test metrics provided in the second embodiment of this application; Figure 3 This is a schematic diagram of the virtual structure of the solid-state drive test index optimization device provided in this application; Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0018] The present application will be further described below with reference to the accompanying drawings and embodiments.
[0019] The following will clearly and completely describe the concept, specific structure, and resulting technical effects of this application in conjunction with embodiments and accompanying drawings, so as to fully understand the purpose, features, and effects of this application. Obviously, the described embodiments are only a part of the embodiments of this application, not all of them. Other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are all within the scope of protection of this application. Furthermore, all connections / linkages involved in the patent do not simply refer to direct contact between components, but rather to the ability to form a better connection structure by adding or reducing connecting accessories according to specific implementation conditions. The various technical features in this application can be combined interactively without contradicting each other.
[0020] Reference Figure 1 , Figure 1 This is a flowchart illustrating the method for optimizing solid-state drive test metrics provided in the first embodiment of this application. It shows several key steps involved in the method for optimizing test metrics provided in this application, which are described in detail below: In step S1, the product to be tested is determined, a portion of the product to be tested is set as a sample group, and a preset first test parameter set and a multi-environment parameter matrix are obtained.
[0021] Among them, the sample group refers to a representative part of all products to be tested, usually accounting for 10%-20%, which is used to obtain key data through small-scale testing in order to infer the overall performance and reliability characteristics of the products and avoid the waste of resources caused by full-scale testing; The first test parameter set is a preset combination of initial test parameters, including basic test conditions such as voltage, frequency, and IO mode (such as random read / write ratio), which serve as the benchmark configuration for screening and verifying sensitive parameters.
[0022] It is worth noting that each test parameter has a preset selection probability as a parameter priority indicator, used to quantify the frequency with which each test parameter (such as voltage, erase count, and I / O mode) is selected in the initial test. For example, a selection probability of 60% for a parameter means that when a test scenario is randomly triggered, this parameter has a 60% probability of being included in the current test combination, reflecting an initial prediction of parameter sensitivity. The selection probability can be based on historical data or experience with similar products, with the "ECC error count" selection probability preset to be higher than the "idle power consumption".
[0023] Among them, the multi-environment parameter matrix is a parameter combination table designed for different application scenarios (such as normal temperature 25℃, high temperature 60℃, and low temperature -20℃), covering variables such as voltage fluctuation range, temperature threshold, and IO load intensity under various environments, forming a systematic multi-scenario test solution.
[0024] Specifically, the scope of SSD products to be tested is first defined, including but not limited to specific models or batches. From this scope, a sample group and a full-volume group are divided. The sample group must be statistically representative, covering different production batches and NAND flash memory chips. Then, a preset first set of test parameters is used as the initial test conditions, combined with a multi-environment parameter matrix to provide a standardized parameter framework for the sample group testing. In one embodiment, the initial test conditions could be a base voltage of 3.3V, a random read / write ratio of 7:3, and the multi-environment parameter matrix could be a voltage increase of 10% at high temperatures and an I / O mode adjustment to 5:5 at low temperatures. By testing the sample group under controllable parameters and multiple environments, preliminary screening data for sensitive parameters can be efficiently obtained, while avoiding high-cost testing of the entire product line. This lays the foundation for subsequent full-cycle modeling and full-volume testing strategy optimization.
[0025] In step S2, based on the first test parameter set and the multi-environment parameter matrix, the sample group is subjected to initial screening test and multiple rounds of verification test, sample test data is collected, sensitive parameters in the first test parameter set are identified and their sensitivity is graded, and the selection probability of the corresponding parameter is updated according to the grading result to obtain the second test parameter set.
[0026] The initial screening test refers to the first round of RDT testing performed on the sample group based on the first set of test parameters. It is a preliminary screening process that identifies potential sensitive parameters by monitoring SMART data (such as the number of ECC errors and erase counts). The aim is to quickly locate key parameters that may affect the reliability of SSDs. The multi-round verification test is a cross-validation and extreme environment test performed on the suspected sensitive parameters identified in the initial screening to eliminate occasional interference and confirm the sensitivity of the parameters. Sensitive parameters refer to test parameters that have a significant impact on the reliability or performance of SSDs. Abnormal changes in these parameters may indicate product failure risks, such as excessive ECC error counts or excessive voltage fluctuations. Sensitivity grading is based on the performance of parameters in multiple rounds of verification, dividing sensitive parameters into core sensitive parameters (confirmed to be stable and sensitive through extreme testing) and secondary sensitive parameters (sensitive only under specific conditions), providing a basis for subsequent resource allocation. The second test parameter set is an optimized parameter set formed by dynamically updating the selection probability based on the sensitivity classification results on the basis of the first test parameter set. The selection probability of core sensitive parameters is significantly improved, and the secondary parameters are moderately improved, so as to achieve precise allocation of test resources.
[0027] Specifically, firstly, based on the first set of test parameters and a multi-environment parameter matrix (scenarios such as normal temperature, high temperature, and low temperature), an initial screening test is performed on the sample group, collecting SMART (Self-Monitoring, Analysis, and Reporting Technology) data and information on abnormal scenarios, such as parameter anomalies during voltage fluctuations, to initially identify suspected sensitive parameters. Subsequently, multiple rounds of verification tests are conducted to eliminate interference and confirm the core and secondary sensitive parameters. Next, the selection probability is adjusted based on the grading results, ultimately forming the second set of test parameters. The entire process, through refined testing of the sample group, ensures the accuracy of parameter identification, laying the foundation for subsequent full-cycle modeling and full-scale testing.
[0028] More specifically, in the embodiments of this application, the test consists of two stages: RDT verification testing and HALT verification testing, which eliminate interference step by step and achieve precise parameter classification. RDT verification testing refers to a secondary reliability verification test performed on the sample group after changing the initial RDT testing strategy (such as adjusting the IO mode and IOPS pressure); HALT verification testing (high-accelerated life testing) is an enhanced verification of parameter sensitivity under extreme environments (such as voltage ±10% of rated value and high temperature 70°C).
[0029] The steps of performing initial screening tests and multiple rounds of verification tests on the sample group based on the first test parameter set and the multi-environment parameter matrix, collecting sample test data, identifying sensitive parameters in the first test parameter set, and performing sensitivity classification include: The initial screening test is performed on the sample group according to the first test parameter set to identify abnormal SMART data and obtain suspected sensitive parameters through initial screening. After changing the RDT testing strategy, the RDT verification test is performed on the sample group. If the suspected sensitive parameter still triggers the SMART anomaly, the corresponding suspected sensitive parameter is updated to a candidate sensitive parameter. Specifically, in the initial screening test, suspected sensitive parameters are identified based on the first set of test parameters. However, at this point, the parameter anomalies may be caused by occasional factors. To eliminate interference, the test strategy is changed during the RDT validation test phase. If the same parameter still triggers SMART data anomalies, it indicates that its sensitivity is not affected by the test strategy, and it is upgraded to a "candidate sensitive parameter," thus initially eliminating strategy-dependent interference.
[0030] The HALT validation test is performed on the sample group based on the multi-environment parameter matrix, and the sensitivity of the candidate sensitive parameters is graded according to the HALT validation results.
[0031] Specifically, the HALT verification test is based on a multi-environmental parameter matrix to further verify the stability of candidate parameters under extreme conditions. Specifically, the extreme environments preset by the multi-environmental parameter matrix include, for example, applying ±10% rated voltage fluctuations, high temperatures of 70°C, or low temperatures of -20°C. If a candidate parameter shows a significant increase in sensitivity under extreme conditions (e.g., a 3-fold increase in the frequency of anomalies), it is identified as a "core sensitive parameter," indicating that it has a critical impact on SSD reliability. If the parameter's sensitivity decreases or disappears, it is downgraded to a "secondary sensitive parameter," indicating that it is only sensitive under normal conditions or may be prone to misjudgment. Through these two rounds of verification, sensitivity grading is finally completed, providing a basis for adjusting the selection probability.
[0032] For example, suppose in the initial screening test, based on the first test parameter set (random read / write 7:3, IOPS 100K), two anomalies are triggered: "ECC error count > 1000" and "erase count > 1000", which are identified as suspected sensitive parameters. Entering the RDT verification test phase, the strategy is changed to sequential read / write. The result is that the "ECC error count" continues to be abnormal (occurring 5 times per hour), upgrading it to a candidate sensitive parameter; the "erase count" anomaly disappears, ruling it out as a pseudo-sensitive parameter.
[0033] Subsequently, HALT verification testing was conducted. Under extreme conditions (voltage 3.795V, temperature 70℃), the "ECC error count" was tested: the frequency of parameter abnormalities increased from 5 times / hour to 15 times / hour, the sensitivity increased by 3 times, and it was determined to be a core sensitive parameter; the "voltage fluctuation ±5%" parameter tested at the same time did not show significant changes in sensitivity under HALT conditions, and was reduced to a secondary sensitive parameter.
[0034] More specifically, the RDT verification test effectively eliminates the interference of the test strategy on parameter sensitivity through policy variable control, avoiding the inclusion of pseudo-sensitive parameters that are abnormal only under specific read and write modes in the core test, thus improving the recognition accuracy of candidate parameters; while the HALT verification test exposes the inherent sensitivity of parameters through extreme environmental stress, solving the deficiency of traditional room temperature testing in not being able to cover industrial / automotive scenarios; in addition, the sensitivity classification provides a clear basis for subsequent resource allocation through the differentiated division of core / minor.
[0035] Further, the step of identifying sensitive parameters in the first test parameter set and classifying their sensitivity, and updating the selection probability of the corresponding parameters based on the classification results, includes: Based on the HALT verification results, the candidate sensitive parameters are divided into core sensitive parameters and secondary sensitive parameters; The selection probability of the core sensitive parameter is adjusted to a first selection probability, and the selection probability of the secondary sensitive parameter is adjusted to a second selection probability, wherein the second selection probability is less than the first selection probability; The adjusted selection probabilities are updated in the first test parameter set to form the second test parameter set.
[0036] Specifically, the first selection probability is the adjusted selection probability of the core sensitive parameter (usually increased by 50%-80%), and the second selection probability is the adjusted selection probability of the secondary sensitive parameter (usually increased by 20%-30%). Both are dynamically updated based on the sensitivity grading results, ultimately forming the optimized second test parameter set.
[0037] More specifically, the core parameters validated through HALT were identified as key factors affecting stability and reliability. Their high selection probability ensures that they are prioritized for high-intensity stress testing in subsequent tests, thereby improving the anomaly detection rate. Moderate increases in secondary parameters cover scenario-dependent risks while controlling resource consumption, avoiding over-testing. Differential adjustments to the selection probability further enable precise resource allocation, improve test resource utilization, and reduce the false positive rate of sensitive parameters.
[0038] In step S3, the sample test data is fitted using a bathtub curve model to obtain a bathtub curve, and the stage features of the initial failure period, accidental failure period, and wear-out failure period of the bathtub curve are extracted as a stage feature set.
[0039] The bathtub curve model is a statistical model describing the change in product failure rate over time. Named for its bathtub-like curve shape, it comprises three phases: initial failure period, random failure period, and wear-out failure period. Sample test data refers to the full-cycle data generated by the sample group of SSDs during multiple rounds of RDT testing (including multi-environment, multi-strategy verification), including SMART anomaly records, failure time distribution, and degradation patterns of sensitive parameters (such as the relationship between erase count and bad block ratio). The phase feature set is a set of key parameters extracted from the three phases through curve fitting, specifically including the initial failure period duration, mean time between failures (MTBF) during the random failure period, and the degradation rate of sensitive parameters during the wear-out failure period, providing a quantitative basis for the full-cycle testing strategy.
[0040] Specifically, statistical methods are used to quantitatively model the failure patterns of the sample group throughout its entire lifecycle. First, based on multiple rounds of RDT test data of the sample group under various environments (e.g., normal temperature, high temperature, low temperature), key indicators are collected. Then, data fitting tools are used to map these data into complete bathtub curves, and features of three stages, including the initial failure period, the accidental failure period, and the wear-out failure period, are extracted to form a stage feature set.
[0041] More specifically, by quantifying the failure pattern throughout the entire lifecycle, the limitations of traditional testing that only focuses on the initial failure period are avoided, enabling phased and specific analysis. Specifically, the step of extracting the phase characteristics of the initial failure period, accidental failure period, and wear-out failure period of the bathtub curve as a phase feature set includes: The bathtub curve is analyzed to extract the duration of the initial failure period, the mean time between failures (MTBF) of the random failure period, and the parameter degradation rate of the wear-out failure period. All extracted stage features are integrated into the stage feature set.
[0042] The initial failure period refers to the duration of the high failure rate phase on the left side of the bathtub curve, reflecting the period of concentrated early defects in SSDs. It is extracted by statistically analyzing the proportion and time distribution of early failure products in multiple rounds of RDT testing. For example, if 5% of the products fail within 24 hours at room temperature, the corresponding duration is 24 hours.
[0043] Among them, the mean time between failures (MTBF) during the period of random failure refers to the average time without failure during the flat phase in the middle of the curve, which characterizes the reliability of the product during the stable operation phase. It is calculated by the distribution of the time without failure in the low failure rate range of the statistical sample group.
[0044] Among them, the degradation rate of wear-out failure period parameters refers to the rate at which core sensitive parameters deteriorate over time during the high failure rate stage on the right side of the curve. For example, for every 1,000 erase counts, the proportion of bad blocks increases by 0.2%. This can be extracted by performing linear regression analysis on long-term monitoring data of core sensitive parameters (such as ECC error count and erase count).
[0045] Specifically, firstly, the fitted complete bathtub curve is segmented and analyzed, and the boundaries of the three stages are identified through the failure rate inflection point. Then, key quantitative indicators are extracted for each stage. For example, the initial failure period is calculated based on the duration when the early failure rate reaches 90%, the random failure period is calculated based on the average time to failure (MTBF), and the wear-out failure period is fitted with the linear relationship between core sensitive parameters and failure indicators, such as the slope of the erase count-bad block percentage curve. Finally, these indicators are integrated into a structured stage feature set, which contains key information such as the duration of each stage, failure rate, and degradation patterns of sensitive parameters, serving as the core input for generating subsequent test strategies.
[0046] More specifically, the initial failure period duration, derived through data fitting, ensures sufficient exposure of early defects; the introduction of the random failure period MTBF enables dynamic optimization of testing frequency, reducing testing resource consumption while ensuring adequate anomaly detection coverage during the random failure period; and the quantification of the degradation rate of the wear-out failure period parameter provides a scientific basis for lifetime testing, reducing lifetime prediction errors by setting targeted testing intensity. The integrated stage feature set transforms the testing strategy from single-stage coverage to precise full-cycle control, significantly improving testing efficiency and reliability verification accuracy.
[0047] In step S4, an environment adaptation test strategy is generated based on the stage feature set and the environment parameter matrix. The second test parameter set and the environment adaptation test strategy are then combined to perform tests on all products to be tested, and full test results are obtained.
[0048] Specifically, the environment adaptation testing strategy is a systematic approach that dynamically adjusts test parameters for different application scenarios, based on a stage feature set and a multi-environment parameter matrix. First, the full-cycle indicators (initial failure duration, MTBF of accidental failure, and degradation rate of wear-out failure) in the stage feature set are correlated with variables (temperature, voltage, I / O mode) in the environmental parameter matrix to establish a three-dimensional mapping relationship of "environment-stage-parameter," and the testing strategy for full-scale testing is set accordingly. Then, resource allocation is performed through a second set of test parameters, with core sensitive parameters given priority sampling and longer test durations. This process enables full-scale testing to achieve dual optimization of environment adaptation and parameter priority, ensuring coverage of multiple scenarios while avoiding resource waste.
[0049] More specifically, the step of generating an environment adaptation testing strategy based on the stage feature set and the environment parameter matrix includes: The multi-environment aging time is determined based on the duration of the initial failure period, the RDT test interval frequency is determined based on the mean time between failures of the accidental failure period, and the key test parameters are determined based on the parameter degradation rate of the wear-out failure period, and integrated to form a basic test strategy. Specifically, the aging time is first determined based on the duration of the initial failure period; secondly, the RDT test interval frequency is optimized based on the mean time between failures (MTBF) during the accidental failure period; and finally, the key parameters for directional lifetime testing are determined based on the parameter degradation rate during the wear-out failure period. These three factors are then integrated to form a basic testing strategy covering the entire lifecycle.
[0050] More specifically, for example, the initial failure period T0=18h is obtained by fitting the bathtub curve under normal temperature conditions, and this is used as the benchmark duration for aging tests; when MTBF=1200h, the RDT test interval is extended from the traditional 1h / time to 4h / time to reduce redundant tests; if the percentage of bad blocks increases by 0.2% for every 1000 erase counts, the erase count is used as the core monitoring indicator, and 5000 cycles of testing are added to verify its long-term stability.
[0051] Based on the multi-environment parameter matrix, the basic test strategy is adjusted for environmental adaptation to generate the environment-adapted test strategy.
[0052] Specifically, based on the basic testing strategy, dynamic adjustments are made in conjunction with a multi-environment parameter matrix. For example, in a high-temperature environment (60℃), since temperature accelerates device aging, the initial failure period aging time is shortened to 0.7×T0=12.6h, while the voltage is increased to +10% of the rated value (e.g., 3.63V), and the IO mode is adjusted to sequential read / write (6:4) to simulate continuous high load in industrial scenarios; At low temperatures (-20℃), the device response speed slows down, the aging time is extended to 1.2 × T0 = 21.6 h, the voltage drops to -10% of the rated value (2.97V), and the IO mode is adjusted to random read / write (5:5) to adapt to the complex data interaction during the vehicle startup phase. Regarding the accidental failure period, the MTBF may be shortened at high temperatures, requiring the RDT test interval to be reduced from 4 h / test to 3 h / test, while the MTBF is extended at low temperatures, allowing the interval to be relaxed to 5 h / test.
[0053] The focus of the wear-out period adjustment is to take into account the differences in parameter sensitivity under different environments. For example, in high-temperature environments, "cache error rate" is added as a monitoring parameter and verified in conjunction with erase count, ultimately forming a differentiated testing strategy that adapts to multiple scenarios.
[0054] More specifically, through this process, the basic strategy ensures the integrity of testing throughout the entire lifecycle, while the environment adaptation adjustment achieves precise scenario-based optimization. The combination of the two significantly improves the fit of test results with actual applications in complex scenarios such as industry and automotive.
[0055] Further, the step of performing tests on all the products under test according to the second test parameter set and the environment adaptation test strategy to obtain full test results includes: Based on the multi-environment aging time in the environment adaptation test strategy, perform multi-environment aging tests on all the products under test. After completing the multi-environment aging test, the product under test is sampled and tested according to the RDT test interval frequency in the environmental adaptation test strategy and the selection probability in the second test parameter set. Based on the key test parameters in the environmental adaptation test strategy, and the test intensity is determined based on the sensitivity level of the key test parameters, targeted lifetime tests are performed on the corresponding parameters and the failure threshold is recorded. All test data are integrated to form the full test results.
[0056] Among them, sampling testing is conducted after aging testing. Based on the RDT test interval frequency in the environmental adaptation test strategy and the selection probability of sensitive parameters in the second test parameter set, samples are drawn from the full product for targeted testing to verify parameter sensitivity and stability. Among them, targeted lifetime testing is to determine the test intensity (including the number of cycle tests and the duration of extreme conditions) based on the key parameters clearly defined in the environmental adaptation test strategy (such as the erase count of wear-out and the cache error rate under high temperature environment) according to their sensitivity level (core or minor), and record the critical value when the parameter fails.
[0057] Specifically, firstly, based on the multi-environment aging time in the environment adaptation testing strategy, all SSDs under test are subjected to batch aging to quickly expose early failure defects. After aging, combined with the sensitivity parameter selection probability of the second test parameter set and the RDT test interval frequency, a sampling test is performed on all products. This avoids the waste of resources in full-scale high-frequency testing and ensures the effective detection of sensitive parameter anomalies. Subsequently, targeted lifetime testing is performed on the key parameters identified in the environment adaptation strategy according to the sensitivity level, and the failure threshold is recorded. Finally, the three types of test data are integrated to form a full-scale test result that includes parameter sensitivity, lifetime characteristics, and failure modes under multiple environments.
[0058] Furthermore, referring to Figure 2 , Figure 2 This is a flowchart illustrating the solid-state drive test metric optimization method provided in the second embodiment of this application. It further proposes that after the step of obtaining full test results, the method also includes dynamic iteration. Dynamic iteration includes the following key steps, which are described in detail below: In step S5, the core sensitive parameters are verified based on the full test results, and the false positive rate is calculated based on the number of misjudged parameters to obtain the false judgment rate of sensitive parameters.
[0059] Among them, the false positive rate refers to the proportion of core sensitive parameters that were incorrectly judged as sensitive but did not actually trigger an anomaly in the test, out of the total number of core sensitive parameters. Specifically, the actual performance data of all core sensitive parameters in multi-environment testing is extracted from the full test results, including whether SMART anomalies are triggered and whether the parameter degradation rate meets expectations. Secondly, core sensitive parameters that do not actually show anomalies are marked as false positives, and the number of false positive parameters is counted. Finally, the false positive rate of sensitive parameters is calculated based on the percentage of false positives to quantify the accuracy of core sensitive parameter identification.
[0060] In step S6, the full test results are compared with the predicted value of the bathtub curve to calculate the deviation and obtain the bathtub curve prediction deviation.
[0061] Among them, the bathtub curve prediction value refers to the prediction data of the characteristics of each stage of the SSD's entire life cycle by the bathtub curve fitted by the full life cycle modeling module, including the predicted aging time of the initial failure period, the predicted mean time between failures of the random failure period, and the predicted parameter degradation rate of the wear-out failure period.
[0062] Specifically, firstly, the actual characteristic values of each stage are extracted from the full test results; secondly, these actual values are compared with the predicted values of the bathtub curve one by one to calculate the absolute or relative deviation; finally, the deviations of each stage are combined to form the predicted deviation of the bathtub curve.
[0063] In step S7, the sensitive parameter verification standard is updated according to the sensitive parameter misjudgment rate, and the model parameters of the bathtub curve model are updated according to the bathtub curve prediction deviation.
[0064] Among them, the sensitive parameter verification standard refers to the process specifications used to screen and confirm core / secondary sensitive parameters, including the RDT testing strategy of the initial screening layer, the cross-testing method (such as HALT extreme environment testing) of the verification layer, and the graded adjustment rules. Specifically, when the false positive rate of sensitive parameters in the full test results exceeds a preset threshold, the system automatically backtracks the multi-dimensional verification process, analyzes the causes of false positive parameters (such as failure to exclude temperature fluctuation interference), and updates the verification standards. For example, the system can analyze the causes of false positive parameters (such as failure to exclude temperature fluctuation interference) and update the verification standards. Regarding the bathtub curve prediction deviation, if the overall deviation exceeds a preset threshold, the model parameters are corrected based on the actual test data. For example, the wear-out failure period degradation rate coefficient is adjusted from 0.2% / 1000 erases to 0.18%, and the initial failure period aging time calculation coefficient is corrected from 1.0 at room temperature to 0.7 at high temperature and 1.2 at low temperature, making the curve prediction more closely match the actual environment. The updated verification standards and model parameters will be applied to the next batch of SSD tests, forming a continuous iteration mechanism.
[0065] More specifically, by responding to test data feedback in real time, it solves the problem that traditional static test models cannot adapt to SSD batch differences and scenario changes, significantly enhancing the scenario adaptability and continuous optimization capability of the test strategy.
[0066] Reference Figure 3 , Figure 3 This is a virtual structural diagram of the solid-state drive (SSD) test performance optimization device provided in this application. A second aspect of this application provides a solid-state drive (SSD) test performance optimization device, comprising: The sample group and parameter determination module 100 is used to determine the products to be tested, set a portion of the products to be tested as a sample group, and obtain a preset first test parameter set and a multi-environment parameter matrix. The parameter sensitivity grading module 200 is used to perform initial screening tests and multiple rounds of verification tests on the sample group based on the first test parameter set and the multi-environment parameter matrix, collect sample test data, identify sensitive parameters in the first test parameter set and perform sensitivity grading, update the selection probability of the corresponding parameters according to the grading results, and obtain the second test parameter set. The stage feature extraction module 300 is used to fit the sample test data through the bathtub curve model to obtain the bathtub curve, and extract the stage features of the initial failure period, accidental failure period and wear-out failure period of the bathtub curve as the stage feature set. The environment adaptation full-scale testing module 400 is used to generate an environment adaptation testing strategy based on the stage feature set and the environment parameter matrix, and to perform tests on all products to be tested in combination with the second test parameter set and the environment adaptation testing strategy to obtain full-scale test results.
[0067] The solid-state drive test index optimization device described in this application embodiment can execute the solid-state drive test index optimization method provided in the above embodiment. The solid-state drive test index optimization device has the corresponding functional steps and beneficial effects of the solid-state drive test index optimization method described in the above embodiment. For details, please refer to the embodiment of the solid-state drive test index optimization method described above. This application embodiment will not be repeated here.
[0068] This application also provides an electronic device, please refer to... Figure 4 , Figure 4This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include a processor and a memory, which can be connected via a bus or other means. The processor may be a Central Processing Unit (CPU). The processor may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or combinations of the above types of chips. The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the solid-state drive test index optimization method in the embodiments of this application. The processor executes various functional applications and data processing by running the non-transitory software programs, instructions, and modules stored in the memory, thereby implementing the solid-state drive test index optimization method in the above method embodiments.
[0069] The memory may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created by the processor, etc. Furthermore, the memory may include high-speed random access memory and non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. One or more modules are stored in the memory and, when executed by the processor, perform the solid-state drive test metric optimization method as described in the above method embodiments. Specific details of the above electronic device can be understood by referring to the corresponding descriptions and effects in the above method embodiments, and will not be repeated here. Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it may include the processes of the embodiments of the above methods. The storage medium may be a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), etc.; the storage medium may also include a combination of the above types of memory.
[0070] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of this application may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0071] Similarly, it should be understood that, in order to streamline this disclosure and aid in understanding one or more of the various inventive aspects, in the above description of exemplary embodiments of this application, various features of this application are sometimes grouped together in a single embodiment, figure, or description thereof. However, this approach to disclosure should not be construed as reflecting an intention that the claimed application requires more features than expressly recited in each claim. Rather, as reflected in the claims, inventive aspects lie in fewer than all features of a single foregoing disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into that detailed description, wherein each claim itself is a separate embodiment of this application.
[0072] It should be noted that the above embodiments are illustrative of this application and not restrictive of this application, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims.
Claims
1. A method for optimizing solid-state drive (SSD) test metrics, characterized in that, The method includes: Identify the products to be tested, set a portion of the products to be tested as a sample group, and obtain a preset first test parameter set and a multi-environment parameter matrix; Based on the first test parameter set and the multi-environment parameter matrix, the sample group is subjected to initial screening test and multiple rounds of verification test, sample test data is collected, sensitive parameters in the first test parameter set are identified and their sensitivity is graded, and the selection probability of the corresponding parameters is updated according to the grading results to obtain the second test parameter set; The sample test data is fitted using a bathtub curve model to obtain a bathtub curve, and the stage features of the initial failure period, accidental failure period, and wear-out failure period of the bathtub curve are extracted as a stage feature set. An environment adaptation test strategy is generated based on the stage feature set and the environment parameter matrix. The second test parameter set and the environment adaptation test strategy are then combined to perform tests on all products to be tested, and full test results are obtained.
2. The method for optimizing solid-state drive test indicators according to claim 1, characterized in that, The multi-round verification test includes RDT verification test and HALT verification test; The steps of performing initial screening tests and multiple rounds of verification tests on the sample group based on the first test parameter set and the multi-environment parameter matrix, collecting sample test data, identifying sensitive parameters in the first test parameter set, and performing sensitivity classification include: The initial screening test is performed on the sample group according to the first test parameter set to identify abnormal SMART data and obtain suspected sensitive parameters through initial screening. After changing the RDT testing strategy, the RDT verification test is performed on the sample group. If the suspected sensitive parameter still triggers the SMART anomaly, the corresponding suspected sensitive parameter is updated to a candidate sensitive parameter. The HALT validation test is performed on the sample group based on the multi-environment parameter matrix, and the sensitivity of the candidate sensitive parameters is graded according to the HALT validation results.
3. The method for optimizing solid-state drive test indicators according to claim 2, characterized in that, The step of identifying sensitive parameters in the first test parameter set and classifying their sensitivity, and updating the selection probability of the corresponding parameters based on the classification results, includes: Based on the HALT verification results, the candidate sensitive parameters are divided into core sensitive parameters and secondary sensitive parameters; The selection probability of the core sensitive parameter is adjusted to a first selection probability, and the selection probability of the secondary sensitive parameter is adjusted to a second selection probability, wherein the second selection probability is less than the first selection probability; The adjusted selection probabilities are updated in the first test parameter set to form the second test parameter set.
4. The method for optimizing solid-state drive test indicators according to claim 3, characterized in that, After the step of obtaining the full test results, the method further includes: The core sensitive parameters are verified based on the full test results. The false positive rate is calculated based on the number of misjudged parameters to obtain the false positive rate of the sensitive parameters. The deviation is calculated by comparing the full test results with the predicted value of the bathtub curve, thus obtaining the bathtub curve prediction deviation. The sensitive parameter verification criteria are updated based on the false positive rate of the sensitive parameters, and the model parameters of the bathtub curve model are updated based on the prediction deviation of the bathtub curve.
5. The method for optimizing solid-state drive test indicators according to claim 1, characterized in that, The step of extracting the stage features of the initial failure period, accidental failure period, and wear-out failure period of the bathtub curve as a stage feature set includes: The bathtub curve is analyzed to extract the duration of the initial failure period, the mean time between failures (MTBF) of the random failure period, and the parameter degradation rate of the wear-out failure period. All extracted stage features are integrated into the stage feature set.
6. The method for optimizing solid-state drive test indicators according to claim 5, characterized in that, The step of generating an environment adaptation testing strategy based on the stage feature set and the environment parameter matrix includes: The multi-environment aging time is determined based on the duration of the initial failure period, the RDT test interval frequency is determined based on the mean time between failures of the accidental failure period, and the key test parameters are determined based on the parameter degradation rate of the wear-out failure period, and integrated to form a basic test strategy. Based on the multi-environment parameter matrix, the basic test strategy is adjusted for environmental adaptation to generate the environment-adapted test strategy.
7. The method for optimizing solid-state drive test indicators according to claim 6, characterized in that, The step of performing tests on all the products to be tested according to the second test parameter set and the environment adaptation test strategy to obtain full test results includes: Based on the multi-environment aging time in the environment adaptation test strategy, perform multi-environment aging tests on all the products under test. After completing the multi-environment aging test, the product under test is sampled and tested according to the RDT test interval frequency in the environmental adaptation test strategy and the selection probability in the second test parameter set. Based on the key test parameters in the environmental adaptation test strategy, and the test intensity is determined based on the sensitivity level of the key test parameters, targeted lifetime tests are performed on the corresponding parameters and the failure threshold is recorded. All test data are integrated to form the full test results.
8. A device for optimizing solid-state drive test indicators, characterized in that, include: The sample group and parameter determination module is used to determine the products to be tested, set a portion of the products to be tested as a sample group, and obtain a preset first test parameter set and a multi-environment parameter matrix. The parameter sensitivity grading module is used to perform initial screening tests and multiple rounds of verification tests on the sample group based on the first test parameter set and the multi-environment parameter matrix, collect sample test data, identify sensitive parameters in the first test parameter set and perform sensitivity grading, update the selection probability of the corresponding parameters according to the grading results, and obtain the second test parameter set. The stage feature extraction module is used to fit the sample test data to obtain the bathtub curve by using the bathtub curve model, and extract the stage features of the initial failure period, accidental failure period and wear-out failure period of the bathtub curve as the stage feature set. The environment adaptation full-scale testing module is used to generate an environment adaptation testing strategy based on the stage feature set and the environment parameter matrix, and to perform tests on all products to be tested in combination with the second test parameter set and the environment adaptation testing strategy to obtain full-scale test results.
9. An electronic device, characterized in that, include: One or more processors; One or more memory units; And one or more computer programs, wherein the one or more computer programs are stored in the one or more memories, the one or more computer programs including instructions that, when executed by the one or more processors, cause the electronic device to perform the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a program or instructions that, when executed, implement the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Storage device firmware and manufacturing software
CN103597443A
Test index optimization method based on test result data
CN117637004A