State monitoring method of solid state disk, electronic equipment and storage medium
By performing weighted summation on multi-source data of solid-state drives and calculating the cross-layer status index, the problem of the inability to accurately assess the health status of SSDs in existing technologies is solved, comprehensive health status assessment and fault prediction of SSDs are achieved, and storage system management is optimized.
Patent Information
- Application Number
- CN202511109192.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-08-08
AI Technical Summary
Existing technologies are unable to effectively integrate and analyze cross-layer data, making it difficult to accurately assess and predict the health status of solid-state drives (SSDs).
By acquiring multi-source data of NAND flash memory, including the current erase and write counts of block storage, average read and write error rate, and real-time temperature, weighted summation is performed to calculate the cross-layer status index, thereby realizing status monitoring of solid-state drives.
It achieves a comprehensive health status assessment of solid-state drives, breaking the limitations of single-dimensional analysis. It can accurately predict potential failures, optimize storage allocation, reduce the risk of data loss, and improve the reliability and efficiency of storage systems.
Smart Images

Figure CN120610871A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data storage technology, and in particular to a state monitoring method, electronic device, and storage medium for a solid-state hard disk. Background Art
[0002] Solid-state drives (SSDs), the core of modern storage technology, rely heavily on the performance and reliability of NAND (Not AND) flash memory media and controller firmware algorithms. As the core storage medium of SSDs, NAND flash memory's reliability is significantly affected by charge leakage, read disturb, and temperature fluctuations.
[0003] Related NAND flash memory management technologies are unable to effectively integrate and analyze cross-layer data (such as temperature, P / E cycles, bad block distribution, etc.), making it difficult to accurately evaluate and predict the health status of SSDs. Summary of the Invention
[0004] This application provides a solid-state drive status monitoring method, electronic device, and storage medium to at least solve the problem in related technologies that cross-layer data cannot be effectively integrated and analyzed, making it difficult to accurately evaluate and predict the health status of the SSD.
[0005] The present application provides a state monitoring method for a solid-state hard disk, comprising: obtaining multi-source data of a NAND-type flash memory in a solid-state hard disk to be tested; the NAND-type flash memory includes multiple block storages; the multi-source data includes current erase and write times, average read and write error rates, and real-time temperatures of the multiple block storages; performing weighted summation on the current erase and write times, average read and write error rates, and real-time temperatures of block storages in the multiple block storages to obtain cross-layer state indexes of the block storages in the multiple block storages; determining a state index of the solid-state hard disk to be tested based on the cross-layer state indexes of the block storages in the multiple block storages, and performing state monitoring on the solid-state hard disk to be tested based on the determined state index of the solid-state hard disk to be tested.
[0006] The present application also provides a state monitoring device for a solid-state hard disk, comprising: a cross-layer acquisition module for acquiring multi-source data of a NAND-type flash memory in a solid-state hard disk to be tested; the NAND-type flash memory includes multiple block storages; the multi-source data includes the current number of erase and write times, average read and write error rates, and real-time temperatures of the multiple block storages; a state prediction module for performing weighted summation of the current number of erase and write times, average read and write error rates, and real-time temperatures of the block storages in the multiple block storages, respectively, to obtain a cross-layer state index of the block storages in the multiple block storages; a state monitoring module for determining a state index of the solid-state hard disk to be tested based on the cross-layer state index of the block storages in the multiple block storages, and performing state monitoring on the solid-state hard disk to be tested based on the determined state index of the solid-state hard disk to be tested.
[0007] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned solid-state hard disk status monitoring methods when executing the computer program.
[0008] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned solid-state hard disk status monitoring methods are implemented.
[0009] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned solid-state hard disk status monitoring methods when executed by a processor.
[0010] Through this application, multi-source data such as temperature, number of erase and write cycles, and average read and write error rate are integrated, and the health status of each block storage in the solid-state drive (SSD) is comprehensively evaluated through weighted summation. This breaks the limitations of isolated analysis of each parameter and single analysis dimension in related monitoring mechanisms, and solves the problem in related technologies that cannot effectively integrate and analyze cross-layer data, resulting in difficulty in accurately evaluating and predicting the health status of SSDs. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0012] Figure 1 This is an application environment diagram of a solid-state hard disk status monitoring method provided in an embodiment of the present application.
[0013] Figure 2 This is a flow chart of a method for monitoring the status of a solid-state drive provided in an embodiment of the present application.
[0014] Figure 3 A schematic diagram of multiple layers of a solid-state drive to be tested provided in an embodiment of the present application.
[0015] Figure 4 Schematic diagram of the voltage threshold curve of the current scanning point provided by this embodiment.
[0016] Figure 5 This is a flowchart provided in this embodiment for drawing a voltage threshold curve corresponding to the current scanning point stored in a specified block.
[0017] Figure 6 4 is a structural diagram of a state monitoring device for a solid state hard disk provided in this embodiment. DETAILED DESCRIPTION
[0018] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0019] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0020] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0021] The terms involved in the embodiments of this application are explained as follows:
[0022] SSD: Solid State Drive, abbreviated as SSD, is a solid-state storage device that includes a main control chip, DRAM (Dynamic Random Access Memory, abbreviated as dynamic random access memory) cache and NAND (Not AND, abbreviated as NAND gate) particle group.
[0023] NAND FLASH: A non-volatile storage technology that uses a "NAND gate" structure and is widely used in devices such as solid-state drives (SSDs) and USB flash drives. Its core features are high density, low cost, and non-volatility.
[0024] P / E Cycle: Program / Erase Cycle, referred to as program / erase cycle, is a core indicator for measuring the degree of wear of NAND blocks.
[0025] FTL: Flash Translation Layer, a flash translation layer that implements logical address to physical address mapping, including wear leveling and garbage collection functions.
[0026] REBR: Raw Bit Error Rate, the raw bit error rate before error correction, reflects the degree of physical degradation of the storage unit.
[0027] SMART: Self-Monitoring, Analysis and Reporting Technology, abbreviated as self-monitoring, analysis and reporting technology, is used to monitor the health status of SSDs.
[0028] QLC: Quadruple Level Cell, referred to as four-layer cell flash memory.
[0029] According to one aspect of the embodiment of the present application, a method for monitoring the state of a solid state hard disk is provided. Optionally, in this embodiment, the method for monitoring the state of a solid state hard disk can be applied to, but is not limited to, Figure 1 In the illustrated computer device, SSD 102 and processor 104 are included. Processor 104 sends commands to SSD 102, requesting status information such as temperature, program / erase cycles (P / E cycles), bit error rate (REBR), and other key metrics. Upon receiving these requests, SSD 102 provides corresponding real-time or historical data to processor 104, which analyzes and displays this data on an interactive visualization platform. This allows users or systems to monitor SSD status and implement timely maintenance measures or optimization strategies.
[0030] The computer device may be, but is not limited to, a personal computer (PC), a mobile phone, a tablet computer, a cloud server, a server cluster, or other server types.
[0031] The state monitoring method of the solid-state hard disk in the embodiment of the present application can be executed by a computer device. Figure 2 FIG. 1 is a flow chart of an optional method for monitoring the state of a solid-state hard disk according to an embodiment of the present application. Figure 2 As shown, the process of the method may include the following steps:
[0032] Step S202, obtaining multi-source data of a NAND flash memory in the solid state drive to be tested; the NAND flash memory includes multiple block storages; the multi-source data includes the current erase and write times, average read and write error rates, and real-time temperatures of the multiple block storages.
[0033] Step S204 , performing weighted summation on the current erase and write times, average read and write error rates, and real-time temperatures of the block storages in the plurality of block storages to obtain a cross-layer status index of the block storages in the plurality of block storages.
[0034] Step S206 , determining a status index of the solid state drive to be tested according to the cross-layer status index of the block storage in the plurality of block storages, and performing status monitoring on the solid state drive to be tested based on the determined status index of the solid state drive to be tested.
[0035] This application provides a method for monitoring the health of solid-state drives (SSDs) for use in data storage technology, specifically in scenarios such as, but not limited to, enterprise servers, data centers, cloud computing storage systems, and high-performance computing environments. By analyzing multi-source SSD data in real time, this method predicts potential failures, optimizes storage allocation, reduces the risk of data loss, and improves the reliability and efficiency of the overall storage system.
[0036] Solid-state drives (SSDs), the core of modern storage technology, rely heavily on the performance and reliability of NAND flash media characteristics and controller firmware algorithms. As the core storage medium of SSDs, NAND flash reliability is significantly impacted by charge leakage, read disturb, and temperature fluctuations. Existing NAND flash management technologies are unable to effectively integrate and analyze cross-layer data (such as temperature, P / E cycles, and bad block distribution), making it difficult to accurately assess and predict SSD health.
[0037] Therefore, in order to solve the above problems, in an embodiment of the present application, multi-source data is collected across layers, and the status of the SSD is monitored based on the multi-source data. Among them, the multi-source data refers to the comprehensive information about the health status of the NAND gate type flash memory collected from multiple layers of the solid-state drive to be tested. The current erase and write times, average read and write error rate and real-time temperature of the block storage in the multiple block storages can be weighted and summed by a pre-trained health index model to obtain the cross-layer status index of the block storage in the multiple block storages. Among them, Figure 3 A schematic diagram of multiple levels of a solid-state drive to be tested provided in an embodiment of the present application, such as Figure 3As shown, the multiple layers of the SSD under test include SMART data, physical layer signals, and FTL metadata. Each data type reflects different aspects of the flash memory status, forming a comprehensive health monitoring framework. SMART data is part of the SSD's self-monitoring, analysis, and reporting technology (SMR) and provides basic health indicators, including real-time temperature, read / write / erase failure counts, and the percentage of remaining life. Physical layer signals relate to the underlying physical characteristics of the NAND flash memory, specifically the voltage threshold (Vth) distribution obtained through offset voltage sweeps. This Vth distribution reveals physical degradation of memory cells, such as charge leakage and read disturb, and provides a valuable window into NAND flash memory health. Combined with FTL metadata, a health index (HI) for each block can be calculated, demonstrating the direct contribution of physical layer data to SSD health assessment. FTL metadata, part of the flash translation layer, acts as a bridge within the SSD. It stores details such as bad blocks, write / erase cycles (P / E cycles), and bit error rate (REBR), which are key indicators of NAND flash wear and degradation. FTL metadata not only includes bad block information at the factory, but also includes bad block records added during operation, as well as detailed P / E and REBR statistics at the block level.
[0038] In the related art, SMART logs and debug interface data are not prioritized, data collection is rough, and resources are seriously wasted. Therefore, in order to solve this problem, in the embodiment of the present application, differentiated collection frequencies are proposed for different data types, such as using different frequencies for collection according to the importance of the data, unifying heterogeneous data, aligning them according to timestamps, and saving them to a local database. This not only ensures the real-time nature of key monitoring indicators, but also avoids resource waste, solving the problem of rough data collection and low efficiency in the related art. The multi-source data of the embodiment of the present application includes but is not limited to the temperature of each block storage, program / erase cycle (P / E Cycle), bad block distribution, raw bit error rate (REBR), etc. The collection strategy of multi-source data is shown in Table 1 below:
[0039] Table 1
[0040]
[0041] Among them, the current erase and write cycle (P / E Cycle) refers to the total number of programming (writing) and erasing operations that a specific block (Block) in NAND Flash has undergone since its production. The average read and write bit error rate (Raw Bit Error Rate, REBR) is an indicator that measures the error ratio between the raw data and the actual stored data in the read or write operation of the NAND Flash storage unit. In this application, it is used to evaluate the degree of physical degradation of the storage block (Block). The higher the average REBR, the lower the reliability of the storage unit. Real-time temperature refers to the real-time temperature reading of the internal NAND flash memory of the solid-state drive during operation. Temperature has a significant impact on the performance and life of NAND Flash. Too high or too low temperature will accelerate the degradation of the storage unit and reduce the reliability of the data.
[0042] NAND flash memory specifically refers to a non-volatile storage technology that uses a "NAND gate" (NAND) circuit architecture. It has high-density storage, fast read and write speeds, and low power consumption, and is widely used in solid-state drives (SSDs). The solid-state drive to be tested contains NAND flash memory, and NAND flash memory consists of multiple block storage units. Block storage is the basic unit of data storage in NAND Flash. Each block consists of multiple pages and can be erased independently. The performance and reliability of solid-state drives are highly dependent on the status of block storage units, especially parameters such as the current number of erase and write cycles, average read and write error rate, and real-time temperature. This application monitors these multi-source data to evaluate the health of the entire solid-state drive and achieve fault prediction and management.
[0043] The Cross-Tier Health Index is a comprehensive health metric used to assess the current state and potential risks of NAND Flash block storage. It intelligently weights and sums key parameters such as the current number of erase / write cycles, average read / write bit error rate, and real-time temperature to generate a quantitative value that intuitively reflects the degree of wear and degradation of the storage block across different dimensions. Weights are dynamically adjusted based on the impact of each parameter on storage reliability, making the assessment more realistic and effective in predicting and diagnosing early failures of storage cells, thereby optimizing SSD management and extending its lifespan.
[0044] The Health Index is a quantitative indicator of the overall health of the solid-state drive (SSD) under test. It relies on a comprehensive analysis of cross-layer health indices. Starting at the NAND Flash block level, it considers key parameters such as temperature, current erase / write cycles (P / E cycles), and average read / write bit error rate (REBR). A weighted summation is then used to derive a health score (HI) for each block. The Health Index further integrates these block-level HI values to reflect the overall health of the SSD, including wear, degradation, and operating environment.
[0045] Through the embodiments of the present application, multi-source data such as temperature, number of erase and write times, and average read and write error rate are integrated, and the health status of each block storage in the solid-state drive (SSD) is comprehensively evaluated through weighted summation. This breaks the limitations of isolated analysis of each parameter and single analysis dimension in the relevant monitoring mechanism, and solves the problem in the relevant technology that it is impossible to effectively integrate and analyze cross-layer data, which makes it difficult to accurately evaluate and predict the health status of the SSD.
[0046] In an exemplary embodiment, a weighted sum is performed on the current number of erase and write times, the average read and write error rate, and the real-time temperature of a block storage in a plurality of block storages to obtain a cross-layer status index of the block storage in the plurality of block storages, including:
[0047] 1. Calculate the ratio between the current erasure count of a block storage in the plurality of block storages and the maximum tolerated erasure count of a block storage in the plurality of block storages to obtain a first ratio corresponding to the block storage in the plurality of block storages.
[0048] 2. Calculate the ratio of the average read and write error rate of the block storage in the multiple block storages to the maximum read and write error rate of the block storage in the multiple block storages to obtain a second ratio corresponding to the block storage in the multiple block storages.
[0049] 3. Calculate the temperature difference between the real-time temperature of the block storage in the multiple block storages and the optimal operating temperature of the block storage in the multiple block storages, and calculate the ratio between the temperature difference of the block storage in the multiple block storages and the maximum allowable temperature of the block storage in the multiple block storages, to obtain a third ratio corresponding to the block storage in the multiple block storages.
[0050] 4. Perform a weighted summation of a first ratio corresponding to a block storage in the multiple block storages, a second ratio corresponding to a block storage in the multiple block storages, and a third ratio corresponding to a block storage in the multiple block storages to obtain a cross-layer state index of the block storage in the multiple block storages.
[0051] In this application, the maximum number of erase and write cycles refers to the maximum number of program / erase (P / E) operations that each block storage unit (Block) in NAND Flash can withstand during its life cycle. If this number is exceeded, the reliability of the unit will be significantly reduced, which may cause data errors or block failures.
[0052] The first ratio is the ratio of the current number of erase / write cycles to the maximum number of cycles a block has endured, and is used to quantify the wear and tear of the block. When calculating the cross-layer health index, a higher first ratio indicates that the block is nearing its lifespan, which in turn has a greater negative impact on the health index.
[0053] The maximum read / write bit error rate (BER) measures the highest acceptable level of data read / write errors in NAND Flash. It reflects the highest possible error rate in data read / write operations without additional error correction measures. This parameter is compared with the average read / write BER of the block storage to calculate a second ratio, which reflects the degree of degradation during data read / write operations.
[0054] The second ratio is the ratio of the block storage's average read and write bit error rate to its maximum read and write bit error rate. This ratio is used to assess the block storage's data integrity status. In the cross-layer health index calculation, the second ratio reflects the impact of bit errors on the block storage's health. A ratio close to 1 indicates that the block storage's bit error rate is approaching its critical threshold, indicating a poor health status.
[0055] The optimal operating temperature refers to the optimal operating environment temperature considered during NAND Flash block storage design. Within this range, the storage device can achieve optimal performance and longest lifespan while maintaining the lowest power consumption and most stable reliability. The optimal operating temperature is selected based on a comprehensive consideration of factors such as material properties, manufacturing process, and workload. It aims to provide an ideal operating environment for the SSD, minimizing temperature-related performance degradation and data errors. For example, the optimal operating temperature could be 45°C.
[0056] The maximum allowable temperature is the maximum temperature limit at which NAND Flash can operate stably without performance degradation or reliability loss. For example, the maximum allowable temperature might be 60°C. Exceeding the maximum allowable temperature can affect the storage medium's charge retention, further impacting data stability and durability. When calculating the third ratio, the ratio of the real-time temperature to the maximum allowable temperature is used to assess the temperature risk of the storage environment.
[0057] The third ratio is the ratio of the temperature difference between the real-time temperature and the optimal operating temperature to the maximum allowable temperature, which reflects the degree to which the temperature deviates from the optimal state. In this embodiment, the third ratio is regarded as a temperature deviation ratio, which characterizes the degree to which the operating temperature of the solid-state drive to be tested deviates from the optimal operating temperature, relative to the standardized ratio of the maximum allowable temperature. The third ratio will increase significantly in a high-temperature environment, triggering an increase in the temperature weight, thereby giving more attention to the temperature factor in the calculation of the cross-layer state index, and providing early warning of potential temperature risks. Among them, the temperature difference refers to the difference between the real-time temperature of the block storage and the optimal operating temperature, which is used to quantify the degree of deviation of the current working environment from the ideal state. In this embodiment, a method is used to calculate the third ratio by calculating the temperature difference between the real-time temperature and the optimal operating temperature and then calculating the ratio with the maximum allowable temperature. This temperature difference can more accurately measure the degree to which the current block storage temperature deviates from its optimal operating temperature. Compared with directly using the real-time temperature, this method can more delicately reflect the potential impact of temperature fluctuations on the health status of the NAND Flash. Furthermore, the ratio of the temperature difference to the maximum allowable temperature is used as the third ratio. The closer the temperature is to the maximum allowable temperature, the larger the ratio of the temperature difference to the maximum allowable temperature, indicating that the block storage faces a higher risk. This method can automatically amplify the impact of temperature when the temperature approaches the limit.
[0058] In this embodiment, a weighted summation is performed by calculating the ratio between the current parameter value and the corresponding maximum value, rather than directly weighting the sum of the current erase / write cycles, average read / write bit error rate, and real-time temperature. This is primarily for the following reasons: 1) Standardized data: By calculating the ratio, parameters with different scales and units (such as erase / write cycles, bit error rate, and temperature) can be converted to a common evaluation scale, typically between 0 and 1. This avoids weight distortion caused by large differences in parameter magnitude when directly adding them together, ensuring fairness and comparability in the evaluation of each parameter. 2) Clear physical meaning: The ratio calculation better reflects the degree of deviation of the parameter from its ideal state (maximum endurable erase / write cycles, maximum read / write bit error rate, and maximum allowable temperature), thereby intuitively reflecting the health of the block storage in three key dimensions: wear, data integrity, and environmental suitability. This approach makes the interpretation of the health index easier to understand and can be directly linked to the block storage's lifespan, reliability, and environmental adaptability. 3) Avoiding the impact of data magnitude: The erase / write cycles may reach thousands or even tens of thousands, while the values of temperature and bit error rate are relatively small. Without normalization, a straightforward weighted summation could cause the erase / write cycle to dominate the health index, while ignoring the effects of temperature and bit error rate. Ratioing ensures that each parameter contributes to the health index based on its deviation from the maximum, avoiding the problem of a single parameter dominating the health index.
[0059] This embodiment introduces a first ratio (i.e., the ratio of the current erase / write count of the block storage to the maximum tolerated erase / write count), a second ratio (i.e., the ratio of the average read / write error rate (BER) to the maximum read / write error rate (BER), and a third ratio (i.e., the ratio of the temperature difference between the real-time temperature and the optimal operating temperature to the maximum allowable temperature). The first ratio overcomes the limitation of related monitoring mechanisms that rely solely on the current erase / write count for evaluation. By quantifying the ratio of the erase / write count to the lifespan limit, it more accurately reflects the wear state of the block storage. The calculation of the second ratio more accurately captures the impact of errors on storage health during data read / write processes, addressing the inefficiency issues in related technologies, providing a means to quickly identify the root cause of error problems, and avoiding lengthy fault location. The third ratio emphasizes the importance of temperature to NAND flash memory health, addressing the issue of data collection strategies in related technologies ignoring temperature sensitivity. The cross-layer status index is derived by weighted summing the first, second, and third ratios, avoiding the dimensional and magnitude inconsistencies that may result from directly using raw data, and ensuring that data from different sources plays an equal role in the evaluation.
[0060] In an exemplary embodiment, the solid-state drive status monitoring method further includes:
[0061] The following weight adjustment operation is performed on each of the multiple block storages as the current block storage: when the real-time temperature of the current block storage is greater than a preset temperature threshold, the initial weight of the third ratio corresponding to the current block storage is adjusted to obtain a first weight; the first weight is positively correlated with the real-time temperature of the current block storage and is less than the preset weight threshold; based on the first weight, the initial weight of the second ratio corresponding to the current block storage is adjusted to obtain a second weight, and based on the first weight, the initial weight of the first ratio corresponding to the current block storage is adjusted to obtain a third weight; the sum of the first weight, the second weight, and the third weight is equal to 1; wherein the weighted summation of the first ratio corresponding to the block storage in the multiple block storages, the second ratio corresponding to the block storage in the multiple block storages, and the third ratio corresponding to the block storage in the multiple block storages is performed based on the first weight, the second weight, and the third weight.
[0062] In this embodiment, the cross-layer state index of block storage can be expressed using the following formula (1):
[0063]
[0064] in, The maximum number of erase and write cycles that can be tolerated for a specified block. The current number of erase and write times for the specified block storage (Block); The average read and write error rate of the pages that have been read in the specified block storage (Block); The maximum read and write error rate in the specified block storage (Block); Stores the real-time temperature of a specified block. The optimal operating temperature for a specified block, for example, 45 degrees. The maximum allowable temperature for a specified block, for example, 60 degrees. represents the third weight; represents the second weight; represents the first weight, and + + =1.
[0065] The current block storage refers to the specific NAND Flash storage block undergoing health assessment. The preset temperature threshold is the warning limit set for the real-time temperature of the block storage within the SSD's health monitoring system. When the real-time temperature of the block storage exceeds this threshold, the system adjusts the initial weight of the third ratio, indicating that the storage device is experiencing an unfavorable temperature environment and may be at risk of performance degradation or shortened lifespan. This threshold is determined based on the specific temperature sensitivity of the NAND Flash material and optimal operating conditions. For example, the preset temperature threshold could be 65 degrees Celsius.
[0066] The initial weight of the third ratio reflects the impact of the temperature dimension on the health of the block storage. This initial weight is preset through experimentation or experience and represents the default importance metric for the REBR parameter when the temperature does not exceed the maximum allowable temperature. In the absence of specific temperature triggers, the initial weight of the third ratio reflects the basic weight of temperature in the cross-layer health index calculation. The first weight is the result of adjusting the initial weight of the third ratio when the real-time temperature changes. The first weight is positively correlated with the real-time temperature but is subject to a preset weight threshold to prevent the temperature parameter from excessively influencing the overall health index. The preset weight threshold serves as an upper limit for adjusting the first weight, preventing significant temperature changes from causing its weight to rise too high, thereby affecting the overall balance of the health index. For example, the preset weight threshold could be 0.4. The preset weight threshold ensures that regardless of real-time temperature fluctuations, the contribution of the temperature parameter to the health index does not exceed a preset limit.
[0067] Optionally, the computer device continuously collects real-time temperature data of the current block storage at a frequency of 2Hz through the SSD physical layer interface. A preset temperature threshold (e.g., 65 degrees Celsius) is set. When the real-time temperature exceeds the preset temperature threshold, a weight adjustment mechanism is triggered. For example, a linear growth model is used, where the first weight increases accordingly with each degree increase in temperature, but the first weight is ensured to be less than a preset weight threshold (e.g., 0.4). Based on the first weight value, the initial weights of the second ratio (e.g., the weight related to the PE cycle) and the third ratio (e.g., the weight related to the REBR) are proportionally adjusted. For example, if the first weight increases, the second and third weights are appropriately reduced to maintain the total weight sum to 1, ensuring a reasonable weight distribution.
[0068] Through this embodiment, when the real-time temperature of the current block storage is greater than the preset temperature threshold, the initial weight of the third ratio corresponding to the current block storage is adjusted to obtain a first weight. The first weight is positively correlated with the real-time temperature of the current block storage and is less than the preset weight threshold. The second weight and the third weight are dynamically adjusted according to the change of the first weight, and the total weight sum is ensured to be 1. This can accurately reflect the physical degradation effect of high temperature on the block storage.
[0069] In an exemplary embodiment, when the real-time temperature of the current block storage is greater than a preset temperature threshold, adjusting the initial weight of the third ratio corresponding to the current block storage to obtain the first weight includes:
[0070] When the real-time temperature of the current block storage is greater than a preset temperature threshold, the temperature difference between the real-time temperature of the current block storage and the maximum allowable temperature is determined as the real-time temperature difference, and the ratio of the real-time temperature difference to the preset weight influence coefficient is determined as the temperature change ratio; according to the temperature change ratio, a first weight adjustment factor is determined, and the product of the first weight adjustment factor and the initial weight of the third ratio of the current block storage is determined as the first weight.
[0071] The real-time temperature difference refers to the difference between the current block's real-time temperature and the maximum allowable temperature. It quantifies the degree to which the temperature deviates from the safe range. For example, when the threshold is 65 degrees and the real-time temperature is 85 degrees, the real-time temperature difference is 20 degrees. This value directly drives the weight adjustment calculation to ensure risk response sensitivity in high-temperature scenarios.
[0072] The temperature change ratio is determined by the ratio of the real-time temperature difference to the preset weight influence coefficient, reflecting the quantitative impact of temperature change on weight adjustment. The preset weight influence coefficient is a fixed parameter used to quantify the impact of the real-time temperature difference on weight adjustment. Its value is set through experimentation or experience and serves as the denominator in the temperature change ratio calculation. For example, if the preset weight influence coefficient is 30 and the real-time temperature difference is 20 degrees, the temperature change ratio is 0.67. The temperature change ratio determines the magnitude of the first weight adjustment factor, achieving a nonlinear relationship between temperature and weight.
[0073] The first weight adjustment factor is a dynamic coefficient calculated based on the temperature change ratio and is used to correct the initial weight. For example, the first weight adjustment factor can be a dynamic adjustment coefficient obtained by adding a constant 1 to the temperature change ratio. In this embodiment, the first weight can be expressed using the following formula (2):
[0074]
[0075] The maximum allowable temperature can be 60 degrees, and the preset weight influence coefficient can be 30. Indicates the initial weight of the third ratio stored in the current block, The weight after the initial weight of the third ratio stored in the current block is adjusted, that is, the first weight.
[0076] According to this embodiment, when the real-time temperature of the current block storage is greater than a preset temperature threshold, the temperature difference between the real-time temperature of the current block storage and the maximum allowable temperature is determined as the real-time temperature difference. The real-time temperature difference directly reflects the physical degradation pressure of the block storage by quantifying the degree of deviation between the current temperature and the maximum allowable temperature. The ratio between the real-time temperature difference and the preset weight influence coefficient is determined as the temperature change ratio. The temperature difference and the preset weight influence coefficient are normalized by the temperature change ratio to eliminate the impact of hardware model differences on the calculation and ensure that the weight adjustment of different SSDs is comparable.
[0077] In an exemplary embodiment, adjusting the initial weight of the second ratio corresponding to the current block storage according to the first weight to obtain the second weight includes:
[0078] Determine, based on the first weight, the current remaining ratio of the first weight relative to the specified weight state, and, based on the initial weight of the third ratio corresponding to the current block storage, determine the original remaining ratio of the initial weight of the third ratio corresponding to the current block storage relative to the specified weight state; determine the ratio between the current remaining ratio and the original remaining ratio as a second weight adjustment factor, and determine the product of the second weight adjustment factor and the initial weight of the second ratio corresponding to the current block storage as the second weight.
[0079] The designated weight state refers to the situation where the sum of all parameter weights is 1, which serves as the baseline for weight adjustments. The designated weight state represents the original importance distribution of each parameter in the health index model under normal or ideal operating conditions and serves as a reference point for subsequent weight changes.
[0080] The current remaining ratio refers to the difference between the first weight and 1 (i.e. 1- ), reflects the change in the proportion of the remaining space occupied by other parameter weights after the temperature weight is adjusted, and is used to calculate the second weight adjustment factor to achieve dynamic balance of the weight system.
[0081] The original residual ratio is the difference between the initial weight of the third ratio and 1 when the current block is stored without being affected by temperature (i.e. 1- ), which shows the proportion of the temperature parameter's relative residual contribution to the health index calculation under normal conditions, used to compare the effect of weight adjustment after temperature changes.
[0082] The second weight adjustment factor is calculated based on the ratio of the current remaining ratio to the original remaining ratio. It is used to dynamically adjust the initial weight of the second ratio (such as the bit error rate weight) corresponding to the current block storage. This second weight adjustment factor ensures that the importance of key parameters such as read and write bit error rates can be automatically adapted under temperature-sensitive conditions, maintaining the accuracy and stability of the health index model.
[0083] In this embodiment, the adjusted second weight can be expressed by the following formula (3):
[0084]
[0085] In an exemplary embodiment, the third weight is adjusted in the same manner as the second weight, and details thereof will not be repeated here.
[0086] In an example, assuming that the optimal operating temperature of the current block storage is 45 degrees and the maximum allowable temperature is 60 degrees, the maximum number of erase and write times of the current block storage is =1000, the current number of erase and write times of the current block storage =400, the maximum read and write error rate in the current block storage =1.1x10-3, third weight =0.5, second weight , first weight .
[0087] If the temperature is normal, the real-time temperature The temperature is 45 degrees, which does not exceed the maximum allowable temperature. The first weight remains unchanged. The average read and write error rate of the pages that have been read in the current block storage is =7x10-5, then the cross-layer state index of the current block storage can be expressed by the following formula (4):
[0088]
[0089] The above formula (4) indicates that the current block storage health status is good (HI=0.221).
[0090] If the temperature is medium, the real-time temperature The temperature is 65 degrees, which does not exceed the maximum allowable temperature. The first weight remains unchanged. The average read and write error rate of the pages that have been read in the current block storage is =4x10-4, then the cross-layer state index of the current block storage can be expressed by the following formula (5):
[0091]
[0092] The above formula (5) indicates that the current block storage is at a medium risk (HI=0.387).
[0093] If in high temperature scenario, real-time temperature The temperature is 85 degrees, which exceeds the maximum allowable temperature. The first weight, the second weight, and the third weight all need to be readjusted. Assuming that the preset weight threshold is 0.4, the first weight can be adjusted using the following formulas (6) and (7), the second weight can be adjusted using the following formula (8), and the third weight can be adjusted using the following formula (9):
[0094]
[0095]
[0096]
[0097]
[0098] Assume that the average read and write error rate of the pages that have been read in the current block storage is =9x10-4, then the cross-layer state index of the current block storage corresponding to the weight adjustment can be calculated using the following formula (10):
[0099]
[0100] The above formula (10) indicates that the current block storage is at high risk (HI=0.616).
[0101] Through this embodiment, the current remaining proportion of the first weight relative to the specified weight state is calculated, and the weight of the temperature parameter in the health index calculation can be dynamically adjusted. When the temperature exceeds the preset temperature threshold, the temperature impact ratio is automatically increased, thereby achieving a rapid response to changes in temperature sensitivity; the initial weight of the third ratio relative to the original remaining proportion of the specified weight state can quantify the difference before and after the temperature weight adjustment; the second weight adjustment factor is multiplied by the initial weight of the second ratio to determine the second weight, and the weights of important indicators such as the read and write error rate can be adaptively adjusted according to real-time temperature changes to ensure that when the temperature rises and the first weight increases, the weights of other parameters are automatically reduced to maintain the stability and accuracy of the overall evaluation system.
[0102] In an exemplary embodiment, fault location in related technologies takes up to several weeks. Therefore, to solve this technical problem, this embodiment provides an interactive visualization platform that visualizes the physical structure of the SSD, supports timeline backtracking and data drill-down, and supports multi-chart linkage analysis, thereby improving fault location speed. The above-mentioned solid-state drive status monitoring method also includes:
[0103] According to the physical structure of the solid-state drive to be tested, a multi-level addressing heat map is constructed; through the multi-level addressing heat map, a group of target block storages is located and displayed; a group of target block storages refers to at least one block storage whose cross-layer health index is greater than a preset health index; wherein, the multi-level addressing heat map is a multi-level view constructed with the cross-layer health index of the block storage as the heat value; the multi-level addressing heat map includes a channel level view, a target level view, a logical unit level view and a block level view; the channel level view displays a summary of the health status of the channel level in the solid-state drive to be tested; the target level view displays a summary of the health status of the target level in the solid-state drive to be tested; the logical unit level view displays the cross-layer health index of the block storage under the logical unit level in the solid-state drive to be tested; the block level view displays the properties of the word line under the block level in the solid-state drive to be tested; the word line is a control line used to select a specific storage unit.
[0104] The multi-level addressing heat map is a visualization tool based on the physical structure of solid-state drives (SSDs). Using a cross-layer health index (HI) as a heat value, it displays the health status of various SSD components through views at different levels (channel, target, logical unit, and block), helping to quickly locate problem areas. Within a solid-state drive (SSD), NAND flash memory is organized in a complex, layered structure. Specifically, the SSD's internal physical structure is organized as Channel -> Target -> LUN (logical unit) -> Plane -> Block. In the SSD's physical architecture, the Channel is the outermost organizational structure. Each Channel contains multiple Targets, each Target manages several LUNs (logical units), each LUN is divided into multiple Planes, and finally, each Plane consists of multiple Blocks.
[0105] The channel level view is part of the multi-level addressing heat map, which is used to display the health status summary of the channel level in the SSD. Among them, the channel level health status summary provides a comprehensive and detailed overview of the health status of the solid state drive at the channel level. For example, the channel level health status summary includes the HI value (the average of all HI values obtained under the channel), the problem target, LUN, the number of blocks, and the number of bad blocks. Assume that the solid state drive to be tested includes 16 channels, then the channel level view Figure 1 It is generally presented in the form of a 16×1 matrix, where each unit reflects the health status summary of a channel, such as the mean HI value, making it easier to grasp the overall health status of the channel at a macro level.
[0106] The target-level view, also part of the multi-level addressing heat map, focuses on a target-level health summary within the SSD. This target-level health summary provides a comprehensive and detailed overview of the SSD's health at the target level. For example, the target-level health summary includes the health indicator (HI) value (the average of all acquired HI values for that target), problematic targets, LUNs, number of blocks, and number of bad blocks. Assuming the SSD under test includes four targets per channel, the health indicators for the four targets are presented in a 4×1 matrix, providing data support for mid-level analysis.
[0107] The logical unit-level view displays the cross-layer health index (HI) of all block storage within each logical unit (LUN) in a multi-level addressing heat map system. All blocks within a LUN are distributed by plane and converted into a two-dimensional grid based on their physical topology. This grid is represented by the plane number (e.g., 0-3) on the x-axis and the block number (e.g., 0-M) on the y-axis. Each grid cell represents a single block (the size dynamically adjusts with the zoom level). The color of the grid cell reflects the HI value. Bad blocks are marked with different colors depending on whether they are currently bad or newly added, facilitating detailed LUN-level health analysis. In addition to displaying the cross-layer health index of block storage within the logical unit, the logical unit-level view also displays each block's attributes, including the physical address, LBA address, HI, PE, and REBR (mean). When the logical unit-level view focuses on a block, all of these attribute values are displayed.
[0108] The block-level view is the lowest level of the heat map. It displays the health attributes of individual blocks using a wordline x 1 matrix, including wordline-level physical addresses, REBR information, and more. It supports wordline-level Vth distribution scanning, enabling precise location and analysis of internal block health issues. A wordline is a control line used to select specific memory cells or pages on the NAND Flash chip in an SSD for read and write operations. It forms the foundation of the block-level view. By monitoring wordline properties, in-depth analysis and location of subtle changes in memory cells can be achieved, ensuring SSD stability and reliability.
[0109] In this embodiment, a target block storage group specifically refers to a set of block storage whose cross-layer health index (HI) exceeds a preset health threshold, and includes at least one block storage. In this embodiment, HI values and REBR values can be displayed using different statistical methods at different layers. A heat map of high-risk block distribution can be displayed at the lun level, and a list of the top 10 degraded blocks ranked by HI values can be output.
[0110] This embodiment can bind real-time and historical data to various view levels of the visualization system, linking relevant information. For example, when locating a specific wordline, a Vth scan is initiated, a curve is drawn, and historical scan records are provided, with the ability to overlay historical curves. Different types of alarm conditions can also be configured based on usage requirements. When triggered, relevant data and information are automatically recorded and corresponding operations (such as Vth scans) are executed, with the corresponding issues popping up or flagged in the visualization system.
[0111] In one example, during a high-temperature stress test on a QLC SSD, the method of this embodiment was used to configure an alarm when the HI value of all blocks exceeded 0.7. When the test temperature reached 85°C, the heat map showed that the HI value of CH2-T1-L0 rose to 0.82, detecting the charge leakage problem 14 days ahead of the SMART warning. The visualization system accurately located the specific location. By analyzing the Vth distribution of a large number of wordlines in this block and overlaying them, it was found that all of them had large offsets. Reading the associated REBR curve revealed that the REBR in this area was high. After enhancing the heat dissipation comparison, the problem was alleviated.
[0112] Through this embodiment, the multi-level addressing heat map is organized into levels of channel, target, logical unit (LUN), block, and wordline, so that fault diagnosis can start from the macro channel level and gradually refine to the specific wordline level. This top-down analysis path greatly accelerates the problem location process; the heat map intuitively represents the health status (HI value) of different levels with color depth or temperature, allowing maintenance personnel to see at a glance which levels or blocks have abnormal health indexes and immediately identify fault points, without having to delve into the code or log files to gain an overview.
[0113] In one exemplary embodiment, as storage density increases (e.g., QLC technology), cell charge capacity decreases, and even a slight shift in the voltage threshold distribution can lead to read errors. Therefore, in order to monitor the status of the SSD under test, in this embodiment, the SSD status monitoring method further includes:
[0114] Initiate a scan instruction packet for the specified block storage in the block-level view; the scan instruction packet includes voltage threshold scan configuration parameters; the scan instruction packet is used to instruct the solid-state drive under test to cyclically scan the specified block storage according to the voltage threshold scan configuration at the current scan point and return the scan data corresponding to the current scan point; based on the scan data, draw the voltage threshold curve corresponding to the specified block storage at the current scan point; the voltage threshold curve corresponding to the current scan point is used to analyze the read status and data reliability of the NAND gate flash memory at the current scan point.
[0115] Among them, this embodiment supports NAND wordline level Vth scanning and generates a Vth distribution curve diagram, supports overlay comparison, difference calculation, etc.
[0116] Designated block storage refers to a block or blocks within the physical storage structure of a solid-state drive (SSD) that are specifically selected for monitoring, analysis, or operation. Designated block storage can be located anywhere within the SSD. For example, a designated block storage can be a target block storage.
[0117] A scan command packet is a series of instructions and configuration parameters sent by the visualization system to the SSD controller. It performs a Vth (voltage threshold) distribution scan on a specified NAND Flash block. For example, the scan command packet includes key information such as the target NAND physical address, target Vth value (generally, all Vth values are scanned by default; QLC has 15), offset step size (e.g., 4), and scan range (e.g., -64 to +64). This ensures the SSD controller accurately performs the Vth scan and returns the required data.
[0118] The current scan point typically refers to the moment when the system issues a Vt distribution scan command and performs a read operation. The voltage threshold curve, or Vth curve, is a graph generated based on the Vth distribution scan data. It visually displays the voltage threshold distribution of each memory cell in the NAND Flash memory. This curve is crucial for analyzing physical cell degradation (such as charge leakage) and assessing the reliability of read operations. By analyzing the curve, the health of the NAND Flash block can be promptly identified and assessed. Figure 4 is a schematic diagram of the voltage threshold curve of the current scanning point provided by this embodiment, such as Figure 4 As shown, the voltage threshold curve of the current scanning point includes a scanning range of -64 to +64.
[0119] Optionally, Figure 5 This is a flow chart provided in this embodiment for drawing a voltage threshold curve corresponding to the current scanning point stored in a specified block, such as Figure 5 As shown, the visualization system in a computer device sends a scan command packet to the solid-state drive (SSD) controller chip. The scan command packet includes a target Vth value (e.g., RL10_QLC, indicating the 10th read level for QLC (quad-level cell) flash memory), an offset step size (e.g., 4), and a scan range (e.g., -64 to +64). The SSD controller loops through the NAND gate (NAND) chip 65 times according to the set offset step size. The NAND chip then returns scan data to the SSD controller. The SSD controller then forwards the scan data to the visualization system. Based on the scan data, the visualization system plots the voltage threshold curve corresponding to the specified block storage at the current scan point.
[0120] Through this embodiment, a Vt distribution scan is performed on a specified block of storage at the current scan point, and direct health status data of the storage cell at that moment, such as charge retention and read stability, can be obtained. This helps to instantly assess data reliability and whether the storage cell has suffered from faults such as read disturbance and programming failure. In a multi-level SSD structure, the ability to initiate a scan directly at the block level means that the specific location of the faulty storage cell can be accurately located.
[0121] In an exemplary embodiment, the solid-state drive status monitoring method further includes:
[0122] Select the voltage threshold curves corresponding to multiple historical scanning points of the specified block storage, superimpose and display the voltage threshold curve corresponding to the current scanning point of the specified block storage and the voltage threshold curves corresponding to multiple historical scanning points, and use the baseline curve as a reference to determine the average offset of the voltage threshold curve; the baseline curve refers to the voltage threshold curve generated based on the first scan of the specified block storage when the solid-state drive to be tested leaves the factory; when the cross-layer health index of the specified block storage is greater than the preset health index, or the average offset of the voltage threshold curve of the specified block storage is greater than the preset offset threshold, an alarm report is generated.
[0123] Average offset refers to the average difference in voltage between data points at the same voltage between the Vt curves of the current scan point and those of historical scan points. Average offset is a quantitative indicator for assessing memory cell aging and performance degradation, helping to dynamically monitor changes in NAND health.
[0124] The baseline curve is the first voltage threshold curve obtained by performing a Vt distribution scan on a specified storage block in the SSD's factory default state or under ideal conditions. It serves as a reference baseline for comparative analysis of subsequent scan results, helping to identify storage cell degradation trends over time.
[0125] The preset health index (HI) is a pre-set threshold used to determine whether the health of a specific block storage has deteriorated to a level requiring attention. When the HI exceeds the preset health index, an alarm is triggered, prompting troubleshooting or remediation measures.
[0126] The preset offset threshold is a pre-set threshold for the average offset of the Vt curve. It is used to determine whether a memory cell's read stability and data retention have significantly changed. If the average offset exceeds this threshold, the system will generate an alarm, indicating that the memory cell may have experienced severe performance degradation.
[0127] Optionally, the computer retrieves voltage threshold curve data for a specified block at multiple historical scan points (up to five time points) from a database, as well as the latest voltage threshold curve data for the current scan point. The Vt curve for the current scan point is overlaid and displayed on a visual interface. The overlaid curve is then compared with the baseline curve generated at the factory, as a reference. The voltage difference between the corresponding points on the historical and current scan points and the baseline curve is calculated, and the average offset is calculated. The offset values at different points are calculated, and the maximum and minimum offsets are recorded. Each configured alarm event, such as high (HI) warning and Vth warning, automatically records the time and triggers a Vth scan upon triggering. This monitors the HI value in real time and checks whether the cross-layer health index of the specified block exceeds a preset health index threshold. Once the threshold is reached, subsequent alarm mechanisms are triggered. The average offset of the historical and current Vt curves is analyzed to see if it exceeds a preset offset threshold. If the HI value of the specified block exceeds the specified threshold or the average offset of the Vt curve is excessive, the system automatically generates an alarm report. The alarm report should include key data such as the block information that triggered the alarm, HI value, maximum offset, minimum offset, proportion of high-wear blocks, average temperature, SAMAT data, and recommended maintenance operations, summarizing and displaying the global health status.
[0128] This embodiment overlays and displays the voltage threshold curves of a specified storage block at different time points (historical scan points and the current scan point), allowing intuitive analysis of the evolution of the Vt distribution over time. This trend analysis can proactively detect signs of storage cell performance deterioration and accurately locate the most severely degraded cells, enabling early warning and preventive measures before failures occur. When the cross-layer health index of a specified storage block exceeds a preset threshold, or the average offset of the Vt curve exceeds a predetermined limit, the system automatically generates an alarm report. This automated monitoring and alarm mechanism significantly improves SSD maintenance efficiency and reduces the burden of manual monitoring.
[0129] In an exemplary embodiment, to more accurately determine the cross-layer status index of the block storage, after determining the cross-layer status index of the specified block storage based on multi-source data, the state monitoring of the solid-state drive further includes the following steps:
[0130] Obtain the voltage threshold curve stored in the specified block at the current scanning point, superimpose the comparison benchmark curve and the voltage threshold curve stored in the specified block at the current scanning point, and obtain the real-time offset of the voltage threshold curve stored in the specified block at the current scanning point and the voltage threshold distribution change rate of the voltage threshold distribution parameters; determine the voltage state index stored in the specified block according to the real-time offset of the voltage threshold curve stored in the specified block at the current scanning point and the voltage threshold distribution change rate of the voltage threshold distribution parameters; fuse the cross-layer state index stored in the specified block and the voltage state index stored in the specified block to obtain the final cross-layer state index stored in the specified block.
[0131] The real-time offset of the voltage threshold curve corresponding to the current scanning point stored in the designated block refers to the offset between the voltage threshold curve corresponding to the current scanning point stored in the designated block and the reference curve.
[0132] Voltage threshold distribution parameters refer to mathematical indicators that describe the characteristics of the voltage threshold distribution (Vth distribution), including but not limited to the distribution mean, standard deviation, distribution width, or distribution shape variation. These parameters can reflect the concentration trend and dispersion of the memory cell Vth distribution.
[0133] The voltage threshold distribution change rate refers to the rate of change of a specific Vth distribution parameter within a certain time interval. It is used to quantify the time-dependent trend of the memory cell Vth distribution stability. For example, in two consecutive scans (one hour apart), the average value of the memory cell Vth distribution changes from 1.2V to 1.3V. The rate of change is (1.3V - 1.2V) / 1.2V = 0.0833 (approximately 8.33%), indicating that the Vth distribution is gradually shifting over time.
[0134] Optionally, after determining the cross-layer state index of the designated block storage based on multi-source data, the computer device obtains the voltage threshold curve corresponding to the designated block storage at the current scanning point, superimposes the comparison benchmark curve and the voltage threshold curve corresponding to the designated block storage at the current scanning point, obtains the real-time offset of the voltage threshold curve corresponding to the current scanning point of the designated block storage, and determines the voltage threshold distribution change rate of the voltage threshold distribution parameters; performs weighted summation of the real-time offset of the voltage threshold curve corresponding to the current scanning point of the designated block storage and the voltage threshold distribution change rate of the voltage threshold distribution parameters according to preset weights to obtain the voltage state index of the designated block storage; performs weighted summation of the cross-layer state index of the designated block storage and the voltage state index of the designated block storage to obtain the final cross-layer state index of the designated block storage.
[0135] Through this embodiment, the cross-layer status index of the specified block storage and the voltage status index of the specified block storage are integrated, and the Vth distribution stability of the physical layer and the cross-layer environmental influences (such as temperature, wear, etc.) are jointly considered, so that the health index not only reflects the immediate status of the storage unit, but also comprehensively considers the long-term and multi-dimensional performance changes, providing a more comprehensive health assessment; the calculation of real-time offset and distribution change rate can accurately locate the specific reasons for the performance degradation of the storage unit, such as the drift of the voltage threshold over time or the excessive distribution width, providing direct evidence for fault diagnosis and accelerating maintenance response speed.
[0136] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0137] The embodiment of the present application also provides a state monitoring device for a solid state hard disk, such as Figure 6 Shown, including:
[0138] The cross-layer acquisition module 602 is used to obtain multi-source data of the NAND flash memory in the solid-state drive to be tested; the NAND flash memory includes multiple block storages; the multi-source data includes the current erase and write counts, average read and write bit error rates, and real-time temperatures of the multiple block storages;
[0139] The state prediction module 604 is configured to perform weighted summation of the current erase and write counts, average read and write error rates, and real-time temperatures of the block storages in the plurality of block storages to obtain a cross-layer state index of the block storages in the plurality of block storages;
[0140] The status monitoring module 606 is configured to determine the status index of the SSD to be tested according to the cross-layer status index of the block storage in the plurality of block storages, and perform status monitoring on the SSD to be tested based on the determined status index of the SSD to be tested.
[0141] In an exemplary example, the state prediction module 604 is further configured to respectively calculate a ratio between a current number of erase and write cycles of a block storage in the plurality of block storages and a maximum number of erase and write cycles tolerable for the block storage in the plurality of block storages, to obtain a first ratio corresponding to the block storage in the plurality of block storages; respectively calculate a ratio between an average read and write error rate of the block storage in the plurality of block storages and a maximum read and write error rate of the block storage in the plurality of block storages, to obtain a second ratio corresponding to the block storage in the plurality of block storages; respectively calculate a temperature difference between a real-time temperature of the block storage in the plurality of block storages and an optimal operating temperature of the block storage in the plurality of block storages, and respectively calculate a ratio between the temperature difference of the block storage in the plurality of block storages and a maximum allowable temperature of the block storage in the plurality of block storages, to obtain a third ratio corresponding to the block storage in the plurality of block storages; and perform a weighted summation of the first ratio corresponding to the block storage in the plurality of block storages, the second ratio corresponding to the block storage in the plurality of block storages, and the third ratio corresponding to the block storage in the plurality of block storages, to obtain a cross-layer state index of the block storage in the plurality of block storages.
[0142] In an illustrative example, the state prediction module 604 is further configured to perform the following weight adjustment operation on a block storage in the multiple block storages as the current block storage: when the real-time temperature of the current block storage is greater than a preset temperature threshold, adjusting the initial weight of the third ratio corresponding to the current block storage to obtain a first weight; the first weight is positively correlated with the real-time temperature of the current block storage and is less than the preset weight threshold; adjusting the initial weight of the second ratio corresponding to the current block storage according to the first weight to obtain a second weight, and adjusting the initial weight of the first ratio corresponding to the current block storage according to the first weight to obtain a third weight; the sum of the first weight, the second weight, and the third weight is equal to 1; wherein the weighted summation of the first ratio corresponding to the block storage in the multiple block storages, the second ratio corresponding to the block storage in the multiple block storages, and the third ratio corresponding to the block storage in the multiple block storages is performed according to the first weight, the second weight, and the third weight.
[0143] In an exemplary example, the state prediction module 604 is further configured to, when the real-time temperature of the current block storage is greater than a preset temperature threshold, determine the temperature difference between the real-time temperature of the current block storage and the maximum allowable temperature of the current block storage as the real-time temperature difference, and determine the ratio of the real-time temperature difference to a preset weight influence coefficient as the temperature change ratio; determine a first weight adjustment factor based on the temperature change ratio, and determine the product of the first weight adjustment factor and the initial weight of the third ratio of the current block storage as the first weight.
[0144] In an illustrative example, the state prediction module 604 is further configured to determine, based on the first weight, a current remaining ratio of the first weight relative to a specified weight state, and, based on the initial weight of the third ratio corresponding to the current block storage, determine an original remaining ratio of the initial weight of the third ratio corresponding to the current block storage relative to the specified weight state; determine the ratio between the current remaining ratio and the original remaining ratio as a second weight adjustment factor, and determine the product of the second weight adjustment factor and the initial weight of the second ratio corresponding to the current block storage as the second weight.
[0145] In an exemplary example, the device also includes a visualization module for constructing a multi-level addressing heat map according to the physical structure of the solid-state drive to be tested; the multi-level addressing heat map is a multi-level view constructed with the cross-layer health index of the block storage as the thermal value; the multi-level addressing heat map includes a channel level view, a target level view, a logical unit level view and a block level view; the channel level view displays a summary of the health status of the channel level in the solid-state drive to be tested; the target level view displays a summary of the health status of the target level in the solid-state drive to be tested; the logical unit level view displays the cross-layer health index of the block storage under the logical unit level in the solid-state drive to be tested; the block level view displays the properties of the word line under the block level in the solid-state drive to be tested; the word line is a control line for selecting a specific storage unit; a group of target block storage is located and displayed through the multi-level addressing heat map; a group of target block storage refers to at least one block storage whose cross-layer health index is greater than a preset health index.
[0146] In an exemplary example, the visualization module is also used to initiate a scan instruction packet for a specified block storage in a block-level view; the scan instruction packet includes voltage threshold scan configuration parameters; the scan instruction packet is used to instruct the solid-state drive to be tested to cyclically scan the specified block storage according to the voltage threshold scan configuration at the current scan point and return the scan data corresponding to the current scan point; based on the scan data, a voltage threshold curve corresponding to the specified block storage at the current scan point is drawn; the voltage threshold curve corresponding to the current scan point is used to analyze the read status and data reliability of the NAND flash memory corresponding to the current scan point.
[0147] In an exemplary example, the visualization module is also used to select voltage threshold curves corresponding to multiple historical scanning points of the specified block storage, superimpose and display the voltage threshold curve corresponding to the current scanning point of the specified block storage and the voltage threshold curves corresponding to multiple historical scanning points, and determine the average offset of the voltage threshold curve with reference to the baseline curve; the baseline curve refers to the voltage threshold curve generated based on the first scan of the specified block storage when the solid-state drive to be tested leaves the factory; when the cross-layer health index of the specified block storage is greater than the preset health index, or the average offset of the voltage threshold curve of the specified block storage is greater than the preset offset threshold, an alarm report is generated.
[0148] For the description of the features in the embodiment corresponding to the state monitoring device for a solid state hard disk, reference can be made to the relevant description of the embodiment corresponding to the state monitoring method for a solid state hard disk, and no further details will be given here.
[0149] An embodiment of the present application further provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps of any of the above-mentioned solid-state hard drive status monitoring method embodiments.
[0150] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned solid-state hard disk status monitoring method embodiments when running.
[0151] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0152] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned solid-state hard disk status monitoring method embodiments are implemented.
[0153] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of any of the above-mentioned solid-state hard disk status monitoring method embodiments.
[0154] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0155] The above is a detailed introduction to the state monitoring method, electronic device and storage medium of a solid-state hard disk provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A method for monitoring the status of a solid-state hard disk, characterized in that: include: Acquire multi-source data of a NAND flash memory in a solid-state drive to be tested; the NAND flash memory includes multiple block storages; the multi-source data includes the current number of erase and write times, average read and write bit error rate, and real-time temperature of the multiple block storages; Performing weighted summation on the current number of erase and write times, average read and write error rates, and real-time temperatures of the block storages in the plurality of block storages to obtain a cross-layer status index of the block storages in the plurality of block storages; The status index of the solid state drive to be tested is determined according to the cross-layer status index of the block storage in the multiple block storages, and the status of the solid state drive to be tested is monitored based on the determined status index of the solid state drive to be tested.
2. The method according to claim 1, characterized in that The weighted summing of the current number of erase and write times, the average read and write error rate, and the real-time temperature of the block storage in the plurality of block storages to obtain the cross-layer status index of the block storage in the plurality of block storages includes: Calculating the ratio between the current erasure count of the block storage in the plurality of block storages and the maximum tolerated erasure count of the block storage in the plurality of block storages respectively, to obtain a first ratio corresponding to the block storage in the plurality of block storages; Calculating the ratio of an average read and write bit error rate of a block storage in the plurality of block storages to a maximum read and write bit error rate of a block storage in the plurality of block storages to obtain a second ratio corresponding to the block storage in the plurality of block storages; Calculating the temperature difference between the real-time temperature of the block storage in the plurality of block storages and the optimal operating temperature of the block storage in the plurality of block storages, and calculating the ratio between the temperature difference of the block storage in the plurality of block storages and the maximum allowable temperature of the block storage in the plurality of block storages, respectively, to obtain a third ratio corresponding to the block storage in the plurality of block storages; A weighted sum is performed on a first ratio corresponding to a block storage in the plurality of block storages, a second ratio corresponding to a block storage in the plurality of block storages, and a third ratio corresponding to a block storage in the plurality of block storages to obtain a cross-layer state index of the block storage in the plurality of block storages.
3. The method according to claim 2, characterized in that The method further comprises: The following weight adjustment operations are performed on each of the multiple block stores as the current block store: When the real-time temperature of the current block storage is greater than a preset temperature threshold, adjusting an initial weight of the third ratio corresponding to the current block storage to obtain a first weight; the first weight is positively correlated with the real-time temperature of the current block storage and is less than the preset weight threshold; adjusting an initial weight of the second ratio corresponding to the current block storage according to the first weight to obtain a second weight, and adjusting the initial weight of the first ratio corresponding to the current block storage according to the first weight to obtain a third weight; the sum of the first weight, the second weight, and the third weight is equal to 1; The weighted summation of the first ratio corresponding to the block storage in the plurality of block storages, the second ratio corresponding to the block storage in the plurality of block storages, and the third ratio corresponding to the block storage in the plurality of block storages is performed according to the first weight, the second weight, and the third weight.
4. The method according to claim 3, characterized in that When the real-time temperature of the current block storage is greater than a preset temperature threshold, adjusting the initial weight of the third ratio corresponding to the current block storage to obtain the first weight includes: When the real-time temperature of the current block storage is greater than a preset temperature threshold, determining a temperature difference between the real-time temperature of the current block storage and a maximum allowable temperature of the current block storage as a real-time temperature difference, and determining a ratio between the real-time temperature difference and a preset weight influence coefficient as a temperature change ratio; A first weight adjustment factor is determined according to the temperature change ratio, and a product of the first weight adjustment factor and an initial weight of the third ratio stored in the current block is determined as the first weight.
5. The method according to claim 3, characterized in that The step of adjusting the initial weight of the second ratio corresponding to the current block storage according to the first weight to obtain the second weight includes: determining, based on the first weight, a current remaining proportion of the first weight relative to a specified weight state, and determining, based on an initial weight of a third ratio corresponding to the current block storage, an original remaining proportion of the initial weight of the third ratio corresponding to the current block storage relative to the specified weight state; The ratio of the current remaining ratio to the original remaining ratio is determined as a second weight adjustment factor, and the product of the second weight adjustment factor and an initial weight of the second ratio corresponding to the current block storage is determined as the second weight.
6. The method according to claim 1, characterized in that The method further comprises: According to the physical structure of the solid-state drive to be tested, a multi-level addressing heat map is constructed; the multi-level addressing heat map is a multi-level view constructed with the cross-layer health index of the block storage as the heat value; the multi-level addressing heat map includes a channel level view, a target level view, a logical unit level view and a block level view; the channel level view displays a summary of the health status of the channel level in the solid-state drive to be tested; the target level view displays a summary of the health status of the target level in the solid-state drive to be tested; the logical unit level view displays the cross-layer health index of the block storage under the logical unit level in the solid-state drive to be tested; the block level view displays the properties of the word line under the block level in the solid-state drive to be tested; the word line is a control line used to select a specific storage unit; A group of target block storages is located and displayed through the multi-level addressing heat map; the group of target block storages refers to at least one block storage whose cross-layer health index is greater than a preset health index.
7. The method according to claim 6, characterized in that The method further comprises: Initiating a scan instruction packet for a specified block storage in the block-level view; the scan instruction packet includes voltage threshold scan configuration parameters; the scan instruction packet is used to instruct the solid-state drive to cyclically scan the specified block storage according to the voltage threshold scan configuration at a current scan point and return scan data corresponding to the current scan point; Based on the scan data, a voltage threshold curve corresponding to the specified block stored at the current scan point is drawn; the voltage threshold curve corresponding to the current scan point is used to analyze the reading state and data reliability of the NAND flash memory corresponding to the current scan point.
8. The method according to claim 7, characterized in that The method further comprises: selecting voltage threshold curves corresponding to multiple historical scan points stored in the designated block, overlaying and displaying the voltage threshold curve corresponding to the current scan point stored in the designated block with the voltage threshold curves corresponding to the multiple historical scan points, and determining an average offset of the voltage threshold curves with reference to a reference curve; the reference curve being a voltage threshold curve generated based on the first scan of the designated block stored when the solid-state drive to be tested leaves the factory; When the cross-layer health index of the designated block storage is greater than the preset health index, or the average offset of the voltage threshold curve stored in the designated block storage is greater than the preset offset threshold, an alarm report is generated.
9. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the solid state hard disk status monitoring method according to any one of claims 1 to 8 when executing the computer program.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the solid-state hard disk status monitoring method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Method and device for improving data security of solid state disk and storage medium
CN112256193A
Solid state disk fault prediction method and system based on artificial intelligence
CN119179598A
Chip testing method and chip testing system
CN119805160A
Solid state disk load test method and system, electronic equipment and medium
CN119847900A
Storage system with flash memory, and storage control method
US20130262749A1
Cited By
Solid state disk detection method and device, electronic equipment, medium and product
CN120895086A
Wear balance processing method of storage device, storage device and storage medium
CN122173412A
Solid state disk grain degradation trend evaluation method based on timing behavior modeling
CN122450726A
Solid state disk grain degradation trend evaluation method based on timing behavior modeling
CN122450726B