Storage device management method based on fault prediction and automatic recovery
By using health score-driven segmented data migration and dynamic timing control, the problem of balancing performance and reliability in solid-state drives is solved. This enables flexible adjustment of data migration strategies and improves system stability, extending module lifespan and reducing the risk of data loss.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JINDA SEMICONDUCTOR CO LTD
- Filing Date
- 2025-12-08
- Publication Date
- 2026-05-05
AI Technical Summary
Existing technologies for solid-state drive (SSD) data storage management struggle to balance performance and reliability, especially when dealing with partially aged or high-risk data. They cannot flexibly adjust the data migration and writing rhythm, resulting in low migration efficiency or excessive interference with normal operations.
Through health assessment and status switching steps, the health status of the storage module is monitored in real time. Segmented data migration and operation timing control are adopted to dynamically adjust data migration and write operations, including health score-driven segmented migration, real-time assessment and permanent isolation mechanisms, to ensure module reliability and extend lifespan.
This technology enables real-time adjustment of data migration strategies within solid-state drives (SSDs), reducing the impact on normal data writing, improving system performance and stability, avoiding excessive migration or premature isolation, extending module lifespan, and enhancing overall system reliability and data security.
Smart Images

Figure CN121979705A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to solid-state drive (SSD) data management technology, and more particularly to a storage device management method based on fault prediction and automatic recovery. This method aims to improve the reliability and lifespan of each storage module within an SSD by using dynamic health score monitoring, segmented data migration, real-time logical mapping updates, and automatic isolation control to reduce the risk of data loss, improve write performance, and ensure stable operation of the storage system. Background Technology
[0002] In the field of solid-state drive (SSD) data storage management, systems need to continuously monitor the health status of each storage unit and take corresponding measures to address potential failures or performance degradation. However, existing technologies do not adequately consider the dynamic changes in data distribution, local aging, and write load within storage units, and have limited overall coordination capabilities for data migration strategies and write operations, potentially leading to a trade-off between performance and reliability. Furthermore, existing methods often rely on static or preset conditions for health assessment, lacking strategies for automatically adjusting to the real-time health status of storage units. This is especially problematic when dealing with locally aged or high-risk data, as they cannot flexibly adjust data migration and write rhythms, potentially resulting in inefficient migration or excessive interference with normal operations. Therefore, when addressing dynamic changes in storage unit health status, localized data aging, and the need for segmented migration, existing technologies still face the challenge of balancing reliability and efficiency. Summary of the Invention
[0003] One objective of this invention is to provide a storage device management method based on fault prediction and automatic recovery, which can assess the health status of each storage module in real time and dynamically adjust data migration and writing operations according to the module health score. Through segmented data migration and operation timing control, the method can improve module reliability and extend service life while reducing the impact on normal data writing and improving the overall performance and stability of the storage system.
[0004] To achieve the above objectives, the present invention provides a storage device management method based on fault prediction and automatic recovery, applicable to solid-state drives, comprising: (A) a health assessment and state switching step: the controller of the computer system continuously monitors the trend changes of multiple storage behavior indicators, calculates a health score for multiple storage modules, and automatically switches the multiple storage modules between three states: a normal mode, a degraded mode, and an isolation mode based on the health score; wherein, the multiple storage behaviors include a write error event frequency, an erase / write latency offset, and a retention error increase, and the multiple storage modules refer to storage units composed of at least one flash memory die or its logical sub-region; wherein, the health score is dynamically calculated based on the distribution of effective data, local aging degree, and changes in write load within the multiple storage modules, so that data migration can affect the health score in real time; (B) a segmented migration step: when any of the storage modules enters the degraded mode, actions b1 to b4 are executed; b1. The controller, based on logical address continuity, physical page clustering, or block correspondence, divides the existing valid data within the storage module entering the degraded mode into multiple independently movable data segments; wherein each data segment corresponds to at least one logical address range and at least one physical page group; b2. For each data segment, the controller determines the risk index of the multiple data segments based on the local aging degree of the physical page group to which it belongs, and sorts them from high to low risk to form a segmented migration sequence; b3. The controller migrates the data segments sequentially according to the segmented migration sequence, starting with the data segment with the highest risk, to any other storage module belonging to the normal mode; b4.After each data segment migration is completed, the controller immediately updates the local logical mapping table corresponding to that data segment, making the data segment effective in the logical location of the new module in real time, without waiting for all the other unmigrated data segments to be migrated; (C) Real-time migration evaluation step: When the abnormal storage module is at the beginning of the migration of the multiple data segments, new data writing is paused, and after each data segment migration, the controller recalculates the health score of the abnormal storage module in real time; if the health score rises above a first preset threshold, the migration of the remaining multiple data segments and the synchronous execution of new data writing continue. The remaining data segments are migrated at their original speed, while the timing of new data write operations is adjusted to twice the original timing to reduce instantaneous write speed. If the health score rises above a second preset threshold, the previously abnormal storage module is switched to normal mode, the migration of the remaining data segments is stopped, and the new data write operation timing is restored to normal write speed. The second preset threshold is higher than the first preset threshold. The remaining data segments after migration stops remain in the original storage module and are subsequently included in the relocation or isolation determination process based on the health score during subsequent state changes. (D) Permanent isolation steps: If the health score of any storage module is initially determined to be in the isolation mode, write operations are permanently disabled and the relevant logical mapping is replaced by another storage module in the normal mode; or if the health score of a storage module is initially determined to be in the degraded mode, and before reaching the first preset threshold, if the health score after moving N data segments is still lower than or equal to the health scores of (N-5) moved segments, the controller switches the storage module to the isolation mode, permanently disables write operations, and replaces the relevant logical mapping by another storage module in the normal mode; where N is an integer greater than or equal to 5.
[0005] In summary, the storage device management method provided by this invention, through health score-driven segmented data migration and dynamic job timing control, can not only adjust the data processing strategy of abnormal modules in real time to avoid data loss or premature module isolation, but also maintain the high-efficiency write performance of normal modules, further improving the overall reliability, lifespan and system performance of solid-state drives, and reducing data access delays and system blockages caused by module aging or local failures. Attached Figure Description
[0006] Figure 1 This is a flowchart of a preferred embodiment of the present invention.
[0007] Explanation of reference numerals in the attached figures: (A)~(D) - steps. Detailed Implementation
[0008] To enable those skilled in the art to clearly understand the technical content of the present invention, the following embodiments are provided in conjunction with the accompanying drawings to further illustrate the storage device management method based on fault prediction and automatic recovery provided by the present invention.
[0009] Please see Figure 1 This is a flowchart of a preferred embodiment of the present invention. This embodiment provides a storage device management method based on fault prediction and automatic recovery, applicable to solid-state drives (SSDs). First, the (A) health assessment and state switching step is executed. In this embodiment, the controller in the computer system configured with SSDs continuously monitors multiple storage behavior indicators for each storage module and calculates a module health score based on the changing trends of these indicators. The storage behavior indicators include write error event frequency, erase / write latency offset, and retention error increase. Write error event frequency refers to the number of write failures occurring per unit time; the controller can count these events using a counter. Erase / write latency offset represents the deviation of the actual erase / write time from the nominal specification; a large deviation may indicate partial aging of the flash memory or resource constraints on the flash controller. Retention error increase refers to the increase in tolerable errors in the retention area, reflecting a weakening of the module's fault tolerance. Each storage module can be composed of a single flash memory die or formed by logical sub-regions, serving as an independent storage unit. A logical sub-region refers to multiple independently manageable page groups divided on a physical die; each page group can have independent erase / write counts and error records. Based on the calculated health score, the controller can automatically switch each storage module to one of three modes: normal mode, degraded mode, or isolation mode. Normal mode indicates that the storage module is in good health and data can be read and written freely; degraded mode indicates that the storage module is in slightly poor health, and the controller may limit the amount of simultaneous writes, adjust the data migration strategy, or extend the write interval to prevent further deterioration; isolation mode indicates that the storage module is in severely degraded health, and the controller will prohibit writes and transfer the original module's logical mapping to other healthy modules.
[0010] Furthermore, the health score comprehensively considers the distribution of valid data within the storage module, the degree of local aging, and changes in write load, allowing data migration or write operations to affect the health score in real time. For example, if a segment of the module is subjected to high-frequency writes for an extended period and its erase / write cycles are approaching the upper limit of its specifications, the controller will include that segment in the health score calculation, reflecting the overall risk of the module. If the data is evenly distributed and the load is low, the health score will remain at a high value, and the module can operate normally. The controller can use sliding windows, weighted averages, or other statistical methods to integrate various indicators and generate a single health score to provide a basis for automatic switching decisions. Through this mechanism, the module's health status can be reflected in the control strategy in real time, providing a foundation for subsequent segmented data migration or isolation determination.
[0011] Next, step (B) is performed: when any storage module enters degrade mode, the controller performs a segmented migration of the existing valid data within that module. Valid data refers to data that is still referenced by the system or has practical use value, rather than deleted or marked as free areas.
[0012] First, in detailed action b1, the controller divides the valid data within the degraded module into multiple independently movable data segments based on logical address continuity, physical page clustering, and block correspondence. Each data segment corresponds to at least one logical address range and at least one physical page group, allowing each data segment to be moved independently without affecting access to other data segments. For example, if a module contains 100 consecutive logical pages, with every 10 pages forming a physical page group, then 10 data segments can be formed, each of which can be independently moved to other healthy modules.
[0013] Next, in detailed action b2, the controller assesses risk indicators for each data segment. Risk indicators are calculated based on the local aging degree, cumulative write count, and error event frequency of the physical page group to which the data segment belongs. Local aging degree can be determined by comparing the page group's erase / write cycle count with a preset durability; a page group nearing its durability limit carries a higher risk. All data segments are sorted from highest to lowest risk, forming a segmented migration sequence to ensure that high-risk data segments are moved first. For example, if data segments A, B, and C have local aging degrees of 90%, 75%, and 60% respectively, the migration order would be A → B → C.
[0014] Next, in detailed action b3, the controller sequentially moves data segments to other storage modules in normal mode according to the segmented migration sequence. The migration can be completed via DMA or built-in instructions of the flash memory controller, ensuring data consistency and integrity during the migration process. Following this, in detailed action b4, after each data segment migration is completed, the controller immediately updates the corresponding local logical mapping table, making the data effective in its logical location in the new module in real time, without waiting for other data segments to be migrated. For example, after the first data segment migration is completed, the system can immediately read and write that segment of data from the new module; subsequent data segments remain in the original module or are being migrated, without affecting existing access. Therefore, through this segmented migration mechanism, the controller can precisely control the data migration order and speed, reducing the impact on normal data write operations, and providing a basis for real-time adjustments when the health status of degraded modules improves or deteriorates. Compared to traditional full-module migration, this strategy can effectively improve migration efficiency and reduce the impact on system performance.
[0015] Following steps (A) and (B), step (C) – real-time migration assessment – is executed. When a faulty storage module begins data segment migration, the controller first suspends new data write operations on that module. This design primarily aims to prevent new write loads from further exacerbating the module's local aging and to ensure data integrity during the initial migration phase. During the pause, the module only migrates existing data segments and does not bear new write pressure. After each data segment migration is completed, the controller immediately recalculates the health score of the faulty module. The health score is based on a combination of factors, including the module's error rate, changes in recent erase / write cycles, reported internal device health parameters, and changes in heat or pressure caused by the migration activity. Real-time calculation after migration accurately reflects whether the module shows signs of stabilization after partial data unloading. Specifically, if the health score rises above a first preset threshold, the controller initiates a synchronization operation, allowing the migration of the remaining data segments to proceed simultaneously with new data write operations. In this state, the migration timing of the remaining data segments remains unchanged, but the timing of new data write jobs is adjusted to twice the original cycle to reduce the instantaneous write speed. For example, if the original write interval is 1ms, it will be extended to 2ms in this stage to prevent the module from being subjected to excessive load immediately after stabilizing and deteriorating again.
[0016] Furthermore, if the health score continues to improve and rises above the second preset threshold (which is higher than the first threshold), the controller will switch the module from abnormal mode back to normal mode. At this point, the remaining unmigrated data segments will stop migrating and remain in the original storage module. Since the module has returned to a normal operating state, the data does not need to be forcibly migrated to other modules. Subsequently, the new data write timing returns to normal speed and is no longer subject to speed reduction control.
[0017] As for the remaining data segments retained after the migration is stopped, the controller will reassess their health scores during subsequent evaluations to determine whether they need to be reinstated into the migration process or returned to isolation. For example, if the module experiences anomalies again during subsequent operation and its health score declines, the previously retained data segments may be reinstated on the migration list; conversely, if the health remains stable, the retained data segments can continue to serve within the original module. Therefore, this design allows for a highly flexible data migration strategy, preventing over-migration and enabling modules to quickly resume normal operation after stabilization, effectively improving the overall system's durability and performance maintenance capabilities.
[0018] Following steps (A), (B), and (C), step (D) permanent isolation is executed. When the controller determines during the initial evaluation phase that a storage module's health score is in isolation mode, that module is considered unable to safely undertake any data write operations. In this case, the controller immediately permanently disables writes to that module and reassigns the entire logical address range that originally belonged to that module to other storage modules that are still in normal mode. This action also includes updating the logical mapping table so that all subsequent data writes and reads no longer touch the abnormal module, thus preventing data from falling into its unreliable area.
[0019] If the storage module is classified as degraded in the initial assessment, the controller will migrate the data in the module segment by segment according to the aforementioned segmented migration process. During this degrade period, the system observes its recovery trend through the health score after each migration. Specifically, after the Nth data segment is migrated, the controller will obtain the health score at that point in time; at the same time, the controller will also backtrack to the health score when the (N-5)th data segment migration was completed and compare the changes between the two.
[0020] If, after moving N data segments, the module's health score remains lower than or equal to the health score at the time of completing the (N-5) data segment relocation, it indicates that the module has not shown any improvement after multiple consecutive relocations and may even continue to deteriorate. This situation means that the module's localized aging is not caused by a single pressure or instantaneous load, but may have comprehensive, irreversible, or unstable deterioration characteristics. Therefore, continuing to relocate the module will not only fail to help with recovery but may also delay the isolation time and increase data risk.
[0021] Based on the above determination, the controller will switch the storage module from degraded mode to isolated mode at that point in time, and also apply a permanent write restriction. At the same time, all logical address mappings belonging to this module will be immediately transferred to other normal mode storage modules, so that this module will no longer participate in subsequent data writing and migration operations.
[0022] Here, parameter N is an integer greater than or equal to 5, used to ensure that the health score comparison has a sufficient observation span to avoid fluctuations caused by a single or very short-term data transfer misleading the isolation decision. By comparing the health scores at two time points, N and (N-5), the overall trend of the module after multiple consecutive data segment transfers can be observed, distinguishing whether the module's health status shows a gradual recovery, remains stable, or continues to decline. The technical significance of this comparison method is that during the data transfer process, the module may temporarily reduce its load due to partial unloading, but a short-term increase in the health score does not mean that the module has fully recovered as a whole; conversely, a decrease may also create a false recovery due to a momentary drop in write load. By setting the observation span between N and (N-5), the interference of such short-term fluctuations on the judgment can be reduced, ensuring that the isolation decision reflects the module's long-term trend and true health status. For example, in a real-world scenario, if N=7, the system compares the health score after moving the 7th segment with the health score after moving the 2nd segment. If no significant improvement is observed, it indicates that the module's health status has not improved over a long period, potentially indicating a comprehensive, irreversible, or unstable aging phenomenon, and therefore it should be permanently isolated. This mechanism effectively distinguishes between "recoverable short-term anomalies" and "irreversible deep degradation," preventing misjudgments of a module's continued usability due to a single, momentary rebound. This improves data security, reduces system risk, and extends the reliable operating time of other healthy modules. Furthermore, this method is applicable to storage systems of different sizes and configurations, providing a consistent and predictable isolation strategy, enabling the system to maintain stability and data integrity under high load or long-term use.
[0023] In summary, the storage device management method based on fault prediction and automatic recovery provided by this invention accurately reflects the health status of each storage module through real-time health score calculation and multi-level mode switching. It also provides a dynamic, segmented data migration and synchronous write strategy, effectively reducing the performance impact caused by full module relocation. Through segmented migration and risk prioritization, the method prioritizes high-risk data segments, ensuring data security and access consistency while maintaining system write performance. Utilizing health score threshold design and real-time monitoring, degraded modules can automatically adjust their data migration pace and new data write sequence based on the recovery rate, balancing module lifespan extension and system performance maintenance. When a module exhibits irreversible or continuously deteriorating trends, it can be permanently isolated in a timely manner to prevent data loss or errors from spreading to normal modules, further improving the overall system reliability. Segmented relocation, real-time health score updates, and isolation judgment mechanisms enable the storage system to react quickly and accurately when faced with local aging, instantaneous high load, or degradation caused by long-term use. This avoids excessive relocation or premature isolation, reduces resource waste, and extends module lifespan, thereby improving the overall durability, data security, and operational stability of the storage system. It provides a reliable and predictable storage operation mode for multi-module, high-load environments.
[0024] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of the present invention. Therefore, all equivalent changes and modifications made without departing from the scope of the present invention should be covered within the protection scope of the present invention.
Claims
1. A storage device management method based on fault prediction and automatic recovery, applicable to solid-state drives, characterized in that, Include: (A) Health Assessment and State Switching Steps: The computer system controller continuously monitors the trend changes of multiple storage behavior indicators, calculates a health score for multiple storage modules, and automatically switches the multiple storage modules between three states: a normal mode, a degraded mode, and an isolation mode based on the health score; wherein, the multiple storage behaviors include a write error event frequency, an erase / write latency offset, and a retention error increase; the multiple storage modules refer to storage units composed of at least one flash memory die or its logical sub-region; wherein, the health score is dynamically calculated based on the distribution of valid data within the multiple storage modules, the degree of local aging, and changes in write load, so that data migration can affect the health score in real time; (B) Segmented migration steps: When any of the storage modules enters the degradation mode, actions b1 to b4 are executed; b1. The controller divides the existing valid data in the storage module that has entered the degradation mode into multiple data segments that can be moved independently, based on logical address continuity, physical page clustering, or block correspondence; wherein each data segment corresponds to at least one logical address range and at least one physical page group. b2. For each data segment, the controller determines the risk index of the multiple data segments based on the local aging degree of the physical page group to which it belongs, and sorts them from high to low risk to form a segmented migration sequence; b3. The controller migrates the data segment in sequence according to the segmented migration sequence, starting with the data segment with the highest risk and moving sequentially to any other storage module belonging to the normal mode; b4. After each data segment is migrated, the controller immediately updates the local logical mapping table corresponding to the data segment, so that the data segment takes effect in the logical location of the new module in real time, without waiting for all other data segments that have not yet been migrated to be moved. (C) Real-time Migration Assessment Steps: When the abnormal storage module is at the beginning of the migration of multiple data segments, new data writing is paused. After each data segment migration, the controller recalculates the health score of the abnormal storage module in real time. If the health score rises above a first preset threshold, the migration of the remaining multiple data segments continues and new data writing is performed synchronously. The migration timing of the remaining data segments remains at the original speed, while the timing of new data writing is adjusted to twice the original timing to reduce instantaneous write speed. If the health score rises above a second preset threshold, the previously abnormal storage module is switched to normal mode, the migration of the remaining multiple data segments is stopped, and the new data writing timing is restored to normal write speed. The second preset threshold is higher than the first preset threshold. The remaining data segments after the migration stops remain in the original storage module and are re-included in the subsequent relocation or isolation determination process based on the health score during subsequent state changes. (D) Permanent Isolation Step: If the health score of any storage module is initially determined to be in the isolation mode, write operations are permanently prohibited and the relevant logic mapping is replaced to other storage modules in the normal mode; or if the health score of the storage module is initially determined to be in the degraded mode, and before the first preset threshold is reached, when the health score after moving N data segments is still lower than or equal to the health scores of (N-5) data segments, the controller switches the storage module to the isolation mode, permanently prohibits write operations, and replaces the relevant logic mapping to other storage modules in the normal mode; where N is an integer greater than or equal to 5.