Performance improvement method and device for all-flash memory

By optimizing the data storage structure, controller and firmware of the all-flash storage system, combined with intelligent monitoring and self-repair, the bad block problem is solved, performance improvement and stability enhancement under high loads is achieved, flash memory life is extended, and maintenance costs are reduced.

CN120371207APending Publication Date: 2025-07-25BEIJING ZHIYE TECH IND CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510452677.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

There are bad block problems in all-flash storage systems, resulting in reduced reliability, untimely manual maintenance, performance bottlenecks under high loads, and lack of intelligent management, which affects system stability and data security.

Method used

By optimizing the data storage structure, controller and firmware, improving parallelism, optimizing write strategies, combining intelligent monitoring and self-healing, load balancing and optimized scheduling are achieved, using embedded sensors and machine learning to predict bad blocks, dynamically adjust data storage, and supporting NVMe protocol.

Benefits of technology

It significantly improves the performance, stability and life of all-flash storage systems, reduces the incidence of bad blocks, shortens system downtime, reduces maintenance costs, and improves the system's self-repair capability and availability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371207A_ABST
    Figure CN120371207A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data storage, and discloses a performance improvement method for all-flash memory, which comprises the following steps: optimizing a data storage structure; optimizing the controller and firmware, including updating the firmware of the storage controller; improving the degree of parallelism, including increasing the parallelism of a flash memory channel and a chip, and supporting an NVMe protocol to optimize a writing strategy, including writing combination and reduction of writing amplification; intelligent maintenance including intelligent monitoring and self-repairing; load balancing and optimal scheduling are carried out; according to the invention, through intelligent maintenance, the performance of the system under a high-load condition is obviously improved; through intelligent maintenance, the system can detect and process potential bad blocks in the flash memory storage unit in real time, and the occurrence rate of the bad blocks is reduced; through intelligent maintenance, various performance indexes can be monitored in real time in the operation process of the storage system, and automatic repair and optimization measures are triggered according to a preset health state; and through intelligent maintenance, the self-repairing capability of the system is improved, and the frequency of manual intervention is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data storage, and specifically to a method and device for improving the performance of all-flash storage. Background Art

[0002] With the rapid development of information technology and the wide popularization of applications such as big data, cloud computing, and artificial intelligence, the demand for data storage is increasing day by day, especially the demand for high-performance and high-reliability storage systems is becoming more urgent; Flash Storage, as a high-speed, low-power, and stable storage medium, has been widely used in various storage systems, such as data centers, enterprise-level storage arrays, cloud storage services, etc. All-flash storage systems have faster read and write speeds, lower latency, higher seismic resistance, and lower energy consumption compared to traditional hard disk drives (HDDs), and have become the core components of modern data storage architectures.

[0003] However, despite the excellent performance of all-flash storage technology, it still faces some technical challenges:

[0004] First of all, the number of write operations of flash memory storage units is limited, and each storage unit has a certain erase-write life. This means that during long-term use, bad blocks may gradually appear in the flash memory or the performance may decline. The bad block problem of flash memory is one of the important factors leading to the decline in the reliability of the storage system. If these bad blocks cannot be detected and processed in time, it will affect the normal operation of the storage system and may cause data loss or system crashes;

[0005] Secondly, in the maintenance of all-flash storage systems in the prior art, it usually relies on manual monitoring and regular detection. Operations such as bad block detection, performance optimization, and fault recovery in the storage system are often carried out manually by operation and maintenance personnel; however, this traditional manual intervention method has many problems. Manual detection cannot respond to system failures in a timely manner, which may lead to bad blocks not being processed in time, thus affecting the reliability and data security of the system. Moreover, manual maintenance and operations increase labor costs and may cause the system to be down for a long time, affecting the continuity and stability of the business;

[0006] Furthermore, most of the existing all-flash storage systems lack intelligent performance management and optimization mechanisms. In high-load or high-concurrency scenarios, it is easy to encounter performance bottlenecks, resulting in problems such as increased latency and decreased throughput in the storage system. Summary of the Invention

[0007] The purpose of the present invention is to provide a method and device for improving the performance of all-flash storage to solve the problems raised in the above background art.

[0008] To achieve the above purpose, the present invention provides the following technical solutions:

[0009] Methods for improving the performance of all-flash storage, including:

[0010] Optimizing the data storage structure, including data compression and deduplication and data tiering;

[0011] Optimizing the controller and firmware, including updating the firmware of the storage controller to optimize data processing and caching algorithms and using high-performance controllers to handle more concurrent operations;

[0012] Increasing parallelism, including increasing flash channel and chip parallelism and supporting the NVMe protocol;

[0013] Optimizing the write strategy, including write merging and reducing write amplification;

[0014] Intelligent maintenance, including intelligent monitoring and self-repair;

[0015] Load balancing and optimized scheduling, including using load balancing algorithms and intelligent scheduling of input / output requests.

[0016] As a further aspect of the present invention: The data compression and deduplication improves storage efficiency and the number of input / output operations per second by reducing the amount of data stored; The data tiering stores data in layers based on the data access pattern to ensure that the most frequently accessed data is stored in the optimal flash region, while cold data is stored in slower tiers.

[0017] As a further aspect of the present invention: The increase in flash channel and chip parallelism improves the system throughput and response speed by processing multiple read / write requests in parallel.

[0018] As a further aspect of the present invention: The write merging reduces the number of writes to the flash by merging small write operations into larger blocks for writing, thereby reducing latency and improving performance; The reduction of write amplification improves performance and extends the flash lifetime by optimizing the data write strategy and using more intelligent garbage collection and block management mechanisms to reduce the write amplification phenomenon.

[0019] As a further aspect of the present invention: The intelligent monitoring analyzes the data of each flash cell through embedded sensors and algorithms, predicts which cells will soon reach the end of their lifespan, then learns the write patterns of different applications through a machine learning model, identifies which areas are most vulnerable to damage, and then dynamically switches these areas with less written areas, thereby extending the overall lifespan of the flash; The self-repair remaps the bad blocks through virtualization technology and replaces the bad blocks with healthy blocks when a bad block is detected in the flash cell, without affecting data access and performance, reducing the need for manual intervention, and improving the continuous stability of the system.

[0020] As a further solution of the present invention: The data includes the number of write cycles, wear level, and temperature.

[0021] As a further solution of the present invention: The embedded sensors include a temperature sensor, a voltage sensor, and a write cycle counter.

[0022] As a further solution of the present invention: The use of a load balancing algorithm can ensure the even distribution of load among different storage channels, controllers, and storage devices, avoid a certain component becoming a bottleneck, and improve the throughput and response speed of the overall system; The intelligent scheduling of input / output requests can optimize the processing order of requests, thereby reducing latency and improving performance.

[0023] The present invention also discloses a device for improving the performance of all-flash storage, based on any one of the methods described above, including:

[0024] A data compression and deduplication unit, which reduces the amount of data stored through compression and deduplication techniques;

[0025] A data hierarchical management unit, which divides data into different levels for storage according to the access frequency and importance of the data;

[0026] A controller upgrade and firmware optimization unit, which regularly updates the firmware of the storage controller to optimize the data processing flow and improve the efficiency of cache management, data processing algorithms, and concurrent operations;

[0027] A high-performance controller unit, which uses a high-performance controller to handle a large number of concurrent operations, especially in high-load situations, to improve the data transfer speed;

[0028] A parallelism improvement unit, which improves the parallel processing ability of the storage system by increasing the number of flash channels; increases the parallel access ability of multiple flash chips, enabling data to be read and written simultaneously on multiple flash chips, improving the bandwidth and throughput of the storage system; supports the NVMe protocol, improves the communication speed with the storage medium, and reduces the latency of the protocol layer;

[0029] A write combining unit, which combines multiple small write requests into one large write request, thereby reducing the number of write cycles of the flash medium and reducing the write amplification phenomenon;

[0030] A write amplification reduction unit, which optimizes the data writing method and reduces the write amplification phenomenon by reducing the writing of redundant data;

[0031] An intelligent detection unit, which monitors various indicators of the storage system in real time, automatically detects potential faults or performance degradation, and issues early warnings;

[0032] The self - repair unit automatically performs operations such as data migration, error repair, or module replacement according to the monitoring data and fault prediction information to restore the normal operation of the system;

[0033] The load - balancing algorithm unit dynamically adjusts the distribution of data read - write requests according to the load condition of the current storage system to ensure uniform load distribution among storage units and avoid local overload;

[0034] The intelligent scheduling unit intelligently schedules I / O requests, optimizes the scheduling order of read - write operations according to the request type, priority, and current storage state, and improves the overall response speed and throughput.

[0035] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0036] The present invention is provided with intelligent maintenance including intelligent monitoring and self - repair. Through intelligent maintenance, the performance of the system under high - load conditions is significantly improved. Through intelligent maintenance, the system can detect and process potential bad blocks in flash memory storage units in real time, reducing the incidence of bad blocks. The system can intelligently allocate the number of erase - write cycles, avoiding over - use of a single storage unit, thereby effectively reducing the wear phenomenon and further extending the reliability and durability of the flash memory. Through intelligent maintenance, various performance indicators can be monitored in real time during the operation of the storage system, and automatic repair and optimization measures can be triggered according to the preset health status, shortening the system downtime, improving the continuous availability and stability of the system, and avoiding major data loss or system crashes caused by hardware failures. Through intelligent maintenance, the self - repair ability of the system is improved, reducing labor costs and maintenance difficulties. In summary, the present invention significantly improves the performance, stability, lifespan, and reliability of the all - flash storage system and reduces the maintenance cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a structural block diagram of a method for improving the performance of all - flash storage;

[0038] Figure 2 It is a structural block diagram of a device for improving the performance of all - flash storage. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] Please refer to Figure 1 and Figure 2 In the embodiments of the present invention, the method for improving the performance of all - flash storage includes:

[0040] Optimizing the data storage structure, including data compression and deduplication and data layering;

[0041] Optimizing the controller and firmware, including updating the firmware of the storage controller to optimize data processing and caching algorithms and using high - performance controllers to handle more concurrent operations;

[0042] Improve parallelism, including increasing flash channels and chip parallelism and supporting the NVMe protocol;

[0043] Optimize write strategies, including write merging and reducing write amplification;

[0044] Intelligent maintenance, including intelligent monitoring and self - repair;

[0045] Load balancing and optimized scheduling, including using load - balancing algorithms and intelligently scheduling input / output requests.

[0046] Preferably, data compression and deduplication improve storage efficiency and the number of input / output operations per second by reducing the amount of stored data; data tiering stores data in layers based on the data access pattern, ensuring that the most frequently accessed data is stored in the optimal flash region, while cold data is stored in slower tiers.

[0047] Preferably, increasing flash channels and chip parallelism improves the system's throughput and response speed by processing multiple read - write requests in parallel.

[0048] Preferably, write merging reduces the number of writes to the flash by merging small write operations into larger blocks for writing, thereby reducing latency and improving performance; reducing write amplification improves performance and extends the flash lifespan by optimizing the data write strategy and using more intelligent garbage collection and block management mechanisms to reduce write amplification.

[0049] Preferably, intelligent monitoring analyzes data including write count, wear level, and temperature of each flash cell through embedded sensors and algorithms including temperature sensors, voltage sensors, and write - count counters, and predicts which cells will soon reach their lifespan limit. Then, by learning the write patterns of different applications through a machine - learning model, it identifies which areas are most vulnerable and dynamically switches these areas with less - written areas, thereby extending the overall lifespan of the flash.

[0050] The above algorithms include: dynamic wear leveling, which dynamically writes data to less - used flash cells according to real - time write conditions to ensure that the write counts of each flash cell are as even as possible.

[0051] Static wear leveling, when a data block is idle, the system migrates the data to less - used blocks to extend the lifespan of the entire storage system;

[0052] The machine - learning model uses supervised learning algorithms (such as support vector machines, random forests, etc.) to train historical data to establish a health prediction model. By predicting the remaining lifespan of flash cells, the machine - learning model can identify which blocks are about to fail and recommend data migration or backup before a failure occurs;

[0053] When self - repair detects a bad block in a flash memory cell, it remaps through virtualization technology and replaces the bad block with a healthy block, which does not affect data access and performance, reduces the need for manual intervention, and improves the continuous stability of the system.

[0054] Virtualization technology utilizes storage virtualization technology, and all physical storage devices (such as flash memory, SSD, HDD, etc.) can be uniformly managed and scheduled. The operating system and application programs only need to interact with the virtual storage interface without caring about the physical implementation of the underlying storage devices. This makes the management of storage more flexible and enables rapid data migration in case of hardware failures, reducing the risk of data loss.

[0055] Preferably, using a load - balancing algorithm can ensure uniform distribution of the load among different storage channels, controllers, and storage devices, avoiding a certain component becoming a bottleneck and improving the throughput and response speed of the overall system; intelligent scheduling of input / output requests can optimize the processing order of requests, thereby reducing latency and improving performance.

[0056] The present invention also discloses a performance - improvement device for all - flash storage, which includes the method of any one of the following:

[0057] A data compression and deduplication unit that reduces the amount of stored data through compression and deduplication techniques;

[0058] A data hierarchical management unit that classifies data into different levels for storage according to the access frequency and importance of the data;

[0059] A controller upgrade and firmware optimization unit that regularly updates the firmware of the storage controller to optimize the data - processing flow and improve the efficiency of cache management, data - processing algorithms, and concurrent operations;

[0060] A high - performance controller unit that uses a high - performance controller to handle a large number of concurrent operations, especially in high - load situations, to improve data - transfer speed;

[0061] A parallelism improvement unit that improves the parallel - processing ability of the storage system by increasing the number of flash channels; increases the parallel access ability of multiple flash chips, enabling data to be read and written simultaneously on multiple flash chips, improving the bandwidth and throughput of the storage system; supports the NVMe protocol to improve the communication speed with the storage medium and reduce the latency of the protocol layer;

[0062] A write - combining unit that combines multiple small write requests into one large write request, thereby reducing the number of writes to the flash medium and reducing the write - amplification phenomenon;

[0063] A write - amplification reduction unit that optimizes the data - writing method and reduces the write - amplification phenomenon by reducing the writing of redundant data;

[0064] Intelligent detection unit, which monitors various indicators of the storage system in real time, automatically detects potential faults or performance degradation, and issues early warnings;

[0065] Self-repair unit, which automatically performs operations such as data migration, error repair, or module replacement according to the monitoring data and fault prediction information to restore the normal operation of the system;

[0066] Load balancing algorithm unit, which dynamically adjusts the distribution of data read and write requests according to the current load situation of the storage system to ensure uniform load distribution among storage units and avoid local overload;

[0067] Intelligent scheduling unit, which intelligently schedules I / O requests, optimizes the scheduling order of read and write operations according to the request type, priority, and current storage status, and improves the overall response speed and throughput.

[0068] To further illustrate the technical effects of the present invention, it is verified through the following experiments:

[0069] Experimental group

[0070] The performance improvement method of the present invention;

[0071] Traditional performance improvement method.

[0072] Test equipment and environment

[0073] Storage medium: all-flash array (SSD or NVMe device);

[0074] Storage controller: high-performance storage controller supporting self-repair mechanism;

[0075] Load testing tool: perform read and write operations with high IOPS load to simulate various access modes (sequential read and write, random read and write, high-load write, etc.);

[0076] Monitoring tool: embedded sensors monitor temperature, voltage, write count, etc., and cooperate with machine learning models for data analysis and dynamic switching.

[0077] Test method

[0078] IOPS and throughput test: measure the throughput, IOPS, and latency of the experimental group and the control group through continuous random read and write tests;

[0079] Lifetime monitoring: record the erase / write count and wear condition of each flash cell, monitor whether there are bad blocks generated, and observe the impact of intelligent monitoring and self-repair mechanism on the lifetime;

[0080] System stability test: test the failure rate of the system, the repair time after bad blocks appear, and the system recovery time under long-term operation.

[0081] Experimental table

[0082]

[0083]

[0084] It can be concluded from the above table that:

[0085] The IOPS of the present invention is improved in both random write and sequential read / write operations; the random write performance is increased from 150,000 to 180,000, and the sequential read / write is increased from 300,000 to 350,000;

[0086] The throughput of the present invention is increased by about 20% (from 1.0 GB / s to 1.2 GB / s), and the latency is reduced from 2.5 ms to 1.8 ms;

[0087] The incidence rate of bad blocks of the present invention is significantly reduced, from 5% to 1%;

[0088] The bad block repair time of the present invention is significantly reduced, and the repair time is reduced from 24 hours to 4 hours;

[0089] The system failure rate of the present invention is significantly reduced, from 3% to 0.5%;

[0090] The number of times of manual intervention in the system of the present invention is reduced, from 10 times per month to 2 times;

[0091] The system flash memory life (number of erase / write cycles) of the present invention is extended, from 5,000 times to 10,000 times.

[0092] It can be analyzed and concluded from this that:

[0093] In terms of performance, the throughput and IOPS of the present invention are both increased, and the latency is significantly reduced; in terms of life and stability, the incidence rate of bad blocks is reduced, the system failure rate is reduced, the life is extended, and the repair time is greatly shortened; the maintenance cost is also effectively reduced.

[0094] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.

[0095] In addition, it should be understood that although this specification is described in terms of embodiments, not every embodiment contains only an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A method for improving the performance of all-flash storage, characterized in that, Including: Optimizing the data storage structure, including data compression and deduplication and data tiering; Optimizing the controller and firmware, including updating the firmware of the storage controller to optimize data processing and caching algorithms and using high-performance controllers to handle more concurrent operations; Increasing parallelism, including increasing flash channel and chip parallelism and supporting the NVMe protocol; Optimizing the write strategy, including write merging and reducing write amplification; Intelligent maintenance, including intelligent monitoring and self-repair; Load balancing and optimized scheduling, including using load balancing algorithms and intelligent scheduling of input / output requests.

2. The performance improvement method for all-flash storage according to claim 1, wherein The data compression and deduplication improves storage efficiency and the number of input / output operations per second by reducing the amount of stored data; the data tiering stores data in different tiers based on the data access pattern, ensuring that the most frequently accessed data is stored in the optimal flash region while cold data is stored in slower tiers.

3. The performance improvement method for all-flash storage according to claim 1, characterized in that The increase in flash channel and chip parallelism improves the system throughput and response speed by processing multiple read / write requests in parallel.

4. The performance improvement method for all-flash storage according to claim 1, wherein The write merging reduces the number of writes to the flash by merging small write operations into larger blocks for writing, thereby reducing latency and improving performance; the reduction of write amplification improves performance and extends the flash lifespan by optimizing the data write strategy and using more intelligent garbage collection and block management mechanisms to reduce the write amplification phenomenon.

5. The performance improvement method for all-flash storage according to claim 1, wherein The intelligent monitoring analyzes the data of each flash cell through embedded sensors and algorithms, predicts which cells will soon reach the end of their lifespan, then learns the write patterns of different applications through a machine learning model, identifies which areas are most vulnerable to damage, and then dynamically switches these areas with less-written areas, thereby extending the overall lifespan of the flash; the self-repair remaps through virtualization technology when a bad block is detected in a flash cell and replaces the bad block with a healthy block, without affecting data access and performance, reducing the need for manual intervention and enhancing the continuous stability of the system.

6. The method for improving the performance of all-flash storage according to claim 5, wherein The data includes the number of writes, wear level, and temperature.

7. The performance improvement method for all-flash storage according to claim 5, characterized in that, The embedded sensors include a temperature sensor, a voltage sensor, and a write count counter.

8. The performance improvement method for all-flash storage according to claim 1, wherein, The use of load balancing algorithms can ensure uniform load distribution among different storage channels, controllers, and storage devices, avoiding a bottleneck in a certain component and improving the overall system throughput and response speed; the intelligent scheduling of input / output requests can optimize the processing order of requests, thereby reducing latency and improving performance.

9. A performance improvement device for all-flash storage, characterized in that, Based on the method according to any one of claims 1-8, including: A data compression and deduplication unit that reduces the amount of stored data through compression and deduplication techniques; A data tiering management unit that divides data into different tiers for storage according to the access frequency and importance of the data; A controller upgrade and firmware optimization unit that regularly updates the firmware of the storage controller to optimize the data processing flow and enhance the efficiency of cache management, data processing algorithms, and concurrent operations; A high-performance controller unit that uses high-performance controllers to handle a large number of concurrent operations, especially in high-load situations, to improve data transfer speed; Parallelism Enhancement Unit, which enhances the parallel processing ability of the storage system by increasing the number of flash channels; increases the parallel access ability of multiple flash chips, enabling data to be read and written simultaneously on multiple flash chips, and improving the bandwidth and throughput of the storage system; supports the NVMe protocol, enhances the communication speed with the storage medium, and reduces the latency of the protocol layer; Write Combining Unit, which combines multiple small write requests into one large write request, thereby reducing the number of writes to the flash medium and reducing the write amplification phenomenon; Write Amplification Reduction Unit, which optimizes the data writing method and reduces the write amplification phenomenon by reducing the writing of redundant data; Intelligent Detection Unit, which monitors various indicators of the storage system in real time, automatically detects potential faults or performance degradation, and issues early warnings; Self-Repair Unit, which automatically performs operations such as data migration, error repair, or module replacement based on the monitoring data and fault prediction information to restore the normal operation of the system; Load Balancing Algorithm Unit, which dynamically adjusts the distribution of data read and write requests according to the current load situation of the storage system, ensures uniform load distribution among storage units, and avoids local overload; Intelligent Scheduling Unit, which intelligently schedules I / O requests, optimizes the scheduling order of read and write operations according to the request type, priority, and current storage status, and improves the overall response speed and throughput.