Memory system, operating method of memory system, and controller

By introducing a data backup mechanism of retained blocks and controllers in the storage system, the problem of low efficiency of failed storage page analysis is solved, efficient and accurate analysis of failure causes is achieved, and the reliability and data integrity of the storage system are ensured.

CN120295550APending Publication Date: 2025-07-11YANGTZE MEMORY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410048170.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-11
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When the prior art detects a failed storage page, it cannot effectively analyze the causes of it, resulting in the risk remaining in the storage system, affecting the reliability and efficiency of data storage.

Method used

It provides a storage system, including multiple storage blocks and at least one reserved block. When the controller detects a failed storage page, its data is backed up to the backup storage page in the reserved block, and ensures data integrity through LDPC decoding, MCRC detection, descrambling, etc., and can subsequently analyze the cause of failure from the backup storage page.

Benefits of technology

Improve the efficiency and accuracy of failed storage page analysis, avoid risk legacy, and ensure the reliability and data integrity of the storage system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295550A_ABST
    Figure CN120295550A_ABST
Patent Text Reader

Abstract

The invention provides a storage system, an operation method of the storage system and a controller, and belongs to the technical field of storage. In the scheme provided by the invention, the memory comprises a plurality of memory blocks and at least one reserved block. The controller can obtain data in a failure storage page when the failure storage page is detected in the storage pages included in the plurality of storage blocks, and backs up the data in the failure storage page to one backup storage page in the at least one reserved block. Therefore, when the failure analysis needs to be carried out on the failure storage page, the data can be directly read from the backup storage page and the failure analysis is carried out, so that the risk is prevented from being left in the storage system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of storage technologies, and particularly to a storage system, an operation method of a storage system, and a controller. Background Art

[0002] A NAND flash memory generally includes a plurality of storage blocks, and each storage block includes a plurality of storage pages. Among them, the smallest unit for data reading and writing is a storage page, and the smallest unit for data erasing is a storage block.

[0003] When the controller of the storage system detects a data read / write error, that is, when a fail page is detected, the controller can mark the storage block to which the fail page belongs as a bad block (BB). Summary of the Invention

[0004] This application provides a storage system, an operation method of a storage system, and a controller, which can analyze the cause of the generation of fail pages in a memory. The technical solutions are as follows:

[0005] In a first aspect, a storage system is provided. The storage system includes a memory and a controller coupled to the memory. The memory includes a plurality of storage blocks and at least one reserved block. Each of the plurality of storage blocks includes a plurality of storage pages, and each of the at least one reserved blocks includes a plurality of backup storage pages. The controller is configured to:

[0006] In response to detecting a fail page in the storage pages included in the plurality of storage blocks, obtain the data in the fail page;

[0007] Back up the data in the fail page to a backup storage page in the at least one reserved block.

[0008] Optionally, the fail page is a storage page that generates a data read error. The controller is configured to: recover the data in the fail page;

[0009] Back up the recovered data in the fail page to a backup storage page in the at least one reserved block.

[0010] Optionally, the controller is configured to:

[0011] Read other data in the redundant arrays of independent disks (RAID) stripe to which the fail page belongs, except for the data in the fail page;

[0012] Restore the data in the failed storage page based on the other data.

[0013] Optionally, the controller is configured to:

[0014] Perform low density parity check (LDPC) decoding, media cyclic redundancy check (MCRC) detection, de-scrambling, and exclusive OR (XOR) processing on the other data in sequence to restore the data in the failed storage page;

[0015] Perform scrambling, generate MCRC and LDPC encoding on the data in the restored failed storage page in sequence, and program it into a backup storage page in the at least one reserved block.

[0016] Optionally, the failed storage page is a storage page that generates a data write error; the controller is configured to:

[0017] Obtain the data to be written to the failed storage page from the memory of the controller;

[0018] Back up the data to be written to the failed storage page to a backup storage page in the at least one reserved block.

[0019] Optionally, the controller is further configured to:

[0020] Mark the storage block to which the failed storage page belongs as a grown bad block (GBB);

[0021] After garbage collection of the memory, read the original data in the failed storage page;

[0022] Compare the original data with the data in a backup storage page in the at least one reserved block.

[0023] Optionally, the controller is configured to:

[0024] Perform LDPC decoding, MCRC detection, and de-scrambling on the data in a backup storage page in the at least one reserved block in sequence;

[0025] Generate a scrambling seed based on the address of the failed storage page;

[0026] Scramble the data in the de-scrambled backup storage page based on the scrambling seed;

[0027] Perform MCRC generation and LDPC encoding on the scrambled data in sequence, and write it into the cache of the controller;

[0028] Compare the original data with the data stored in the cache.

[0029] Optionally, the controller is configured to:

[0030] Back up the data in the failed storage page to a backup storage page in one of the at least one reserved block in single level cell (SLC) mode or triple level cell (TLC) mode.

[0031] Optionally, the at least one reserved block belongs to a target stripe, and the target stripe includes at least one factor bad block (FBB).

[0032] In a second aspect, an operation method of a storage system is provided. The storage system includes a controller and a memory. The memory includes a plurality of storage blocks and at least one reserved block. Each storage block in the plurality of storage blocks includes a plurality of storage pages, and each reserved block in the at least one reserved block includes a plurality of backup storage pages. The method includes:

[0033] In response to detecting a failed storage page in the storage pages included in the plurality of storage blocks, the controller obtains the data in the failed storage page;

[0034] The controller backs up the data in the failed storage page to a backup storage page in one of the at least one reserved block.

[0035] Optionally, the failed storage page is a storage page that generates a data read error. The controller backing up the data in the failed storage page to a backup storage page in one of the at least one reserved block includes: the controller recovers the data in the failed storage page;

[0036] The controller backs up the recovered data in the failed storage page to a backup storage page in one of the at least one reserved block.

[0037] Optionally, the controller recovering the data in the failed storage page includes:

[0038] The controller reads other data in the RAID stripe to which the failed storage page belongs except the data in the failed storage page;

[0039] The controller recovers the data in the failed storage page based on the other data.

[0040] Optionally, the controller recovering the data in the failed storage page based on the other data includes:

[0041] The controller sequentially performs LDPC decoding, MCRC detection, descrambling, and exclusive OR processing on the other data to recover the data in the failed storage page;

[0042] The controller backs up the data in the recovered failed storage page to a backup storage page in at least one of the reserved blocks, including:

[0043] The controller sequentially performs scrambling, MCRC generation, and LDPC encoding on the data in the recovered failed storage page, and programs it into a backup storage page in at least one of the reserved blocks.

[0044] Optionally, the failed storage page is a storage page with a data write error; the controller backs up the data in the failed storage page to a backup storage page in at least one of the reserved blocks, including: the controller obtains the data to be written to the failed storage page from the memory of the controller;

[0045] The controller backs up the data to be written to the failed storage page to a backup storage page in at least one of the reserved blocks.

[0046] Optionally, the method further includes:

[0047] The controller marks the storage block to which the failed storage page belongs as GBB;

[0048] After garbage collection of the memory, the controller reads the original data in the failed storage page;

[0049] Compare the original data with the data in a backup storage page in at least one of the reserved blocks.

[0050] Optionally, the comparing the original data with the data in a backup storage page in at least one of the reserved blocks includes:

[0051] The controller sequentially performs LDPC decoding, MCRC detection, and descrambling on the data in a backup storage page in at least one of the reserved blocks;

[0052] The controller generates a scrambling seed based on the address of the failed storage page;

[0053] The controller scrambles the data in the descrambled backup storage page based on the scrambling seed;

[0054] The controller sequentially performs MCRC generation and LDPC encoding on the scrambled data, and writes it into the cache of the controller;

[0055] Compare the original data with the data stored in the cache.

[0056] Optionally, the controller backs up the data in the failed storage page to a backup storage page in one of the at least one reserved block, including:

[0057] The controller backs up the data in the failed storage page to a backup storage page in one of the at least one reserved block in SLC mode or TLC mode.

[0058] In a third aspect, a controller is provided, which includes: a processor and a memory; the memory is used to store computer instructions, and the processor is used to execute the computer instructions to implement the operation method of the memory provided in the second aspect above.

[0059] The technical solution provided by this application can at least include the following beneficial effects:

[0060] This application provides a storage system, an operation method of the storage system, and a controller. In the storage system provided by this application, the memory includes a plurality of storage blocks and at least one reserved block. When the controller of the storage system detects a failed storage page in the storage pages included in the plurality of storage blocks, it can obtain the data in the failed storage page and back up the data in the failed storage page to a backup storage page in one of the at least one reserved block. Thus, when it is necessary to perform a failure analysis on the failed storage page, the data can be directly read from the backup storage page and the failure analysis can be performed to avoid leaving risks in the storage system. Description of the Drawings

[0061] In order to more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0062] Figure 1 is a schematic structural diagram of an electronic device provided by an embodiment of this application;

[0063] Figure 2 is a schematic structural diagram of a memory card provided by an embodiment of this application;

[0064] Figure 3 is a schematic structural diagram of a solid-state drive provided by an embodiment of this application;

[0065] Figure 4 is a schematic diagram of the data stored in a memory provided by an embodiment of this application;

[0066] Figure 5 It is a schematic diagram of data stored in another memory provided by an embodiment of the present application;

[0067] Figure 6 It is a schematic structural diagram of a memory provided by an embodiment of the present application;

[0068] Figure 7 It is a schematic flow diagram of an operation method of a storage system provided by an embodiment of the present application;

[0069] Figure 8 It is a schematic diagram of a data recovery and data backup process provided by an embodiment of the present application;

[0070] Figure 9 It is a schematic flow diagram of another operation method of a storage system provided by an embodiment of the present application;

[0071] Figure 10 It is a schematic flow diagram of yet another operation method of a storage system provided by an embodiment of the present application;

[0072] Figure 11 It is a schematic flow diagram of a data comparison method provided by an embodiment of the present application;

[0073] Figure 12 It is a schematic diagram of a data recovery and data comparison analysis process provided by an embodiment of the present application;

[0074] Figure 13 It is a schematic structural diagram of a controller provided by an embodiment of the present application. Detailed implementation manners

[0075] The following further describes the embodiments of the present application in detail with reference to the accompanying drawings.

[0076] The solution provided by the embodiment of the present application can be applied to an electronic device. The electronic device can be a mobile terminal, a desktop computer, a laptop computer, a tablet computer, a vehicle computer, a game console, a printer, a positioning device, a wearable electronic device, a smart sensor, a virtual reality (VR) device, an augmented reality (AR) device, or any other suitable electronic device having a memory.

[0077] Figure 1 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application, as Figure 1As shown, the electronic device includes a storage system 10 and a host 20. Among them, the host 20 can be the central processing unit (CPU) or the system on chip (SOC) of the electronic device. The host 20 is used to send data to the storage system 10 for storage or read data from the storage system 10.

[0078] An embodiment of this application also provides a storage system. Continuing to refer to Figure 1 , the storage system 10 includes: one or more memories 100, and a controller 200. For example Figure 1 shows a plurality of memories 100. Among them, each memory 100 can be a three-dimensional (3D) memory, such as 3D NAND flash. Each memory 100 may include at least one storage plane, each storage plane includes a plurality of storage blocks, and each storage block includes a plurality of storage pages. The controller 200 is respectively connected to the memory 100 and the host 20. The controller 200 is used to manage the data stored in the memory 100 and to communicate with the host 20.

[0079] In the embodiment of this application, the controller 200 can be configured to control the operations performed by the memory 100. Such as read, erase, and program operations. The controller 200 can also be configured to manage various functions regarding the data stored in or to be stored in the memory 100, including but not limited to bad block management, garbage collection (GC), logical address to physical address conversion, and wear leveling. Optionally, the controller 200 can also be configured to process the error correcting code (ECC) for the data read from or written to the memory 100. The controller 200 can also perform any other suitable functions. For example, formatting the memory 100.

[0080] The controller 200 can also communicate with external devices according to specific communication protocols. By way of example, the controller 200 can communicate with external devices through at least one of various interface protocols. The interface protocols can include the universal serial bus (USB) protocol, the Multi-Media card (MMC) protocol, the peripheral component interconnect (PCI) protocol, the PCI Express (PCI-E) protocol, the advanced technology attachment (ATA) protocol, the serial ATA protocol, the parallel ATA protocol, the small computer system interface (SCSI) protocol, the enhanced small drive interface (ESDI) protocol, the integrated drive electronics (IDE) protocol, and the fire wire protocol, etc.

[0081] In some embodiments, the controller 200, and one or more memories 100 can be integrated into various types of storage devices.

[0082] As an example, as Figure 2 shown, the controller 200 and a single memory 100 can be integrated into a memory card 300. The memory card 300 can include a personal computer memory card international association (PCMCIA) card, a compact flash (CF) card, a smart media (SM) card, a memory stick, a multi-media card (MMC), a secure digital (SD) card, and a universal flash storage (UFS), etc. As Figure 2 shown, the memory card 300 can also include a connector 310 for coupling with the host 20.

[0083] As another example, as Figure 3As shown, the controller 200 and multiple memories 100 can be integrated into a solid state disk (SSD) 400. The solid state drive 400 can also include a connector 410 for coupling with a host 20. Among them, the storage capacity and / or operating speed of the solid state drive 400 is greater than that of the memory card 300.

[0084] To improve the reliability of data storage, storage systems typically use RAID striping technology to protect and recover data. Figure 4 It is a schematic diagram of data stored in a memory provided by an embodiment of the present application. Figure 4 Each row in it can be a RAID stripe, and each column can be a storage plane of the memory. As Figure 4 shown, each RAID stripe can include multiple data parts and a parity part. The multiple data parts and a parity part can be located in different storage planes of the memory. Among them, each data part is used to store a piece of data DATA. The parity part is used to store the parity data obtained by performing an exclusive OR on the multiple pieces of data DATA stored in the multiple data parts. The parity data can be a parity block. Optionally, a piece of data DATA stored in each data part can also be called a code word, and its size can be 2 kilobytes (K) or 4K bytes.

[0085] Optionally, referring to Figure 5 , the multiple data parts and the parity part included in each RAID stripe can also be distributed in different memories. For example, Figure 5 in the scenario shown, the multiple data parts and the parity part in each RAID stripe can be distributed in n different memories. Among them, n can be an integer greater than 1.

[0086] When a data read error occurs in any storage page in the memory, that is, when a fail page appears, the controller can read the data stored in the other storage pages in the RAID stripe to which the fail page belongs, except for the fail page, and can recover the data of the fail page based on the exclusive OR relationship between the data. Among them, the data read error can refer to an uncorrectable ECC (UECC) error. It can be understood that since the size of each storage page in the memory can be greater than the size of the code word, the RAID stripe to which the fail page belongs can include multiple ones, that is, a fail page can belong to multiple RAID stripes. For example, assuming that a storage page can store 4 code words, then a fail page can belong to 4 RAID stripes.

[0087] After detecting a failed storage page, the controller can also mark the storage block to which the failed storage page belongs as a GBB and trigger the garbage collection process. In this garbage collection process, the controller can rewrite the valid data within the RAID stripe to which the GBB belongs to a new storage location, and can erase the other storage blocks in the RAID stripe to which the GBB belongs except the GBB, that is, the GBB does not participate in the garbage collection process.

[0088] Since the above garbage collection process erases the data in the other storage blocks in the RAID stripe to which the GBB belongs except the GBB, the XOR relationship between the data in the RAID stripe to which the GBB belongs is destroyed, and further the data in the failed storage page in the GBB cannot be restored, that is, the failure scene is lost. After the failure scene is lost, the cause of the GBB cannot be effectively analyzed, that is, the cause of the failed storage page cannot be effectively analyzed.

[0089] The embodiment of the present application provides an operation method for a storage system, which can back up the data in a failed storage page for subsequent analysis of the cause of the failed storage page, effectively improving the efficiency and accuracy of failure analysis. This method can be applied to, for example Figures 1 to 3 the storage system shown in any of the accompanying drawings. And, as Figure 6 shown, the memory 100 in this storage system may include multiple storage blocks 101 and at least one reserved (RSV) block 102. Each storage block 101 includes multiple storage pages, and each reserved block 102 includes multiple backup storage pages. Among them, the multiple storage blocks 101 can be used to store user data, that is, the multiple storage blocks 101 can be user-addressable blocks. The at least one reserved block 102 can be dedicated to backing up the data of the failed storage page.

[0090] As a possible implementation, as Figure 4 and Figure 5 shown, the at least one reserved block 102 ( Figure 4 and Figure 5 represented by RSV in and ) can belong to the target stripe, and the target stripe can be a stripe including an FBB. Among them, the FBB can refer to a bad block generated due to manufacturing processes and other reasons during the production process of the memory, that is, a bad block existing before leaving the factory. The other available storage blocks in the target stripe except the FBB are usually used to replace the GBB, that is, the controller 200 can remap (RMP) the GBB to the available storage block. In the embodiment of the present application, one or more of the available storage blocks in the target stripe can be selected as the reserved block 102 to be dedicated to backing up the data in the failed storage page. For example, when forming the remap table, one or more available storage blocks can be reserved as the reserved block 102.

[0091] As another possible implementation, the stripe to which the at least one reserved block 102 belongs may not include an FBB either, that is, the at least one reserved block 102 may be reserved by other means before leaving the factory. For example, the at least one reserved block 102 may be randomly selected from all available storage blocks of the memory.

[0092] The operation method of the storage system provided in the embodiments of the present application is introduced below. As Figure 7 shown, the method may include:

[0093] Step 501, in response to detecting a failed storage page in the storage pages included in multiple storage blocks, the controller obtains the data in the failed storage page.

[0094] In the embodiments of the present application, when the controller detects a data read error or a data write error in any storage page in the memory, it may determine that the storage page is a failed storage page.

[0095] In the first possible implementation, when the controller writes data to any storage page, if a data write error occurs, that is, a programming status failed (PSF) occurs, the controller may determine that the storage page is a failed storage page. And, the controller may directly obtain the data to be written to the failed storage page from its memory. Wherein, the memory of the controller may be a static random access memory (SRAM).

[0096] In the second possible implementation, when the controller reads data from any storage page, if a data read error occurs, that is, a data read UECC occurs, the controller may determine that the storage page is a failed storage page. And, the controller may recover the data in the failed storage page to obtain the data in the failed storage page.

[0097] As an example of the second implementation, the data in the failed storage page may be data protected by RAID technology. Correspondingly, after the controller detects the failed storage page, it may read the data other than the data in the failed storage page in the RAID stripe to which the failed storage page belongs. Then, based on the other data read, the data in the failed storage page can be recovered.

[0098] Optionally, referring to Figure 8 , the controller may perform LDPC decoding, MCRC detection, descrambling, and exclusive OR processing on the other data in the RAID stripe in sequence to recover the data in the failed storage page.

[0099] It can be understood that a failed storage page may include multiple codewords. If only some of the codewords in the failed storage page have UECC during data reading, the controller can read the data of the RAID stripe to which the part of the codewords with UECC belongs and recover the part of the codewords with UECC. That is to say, for the other codewords in the failed storage page that can be correctly read, there is no need to perform data recovery through the RAID stripe.

[0100] As another example of this second implementation, the data in the failed storage page can be data protected by a backup strategy. The backup strategy can refer to storing the same data in multiple different storage locations (such as multiple different storage planes) of the memory to ensure that after the data at any storage location is damaged, the data can still be read from other storage locations. For example, for relatively important system data in the storage system, such as the logical-to-physical mapping (L2P) table, this backup strategy can be used for data protection. In this example, the controller can directly obtain the data from the backup storage location corresponding to the failed storage page, and the obtained data is the data in the failed storage page.

[0101] Step 502, the controller backs up the data in the failed storage page to a backup storage page in at least one reserved block of the memory.

[0102] After the controller obtains the data in the failed storage page, it can back up the data to a backup storage page in the at least one reserved block. Thus, the data of the failure scene can be effectively retained, thereby improving the efficiency of subsequent failure analysis.

[0103] Exemplarily, referring to Figure 8 , the controller can first determine the target storage location, that is, determine a backup storage page for storing data from at least one reserved block. Then, the controller can program the data obtained from the failed storage page into this backup storage page.

[0104] It can be understood that for each reserved block, the controller can back up data in each backup storage page in sequence according to the order of the backup storage pages in the reserved block. That is to say, when the controller writes data into a certain backup storage page and then detects a failed storage page again, it can write the data of the failed storage page into the next backup storage page. Correspondingly, the above-mentioned target storage location can refer to the next backup storage page after the last written backup storage page. It can also be understood that if the storage space of a certain reserved block is full, the controller can continue to back up data in the next reserved block.

[0105] On the one hand, if the data in the failed storage page is the data to be written that the controller obtains from its memory, or the data recovered based on the RAID stripe, then as Figure 8 shown, the controller can successively scramble the obtained data, generate MCRC and LDPC codes, and program them into the backup storage page.

[0106] On the other hand, if the data in the failed storage page is directly obtained by the controller from the backup storage location corresponding to the failed storage page based on the backup policy, the controller can first perform LDPC decoding, MCRC detection, and descrambling on the obtained data. After that, the controller can successively scramble the descrambled data, generate MCRC and LDPC codes, and program them into the backup storage page.

[0107] Optionally, in the embodiments of the present application, the controller can adopt the SLC mode or the TLC mode to back up the data in the failed storage page to a backup storage page in one of the at least one reserved block. Among them, the SLC mode can effectively reduce the probability of bit flipping and ensure the security of data storage. The TLC mode can effectively improve the utilization rate of the storage space and ensure that more backup data is written in one reserved block.

[0108] Continuing to refer to Figure 9 , the operation method of the storage system provided by the embodiments of the present application may further include:

[0109] Step 503, the controller marks the storage block to which the failed storage page belongs as GBB.

[0110] In the embodiments of the present application, after the controller detects the failed storage page, it can mark the storage block to which the failed storage page belongs as GBB.

[0111] Step 504, after performing garbage collection on the memory, the controller reads the original data in the failed storage page.

[0112] It can be understood that, as Figure 10 shown, after the controller marks the bad blocks in the storage blocks in the memory, the garbage collection process is usually triggered. For example, if the storage system uses RAID technology to protect data, the controller can move the data in the RAID stripe to which the GBB belongs through the garbage collection process, and erase the data in the other storage blocks in the RAID stripe to which the GBB belongs except the GBB.

[0113] After the controller performs the garbage collection operation, if it is necessary to analyze the cause of GBB, that is, the failure cause of the failed storage page, the raw data in the failed storage page can be read. It can be understood that when the controller reads the raw data in the failed storage page, it does not need to go through processes such as LDPC decoding, MCRC detection, and descrambling, so there will be no data reading errors.

[0114] Step 505: Compare the raw data with the data in a backup storage page in at least one reserved block.

[0115] In the embodiment of the present application, the controller can read the data in the backup storage page of the failed storage page it backs up. Then, the data can be compared and analyzed with the raw data in the failed storage page to clarify the failure cause.

[0116] Optionally, referring to Figure 11 , the implementation process of the above step 505 may include:

[0117] Step 5051: The controller reads the data in a backup storage page in at least one reserved block.

[0118] For the GBB to be analyzed, the controller can determine the backup storage page for storing the data of the failed storage page in the GBB from at least one reserved block and read the data in the backup storage page.

[0119] Step 5052: The controller performs LDPC decoding, MCRC detection, and descrambling on the read data in sequence.

[0120] Referring to Figure 12 , after the controller reads the data in the backup storage page, it can perform LDPC decoding, MCRC detection, and descrambling on the read data in sequence to recover the initial data, which can also be called the target write buffer data.

[0121] Step 5053: The controller generates a scrambling seed based on the address of the failed storage page.

[0122] It can be understood that when the controller writes the initial data to the storage page, it can randomly generate a scrambling seed based on the position of the storage page (i.e., the storage position of the initial data), and scramble the initial data based on the scrambling seed. Since the scrambling seeds corresponding to different storage positions are different, the scrambling results of the same initial data are also different.

[0123] It can be seen from this that the stored contents of the same initial data at different storage locations in the memory are different. That is, the stored contents of the same initial data in the failed storage page are different from those in the backup storage page. Based on this, the controller needs to match the failed storage page in the GBB and generate a scrambling seed based on the address of the failed storage page.

[0124] Step 5054: The controller scrambles the data in the descrambled backup storage page based on the scrambling seed.

[0125] After the controller generates the scrambling seed, it can scramble the data (i.e., the initial data or the target write buffer data) in the descrambled backup storage page based on the scrambling seed.

[0126] Step 5055: The controller sequentially performs MCRC generation and LDPC encoding on the scrambled data and writes it into the cache of the controller.

[0127] The controller can also sequentially perform MCRC generation and LDPC encoding on the scrambled data to obtain the data to be written into the failed storage page. However, since the failed storage page has failed, as Figure 12 shown, the controller can write the scrambled, MCRC-generated, and LDPC-encoded data into the cache of the controller and save the data in the cache. Thus, the controller can simulate the process of writing the initial data into the failed storage page and obtain the correct data that should have been stored in the failed storage page.

[0128] Step 5056: Compare the original data with the data stored in the cache.

[0129] Since the data stored in the cache is the correct data that should have been stored in the failed storage page, by comparing the data stored in the cache and the original data in the failed storage page, the failure reason of the failed storage page, that is, the generation reason of the GBB, can be analyzed. Optionally, this comparison and analysis process can be performed by a tester, or it can also be implemented through a comparison and analysis script. The embodiments of the present application do not make any limitations in this regard.

[0130] It can be understood that in most cases, there is only one failed storage page inside a GBB, for example, only one UECC storage page. Assuming that each storage block in the memory includes 1392 SLC storage pages, one reserved block can store the data of the failed storage pages of 1392 GBBs, and this storage capacity is much larger than the normal demand. It can also be understood that if the storage space of at least one reserved block has been filled, the controller can stop continuing to write and perform debug analysis in a timely manner.

[0131] For example, for a failed storage page with UECC, based on the data recovery and comparison process shown in the above steps 5051 to 5056, it is possible to detect a data shift error in the failed storage page, thereby further leading to the analysis that an error has occurred in the logic domain of the controller. Thus, the efficiency of failure analysis is effectively improved.

[0132] In summary, the embodiments of the present application provide an operation method for a memory. The memory of the storage system includes multiple storage blocks and at least one reserved block. When the controller of the storage system can detect a failed storage page in the storage pages included in the multiple storage blocks, it can obtain the data in the failed storage page and back up the data in the failed storage page to a backup storage page in the at least one reserved block. Thus, when it is necessary to perform a failure analysis on the failed storage page, the data can be directly read from the backup storage page and the failure analysis can be performed to avoid leaving risks in the storage system.

[0133] Figure 13 is a schematic structural diagram of a controller provided by the embodiments of the present application. The controller can be applied to scenarios such as Figures 1 to 3 shown. As Figure 13 shown, the controller 200 includes a processor 201 and a memory 202. The memory 202 is used to store computer instructions, and the processor 201 is used to execute the computer instructions to implement the operation method of the storage system provided by the above embodiments. The processor 201 can be, for example, a micro - controller unit (MCU), etc.

[0134] Among them, the controller 200 can be used to implement the functions of the controller in the foregoing embodiments to implement the functions of the storage system provided by the embodiments of the present application. The specific implementation manner can refer to the embodiments shown in Figure 7 、 Figure 9 and Figure 11 shown, and will not be elaborated here.

[0135] The embodiments of the present application also provide a computer - readable storage medium. Instructions are stored on the computer - readable storage medium. When the instructions are executed by a processor in the controller, any step in the operation method of the storage system provided by the above embodiments can be implemented.

[0136] The embodiments of the present application also provide a computer program product containing instructions. When the instructions are executed by a processor in the controller, any step in the operation method of the storage system provided by the above embodiments can be implemented.

[0137] In this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance. The term "at least one" means one or more, and the term "a plurality" means two or more, unless otherwise clearly defined.

[0138] The above are only exemplary embodiments of this application and are not intended to limit this application. The protection scope of this application shall be subject to the protection scope of the claims.

Claims

1. A storage system, characterized in that, The storage system includes a memory and a controller coupled to the memory. The memory includes a plurality of storage blocks and at least one reserved block. Each of the plurality of storage blocks includes a plurality of storage pages, and each of the at least one reserved blocks includes a plurality of backup storage pages. The controller is configured to: In response to detecting a failed storage page in the storage pages included in the plurality of storage blocks, obtain the data in the failed storage page; Back up the data in the failed storage page to a backup storage page in the at least one reserved block.

2. The storage system according to claim 1, wherein The failed storage page is a storage page that generates a data read error. The controller is configured to: Recover the data in the failed storage page; Back up the recovered data in the failed storage page to a backup storage page in the at least one reserved block.

3. The storage system according to claim 2, wherein The controller is configured to: Read the other data in the redundant array of independent disks RAID stripe to which the failed storage page belongs, except for the data in the failed storage page; Based on the other data, recover the data in the failed storage page.

4. The storage system according to claim 3, wherein The controller is configured to: Perform low-density parity-check LDPC decoding, medium cyclic redundancy check code MCRC detection, descrambling, and exclusive OR processing on the other data in sequence to recover the data in the failed storage page; Perform scrambling, generate MCRC, and LDPC encoding on the data in the recovered failed storage page in sequence, and program it to a backup storage page in the at least one reserved block.

5. The storage system according to claim 1, wherein The failed storage page is a storage page that generates a data write error. The controller is configured to: Obtain the data to be written to the failed storage page from the memory of the controller; Back up the data to be written to the failed storage page to a backup storage page in the at least one reserved block.

6. The storage system according to any one of claims 1 to 5, characterized in that, The controller is further configured to: Mark the storage block to which the failed storage page belongs as a growing bad block GBB; After garbage collection is performed on the memory, read the original data in the failed storage page; Compare the original data with the data in a backup storage page in the at least one reserved block.

7. The storage system according to claim 6, characterized in that The controller is configured to: Perform LDPC decoding, MCRC detection, and descrambling on the data in a backup storage page in the at least one reserved block in sequence; Generate a scrambling seed based on the address of the failed storage page; Scramble the data in the descrambled backup storage page based on the scrambling seed; Perform MCRC generation and LDPC encoding on the scrambled data in sequence, and write it to the cache of the controller; Compare the original data with the data stored in the cache.

8. The storage system according to any one of claims 1 to 5, characterized in that, The controller is configured to: Back up the data in the failed storage page to a backup storage page in the at least one reserved block in a single-level cell SLC mode or a triple-level cell TLC mode.

9. The storage system according to any one of claims 1 to 5, characterized in that, The at least one reserved block belongs to a target stripe, and the target stripe includes at least one factory bad block FBB.

10. A method for operating a storage system, characterized in that, The storage system includes a controller and a memory. The memory includes a plurality of storage blocks and at least one reserved block. Each of the plurality of storage blocks includes a plurality of storage pages, and each of the at least one reserved blocks includes a plurality of backup storage pages. The method includes: In response to detecting a failed storage page in the storage pages included in the plurality of storage blocks, the controller acquires the data in the failed storage page. The controller backs up the data in the failed storage page to a backup storage page in the at least one reserved block.

11. The method according to claim 10, wherein The failed storage page is a storage page that generates a data read error. The controller backing up the data in the failed storage page to a backup storage page in the at least one reserved block includes: The controller recovers the data in the failed storage page. The controller backs up the recovered data in the failed storage page to a backup storage page in the at least one reserved block.

12. The method according to claim 11, wherein The controller recovering the data in the failed storage page includes: The controller reads the other data in the RAID stripe to which the failed storage page belongs, except for the data in the failed storage page. The controller recovers the data in the failed storage page based on the other data.

13. The method according to claim 12, characterized in that, The controller recovering the data in the failed storage page based on the other data includes: The controller sequentially performs LDPC decoding, MCRC detection, descrambling, and exclusive OR processing on the other data to recover the data in the failed storage page. The controller backing up the recovered data in the failed storage page to a backup storage page in the at least one reserved block includes: The controller sequentially performs scrambling, generating MCRC, and LDPC encoding on the recovered data in the failed storage page, and programs it to a backup storage page in the at least one reserved block.

14. The method according to claim 10, wherein The failed storage page is a storage page that generates a data write error. The controller backing up the data in the failed storage page to a backup storage page in the at least one reserved block includes: The controller acquires the data to be written to the failed storage page from the memory of the controller. The controller backs up the data to be written to the failed storage page to a backup storage page in the at least one reserved block.

15. The method according to any one of claims 10 to 14, characterized in that The method further includes: The controller marks the storage block to which the failed storage page belongs as a growing bad block (GBB). After garbage collection of the memory, the controller reads the original data in the failed storage page. Compare the original data with the data in a backup storage page in the at least one reserved block.

16. The method according to claim 15, wherein The comparing the original data with the data in a backup storage page in the at least one reserved block includes: The controller sequentially performs LDPC decoding, MCRC detection, and descrambling on the data in a backup storage page in the at least one reserved block. The controller generates a scrambling seed based on the address of the failed storage page. The controller scrambles the descrambled data in the backup storage page based on the scrambling seed. The controller sequentially performs MCRC generation and LDPC encoding on the scrambled data, and writes the data into the cache of the controller; Compare the original data with the data stored in the cache.

17. The method according to any one of claims 10 to 14, characterized in that, The controller backs up the data in the failed storage page to a backup storage page in one of the at least one reserved block, including: The controller backs up the data in the failed storage page to a backup storage page in one of the at least one reserved block in SLC mode or TLC mode.

18. A controller, characterized in that, The controller includes: a processor and a memory; the memory is used to store computer instructions, and the processor is used to execute the computer instructions to implement the method according to any one of claims 10 to 17.