Fast processing method and device for SSD flash block write errors and SSD device
By flexibly forming RG Blocks in SSDs and migrating data using cache space, the problem of limited number of RG Blocks is solved, the performance and life of SSDs are improved, and the resource usage of write error processing is reduced.
Patent Information
- Application Number
- CN202411417905.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-11
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-10-11
AI Technical Summary
The organizational form of RG Block in existing SSDs has resulted in limited number of RG Blocks that can actually be formed, affecting the service life and performance of SSDs, and the write error processing process occupies resources and affects performance.
通过在SSD中组建RG Block时,选取每个Die中的部分或全部plane上的好块,利用缓存空间进行数据迁移和重新组织,避免写错误块的替换,减少GC处理。
Improves the operation performance of SSD, extends service life, and maintains performance stability during write errors, reducing resource usage.
Smart Images

Figure CN119322589B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data storage, and particularly to a method and apparatus for quickly processing write errors of SSD flash blocks and an SSD device. Background Art
[0002] In the prior art, during the SSD firmware processing, the RG Block is organized in units of Dies. In each Die, one good block located on different planes is selected to form a group of multi-plane Blocks.
[0003] For the organization form of the prior art RG Block, since the number of bad blocks in each Die is inconsistent, the actual number of good blocks on each Plane is inconsistent. The number of RG Blocks that can be actually formed in an SSD is affected by the Plane with the least number of good blocks among all Dies. Then, some extra Blocks in all other Planes cannot participate in the formation of RG Blocks, resulting in a small SSD OP, which will affect the service life and performance of the SSD.
[0004] When a write error occurs in one or more Blocks in a group of multi-plane Blocks in one or more Dies during the data writing process of an RG Block, the data that has not been successfully stored in this group of RG Blocks will be written into another group of intact RG Blocks, and the valid data that has been successfully written in the RG Block with the write error will be subjected to GC. After the valid data recovery is completed, the Block with the write error will be replaced with a spare Block that is good but not used to form an RG Block. If there is no good spare Block available for replacement, this group of RG Blocks will be downgraded. When they cannot be used after downgrading, the firmware will invalidate this group of RG Blocks, and the good Blocks in this group of RG Blocks will be put into the spare Block pool for replacing bad Blocks during subsequent operation.
[0005] This method makes the organization form of RG Blocks too strict. Blocks that do not meet this organization form cannot be organized into RG Blocks, resulting in a small OP and affecting the life and performance of the SSD. Moreover, the processing flow when a write error occurs is to first select a good RG Block to rewrite the data that has not been successfully written, and perform GC processing on the valid data in the RG Block with the write error. The GC processing is usually a background operation, and these two operations will occupy more resources during operation, affecting the performance of the SSD. Summary of the Invention
[0006] To solve the above technical problems or at least partially solve the above technical problems, the present invention provides a fast processing method, device and SSD device for write errors of SSD flash blocks.
[0007] In one aspect of the present invention, a fast processing method for write errors of SSD flash blocks is provided. The RGBlock of the SSD is composed of one good block selected from each part or all planes of each Die constituting the RG Block. The method includes:
[0008] When performing a write operation on the first RG Block, buffer the data to be written into the buffer units of a preset first buffer space according to the position layout of each Block in the first RG Block, so that the position layout of the buffer units storing the data to be written in the first buffer space is the same as the position layout of the Blocks in the target RG Block;
[0009] When sequentially writing the data to be written in the first buffer space into the first RG Page of the first RG Block, if a write error occurs, obtain the position information where the write error occurs. The position information where the write error occurs includes the position information of the first Block where the write error occurs and the information of the first RG Page;
[0010] Determine the data migration rule of the data to be written pre-stored in the first buffer space according to the position information where the write error occurs, and migrate the data to be written in the first buffer space within the first buffer space, and / or migrate the data to be written in the first buffer space between the first buffer space and a preset second buffer space;
[0011] Write the data to be written after data migration in the first buffer space into the second RGPage adjacent to the first RG Page; if the data to be written in the first buffer space is migrated between the first buffer space and a preset second buffer space, write the data to be written after data migration in the first buffer space into the second RG Page adjacent to the first RG Page, and write the data to be written migrated to the second buffer space into the third RG Page adjacent to the second RG Page.
[0012] Further, the determining the data migration rule of the data to be written pre-stored in the first buffer space according to the position information where the write error occurs includes:
[0013] Judge whether the first RG Page is the last page in the first RG Block;
[0014] If the first RG Page is not the last page in the first RG Block, the data migration rule is to migrate the data to be written in the first cache space between the first cache space and a preset second cache space;
[0015] If the first RG Page is the last page in the first RG Block, obtain the position layout of the Block of the second RG Block to which the second RG Page belongs, and determine the data migration rule according to the position layout of the Block of the second RG Block;
[0016] Among them, the determining the data migration rule according to the position layout of the Block of the second RG Block includes:
[0017] Migrate the data to be written in the first cache space within the first cache space; or,
[0018] Migrate the data to be written in the first cache space within the first cache space, and migrate the data to be written in the first cache space between the first cache space and a preset second cache space.
[0019] Further, if the first RG Page is the last page in the first RG Block, the data migration rule includes:
[0020] Take the data to be written in the cache unit corresponding to the first Block position in the first cache space as the redundant data of the second RG Page, and migrate the redundant data of the second RG Page to the second cache space.
[0021] Further, the determining the data migration rule according to the Block distribution of the second RG Block includes:
[0022] Obtain the correspondence between the position of the cache unit pre-storing the data to be written in the first cache space and the position of the Block of the second RG Block;
[0023] Take the cache unit in the first cache space corresponding to the Block position in the second RG Block and not storing the data to be written as the migration-in position; take the cache unit in the first cache space not corresponding to the Block position in the second RG Block and storing the data to be written as the migration-out position;
[0024] When the number of migration-in positions is greater than or equal to the number of migration-out positions, migrate the data to be written in the cache unit at the migration-out position in the first cache space to the cache unit at the migration-in position in the first cache space.
[0025] When the number of relocation positions is less than the number of eviction positions, preferentially migrate the data to be written in each cache unit at the eviction positions in the first cache space to each cache unit at the relocation positions in the first cache space; after the data to be written is pre-stored in each cache unit at the relocation positions in the first cache space, use the data to be written that has not been migrated at the eviction positions as the redundant data of the second RG Page, and migrate the redundant data of the second RG Page to the second cache space.
[0026] Further, the migrating the redundant data of the second RG Page to the second cache space includes:
[0027] Obtain the position layout of the Block of the third RG Block to which the third RG Page belongs;
[0028] Migrate the redundant data of the second RG Page to the cache units in the second cache space that have the same position layout as the Block of the third RG Block.
[0029] Further, after migrating the redundant data of the second RG Page to the second cache space, the method further includes:
[0030] Continue to pre-store the data to be written in the second cache space, so that each cache unit in the second cache space that has the same position layout as the Block of the third RG Block pre-stores the data to be written.
[0031] Further, after obtaining the position information of the write error, the method further includes:
[0032] Mark the first Block as a bad block and remove it from the first RG Block;
[0033] Mark the data in the first RG Page as invalid data, and retain the data in each RG Page in the first RG Block where the write operation has been successfully performed.
[0034] On the other hand, the present invention also provides a fast processing device for write errors of SSD flash blocks. The RG Block of the SSD is composed of selecting one good block from some planes or all planes of each Die that makes up the RG Block. The device includes:
[0035] An arrangement module, configured to, when performing a write operation on the first RG Block, cache the data to be written into the cache units of a preset first cache space according to the position layout of each Block in the first RG Block, so that the position layout of the cache units storing the data to be written in the first cache space is the same as the position layout of the Block of the target RG Block;
[0036] An acquisition module, configured to obtain position information where a write error occurs when sequentially writing data to be written in a first cache space into a first RG Page of a first RG Block. The position information where the write error occurs includes the position information of the first Block where the write error occurs and the information of the first RG Page.
[0037] A migration module, configured to determine a data migration rule for the data to be written pre-stored in the first cache space according to the position information where the write error occurs, and migrate the data to be written in the first cache space within the first cache space, and / or migrate the data to be written in the first cache space between the first cache space and a preset second cache space.
[0038] The write operation module is further configured to write the data to be written after data migration in the first cache space into a second RG Page adjacent to the first RG Page; if the data to be written in the first cache space is migrated between the first cache space and a preset second cache space, write the data to be written after data migration in the first cache space into a second RG Page adjacent to the first RG Page, and write the data to be written migrated to the second cache space into a third RG Page adjacent to the second RG Page.
[0039] Further, the migration module includes:
[0040] A judgment sub-module, configured to judge whether the first RG Page is the last page in the first RG Block;
[0041] A migration determination sub-module, configured to, if the first RG Page is not the last page in the first RG Block, the data migration rule is to migrate the data to be written in the first cache space between the first cache space and a preset second cache space;
[0042] The migration determination sub-module is further configured to, if the first RG Page is not the last page in the first RG Block, obtain the position layout of the Block of the second RG Block to which the second RG Page belongs, and determine the data migration rule according to the position layout of the Block of the second RG Block; wherein, determining the data migration rule according to the position layout of the Block of the second RG Block includes:
[0043] Migrating the data to be written in the first cache space within the first cache space; or,
[0044] Migrate the data to be written in the first cache space within the first cache space, and migrate the data to be written in the first cache space between the first cache space and a preset second cache space.
[0045] Another aspect of the present invention also provides an SSD device, which includes a storage controller. The storage controller includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the above method are implemented.
[0046] The fast processing method, device and SSD device for write errors of SSD flash blocks provided by the present invention. The RG Block of the SSD is composed of one good block selected from some or all of the planes in each Die that makes up the RG Block. The formation method of the RG Block is more flexible, maximizing the OP of the SSD, extending the service life of the SSD and improving the performance of the SSD. And when a write error message is received, the preset first cache space and second cache space are used for data migration and reorganization, so that the first RG Block with a write error can continue to be used without replacing the bad block, reducing the GC processing process and ensuring the stability of the SSD performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0048] Figure 1 It is a schematic flowchart of the fast processing method for write errors of SSD flash blocks in an embodiment of the present invention;
[0049] Figure 2 It is the fast processing method for write errors of SSD flash blocks in a specific embodiment of the present invention;
[0050] Figure 3 It is a schematic structural diagram of the fast processing device for write errors of SSD flash blocks in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0052] Before introducing the fast processing method for write errors of SSD flash memory blocks in the embodiments of the present invention, a brief introduction to the background technology in the present invention is given first:
[0053] Die: Usually refers to an independent chip unit formed by cutting from a wafer and encapsulating, also known as a bare chip. It is the smallest independently operable component of a NAND Flash memory particle.
[0054] Plane: In a NAND flash memory chip, a plane is an organization form of a large number of storage blocks, mainly used to improve the parallelism of storing and managing data. It is a part of the internal logic structure of the flash memory chip, and together with Block and Page, it constitutes the basic storage unit of the flash memory.
[0055] RG Block: A logical concept at the SSD level, the smallest parallel erasure unit in SSD firmware management. A group of RG Blocks is an operation unit formed by organizing one Block on each Plane in each Die of some or all Dies in the SSD. In fact, the Blocks in each Die can be flexibly organized, such as selecting one to less than or equal to the number of Planes of good Blocks located on different Planes in the corresponding Die.
[0056] Good block: Refers to a Block in the SSD that can perform normal read and write operations. These Blocks maintain a good state during manufacturing and use, can reliably store and read data, and ensure that users can use the SSD normally.
[0057] Bad block: Refers to a Block in the SSD that cannot perform normal read or write operations. These Blocks may be damaged due to defects in the manufacturing process, wear during use, or other factors, resulting in incorrect storage or reading of data.
[0058] Page: The smallest unit for storing data in NAND flash memory, which allows data to be read and written in units of pages.
[0059] GC (Garbage Collection): Also known as garbage collection, it refers to the process of transferring the valid written data on a certain flash block (Block) in the SSD to other flash blocks, and completely erasing the original flash block to obtain a new flash block that can be used for subsequent writing.
[0060] OP (Over-Provision, reserved space): The reserved space in the SSD is usually the actual capacity of the SSD minus the user-available capacity. The user space is the space where users store their data, and the reserved space is the capacity space that users cannot operate. The reserved space is mainly used to improve the performance of the SSD and extend its service life.
[0061] Firmware: Firmware is the program in the chip that drives the chip to work. The firmware in the SSD is deeply developed according to the characteristics of NandFlash and the usage scenarios of the SSD. There are many algorithms in the firmware, and the firmware development process is complex. The entire firmware in the SSD can ensure the normal and stable operation of the SSD.
[0062] Buffer (buffer): A buffer is a memory area used to temporarily store data. Specifically in the SSD, the data buffer area is mainly used to receive and temporarily store the user data sent from the host to the SSD, and these data will be stored in the NAND Flash later.
[0063] Write error: When the firmware in the SSD is running, when the controller writes the data in the buffer to the Block in the specific Die of the NandFlash, if there is a problem with the corresponding Block, the corresponding Die will report a write error to the controller, indicating that the data has not been successfully stored in the Block of the NAND Flash. For this situation, the firmware will make special processing to ensure data integrity.
[0064] In the fast processing method for SSD flash block write errors provided by the embodiments of the present invention, the RG Block of the SSD is composed of one good block selected from some or all of the planes in each Die that makes up the RGBlock. There is no strict regulation that one Block must be selected on each Plane for the formation method of the RG Block. One or more Planes can not select Blocks to form the RG Block. In this way, the good blocks on each Plane in each Die can be fully utilized, increasing the OP.
[0065] The following combines the attached Figure 1 and the attached Figure 2 A detailed introduction to the fast processing method for SSD flash block write errors proposed by the embodiments of the present invention is given below. The method includes the following steps:
[0066] S1. When performing a write operation on the first RG Block, buffer the data to be written into the buffer units of a preset first buffer space according to the position layout of each Block in the first RG Block, so that the position layout of the buffer units storing the data to be written in the first buffer space is the same as the position layout of the Blocks in the target RG Block;
[0067] As Figure 2 In the specific embodiment shown, the present invention presets Buffer0 as the first buffer space and presets Buffer1 as the second buffer space in the Buffer area, and pre-stores the data to be written through the first buffer space and the second buffer space. The arrangement of each buffer unit in the first buffer space and the second buffer space corresponds to the arrangement of all dies and all planes in the die used for the component RG Block, and the storage capacity of each buffer unit in the first buffer space or the second buffer space is the storage capacity of one page, so that the data to be written stored in the first buffer space or the second buffer space at one time can perform a write operation on all pages in a group of RG Blocks.
[0068] S2. When sequentially writing the data to be written in the first buffer space into the first RG Page of the first RG Block, if a write error occurs, obtain the position information where the write error occurs. The position information where the write error occurs includes the position information of the first Block where the write error occurs and the information of the first RG Page;
[0069] It should be noted that after the present invention pre-stores the data to be written in the first buffer space, the data to be written in the first buffer space is written into RG BlockM Page N. Here, for the sake of simplicity of expression, the Nth RGPage of RG BlockM is expressed as RG BlockM Page N. Then, the data to be written is pre-stored in the first buffer space again, and the data to be written in the first buffer space is written into RG BlockM Page N + 1. After all the RG Pages in RG BlockM have completed the write operation, the write operation is performed on each RG Page in RG Block(M + 1).
[0070] S3. Determine the data migration rule of the data to be written pre-stored in the first buffer space according to the position information where the write error occurs, and migrate the data to be written in the first buffer space within the first buffer space, and / or migrate the data to be written in the first buffer space between the first buffer space and a preset second buffer space;
[0071] S4. Write the data to be written after data migration in the first cache space into the second RG Page adjacent to the first RG Page; if the data to be written in the first cache space is migrated between the first cache space and a preset second cache space, write the data to be written after data migration in the first cache space into the second RG Page adjacent to the first RG Page, and write the data to be written migrated to the second cache space into the third RG Page adjacent to the second RG Page.
[0072] In the embodiments of the present invention, the second RG Page may belong to the same RG Block as the first RG Page, or may belong to a different RG Block from the first RG Page. With different position correspondence relationships, the data migration rules for the data to be written pre-stored in the first cache space are different. Correspondingly, the third RG Page may belong to the same RG Block as the second RG Page, or may belong to a different RG Block from the second RG Page. The present invention uses the preset first cache space and second cache space for data migration and reorganization, enabling the first RG Block with a write error to continue performing write operations in the second RG Page without replacing the bad Block, reducing the GC processing flow, and ensuring the stability of the SSD performance.
[0073] Specifically, the step of determining the data migration rule for the data to be written pre-stored in the first cache space according to the position information where the write error occurs in the embodiments of the present invention further includes the following steps:
[0074] S01. Determine whether the first RG Page is the last page in the first RG Block;
[0075] S02. If the first RG Page is not the last page in the first RG Block, the data migration rule is to migrate the data to be written in the first cache space between the first cache space and a preset second cache space;
[0076] S03. If the first RG Page is the last page in the first RG Block, obtain the position layout of the Block of the second RG Block to which the second RG Page belongs, and determine the data migration rule according to the position layout of the Block of the second RG Block; wherein, determining the data migration rule according to the position layout of the Block of the second RG Block includes: migrating the data to be written in the first cache space within the first cache space; or, migrating the data to be written in the first cache space within the first cache space and migrating the data to be written in the first cache space between the first cache space and a preset second cache space.
[0077] Further, in the embodiment of the present invention, if the first RG Page is not the last page in the first RG Block, the first Block that cannot perform a write operation in the second RG Page is the first Block where a write error occurs. At this time, the data to be written in the cache unit corresponding to the position of the first Block in the first cache space cannot be successfully written into the second RG Page. Therefore, the data migration rule includes: taking the data to be written in the cache unit corresponding to the position of the first Block in the first cache space as redundant data of the second RG Page, and migrating the redundant data of the second RG Page to the second cache space.
[0078] Further, if the first RG Page is the last page in the first RG Block, the position layout of the Block of the second RG Block will be different from the position layout of the first RG Block. Therefore, it is necessary to determine the specific data migration rule according to the position layout of the Block of the second RG Block.
[0079] In the embodiment of the present invention, determining the data migration rule according to the position layout of the Block of the second RG Block includes the following steps:
[0080] S031. Obtain the correspondence between the position of the cache unit pre-storing the data to be written in the first cache space and the position of the Block of the second RG Block; take the cache unit in the first cache space corresponding to the Block position of the second RG Block and not storing the data to be written as the migration-in position; take the cache unit in the first cache space not corresponding to the Block position of the second RG Block and storing the data to be written as the migration-out position;
[0081] Understandably, cache units in the first cache space corresponding to the Block positions in the second RG Block and not storing data to be written can cache the data to be written and then write the data into the second RG Page. Therefore, the data to be written can be cached again during the data migration process. Correspondingly, cache units in the first cache space not corresponding to the Block positions in the second RG Block and storing data to be written cannot write the data into the second RG Page. Therefore, during the data migration process, the data to be written in such cache units needs to be migrated away to avoid writing failure again.
[0082] S032. When the number of migration positions is greater than or equal to the number of migration-out positions, migrate the data to be written in the cache units at the migration-out positions in the first cache space to the cache units at the migration positions in the first cache space.
[0083] Understandably, when the number of migration positions is greater than or equal to the number of migration-out positions, the first cache space can successfully cache all the data to be written, and each cache unit storing the data to be written will not fail to write due to the difference in its arrangement with the Block arrangement in the second RG Block during the write operation.
[0084] S033. When the number of migration positions is less than the number of migration-out positions, first migrate the data to be written in each cache unit at the migration-out positions in the first cache space to each cache unit at the migration positions in the first cache space; after pre-storing the data to be written in each cache unit at the migration positions in the first cache space, use the data to be written that has not been migrated at the migration-out positions as the redundant data of the second RG Page, and migrate the redundant data of the second RG Page to the second cache space.
[0085] Understandably, when the number of migration positions is less than the number of migration-out positions, the data to be written cached in the first cache space has exceeded the storage capacity of the second RG Page. Therefore, part of the data needs to be migrated to the second cache space. To minimize the operation steps of data migration, the present invention only migrates the data at positions with different Block position layouts, while retaining the data in most storage units with the same Block position layout as in the second RG Block, saving the resource overhead of data migration.
[0086] Further, since both the data storage order and the data storage location will change after data migration, in the embodiments of the present invention, the data to be written after data migration in the first cache space is written into the second RG Page adjacent to the first RG Page; if the data to be written in the first cache space migrates between the first cache space and a preset second cache space, after writing the data after data migration in the first cache space into the second RG Page adjacent to the first RG Page and writing the data migrated to the second cache space into the third RG Page adjacent to the second RG Page, the method further includes: recording the distribution of the data to be written after the write operation is completed.
[0087] In addition, it should be noted that in the embodiments of the present invention, the migration of the redundant data of the second RG Page to the second cache space includes: obtaining the position layout of the Block of the third RG Block to which the third RG Page belongs; migrating the redundant data of the second RG Page to the cache unit in the second cache space with the same position layout as the Block of the third RG Block.
[0088] Among them, the third RG Block may belong to the same RG Block as the second RG Block, or may belong to a different RG Block from the second RG Block. Therefore, when caching the data to be written in the second cache space, it is necessary to determine the specific data migration location according to the position layout of the Block of the RG Block to which the third RG Page belongs. At this time, it can be for the second cache space and
[0089] Further, the redundant data of the second RG Page can be sequentially migrated to each cache unit corresponding to the Block of the third RG Block in the second cache space, or any cache unit corresponding to the Block of the third RG Block can be selected for storage, and the present invention does not limit this.
[0090] In addition, in order to make more full use of the storage space of the second cache space and the third RG Page, after migrating the redundant data of the second RG Page to the second cache space, the method further includes: pre-depositing the data to be written into the second cache space continuously, so that each cache unit in the second cache space with the same position layout as the Block of the third RG Block pre-stores the data to be written.
[0091] The following combines with the attached Figure 2, the data migration rules in the case of invention writing errors in the embodiments of the present invention will be described in detail. For the convenience of describing the present invention, the data volume of one Page is represented by one unit. In this specific embodiment, it is assumed that one die contains two planes. Therefore, the maximum data volume of one RG Page is 2(N + 1)·unit. The positions marked with unit* in the attached drawings are the buffer units pre-storing the data to be written, and the blue positions are the buffer units pre-storing the invalid data corresponding to the bad block positions.
[0092] Further, Scenario 1 describes the situation when the first RG Page is not the last page in the first RG Block. When performing write operations on each RG Page in RG BlockM, the data to be written is first pre-stored in buffer0 in batches. The positions of the buffer units pre-storing the data to be written in buffer0 correspond one-to-one with the positions of the Blocks included in RG BlockM.
[0093] Further, when writing the data into RG BlockM Page(N + 1), the data in buffer0 is data2. When the firmware sends data2 to the NAND Flash, the Block on Plane1 of Die2 reports a write error. The firmware receives the write error message, indicating that data2 has not been successfully stored in the NAND Flash. The data2 in buffer0 remains. The Block on Plane1 of Die2 reporting a write error means that this Block becomes a bad block during operation and cannot continue to write data. The data in unit5 of data2 cannot be directly written into RG BlockM in the original way. At this time, first mark the data in RG BlockM Page(N + 1) as invalid, and then perform the operation of flow1 to move the data in unit5 from buffer0 to the corresponding position of Die0, Plane0 in buffer1.
[0094] Further, after migrating the unit5 data in data2, fill the original position of unit5 in buffer0 with Dummy Data (invalid data). At this time, the data in buffer0 Figure 1 shows "data2 has moved the data on the write error unit" in the figure. Write this data into RG BlockM Page(N + 2). In addition, the data in buffer1 is directly written into RG BlockM Page(N + 3) according to the operation of flow2 in the figure. Thus, the entire write error handling process is completed.
[0095] It should be noted that at this time, the data to be written in Block PageN to 0 on Die2 Plane1 is not subject to GC processing. Generally, this data can be read correctly. If it cannot be read correctly, the data can also be read correctly through the RAID error correction method.
[0096] Therefore, in the embodiment of the present invention, after obtaining the position information of the write error, the method further includes: marking the first Block as a bad block and removing it from the first RG Block; marking the data in the first RG Page as invalid data, and retaining the data in each RG Page in the first RG Block where the write operation has been successful. This embodiment is applicable to other embodiments.
[0097] In the above specific embodiment, the firmware needs to save the distribution of valid data in the RG Block at the same time. For example, the data on Die1, Plane0 in RGBlockM PageN to 0 is invalid, and the data on Die1, Plane0, Die2, Plane1 in RGBlockM Page(N + 2) is invalid.
[0098] Scenario 2 describes the situation where the first RG Page is the last page in the first RG Block and the number of relocation positions is less than the number of eviction positions. In Scenario 2, a write error occurs when the firmware writes data3 to RG BlockM Page(N + 4). Assume that RG BlockM Page(N + 4) is the last page of the RG Block. At this time, it is necessary to consider switching to another valid RG Block to continue writing during the write error handling process.
[0099] Further, the write error is located on Plane0 of Die(N - 1) and DieN. The firmware will first mark the data in RGBlockM Page(N + 4) as invalid, and then allocate RGBlock(M + 1) for subsequent data writing positions. As shown in the figure, the positions where data can be stored in RGBlock(M + 1) (such as Page0) and the positions where data can be stored in the position of RG BlockM where the data to be written is stored (such as Page(N + 3)) have different layouts, thus involving the "sorting" process of data3.
[0100] Furthermore, the data sorting process is performed based on the principle of minimum number of operations. First, the location where data3 is written to RGBlockM Page (N+4) and the write error occurs is valid in RG Block (M+1) Page0, and no special processing is required for these two data locations. At this time, the data in unit0 is first moved to the location where Die1 and Plane0 were originally filled with Dummy Data, and then the data in unit3 is moved to buffer1. As shown in the figure, after the data in unit3 in buffer0 is moved, it is filled with Dummy Data. Then the data unit (2N+1) in data3 is moved to buffer1. As shown in the figure, the original location of unit (2N+1) in buffer0 is filled with Dummy Data. The final sorted data is shown in "data3 moved data exceeding the write length", and then the data is written to RG Block (M+1) Page0, and the remaining data is written from buffer1 to RG Block (M+1) Page1. At this point, the write error processing process in scenario 2 is completed, without GC and frequent bad block replacement operations, ensuring the stability of SSD performance.
[0101] Scenario 3 describes the situation when the first RGPage is the last page in the first RG Block, and the number of migration-in locations is greater than the number of migration-out locations. In Scenario 3, the RG Block that is subsequently switched can store more data locations relative to the RG Block where the write error occurs. The error handling process for this scenario will be simplified, involving only the movement of data in buffer0, where the write error is located on Plane0 of Die1. The firmware will first mark the data in RG Block (M+1) Page (N+4) as invalid, and then allocate RG Block (M+2) for subsequent data write locations. At this time, since Plane0 of Die1 of RGBlock (M+2) belongs to the block of RG Block (M+2), data migration is not required here. However, the Block on Plane1 of Die0 does not belong to RG Block (M+2). At this time, it is only necessary to migrate the data to be written on Plane1 of Die0 in the first cache space to Plane0 of Die0.
[0102] For method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described action sequence, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present invention.
[0103] Figure 3 The structural schematic diagram of a fast processing device for SSD flash block write errors according to an embodiment of the present invention is schematically shown. Referring to Figure 3 , the fast processing device for SSD flash block write errors according to the embodiment of the present invention specifically includes an arrangement module 301, a write operation module 302, and a migration module 303, where:
[0104] The arrangement module 301 is configured to, when performing a write operation on the first RG Block, cache the data to be written into the cache units of a preset first cache space according to the position layout of each Block in the first RG Block, so that the position layout of the cache units storing the data to be written in the first cache space is the same as the position layout of the Blocks in the target RG Block;
[0105] The write operation module 302 is configured to, when sequentially writing the data to be written in the first cache space into the first RG Pages of the first RG Block, if a write error occurs, obtain the position information where the write error occurs, and the position information where the write error occurs includes the position information of the first Block where the write error occurs and the information of the first RG Page;
[0106] The migration module 303 is configured to determine the data migration rule of the data to be written pre-stored in the first cache space according to the position information where the write error occurs, and migrate the data to be written in the first cache space within the first cache space, and / or migrate the data to be written in the first cache space between the first cache space and a preset second cache space;
[0107] The write operation module 302 is further configured to write the data to be written after data migration in the first cache space into the second RG Page adjacent to the first RG Page; if the data to be written in the first cache space is migrated between the first cache space and a preset second cache space, write the data to be written after data migration in the first cache space into the second RG Page adjacent to the first RG Page, and write the data to be written migrated to the second cache space into the third RG Page adjacent to the second RG Page.
[0108] Further, the migration module includes:
[0109] A judgment sub-module, configured to judge whether the first RG Page is the last page in the first RG Block;
[0110] A migration determination sub-module, configured to, if the first RG Page is not the last page in the first RG Block, the data migration rule is to migrate the data to be written in the first cache space between the first cache space and a preset second cache space;
[0111] The migration determination sub-module is further configured to, if the first RG Page is the last page in the first RG Block, obtain the position layout of the Block of the second RG Block to which the second RG Page belongs, and determine the data migration rule according to the position layout of the Block of the second RG Block.
[0112] Further, the migration determination sub-module is specifically configured to, if the first RG Page is not the last page in the first RG Block, the data migration rule includes: regarding the data to be written in the cache unit corresponding to the first Block position in the first cache space as redundant data of the second RG Page, and migrating the redundant data of the second RG Page to the second cache space.
[0113] Further, the migration determination sub-module determining the data migration rule according to the position layout of the Block of the second RG Block specifically includes:
[0114] An acquisition sub-module, configured to acquire the correspondence between the position of the cache unit pre-storing the data to be written in the first cache space and the position of the Block of the second RG Block;
[0115] A migration determination sub-module, configured to use the cache unit corresponding to the Block position in the second RG Block in the first cache space and not storing the data to be written as the migration-in position; and use the cache unit corresponding to the Block in the second RG Block in the first cache space and storing the data to be written as the migration-out position;
[0116] When the number of migration-in positions is greater than or equal to the number of migration-out positions, migrate the data to be written in the cache unit at the migration-out position in the first cache space to the cache unit at the migration-in position in the first cache space.
[0117] When the number of relocation positions is less than the number of eviction positions, the data to be written in each cache unit at the eviction positions in the first cache space is preferentially migrated to each cache unit at the relocation positions in the first cache space; after the data to be written is pre-stored in each cache unit at the relocation positions in the first cache space, the data to be written that has not been migrated in the eviction positions is used as the redundant data of the second RG Page, and the redundant data of the second RG Page is migrated to the second cache space.
[0118] Further, the obtaining sub-module is further configured to obtain the position layout of the Block of the third RG Block to which the third RG Page belongs;
[0119] The migration sub-module of the migration module 304 is further configured to migrate the redundant data of the second RG Page to the cache units in the second cache space that have the same position layout as the Block of the third RG Block.
[0120] Further, the device further includes a management module, configured to, after obtaining the position information where a write error occurs, mark the first Block as a bad block and remove it from the first RG Block; mark the data in the first RG Page as invalid data, and retain the data in each RG Page in the first RG Block where the write operation has been successful.
[0121] Further, the device further includes a recording module, configured to record the distribution of the written data.
[0122] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the related parts, refer to the partial description of the method embodiment.
[0123] In addition, an embodiment of the present invention further provides an SSD device, which includes a storage controller. The storage controller includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method described above are implemented. For example Figure 1 the steps S1 to S4 shown. Or, when the processor executes the computer program, the functions of each module / unit in the above embodiment of the fast processing device for write errors in the SSD flash block are implemented, for example Figure 3 the layout module 301, the write operation module 302, and the migration module 303 shown.
[0124] In this embodiment, if the modules / units integrated in the SSD device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such understanding, to implement all or part of the processes in the method of the above embodiment, the present invention can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0125] The fast processing method, device and SSD device for write errors of SSD flash blocks provided in this embodiment. The RGBlock of the SSD is composed of one good block selected from some planes or all planes of each Die that makes up the RG Block. The construction method of the RGBlock is more flexible, maximizing the OP of the SSD, extending the service life of the SSD and improving the performance of the SSD. And when a write error message is received, a preset first buffer space and a second buffer space are used for data migration and reorganization, so that the first RG Block with a write error can continue to be used without replacing the bad block, reducing the GC processing flow and ensuring the stability of the SSD performance.
[0126] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.
[0127] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solutions, in essence, or the part that contributes to the prior art can be embodied in the form of a software product, and this computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0128] In addition, those skilled in the art can understand that although some embodiments herein include certain features included in other embodiments rather than other features, the combination of the features of different embodiments means that it is within the scope of the present invention and forms different embodiments. For example, any one of the claimed embodiments can be used in any combination.
[0129] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A fast processing method for SSD write errors, characterized in that, The RG Block of the SSD is composed of selecting a good block from each part or all planes in each Die that makes up the RG Block. The method includes: When performing a write operation on the first RG Block, cache the data to be written into the cache units of a preset first cache space according to the position layout of each Block in the first RG Block, so that the position layout of the cache units storing the data to be written in the first cache space is the same as the position layout of the Blocks in the target RG Block; When sequentially writing the data to be written in the first cache space into the first RG Page of the first RG Block, if a write error occurs, obtain the position information where the write error occurs. The position information where the write error occurs includes the position information of the first Block where the write error occurs and the position information of the first RG Page. After obtaining the position information where the write error occurs, the method further includes: marking the first Block as a bad block and removing it from the first RG Block; marking the data in the first RG Page as invalid data and retaining the data in each RG Page in the first RG Block where the write operation has been successful; Determine the data migration rule of the data to be written pre-stored in the first cache space according to the position information where the write error occurs, and migrate the data to be written in the first cache space within the first cache space, and / or migrate the data to be written in the first cache space between the first cache space and a preset second cache space; Write the data to be written after data migration in the first cache space into the second RG Page adjacent to the first RG Page; if the data to be written in the first cache space is migrated between the first cache space and a preset second cache space, write the data to be written after data migration in the first cache space into the second RG Page adjacent to the first RG Page, and write the data to be written migrated to the second cache space into the third RG Page adjacent to the second RG Page.
2. The method according to claim 1, wherein The determining the data migration rule of the data to be written pre-stored in the first cache space according to the position information where the write error occurs includes: Judging whether the first RG Page is the last page in the first RG Block; If the first RG Page is not the last page in the first RG Block, the data migration rule is to migrate the data to be written in the first cache space between the first cache space and a preset second cache space; If the first RG Page is the last page in the first RG Block, obtain the position layout of the Blocks in the second RG Block to which the second RG Page belongs, and determine the data migration rule according to the position layout of the Blocks in the second RG Block; Among them, determining the data migration rule according to the position layout of the Block of the second RG Block includes: Migrate the data to be written in the first cache space within the first cache space; or, Migrate the data to be written in the first cache space within the first cache space, and migrate the data to be written in the first cache space between the first cache space and a preset second cache space.
3. The method according to claim 2, characterized in that If the first RG Page is not the last page in the first RG Block, the data migration rule includes: Use the data to be written in the cache unit corresponding to the first Block position in the first cache space as the redundant data of the second RG Page, and migrate the redundant data of the second RG Page to the second cache space.
4. The method according to claim 2, wherein Determining the data migration rule according to the position layout of the Block of the second RG Block includes: Obtain the correspondence between the position of the cache unit pre-storing the data to be written in the first cache space and the position of the Block of the second RG Block; Use the cache unit in the first cache space corresponding to the Block position in the second RG Block and not storing the data to be written as the migration-in position; use the cache unit in the first cache space not corresponding to the Block position in the second RG Block and storing the data to be written as the migration-out position; When the number of migration-in positions is greater than or equal to the number of migration-out positions, migrate the data to be written in the cache unit at the migration-out position in the first cache space to the cache unit at the migration-in position in the first cache space; When the number of migration-in positions is less than the number of migration-out positions, first migrate the data to be written in each cache unit at the migration-out position in the first cache space to each cache unit at the migration-in position in the first cache space; after pre-storing the data to be written in each cache unit at the migration-in position in the first cache space, use the data to be written not migrated at the migration-out position as the redundant data of the second RG Page, and migrate the redundant data of the second RG Page to the second cache space.
5. The method according to any one of claims 2-4, characterized in that, Migrating the redundant data of the second RG Page to the second cache space includes: Obtain the position layout of the Block of the third RG Block to which the third RG Page belongs; Migrate the redundant data of the second RG Page to the cache unit in the second cache space with the same position layout as the Block of the third RG Block.
6. The method according to claim 5, wherein After migrating the redundant data of the second RG Page to the second cache space, the method further includes: Continue to pre-store the data to be written in the second cache space, so that each cache unit in the second cache space with the same position layout as the Block of the third RG Block pre-stores the data to be written.
7. A fast processing device for write errors in an SSD flash memory block, characterized in that, The RG Block of the SSD is composed of selecting one good block from some or all planes in each Die constituting the RGBlock, and the device includes: The arrangement module is used to cache the data to be written into the cache units of the preset first cache space according to the position layout of each block in the first RG Block when performing a write operation on the first RG Block, so that the position layout of the cache units storing the data to be written in the first cache space is the same as the position layout of the blocks in the target RG Block; The write operation module is used to obtain the position information where a write error occurs when sequentially writing the data to be written in the first cache space into the first RGPage of the first RG Block. The position information where the write error occurs includes the position information of the first block where the write error occurs and the position information of the first RG Page; The management module is used to mark the first block as a bad block and remove it from the first RG Block after obtaining the position information where the write error occurs; mark the data in the first RG Page as invalid data, and retain the data in each RG Page in the first RG Block where the write operation has been successful; The migration module is used to determine the data migration rule of the data to be written pre-stored in the first cache space according to the position information where the write error occurs, and migrate the data to be written in the first cache space within the first cache space, and / or migrate the data to be written in the first cache space between the first cache space and the preset second cache space; The write operation module is further used to write the data to be written after data migration in the first cache space into the second RG Page adjacent to the first RGPage; if the data to be written in the first cache space is migrated between the first cache space and the preset second cache space, write the data to be written after data migration in the first cache space into the second RG Page adjacent to the first RG Page, and write the data to be written migrated to the second cache space into the third RG Page adjacent to the second RG Page.
8. The device according to claim 7, characterized in that The migration module includes: A judgment sub-module, used to judge whether the first RG Page is the last page in the first RG Block; A migration determination sub-module, used to if the first RG Page is not the last page in the first RG Block, the data migration rule is to migrate the data to be written in the first cache space between the first cache space and the preset second cache space; The migration determination sub-module is further used to if the first RG Page is the last page in the first RG Block, obtain the position layout of the blocks of the second RG Block to which the second RG Page belongs, and determine the data migration rule according to the position layout of the blocks of the second RG Block; wherein, the determining the data migration rule according to the position layout of the blocks of the second RG Block includes: Migrate the data to be written in the first cache space within the first cache space; or, Migrate the data to be written in the first cache space within the first cache space, and migrate the data to be written in the first cache space between the first cache space and a preset second cache space.
9. An SSD device, characterized in that, The SSD device includes a storage controller, the storage controller includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the method according to any one of claims 1-6.
Citation Information
Patent Citations
Management method of memory block, write operation method of memory, and memory
CN113220508A
Bad block management method based on NAND Flash
CN113921071A