A High-Reliability Data Access Method for Spaceborne Solid-State Memories Based on NAND Flash

The method addresses unreliable data storage in NAND Flash by identifying and managing bad blocks and bit flips in space radiation using ECC and RAID5, ensuring reliable and efficient data recovery in star-embedded systems.

CN115357185BActive Publication Date: 2025-07-15XIAN INSTITUE OF SPACE RADIO TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210910171.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-29
Publication Date
2025-07-15
Estimated Expiration
2042-07-29

AI Technical Summary

Technical Problem

In the prior art, NAND Flash chips are prone to data errors due to initial bad blocks and space radiation in satellite-borne aerospace applications, and traditional error correction measures cannot effectively handle chip LUN-level data errors, resulting in data loss and a lower write rate.

Method used

The high-reliable data access method based on NAND Flash is adopted. By establishing a static bad block table, building a LUN channel strip page, performing ECC encoding and RAID5 writing, identifying potential bad blocks, RAID5 recovery data, and combining ECC error correction algorithm to achieve reliable access to data.

Benefits of technology

It realizes the reliability of data storage in different error scenarios, reduces resource usage, improves storage space utilization, simplifies processing flow, and ensures data accuracy and write rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115357185B_ABST
    Figure CN115357185B_ABST
Patent Text Reader

Abstract

The present invention designs a high-reliability data access method for on-board solid-state storage based on NAND Flash. By establishing a static bad block linked list, constructing LUN channel stripe pages, page segmentation and in-page ECC encoding, writing stripe page data with RAID5, handling page programming errors, data decoding and error correction, RAID5 data recovery, and pre-identifying potential bad LUNs or bad blocks through decoding threshold judgment, and storing and backing up key information remotely and synchronizing it regularly, it can pre-identify and isolate and eliminate static bad blocks, dynamic bad blocks, and bad LUNs that randomly appear during the read and write operations of the NAND Flash chip, and achieve flexible and reliable data access when data errors occur under different granularity dimensions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a highly reliable data access method for on-board solid state storage based on NAND Flash, and belongs to the field of on-orbit data intelligent processing. Background Art

[0002] With the rapid development of satellite remote sensing technology, the original data rate of payloads is as high as several Gbps or even dozens of Gbps. The development of on-board computing and processing systems has always been accompanied by the continuous update of storage technology. As the direct carrier connecting the storage of payload data and user mode interaction, storage plays an increasingly important role in on-board computing systems. Especially with the rise of on-orbit intelligent processing, cloud computing, and cloud storage technologies, the performance requirements for storage are getting higher and higher: not only a very high throughput rate is required, but also efficient and reliable management and maintenance of stored data are needed. There are various carriers that can be used as storage media. As an outstanding one with excellent performance among storage media, NAND Flash has the advantages of relatively low production and manufacturing costs, large storage capacity, good earthquake and magnetic resistance, fast read and write speeds, and non-volatile data when power is off, and is favored in both on-board aerospace and civilian consumer fields. However, due to the technological characteristics of NAND Flash chips, there will be a certain number of initial bad blocks when the chips leave the factory, and bad blocks will also be generated during use. In addition, individual bit flips will occur when the storage medium is used. Especially in the harsh space radiation environment of on-board aerospace applications, single event upsets will also cause errors in the data stored in NAND Flash chips. For the data storage of on-board aerospace applications, the detected and acquired data is precious, and the correctness of the data is particularly important. It is necessary to take targeted and reliable management measures to ensure the correct storage of data. Even if new bad blocks are generated during use or data errors occur due to factors such as space radiation and the number of uses, the stored data can be restored to correctness through certain management and maintenance strategies and error correction and detection algorithms, so as to achieve the reliability and correctness of stored data.

[0003] In the design of traditional spaceborne solid-state storage, simple bad block management and error correction and detection algorithms are generally adopted to manage and maintain data; and for the identification of potential bad blocks, it is realized through regular self-check operations that have nothing to do with service data. However, for the data error situation at the chip LUN level, no corresponding treatment measures are basically taken, resulting in data errors in the stored data and inability to recover when the chip LUN fails; either multiple groups of chips are used for redundant backup, which consumes a large amount of resources and costs. In addition, for data errors caused by page programming failure during high-speed data writing, data rewriting operations need to be carried out, which will significantly reduce the effective data rate actually written into the NAND Flash, and thus data loss may occur due to insufficient cache resources or untimely data processing. Therefore, a spaceborne solid-state storage high-reliability access method that occupies less resources and can be applied to different levels and different error scenarios is needed. Summary of the Invention

[0004] The technical problem solved by the present invention is: aiming at the problem of poor reliability and correctness of stored data that easily occurs in the existing traditional data access technology, a spaceborne solid-state storage high-reliability data access method based on NAND Flash is proposed.

[0005] The present invention solves the above technical problems through the following technical solutions:

[0006] A spaceborne solid-state storage high-reliability data access method based on NAND Flash, comprising:

[0007] Determine all static bad blocks in the storage array and establish a static bad block table;

[0008] Perform strip page construction on the storage array, obtain all LUN channel strip pages and sort them;

[0009] Write the original data and encoded check information to each LUN channel strip page, split the original data for ECC encoding, perform exclusive OR check on each LUN channel strip page to obtain exclusive OR checksum data, and at the same time perform ECC encoding on the checksum data;

[0010] According to the LUN channel strip pages, for all data blocks except static bad blocks, write the encoded data and exclusive OR checksum data in the order sorted by the LUN channel strip pages;

[0011] Mark the block addresses where the writing fails or the programming status word of the LUN channel strip page is incorrect;

[0012] Perform data read-back recovery on the marked block addresses;

[0013] Statistics are collected based on the data readback recovery results of each LUN channel stripe page to determine the data decoding error rate of each LUN channel stripe page, and the error rate threshold corresponding to the dynamic bad block is determined. Potential bad blocks are identified in advance based on the error rate threshold.

[0014] The addresses of potential bad blocks and dynamic bad blocks are backed up and stored for the next storage identification.

[0015] The method for establishing a static bad block table is:

[0016] The first byte of the first page or second page preparation area of each data block of the storage array is read to determine the initial bad block, the storage chips except the initial bad block are erased, the failed blocks that fail to be erased are recorded, and the erased failed blocks and the initial bad blocks are regarded as static bad blocks to establish a static bad block table.

[0017] The specific steps to obtain the LUN channel stripe page based on the stripe page construction are:

[0018] Each storage chip is composed of k LUN channels, and n storage arrays are equivalent to k*n LUN channels. A page in each LUN channel is extracted, and a total of k*n pages form a stripe page for striping construction. Each stripe page includes a data area, a check area, and a reserved area. Among them, the data area and the check area of each stripe page constitute RAID5. The data area is composed of i subpages of the stripe page, the check area is composed of 1 subpage of the stripe page, and the reserved area is composed of j subpages in the stripe page, i+1+j=k*n.

[0019] After the original data is segmented, it is segmented into segments with the same number of LUN channel stripe pages, and ECC encoding is performed. XOR verification is performed according to the LUN channel stripe pages. When the original data is less than one stripe page, fixed data is used for padding and completion processing, wherein:

[0020] The first i pages in each LUN channel stripe page are used to store the encoded data, and the i+1th subpage in each LUN channel stripe page is used to store the encoded XOR checksum data corresponding to the RAID5 of the stripe.

[0021] The specific process of encoding data, XOR check and writing encoded data is as follows:

[0022] Data is first written to the first LUN channel stripe page, and then data is first written to the second LUN channel stripe page, and so on, until data writing is completed to all LUN channel stripe pages in the data area and the check area.

[0023] During the data readback process, data recovery is performed on the LUN channel stripe sub-pages within the data decoding and error correction capability range, and the data blocks where the LUN channel stripe sub-pages exceeding the data readback error correction capability range are marked as dynamic bad blocks;

[0024] Among them, the data readback error correction capability range is determined according to the lowest error correction capability, and is one to three times the lowest error correction capability.

[0025] The specific steps of data readback recovery are as follows:

[0026] Read the data of each LUN channel stripe page in sequence and perform ECC decoding. When the decoded error data of any LUN channel stripe page is within the error correction capability range, only the correct data within the page is recovered through the error correction algorithm. When the decoded error data of any LUN channel stripe page exceeds the error correction capability range, recovery processing is performed through stripe page RAID5; if the data error of any sub-page in the stripe page exceeds the error correction capability, the error sub-page data is recovered through RAID5 according to other LUN channel stripe pages without errors or within the error correction capability range, and the recovered correct data is output and moved to the spare area LUN channel stripe page, and the address mapping table is updated; if there are two or more LUN channel data errors in the current faulty LUN channel stripe page and cannot be recovered, no recovery is performed and it is marked as a dynamic bad block.

[0027] The error rate thresholds include the in-page error correction threshold, the number of weak pages within a block threshold, and the number of weak blocks within a LUN threshold. When data readback recovery is performed again, potential bad blocks are directly excluded and no subsequent write operation is performed.

[0028] The dynamic bad blocks also include dynamic weak blocks. After decoding the data of the LUN channel stripe page, when the data decoding error rate exceeds twice the lowest error correction capability but does not exceed three times the lowest error correction capability, the LUN channel stripe page is defined as a weak page, and the specific data block is defined as a dynamic weak block, and the dynamic weak blocks and dynamic bad blocks are uniformly processed.

[0029] For dynamic bad blocks with write failures or incorrect programming status words of LUN channel stripe pages, dynamic bad blocks exceeding the data readback error correction capability range, and dynamic bad blocks with two or more LUN channel data errors in the LUN channel stripe page and cannot be recovered, filter out the block addresses with incorrect programming status words caused by space irradiation, and perform AND operation processing. Mark all processed bad blocks and store their addresses together with static bad blocks, potential bad blocks, and weak blocks.

[0030] The advantages of the present invention compared with the prior art are as follows:

[0031] (1) A highly reliable data access method for on-board solid-state storage based on NAND Flash provided by the present invention adopts a unified processing architecture, realizing reliable data storage in case of errors at the chip LUN level, block level within the LUN, strip page level between RAID5 stripes, and sub-page encoding and decoding level; at the same time, it can achieve closed-loop self-recovery processing in case of different data errors.

[0032] (2) Combining the characteristics of NAND Flash, the present invention divides the storage area into a data storage area, a check information storage area, and a reserve replacement area in units of LUN. Through redundant replacement at the LUN channel level, compared with chip-level redundant replacement, it can greatly reduce the layout area and improve the storage space utilization rate.

[0033] (3) The present invention adopts a hierarchical and classified management strategy. By combining in-page encoding and decoding with strip page RAID5 between LUNs, it can correctly recover data errors caused by newly added bad blocks, bit flips, and single-event upsets in the storage operation process.

[0034] (4) According to the ECC error correction algorithm capabilities, the present invention can flexibly set the in-page error correction threshold, the number of weak pages within a block threshold, and the number of weak blocks within a LUN threshold. During the current read operation, potential bad blocks or bad LUNs can be identified, isolated, and replaced in advance, thus greatly reducing the error probability during subsequent data reading and writing operations by users.

[0035] (5) The present invention constructs a RAID5 write stripe through LUNs between chip groups. For sub-pages with programming errors, there is no need to rewrite data and handle pipeline interrupts, which can greatly simplify the processing flow. Through RAID5 read operations, page programming error data can be correctly recovered. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 Flow chart of the highly reliable data access method for on-board solid-state storage provided by the invention;

[0037] Figure 2 Typical organizational structure diagram of a single NAND Flash chip provided by the invention;

[0038] Figure 3 Strip page construction diagram of the NAND Flash storage array provided by the invention;

[0039] Figure 4 Flow chart of in-page ECC and RAID5 data writing in NAND Flash provided by the invention;

[0040] Figure 5 Hardware composition block diagram of the solid-state storage single board provided by the invention;

[0041] Figure 6Comparison diagram of page data writing methods provided for the invention;

[0042] Figure 7 Page programming write timing diagram of strip page internal pipeline operation provided for the invention;

[0043] Figure 8 Comparison diagram of processing flowcharts for errors occurring during the page programming process provided for the invention;

[0044] Figure 9 Flowchart of in-page error correction and inter-page RAID5 recovery during data read operation provided for the invention;

[0045] Figure 10 Flowchart of early identification of potential bad blocks and bad LUNs provided for the invention; Detailed implementation manners

[0046] A highly reliable data access method for on-board solid-state storage based on NAND Flash, which completes the early identification and isolation of static bad blocks, dynamic bad blocks, and bad LUNs that randomly appear during the read and write operations of the NAND Flash chip through the establishment of a static bad block list, the construction of LUN channel strip pages, page segmentation and in-page ECC encoding, strip page data RAID5 writing, page programming error handling, data decoding and error correction, RAID5 data recovery, the early identification of potential bad LUNs or bad blocks through decoding threshold judgment, and the off-site storage backup and regular update synchronization of key information, and realizes flexible and reliable data access when data errors occur under different granularity dimensions. The specific steps are as follows:

[0047] Determine all static bad blocks in the storage array and establish a static bad block table;

[0048] Conduct strip page construction on the storage array, obtain all LUN channel strip pages and sort them;

[0049] Write the original data to each LUN channel strip page, split and process the original data and then perform ECC encoding, perform exclusive OR checksum on each LUN channel strip page, and obtain the exclusive OR checksum data;

[0050] According to the LUN channel strip pages, for all data blocks except static bad blocks, write the encoded data and the exclusive OR checksum data in the order of the LUN channel strip page sorting;

[0051] Mark the block addresses where the writing fails or the programming status word of the LUN channel strip page is incorrect;

[0052] Perform data read-back recovery on the marked block addresses;

[0053] Statistically analyze the data read-back recovery results of each LUN channel strip page, determine the data decoding error rate of each LUN channel strip page, and determine the error rate threshold corresponding to dynamic bad blocks. Identify potential bad blocks in advance based on the error rate threshold;

[0054] Back up and store the addresses of potential bad blocks and dynamic bad blocks for next storage identification.

[0055] The method for establishing a static bad block table is as follows:

[0056] Read the first byte of the first page or the second page pre-reserved area of each data block in the storage array to determine the initial bad blocks. Erase the storage chips of the initial bad blocks, record the erase failure blocks that fail to be erased, and use both the erase failure blocks and the initial bad blocks as static bad blocks to establish a static bad block table;

[0057] The specific steps for constructing and obtaining LUN channel strip pages according to strip pages are as follows:

[0058] Equivalent n storage arrays to k*n LUN channels, perform striping construction, extract one page from each LUN channel, and a total of k pages form a strip page. Each strip page includes a data area, a parity area, and a reserved area. Among them, the data area and the parity area of each strip page form RAID5. The data area is composed of i sub-pages of the strip page where it is located, the parity area is composed of the sub-pages of the strip page where it is located, and the reserved area is composed of j sub-pages in the strip page, and i + 1 + j = k*n;

[0059] After the original data is segmented, it is segmented into the same number of segments as the number of LUN channel strip pages and ECC encoded, and XOR verification is performed according to the LUN channel strip pages, where:

[0060] Store the encoded data in the first i pages of each LUN channel strip page, and use the (i + 1)-th sub-page of each LUN channel strip page to store the XOR checksum data corresponding to the RAID5 of the strip;

[0061] The specific process of writing the encoded data and the XOR checksum data is as follows:

[0062] First write data to the first LUN channel strip page, then write data to the second LUN channel strip page, and so on, until all LUN channel strip pages have completed data writing;

[0063] During the data read-back process, recover the data of the LUN channel strip pages within the data read-back error correction capability, and mark the data blocks where the LUN channel strip pages exceeding the data read-back error correction capability are located as dynamic bad blocks;

[0064] The specific steps for data read-back recovery are as follows:

[0065] Read the data of each LUN channel strip page in sequence and perform ECC decoding. When the decoded error data of any LUN channel strip page is within the error correction capability range, only the correct recovery of the in-page data is performed through the error correction algorithm. When the decoded error data of any LUN channel strip page exceeds the error correction capability range, the error correction capability range is passed through;

[0066] Output the correct data of RAID5 of other error-free LUN channel strip pages to the spare area LUN channel strip page. If two or more LUN channel data in the current faulty LUN channel strip page are in error and cannot be recovered, no recovery is performed and it is marked as a dynamic bad block. Otherwise, the correct data of the spare area LUN channel strip page is used to recover the data of the current faulty LUN channel strip page;

[0067] The error rate thresholds include the in-page error correction threshold, the number of weak pages within a block threshold, and the number of weak blocks within a LUN threshold. When data is read back and recovered again, potential bad blocks are directly excluded and no subsequent write operation is performed;

[0068] Dynamic bad blocks also include dynamic weak blocks. After decoding the data of the LUN channel strip page, when the data decoding error rate exceeds twice the minimum error correction capability but is not greater than three times the minimum error correction capability, the LUN channel strip page is defined as a weak page, and the specific data block is defined as a dynamic weak block. The dynamic weak blocks and dynamic bad blocks are uniformly processed;

[0069] For dynamic bad blocks with write failures or incorrect programming status words of LUN channel strip pages, dynamic bad blocks beyond the data read-back error correction capability range, and dynamic bad blocks with two or more LUN channel data errors in the LUN channel strip page and unable to be recovered, filter out the block addresses with incorrect programming status words caused by space irradiation and perform AND operation processing. Mark all processed bad blocks and store their addresses together with static bad blocks, potential bad blocks, and weak blocks.

[0070] The following is a further description based on the specification drawings and preferred embodiments:

[0071] When designing on-board solid-state memory, multiple NAND Flash chips are generally used to form a storage array to achieve improvements in capacity and rate bandwidth. Its typical hardware block diagram is as Figure 5 shown, mainly consisting of an FPGA chip or other controller, external non-volatile MRAM / NOR Flash, RAM cache (SRAM, DRAM, DDR, etc.), NAND Flash storage array, and input / output interface. If the number of NAND Flash chips used on a single board is n (n is a natural number), and the number of LUNs in each chip is k (k is a natural number), then the total number of LUNs on the single board is k * n.

[0072] The specific operation process of the high-reliability data access method for on-board solid-state storage based on NAND Flash is as follows:

[0073] (1) Since the invalid blocks in the NAND Flash at the time of factory shipment are pre-marked by the manufacturer, the initial bad block table, denoted as T, can be obtained by reading the first byte of the first page or the second page of the preparation area of each block in the storage array. init Erase the blocks of the storage chip except for the initial bad blocks, and count the erased invalid blocks, denoted as T er . Thus, the static bad blocks T st of the storage array = T init + T er ;

[0074] (2) For all n NAND Flash chips, each chip consists of k LUNs. According to the organizational structure characteristics of the NAND Flash chips, k*n LUNs are constructed into RAID stripes, as Figure 3 shown. The corresponding one page of each LUN, a total of k*n pages, forms a stripe page. Each stripe page is divided into three parts: a data area, a parity area, and a reserved area. The data area and the parity area in the stripe page form RAID5. The data area consists of i sub-pages in the stripe page, the parity area consists of 1 sub-page in the stripe page, and the reserved area consists of j sub-pages in the stripe page. According to the organizational structure characteristics of the stripe page, i + 1 + j = k*n. Since the probability of chip LUN failure is extremely low, in order to make the most of all the storage space of the NAND Flash as much as possible, generally j is taken as 1, and the value of j can also be flexibly adjusted according to the actual situation;

[0075] (3) Each manufacturer's NAND Flash chip is required to have a certain ECC error correction ability. Define the minimum required ECC error correction ability of the chip as E min . Taking the 256Gbit Micron SLC chip as an example, the minimum ECC error correction ability required by the manual is to correct 8-bit data errors per 540 bytes, that is, E min = 8 / (540*8)). Commonly used ECC algorithms include Hamming code, RS code, BCH code, LDPC code, etc. The complexities of different algorithms vary greatly; for the same algorithm, the stronger the error correction ability, the more resources are consumed. The designed error correction ability is required to be greater than E min . Combining with the actual resources, generally the designed error correction ability E d is considered according to 3 times of E min (E d can be flexibly adjusted according to the actual situation).

[0076] Since the basic unit of NAND Flash read and write is a page, in order to improve the write speed of the NAND Flash chip array, after caching the original data to be written, a request is sent from the request cache queue each time, and the request is split into several sub-requests in units of pages. First, the page data of these sub-requests is encoded by the ECC algorithm, and then a LUN channel strip page is constructed as shown in Figure 3 Finally, the exclusive-OR checksum and data of each strip sub-page are generated respectively, that is, P = D1 xor D2 xor D3 xor…xor D i . The first i pages in each strip store user data and ECC encoding, and the (i + 1)th sub-page in each strip is used to store the exclusive-OR sum corresponding to the RAID5 strip (generally using exclusive-OR or parity check, this patent uses the exclusive-OR method). The flowchart of page splitting and in-page ECC encoding is shown in Figure 4 ;

[0077] (4) Based on the strip page constructed in the above step (2), for all storage spaces except the static bad block T st , the user page data and encoding information in Figure 4 and the strip page checksum information and encoding information are in the order of the strip pages in Figure 3. First, data is written to strip page 0, that is, the first page Page 10 of LUN1 is programmed and written, then the first page Page 20 of LUN2 is programmed and written, …, and then the first page Page (i) of LUN (i)0 is programmed and written. The data of the first (i) pages is XORed to obtain the checksum information and encoding information of this strip page, and finally the checksum and encoding information are written to the first page Page (i+1) of LUN (i+1)0 ; At this point, the writing of the first strip page data is completed. The writing of other strip page data is carried out in turn according to the operation sequence of strip page 0. The comparison of the page data writing method of this patent with other methods is shown in Figure 6 ;

[0078] (5) Although the read and write clock cycles of NAND Flash chips are very fast, a certain waiting time is required during the read, write, and erase operations. If the next operation is executed after waiting for the completion of the current operation each time, the long waiting time will greatly affect the average rate. Usually, the write operation of NAND Flash chips mainly includes three steps:

[0079] Loading operation, mainly completing the loading of programming commands, addresses, and data;

[0080] The automatic programming operation, that is, the programming operation of writing the data loaded into the page data register into the internal storage unit (Cell) is actively completed by the Flash chip;

[0081] The detection operation. After the automatic programming is completed, it is detected whether the written data is programmed correctly. If it is correct, subsequent operations can be carried out; if it is incorrect, other processing measures need to be taken;

[0082] According to the page programming working principle of the storage chip, its automatic programming is automatically completed inside the chip. At this time, the chip cannot respond to new commands and does not require external intervention. Through the pipeline technology, the internal programming time of the chip is fully utilized to perform loading operations on other chips. Since the LUN of each chip is an independent logical operation unit, it can independently respond to instructions and operations respectively. The pipeline timing diagram of the data writing operation within the strip page is as shown in Figure 7 shown. When the (i + 1) LUN pipeline starts running, in any time period, there are always several small operations being executed simultaneously, thus realizing data multiplexing on the time slice and greatly improving the data writing rate. Due to the differences of NAND Flash chips with different specifications, their internal programming times are different, and different pipeline operation levels can be set according to different internal programming times.

[0083] According to Figure 7 shown, after continuously loading the (i + 1) page data of the pipeline, generally the internal automatic programming of the data in the first LUN should have been completed, and it can be confirmed by detecting the R / B signal of its LUN. In this patent, the continuous (i) page user data, ECC encoding, the exclusive OR sum of the first (i) page user data, and the EEC encoding data are distributed and written into LUN1 to LUN (i+1) In this way, when a page programming fails, only the block address where the page is located needs to be marked. wr For the incorrect data of the page programming failure, it can be correctly recovered through the subsequent step (6) data error correction and detection and step (7) RAID5 read data.

[0084] In other processing methods, when writing data, the data is first continuously written into all the blocks corresponding to a certain LUN1, and then other data is sequentially written into LUN2, and so on. When page programming fails, generally, the backup replacement Flash chipset is reloaded with commands and addresses, and at the same time, the data of the corresponding LUN in the cache is rewritten and loaded into the storage area. After the data page programming is successful, it is restored to the normal pipeline operation sequence. Due to the rewrite of page data, the actual effective writing rate into the solid state memory will be significantly reduced; at the same time, due to the interruption and recovery of the pipeline, the control logic will also be more complex. For spaceborne aerospace applications, if a single event effect occurs in space and a single bit flip occurs when reading the status word after page programming, it will also cause incorrect rewrite operations. Especially when there are consecutive bad blocks in a certain LUN, the required cache space is very large and the consumed resources are huge.

[0085] In this patent, since there is no need to rewrite the data of the page programming failure to the spare area during the writing process, the complex processing flow during page programming failure is greatly simplified, and the cache space will not overflow or the effective data writing rate into the solid state memory will not be significantly reduced due to data rewrite. As described above, since RAID5 is formed within the stripe page between different LUNs, even if all the pages in a block within a certain LUN fail due to the failure of all the blocks, the correct data can be restored through RAID5. The method used in this patent does not require rewrite operation processing for the page data of programming failure, so the writing pipeline operation will not be interrupted, and thus the data writing rate into the Flash can be maintained at a high level. The processing flow of page programming error in this patent is as follows Figure 8 shown;

[0086] (6) When reading back data, according to the error correction ability of the selected algorithm, errors within the designed error correction ability E d can be corrected and recovered, and errors beyond the algorithm's ability cannot be corrected. In use, the block where the data page beyond the error correction ability is located is marked as a bad block. Limited by the error correction algorithm ability, the ECC algorithm can only correct errors when the data within its ability range is incorrect. For situations beyond the ECC error correction ability, the correct data can be obtained through step (7) RAID5 data recovery;

[0087] (7) According to the above description, through the construction of LUN channel striping, page segmentation, and writing of RAID5 striped data pages, the data written in the first (i) sub-pages of each strip is user data and coding information (collectively referred to as data information). The check information P of the (i + 1)th sub-page is obtained by performing an exclusive OR calculation on the corresponding bytes of each page in the first (i) columns, and ECC coding is performed on the check information P. Therefore, in the storage array strip, not only data information is written, but also check bit information is written. The pages corresponding to each LUN form RAID5 striped pages. The schematic diagram of data storage in the striped pages is shown in Table 1;

[0088] Table 1 Schematic Diagram of Data Storage in a Striped Page

[0089]

[0090]

[0091] The process of reading data stored in NAND Flash is as follows:

[0092] According to the playback instruction, the data stored in each page strip is read sequentially;

[0093] Perform ECC decoding on the page data in each LUN read. When the decoding errors of all bytes in a page are within the E d error correction capability range, the correct recovery of the data within the page can be achieved only through the error correction algorithm. When the data of a certain page in the striped page has more errors than E d error correction capability, the error in the current page area cannot be corrected through the error correction algorithm itself, and further RAID5 is adopted for data recovery;

[0094] When the error of the page data exceeds the error correction capability range of the ECC algorithm, the correctness of the page data is restored through RAID5. The incorrect page data is restored based on the other correct page data in the striped page, and the restored correct data is output. At the same time, the correct data is moved to a new spare area page, and then the address mapping table is updated. The error recovery principle is shown in Table 2 below; if more than two pages of data in the striped page have decoding errors (the probability of this situation is extremely low), it exceeds the data recovery capability of RAID5, and no data recovery operation is performed. The error correction and RAID5 recovery flow chart of data readback is as Figure 9 shown;

[0095] Table 2 Schematic Diagram of Data Error in a Striped Page

[0096]

[0097] Sub - page 3 Data Recovery Principle: Data to be recovered = Check data (i + 1) xor Correct data 1 xor Correct data 2 xor Correct data 4 xor... xor Correct data (i - 1) xor Correct data (i);

[0098] (8) Early identification of potential bad LUNs or bad blocks

[0099] Early identification of potential bad LUNs or bad blocks is an important application. When each chip manufacturer conducts stacked integrated packaging of multiple LUNs and designs the hardware layout and wiring for engineering applications, there will be some subtle differences in the characteristics between different LUNs. As a result, some LUNs or Blocks of NAND Flash chips soldered on the physical single - board are not marked as bad blocks at the factory. However, with the high - low temperature changes during aerospace applications or multiple erase - write operations in a thermal - vacuum environment, weak LUNs or weak blocks appear. A weak LUN is between a bad LUN and a good LUN, and tends to change towards a bad LUN as the number of uses increases or the environment deteriorates; a weak block is also between a bad block and a good block and tends to develop towards a bad block. If potential weak LUNs and weak blocks are not identified and removed in advance, it will lead to serious errors in the stored data and affect the effective throughput rate during read - write operations.

[0100] Since the designed error - correction algorithms are all greater than the minimum error - correction capability range required by the storage chip, during the data read - back process, the decoding situation of the LUN channel stripe pages within the error - correction capability range of the ECC algorithm is statistically analyzed, and the data decoding error rate E of the stripe pages is obtained. c Judgment is made according to different thresholds. Based on the set intra - page error - correction threshold, the number threshold of weak pages within a block, and the number threshold of weak blocks within a LUN, during the current read - back operation, potential bad blocks or bad LUNs can be identified in advance and marked as T. weak Subsequent write operations will no longer be performed;

[0101] The definition of a weak block is as follows: The data decoding error rate Ec of a certain page in the NAND Flash chip exceeds 2 times the minimum error - correction capability required by the chip but is within the error - correction algorithm capability range actually designed (i.e., 2E min <E c ≤E d) If so, define this page as a weak page; when the number of weak pages a of a certain block exceeds a certain threshold (define the threshold as b, and this threshold can be adjusted as needed), that is, when a ≥ b, the block where it is located can be marked as a "weak block", and the block will be listed in the bad block table during subsequent use, and no further write operations will be performed. Similarly, when multiple "weak blocks" (the number of weak blocks is denoted as f) accumulate in an LUN, when the total number of weak blocks f in the LUN exceeds the threshold (define the threshold as g, and this threshold can be set according to the needs), that is, when f ≥ g, the current LUN can be marked as a "weak LUN", and no further write operations will be performed on this LUN, and a reserved area LUN as shown in Figure 4 is used for replacement. When among the k*n LUNs of all n NAND Flash chips, if there is a weak LUN, by selecting one from the reserved j LUNs for replacement, the number of all good LUNs used remains i unchanged, so that the available storage capacity and the read / write speed will not be reduced due to the weak LUN. The flow chart for early identification of potential bad blocks and bad LUNs is as shown in Figure 10 ;

[0102] (9) Off-site storage backup and regular update synchronization of key information

[0103] Perform an "AND operation" on the block address T where page programming fails in step (5) wr and the marked super ECC error correction ability T rd address, so as to filter out the block addresses where the page programming status word is misjudged due to space irradiation. Denote the bad block after performing the "AND operation" on T wr and T rd as a dynamic bad block, denoted as T dyn . Store the static bad block T st , the dynamic bad block T dyn , the weak block T weak where potential bad blocks or bad LUNs are located, and the address mapping table information not only in the storage array of the NAND Flash itself, but also in off-site backup storage in external devices such as non-volatile MRAM or NOR FLASH externally attached to the FPGA. To ensure the correctness and reliability of the storage of this information, it is stored in a way that goes through EDAC and the two-out-of-three redundancy method. At the same time, for the high reliability of on-board use, when the solid state memory is idle during each operation, the relevant information stored in the solid state memory is regularly updated and stored in the external MRAM or NORFLASH. In this way, even if the information in the storage array is incorrect, the correct information can be restored through the external non-volatile storage.

[0104] In the above operation process, steps (2) to (5) are the recording operations of NAND Flash; steps (6) to (8) are the playback operations of NAND Flash. The present invention constructs RAID5 data writing based on an independent LUN channel, distributes and writes the consecutive i-page data into strip pages after ECC encoding, performs an exclusive OR operation on the first i-page data to obtain the parity information of the strip page data, and writes the parity information into the strip after ECC encoding. By combining ECC error correction with RAID5 and according to a custom threshold, reliable storage of user data can be achieved, and potential bad blocks and bad LUNs can be identified and removed in advance. The present invention adopts a unified processing architecture to achieve reliable storage management of data in case of errors at the LUN level, block level within the LUN, strip page level between RAID5, and sub-page encoding and decoding level of the on-board solid-state storage chip; at the same time, closed-loop self-recovery processing can be realized in case of different data errors.

[0105] Although the present invention has been disclosed above with preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make possible changes and modifications to the technical solution of the present invention by using the methods and technical contents disclosed above without departing from the spirit and scope of the present invention. Therefore, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention without departing from the technical solution of the present invention shall fall within the protection scope of the technical solution of the present invention.

[0106] The content not described in detail in the specification of the present invention belongs to the well-known technology of those skilled in the art.

Claims

1. A high-reliability data access method for on-board solid-state storage based on NAND Flash, characterized in that include: Determine all static bad blocks in the storage array and establish a static bad block table; Build stripe pages for the storage array, obtain all LUN channel stripe pages and sort them; Write the original data and coding check information to each LUN channel stripe page, perform ECC encoding after segmenting the original data, perform XOR check on each LUN channel stripe page, obtain XOR checksum data, and perform ECC encoding on the checksum data; According to the LUN channel stripe page, for all data blocks except static bad blocks, the encoded data and XOR checksum data are written in the order of the LUN channel stripe page; Mark the block addresses where writing fails or the programming status word of the LUN channel stripe page is incorrect; Read back and recover the data of the marked block address; Statistics are collected based on the data readback recovery results of each LUN channel stripe page to determine the data decoding error rate of each LUN channel stripe page, and the error rate threshold corresponding to the dynamic bad block is determined. Potential bad blocks are identified in advance based on the error rate threshold. The addresses of potential bad blocks and dynamic bad blocks are backed up and stored for the next storage identification.

2. The method for high-reliability data access based on NAND Flash on-board storage according to claim 1, characterized in that: The method for establishing a static bad block table is: The first byte of the first page or second page preparation area of each data block of the storage array is read to determine the initial bad block, the storage chips except the initial bad block are erased, the failed blocks that fail to be erased are recorded, and the erased failed blocks and the initial bad blocks are regarded as static bad blocks to establish a static bad block table.

3. The method for high-reliability data access based on NAND Flash on-board storage according to claim 1, characterized in that: The specific steps to obtain the LUN channel stripe page based on the stripe page construction are: Each storage chip is composed of k LUN channels, and n storage arrays are equivalent to k*n LUN channels. A page in each LUN channel is extracted, and a total of k*n pages form a stripe page for striping construction. Each stripe page includes a data area, a check area, and a reserved area. Among them, the data area and the check area of each stripe page constitute RAID5. The data area is composed of i subpages of the stripe page, the check area is composed of 1 subpage of the stripe page, and the reserved area is composed of j subpages in the stripe page, i+1+j=k*n.

4. The high-reliability data access method based on NAND Flash on-board storage according to claim 3 is characterized in that: After the original data is segmented, it is segmented into segments with the same number of LUN channel stripe pages, and ECC encoding is performed. XOR verification is performed according to the LUN channel stripe pages. When the original data is less than one stripe page, fixed data is used for padding and completion processing, wherein: The first i pages in each LUN channel stripe page are used to store the encoded data, and the i+1th subpage in each LUN channel stripe page is used to store the encoded XOR checksum data corresponding to the RAID5 of the stripe.

5. A spaceborne fixed storage high-reliability data access method based on NAND Flash according to claim 4, characterized in that: The specific process of writing the encoded data, the XOR checksum, and the encoded data is as follows: First write data to the strip pages of the first LUN channel, then write data to the strip pages of the second LUN channel, and so on, until all the strip pages of the data area and the check area LUN channels have completed data writing.

6. A spaceborne fixed storage high-reliability data access method based on NAND Flash according to claim 5, characterized in that: During the data readback process, data recovery is performed on the sub-pages of the LUN channel strip within the data decoding and error correction capability range, and the data block where the sub-page of the LUN channel strip exceeds the data readback error correction capability range is marked as a dynamic bad block; Among them, the data readback error correction capability range is determined according to the lowest error correction capability, and is one to three times the lowest error correction capability.

7. A spaceborne fixed storage high-reliability data access method based on NAND Flash according to claim 6, characterized in that: The specific steps of data readback and recovery are as follows: Read the data of each LUN channel strip page in sequence and perform ECC decoding. When the decoded error data of any LUN channel strip page is within the error correction capability range, only the correct data within the page is recovered through the error correction algorithm. When the decoded error data of any LUN channel strip page exceeds the error correction capability range, recovery processing is performed through strip page RAID5; if the data error of any sub-page in the strip page exceeds the error correction capability, the error sub-page data is recovered through RAID5 according to other error-free or LUN channel strip pages within the error correction capability range, and the recovered correct data is output and moved to the strip page of the spare area LUN channel, and the address mapping table is updated; if two or more LUN channel data errors occur in the current error LUN channel strip page and cannot be recovered, no recovery is performed and it is marked as a dynamic bad block.

8. A spaceborne fixed storage high-reliability data access method based on NAND Flash according to claim 7, characterized in that: The error rate thresholds include the in-page error correction threshold, the number of weak pages within a block threshold, and the number of weak blocks within a LUN threshold. When data readback and recovery are performed again, potential bad blocks are directly excluded and no subsequent write operation is performed.

9. A spaceborne fixed storage high-reliability data access method based on NAND Flash according to claim 6, characterized in that: The dynamic bad block also includes a dynamic weak block. After decoding the data of the LUN channel strip page, when the data decoding error rate exceeds twice the lowest error correction capability but is not greater than three times the lowest error correction capability, this LUN channel strip page is defined as a weak page, and the specific data block is defined as a dynamic weak block, and the dynamic weak block and the dynamic bad block are uniformly processed.

10. A spaceborne fixed storage high-reliability data access method based on NAND Flash according to claim 9, characterized in that: For dynamic bad blocks with write failures or incorrect programming status words in LUN channel stripe pages, dynamic bad blocks beyond the data read-back error correction capability range, and dynamic bad blocks with two or more LUN channel data errors in the LUN channel stripe pages and unable to be recovered, filter out the block addresses with incorrect programming status words caused by space irradiation, perform AND operation processing, mark all the processed bad blocks, and store their addresses together with the static bad blocks, potential bad blocks, and weak blocks.

Citation Information

Patent Citations

  • A memory management method and apparatus

    CN102298543A

  • Data storage device, host device, and data writing method

    CN109885506A