Flash original data analysis method and device and medium
By using software descrambling and decoding methods, user data and metadata plaintext can be recovered when flash memory hardware ECC error correction fails. This solves the data analysis problem after ECC error correction failure and improves the efficiency of development and testing.
Patent Information
- Application Number
- CN202511033398.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-11-07
Smart Images

Figure CN120909859A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of solid-state storage technology, in particular to a flash memory raw data analysis method, device and medium, mainly suitable for the development test and problem positioning scene of a solid-state drive (SSD) with NAND Flash as the storage medium. BACKGROUND
[0002] Generally, user data from the host needs to be written into the SSD, which needs to be converted by the algorithm layer (FTL) to convert the logical address (LBA) of the host end into the physical address (PBA) in the flash memory, and additional metadata (Meta Data) and other information are attached on the basis of the user data, which is used to record the parameters related to the NAND storage block itself (such as physical address PBA, block erase count P / E Cycle, write timestamp, etc.). After the user data and the metadata are prepared, due to a inherent characteristic of the NAND flash memory, the difference in the number of 0 / 1 bits directly written into it is too large, which will cause the charge distribution of the storage unit to be unbalanced, the probability of bit flipping during reading data to increase, and the difficulty of data correction to become larger. Therefore, the data needs to be scrambled by the random number scrambling module (Scrambler) in the flash memory controller hardware to re-encode and "shuffle" it, so that the number of 0 / 1 bits in it maintains a roughly uniform spatial distribution in the local and overall, so that the written data is more stable and reduces the probability of random bit flipping in the flash memory. When scrambling the data, the algorithm layer needs to provide a 32-bit length random seed (Scramble Seed) for each 4KB size data block as the input stimulus of the hardware scrambling module, and a copy of the random seed will also be written into the flash memory along with other scrambled data. The next step is to calculate the checksum of the scrambled data using the CRC32 hash algorithm, and generate an ECC error correction code sequence (Error Correction Code) using the LDPC error correction algorithm. These newly generated additional check data and the previous scrambled effective data are placed in the memory space according to certain order and position rules, forming the raw data (Rawdata) to be stored in the NAND flash memory particles. Its composition can be represented as: scrambled user data + scrambled metadata + hash check code + random seed + ECC error correction code + a certain length of reserved segment data (filling "0" and other meaningless data to meet the address alignment and the minimum operation unit size of the flash memory).
[0003] The process of reading data from the flash memory particle by the master control is the reverse process of the above operation. After the target physical address is given at the algorithm layer, the master control first moves the original data in the flash memory from the inside of the particle to the specified location in the memory of the master control through the flash memory physical layer interface (PHY) and the data transfer module (DMA), and then restores and separates the user data and the metadata through the processing steps of the hardware modules such as ECC error correction, hash verification, and random number de-scrambling. However, when the flash memory particle is aged and worn, the data bus transmission signal is disturbed, or the internal software and hardware of the master control chip fail to cause the bit flip to exceed the upper limit of the error correction capability of the ECC algorithm, the ECC error correction module will report decoding failure. In this case, only the original data without error correction and de-scrambling will be returned in the memory. The original data with error correction failure is in the de-scrambled state, and its binary content is approximately uniformly distributed random numbers without repetition, and there is almost no regularity, which is greatly different from the standard data (Pattern) written, and cannot be directly compared, which is not conducive to the developer to analyze the content and error cause of the error data. Therefore, a debugging method is needed for this case to restore the user data and the metadata to plaintext as much as possible to facilitate comparison with the standard data. SUMMARY
[0004] In view of the defects of the prior art, the present application provides a flash memory original data analysis method, device and medium, which uses software to decode the original data to restore the user data and the metadata to plaintext as much as possible to facilitate comparison with the standard data.
[0005] To solve the technical problem, the technical scheme adopted by the present application is as follows: a flash memory original data analysis method, comprising the following steps: S01. When the read command error correction fails, the master control issues a data analysis command to count the number of 0 / 1 bits in the valid part of the original data block corresponding to the read command, and outputs the statistical result to the analysis log; S02. The storage offset address of the randomization seed in the original data block is located, and the randomization seed combination is extracted based on the storage offset address; S03. Determine whether the randomization seed combination is valid. If it is valid, execute step S04. If it is not valid, determine whether the data analysis command parameter contains information that can calculate the randomization seed combination. If it does, calculate the randomization seed combination and execute step S04. If it does not, execute step S06; S04. Perform software de-scrambling operation on the user data and the metadata using the randomization seed combination. The de-scrambling operation process is: generating a scrambling code restoration sequence based on the randomization seed combination, and then performing XOR operation on the scrambled data to be processed and the scrambling code restoration sequence, so as to restore the scrambled data to plaintext data; S05, judging whether there is reference standard data available for data plaintext comparison, if yes, comparing data plaintext with reference standard data, recording difference byte offset position, bit error mode and total number of bit flips and outputting to analysis log; if no, data plaintext and randomization seed initial value are recorded to analysis log; S06, outputting analysis log.
[0006] Further, data configuration is performed before step S01, and the configuration information includes basic command parameters and optional additional information. The basic command parameters include the memory address of the original data block to be parsed, the length format of user data and metadata, the total length of the original data to be parsed and the check bit field, the number and alignment interval of the original data block to be parsed, and the target address of the output data. The optional additional information includes the starting physical address of the original data to be parsed in the memory, the number of erase and write of the storage block, and the reference standard data before writing.
[0007] Further, after the data configuration is performed, the original data block is divided into basic units of the same length according to the number and alignment interval of the original data block to be parsed, and the basic units are cached into a message queue for sequential processing.
[0008] Further, the storage address of the randomization seed in the original data block is located according to the length format of the user data and the metadata and the length of the check value, and a copy of the randomization seed is extracted based on the length format of the user data and the metadata.
[0009] Further, the process of judging whether the randomization seed is valid is as follows: searching the random seed library, which is a set table composed of all randomization seed values corresponding to the flash memory particles of the same type, to judge whether the randomization seed combination extracted from the original data block satisfies accurate matching or approximate matching. If yes, it is considered that the randomization seed combination extracted from the original database is valid, and if not, it is considered that the randomization seed combination extracted from the original database is invalid.
[0010] Further, the randomization seed combination includes the randomization seed and its multiple copies. The accurate matching means that there is a randomization seed value in the random seed library that is exactly the same as the randomization seed combination. The approximate matching means that there is one or more randomization seed values in the random seed library that has a Hamming distance of no more than 2 from the randomization seed combination extracted from the original database.
[0011] Further, the linear feedback shift register method is used to generate the scrambling code restoration sequence, and the value of each byte in the scrambling code restoration sequence is determined according to the 0 / 1 state of 8 specific positions Bit in the 32-bit random seed. After each byte of the scrambling code restoration sequence is generated, the random seed is updated, and the original value is replaced with a new seed to continue participating in the iterative calculation of the subsequent scrambling code restoration sequence byte generation. The update rule of the random seed is as follows: check the highest bit of the current 32-bit random seed value. If it is 1, then the current seed value is left shifted by one bit and then XORed with a fixed generating polynomial. Otherwise, only the current seed value is left shifted by one bit, and no XOR is performed. The newly obtained 32-bit operation result is the new seed value required for generating the next scrambling code restoration sequence byte.
[0012] Further, in step S01, the valid part of the original data block corresponding to the read command includes scrambled user data, scrambled metadata, check code, random seed and ECC error correction code.
[0013] The application further discloses a flash memory original data analysis device, which comprises a processor and a memory storing program instructions, and the processor is configured to execute the flash memory original data analysis method as described above when running the program instructions.
[0014] The application further discloses a storage medium storing program instructions, which execute the flash memory original data analysis method as described above when running.
[0015] The application has the following beneficial effects: The original data analysis tool and method are suitable for restoring the states of the user data and the metadata before scrambling as much as possible in the case of hardware ECC correction failure in flash memory reading and unscrambling of the original data, and counting the distribution of 0 / 1 bit positions in the original data, comparing with reference standard data to determine the difference position and bit error mode. This is helpful for developers to locate the problem site, observe and analyze the error law, and infer the possible cause of correction failure.
[0016] In addition, the debugging tool is integrated into the normal firmware of the flash memory controller chip, without the need of power-off or replacement of other special test firmware. The developers can conveniently call the debugging tool through a serial port or other debugging interface. After the command parameters are configured, the operation process is simplified through automatic execution steps, the time cost of manual operation is saved, and the efficiency of analyzing problems is improved. BRIEF DESCRIPTION OF DRAWINGS
[0017] Figure 1 Flowchart for the method described in Example 1 Figure 2 Schematic diagram of the device described in Example 2. DETAILED DESCRIPTION
[0018] The application will be further described below in conjunction with the drawings and specific embodiments.
[0019] Embodiment 1 This embodiment 1 discloses a flash memory raw data analysis method, which is a test program embedded in the firmware of a flash memory controller chip, used for debugging and analysis by developers after the occurrence of read data check error and other abnormal conditions in the flash memory. The method is integrated in the form of an executable program inside the flash memory controller, and interacts with the outside through a UART physical serial port or a virtual serial port command line tool (CLI) and other means, and is activated by inputting specific command fields. When called, the process shown in the figure is performed: Figure 1 S01, data configuration, the configuration information includes basic command parameters and optional additional information, the basic command parameters include the memory address of the raw data block to be analyzed, the length format of user data and metadata, the total length of the raw data to be analyzed and the check bit field, the number and alignment interval of the raw data block to be analyzed, and the target address of the output data, the optional additional information includes the memory start physical address (PBA) of the raw data to be analyzed, the storage block erase count (P / E), and the reference standard data (Pattern) before writing.
[0020] S02, when read command error correction fails, the host issues a data analysis command, and the raw data blocks read out from the flash memory are sequentially placed in a known position in the memory by the data transfer module (DMA). After setting the configuration information and executing the raw data analysis command, the program automatically divides the multiple raw data blocks placed continuously in the memory into basic units of the same format (the size is determined by the total length of the raw data and the check bit field, generally 4KB~5KB) according to the number and alignment interval of the raw data blocks to be processed, and caches them in a message queue for processing in turn.
[0021] S03, when processing each segmented raw data block, first, the valid part in the raw data block is counted, and the statistical result is output to the analysis log; whether the 0 / 1 proportion is balanced or not can be used as a basis for the developer to analyze the quality of the data channel signal between the host and the flash memory.
[0022] In this embodiment, the valid part in the raw data block refers to the part other than the padding reserved data segment, i.e. the scrambled user data, scrambled metadata, check code, randomization seed and ECC error correction code. According to the user data and metadata length format, check value and error correction code length and other information in the program call parameters, the length and position of the valid part in the raw data block can be calculated, and the total length of the raw data block is determined by the type of the flash memory used, so the remaining part after removing the valid part is all padding reserved data segment.
[0023] S04, according to the length format of user data and metadata, the length of the check value and the error correction code, etc. Parameters locate the storage offset address of the random seed in the original data block, and extract it out (for example, 4096 bytes of user data + 32 bytes of metadata + 4 bytes of CRC checksum, the first random seed may appear at an offset position 4132 bytes from the beginning of the data block). In order to ensure reliability, the random seed will be periodically repeated in multiple positions in the original data generated by the hardware during writing to construct multiple backup copies, and all of them need to be found during analysis to improve accuracy.
[0024] For different data length formats, the position and number of random seed repetitions are different, such as 4096 bytes of user data + 32 bytes of metadata, the second random seed appears 64 bytes after the first seed, and repeats 8 times at an interval of 64 bytes. Other data length formats also have their own corresponding intervals and repetition numbers, and the program determines the search rule according to the call parameters during analysis.
[0025] S05, determine whether the random seed combination is valid, if valid, execute step S06, if not, determine whether the data analysis command parameters contain information that can calculate the random seed combination, if yes, calculate the random seed combination and execute step S06, if no, execute step S08.
[0026] The judgment process of whether the random seed combination is valid is as follows: after the random seed is extracted, the program will search the random seed library provided by the search algorithm layer (i.e. the set table of all possible random seed values corresponding to the flash memory particles of this type), and query whether the random seed extracted from the original data exists in the random seed library, if there is an exact match, then proceed to the next step, otherwise continue to confirm whether the remaining other random seed copies are exactly matched, when all random seed copies are queried and still no exact match but there is an approximate match with a hamming distance of not more than 2 (i.e. the binary bit difference between one of the random seed copies and a value in the seed library is ≤2), then choose the closest value in the seed library. If the approximate match is not satisfied, when the command parameters contain the flash memory starting physical address (PBA), storage block erase count (P / E) and other parameters, calculate the expected random seed value according to these parameters, if the command does not give these parameters, since the correct random seed cannot be known, skip the subsequent software descrambling operation step and end (only return the 0 / 1 bit number statistics as the analysis result in the log).
[0027] The process of calculating the expected random seed value according to the flash memory starting physical address (PBA), storage block erase count (P / E) is as follows: When the flash physical address (PBA) and the number of program-erase (P / E) cycles of a memory block are known, the expected offset position of a randomization seed in the random seed library is calculated according to the following formula combination: 1. logic_page = pba.wl * WL_PG_NUM + pba.page, 2. seed_index = (pe_cnt + logic_page * PLN_GU_NUM) mod SSEED_MAX + (logic_page / AGI_STEP) * (SSEED_MAX + WL_PG_NUM * PLN_GU_NUM).
[0028] Where pba.wl and pba.page are two different components in the flash physical address (PBA), representing the word line (WL) and page (Page) number respectively, and pe_cnt represents the number of program-erase (P / E) cycles of a memory block, which are input values.
[0029] The remaining parameters in the formula are constants closely related to the specific model of the flash particle, such as for a certain TLC (three-layer storage unit type) memory particle of a certain manufacturer, i.e., WL_PG_NUM = 3, PLN_GU_NUM = 16, SSEED_MAX = 67, and AGI_STEP = 72, which are determined according to the type of flash.
[0030] seed_index is the offset serial number of the starting randomization seed in the seed library, which is a 32-bit variable (4 bytes), and needs to be multiplied by 4 again when actually addressing.
[0031] S06. Use a randomized seed combination to perform software descrambling on user data and metadata, and store the result at the target memory address of the output data given in the command parameters. The scrambling and descrambling algorithms are symmetrical, so the same randomized seed can be used for descrambling to restore the plaintext data. The process can be briefly described as follows: generate a scrambling code restoration sequence using the randomized seed, and then perform an XOR operation on the scrambled data to be processed and the restoration sequence byte by byte to complete the descrambling process. The restoration sequence is usually generated using the linear feedback shift register (LFSR) method, that is, the value of each byte in the restoration sequence is determined by the 0 / 1 state of 8 specific bits in the 32-bit randomized seed. After each byte of the restoration sequence is generated, the randomized seed needs to be updated, and the new seed replaces the original value to continue participating in the iterative calculation of subsequent restoration sequence byte generation. The randomization seed update rule is as follows: check the most significant bit (MSB, i.e., the 31st bit) of the current 32-bit randomization seed value. If it is 1, shift the current seed value left by one bit and then XOR it with a fixed generator polynomial (a 32-bit integer); otherwise, simply shift the current seed value left by one bit without performing an XOR operation. The newly obtained 32-bit operation result is the new seed value required to generate the next restored sequence byte.
[0032] S07. Determine if there is reference standard data for comparison with plaintext data. If so, compare plaintext data with reference standard data, record the offset position of the difference byte, bit error mode, and total number of bit flips, and output them to the parsing log. If not, record plaintext data (user data plaintext and metadata plaintext) and initial value of randomization seed in hexadecimal character form to the parsing log. S08. After the current raw data block is parsed, continue to retrieve other cached raw data blocks from the message queue and repeat the above parsing steps until all raw data blocks to be processed are parsed and the parsing log is output.
[0033] Example 2 This disclosure provides a flash memory raw data parsing device 300, such as... Figure 2 As shown, the device includes a processor 304 and a memory 301. Optionally, the device may further include a communication interface 302 and a bus 303. The processor 304, communication interface 302, and memory 301 can communicate with each other via the bus 303. The communication interface 302 can be used for information transmission. The processor 304 can invoke logical instructions in the memory 301 to execute the flash memory raw data parsing method of the above embodiment.
[0034] In addition, the logical instructions in the memory 301 described above can be implemented in the form of a software function unit and sold or used as an independent product, and can be stored in a computer readable storage medium.
[0035] The memory 301 as a computer readable storage medium can be used to store software programs, computer executable programs, such as program instructions / modules corresponding to the method in the embodiments of the present disclosure. The processor 304 executes the function application and data processing by running the program instructions / modules stored in the memory 301, that is, implements the flash raw data parsing method in the above embodiments.
[0036] The memory 301 can include a program storage area and a data storage area, wherein the program storage area can store an operating system and application programs required by at least one function; the data storage area can store data created according to the use of the terminal device, etc. In addition, the memory 301 can include a high-speed random access memory, and can also include a non-volatile memory.
[0037] Embodiment 3 The embodiments of the present disclosure provide a computer readable storage medium, which stores computer executable instructions, and the computer executable instructions are configured to execute the flash raw data parsing method.
[0038] The computer readable storage medium described above can be a transitory computer readable storage medium or a non-transitory computer readable storage medium.
[0039] The technical solutions of the embodiments of the present disclosure can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes one or more instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in the embodiments of the present disclosure. The aforementioned storage medium can be a non-transitory storage medium, including: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, etc. various media that can store program codes, or can be a transitory storage medium.
[0040] The above description and drawings are illustrative of embodiments of the present disclosure and are not intended to be limiting. Other embodiments can include structural, logical, electrical, process, and other changes. Embodiments are illustrative only. Individual components and functions are optional, and the order of operations can vary. Parts and features of some embodiments can be included or substituted in other embodiments. Also, words used in this document are words of description, not limitation. As used in the description, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. Similarly, the term "and / or" as used in the application refers to any and all possible combinations of one or more elements, i.e., it represents a disjunctive, and the conjunction of separate events that can or can not be mutually exclusive of each other. Additionally, the term "comprising" as used in the application means the open inclusion of the stated features, integers, steps, operations, elements, and / or components, but does not preclude the addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Without more limitations, an element defined by an indefinite article such as "a" or "an" does not exclude the existence of additional identical elements in the process, method, or device including the recited element. In this document, each embodiment focuses on the differences from other embodiments, and the same or similar parts between embodiments can be referred to each other. For the method, product, etc. disclosed by the embodiments, if it corresponds to the method part disclosed by the embodiments, the relevant part can be referred to the description of the method part.
[0041] Those skilled in the art can understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods for each specific application to implement the described functions, but such implementation should not be considered beyond the scope of the embodiments of the present disclosure. The skilled person can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described system, device and unit can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.
[0042] In the embodiments disclosed herein, the disclosed methods, products (including but not limited to devices, apparatuses, etc.) can be implemented in other manners. For example, the described device embodiments are merely illustrative. For example, the division of the units is merely a logical function division. There can be another division manner for the actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms. The units described as separated components can or can not be physical separate units. The units shown as separate components can or can not be physical separate units, i.e., can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to the actual needs to implement the embodiments. In addition, the units in the embodiments disclosed herein can be integrated in a processing unit, or each unit can exist alone physically, or two or more units can be integrated in a unit.
Claims
1. A method for parsing raw data of a flash memory, the method comprising: The method comprises the following steps: S01, when a read command correction fails, the host issues a data analysis command to count the number of 0 / 1 bits in the valid part of the original data block corresponding to the read command, and outputs the statistical result to the analysis log; S02, locate the storage offset address of the randomization seed in the original data block, and extract the randomization seed combination based on the storage offset address; S03, determine whether the randomization seed combination is valid, if valid, execute step S04, if invalid, determine whether the data analysis command parameter contains information that can calculate the randomization seed combination, if yes, calculate the randomization seed combination and execute step S04, if no, execute step S06; S04, use the randomization seed combination to perform software descrambling operation on user data and metadata, the descrambling operation process is: generating a descrambling sequence based on the randomization seed combination, then performing XOR operation between the scrambled data to be processed and the descrambling sequence, so as to restore the scrambled data to data plaintext; S05, determine whether there is reference standard data available for data plaintext comparison, if yes, compare the data plaintext with the reference standard data, record the difference byte offset position, bit error mode and the total number of bit flips and output to the analysis log; if no, record the data plaintext and the randomization seed initial value to the analysis log; S06, output the analysis log.
2. The method of claim 1, wherein: Before executing step S01, data configuration is performed, the configuration information includes basic command parameters and optional additional information, the basic command parameters include the memory address of the original data block to be analyzed, the length format of user data and metadata, the total length of the original data to be analyzed and the check bit field, the number and alignment interval of the original data block to be analyzed, the output data target address, the optional additional information includes the memory starting physical address of the original data to be analyzed, the storage block erase count, and the reference standard data before writing.
3. The method of claim 2, wherein: After the data configuration is executed, the original data block is divided into basic units of the same length according to the number and alignment interval of the original data block to be analyzed, and the basic units are cached to the message queue for processing in turn.
4. The method of claim 2, wherein: The storage offset address of the randomization seed in the original data block is located according to the length format of the user data and the metadata, and the copy of the randomization seed is extracted based on the length format of the user data and the metadata.
5. The method of claim 1, wherein: The process of determining whether the randomization seed is valid is: searching the random seed library, the random seed library is a set of number tables composed of all randomization seed values corresponding to the flash memory particles of the model, determining whether the randomization seed combination extracted from the original data block satisfies the exact match or approximate match, if yes, considering that the randomization seed combination extracted from the original database is valid, if not, considering that the randomization seed combination extracted from the original database is invalid.
6. The method of claim 5, wherein: The randomized seed combination includes randomized seeds and multiple copies thereof, the exact match refers to that there is a randomized seed in the randomized seed library that is exactly the same as the randomized seed combination, and the approximate match refers to that there is one or more randomized seed values in the randomized seed library that have a Hamming distance of no more than 2 from the randomized seed combination extracted from the original database.
7. The method of claim 1, wherein: The linear feedback shift register method is used to generate the scrambling reduction sequence, and the process is as follows: according to the state of 0 / 1 of 8 specific positions Bit in the 32-bit randomized seed, the value of each byte in the scrambling reduction sequence is determined, after each byte of the scrambling reduction sequence is generated, the randomized seed is updated, and the original value is replaced with a new seed to continue participating in the iterative calculation of the subsequent scrambling reduction sequence byte generation; the update rule of the randomized seed is as follows: check the highest bit of the current 32-bit randomized seed value, if it is 1, then perform XOR operation on the left shifted seed value by one bit and a fixed polynomial; otherwise, only left shift the current seed value by one bit, and do not do XOR; the newly obtained 32-bit operation result is the new seed value required for generating the next scrambling reduction sequence byte.
8. The method of claim 1, wherein: In step S01, the valid part of the original data block corresponding to the read command includes scrambled user data, scrambled metadata, check code, randomized seed and ECC error correction code.
9. A flash raw data parsing apparatus comprising a processor and a memory storing program instructions, characterized in that: The processor is configured to execute the flash memory original data parsing method according to any one of claims 1 to 8 when the program instructions are executed.
10. A storage medium storing program instructions, characterized in that: The program instructions, when executed, perform the flash memory original data parsing method according to any one of claims 1 to 8.