Error correction coding method and device, and storage medium
By monitoring the operating status of the 3D NAND flash storage device, obtaining the target physical address and determining the error correction coding scheme, efficient error correction coding of the storage device is achieved, and the problem of high error rate caused by poor durability of the storage device is solved, and storage efficiency and performance are improved.
Patent Information
- Application Number
- CN202411923135.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-25
AI Technical Summary
During use, 3D NAND flash storage devices have different durability of storage units in different layers or regions due to differences in manufacturing processes and physical characteristics, and are prone to wear and error. The prior art reduces the error rate through data migration, but consumes additional resources and may introduce error risks.
An error correction encoding method is proposed. By monitoring the operating status of the storage device, a target physical address of the storage area that meets the preset conditions is obtained, the target error correction encoding scheme is determined based on the target physical address, and error correction encoding is performed to reduce the error rate of the storage device.
By triggering error correction coding according to the degree of use of the storage area, it can avoid error correction failures and data migration caused by excessive use, improve overall storage efficiency, and configure different encoding error correction solutions for different regions to optimize storage performance.
Smart Images

Figure CN119943126A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of memory technology, and in particular to an error correction coding method, device and storage medium. Background Art
[0002] In 3D NAND flash, memory cells in different layers or regions may have different endurance due to differences in manufacturing processes and physical properties. As a result, edge or corner areas with poorer endurance are more susceptible to wear and errors when subjected to the same level of use.
[0003] At present, in order to reduce the error rate of 3D NAND flash during use, data is generally migrated from areas that are more prone to wear and errors to more stable areas through intelligent data allocation algorithms. However, in the process of data migration, additional storage and I / O resources are consumed, and the risk of data migration errors may be introduced. In addition, frequent data migration may also affect the overall performance and stability of the system.
[0004] The above contents are only used to assist in understanding the technical solution of the present application and do not constitute an admission that the above contents are prior art. Summary of the invention
[0005] The main purpose of this application is to provide an error correction coding method, device and storage medium, aiming to solve the technical problem of how to reduce the error rate of storage devices during use.
[0006] To achieve the above object, the present application proposes an error correction coding method, which includes:
[0007] Monitor the operating status of the storage device and obtain the operating status information of the storage device;
[0008] When the monitored operation status information reaches a preset condition, a target physical address of a storage area in the storage device that meets the preset condition is obtained;
[0009] When a data storage request is received, determining a target error correction coding scheme based on a target physical address;
[0010] Perform error correction coding according to a target error correction coding scheme.
[0011] In one embodiment, the operation status information includes the number of erase and write times. When the operation status information is monitored to reach a preset condition, the step of obtaining a target physical address of a storage area in the storage device that reaches the preset condition includes:
[0012] Get the number of erase and write times of the storage device;
[0013] When the number of erasure and writing times reaches a preset first threshold, the target physical address corresponding to each storage area in the storage device whose number of erasure and writing times reaches the first threshold is obtained.
[0014] In one embodiment, before the step of monitoring the operating status of the storage device and obtaining the operating status information of the storage device, the method further includes:
[0015] Perform erase and write tests on storage devices and generate an error distribution matrix based on the test results;
[0016] According to the error distribution matrix, the weight corresponding to each error storage area where an error occurs is calculated;
[0017] Determine a target error correction coding scheme according to the target physical address, including:
[0018] According to the weight, an error correction coding scheme corresponding to the error storage area is determined.
[0019] In one embodiment, the steps of performing an erase test on a storage device and generating an error distribution matrix according to the test result include:
[0020] Obtain the number of errors detected by the error correction coding algorithm and the logical address where the error occurred;
[0021] Mapping the logical address to the physical structure of the storage device to obtain the physical address of the error storage area;
[0022] Generate an error distribution matrix based on the number of errors and physical addresses.
[0023] In one embodiment, the step of calculating the weight corresponding to each error storage area where an error occurs according to the error distribution matrix includes:
[0024] Input the error distribution matrix into a preset mapping function;
[0025] Compute the weight corresponding to each element in the error distribution matrix.
[0026] In one embodiment, the step of determining the error correction coding scheme corresponding to the error storage area according to the weight includes:
[0027] Group the error storage areas according to the preset classification criteria and weights;
[0028] According to the grouping result, the error correction coding algorithm and the length of the error correction coding check area corresponding to each group weight are determined;
[0029] An error correction coding scheme is generated according to the error correction coding algorithm and the length of the error correction coding check area, and there is a mapping relationship between the error correction coding scheme and the corresponding physical address.
[0030] In one embodiment, the step of performing error correction coding according to the target error correction coding scheme includes:
[0031] When the target physical address receives a test storage request, the test data corresponding to the test storage request is encoded and stored according to the target error correction scheme.
[0032] In one embodiment, after the step of performing error correction coding according to the target error correction coding scheme, the method further includes:
[0033] After reading the test data, the test data is parsed according to the target error correction scheme to obtain comparison data;
[0034] If the test data and the comparison data are inconsistent, a first error message is generated and an error log is updated;
[0035] If the test data and the comparison data are consistent, determine whether the generated error cause matches the preset error pattern;
[0036] If there is no match, a second error message is generated and the error log is updated.
[0037] In addition, to achieve the above-mentioned purpose, the present application also proposes an error correction coding device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the error correction coding method as described above.
[0038] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the error correction coding method as described above are implemented.
[0039] The present application provides an error correction coding method, which continuously monitors the operating status of a storage device to obtain the operating status information of the storage device. When the operating status information is monitored to meet a preset condition, the target physical address of the storage area that meets the preset condition in the storage device is obtained, and the storage area where the problem occurs or meets a specific condition can be accurately located through the physical address. When a data storage request is received, the target error correction coding scheme is determined according to the obtained target physical address, and the data to be stored is error-corrected according to the scheme.
[0040] In this application, by triggering error correction coding according to the usage level of the storage area, it is possible to avoid the current error correction coding scheme being unable to correctly correct errors due to excessive usage, resulting in damaged blocks and data migration. In addition, the target error correction coding scheme is determined based on the target physical address of the storage area, and different coding error correction schemes are configured for different storage areas, which can be optimized according to the specific needs of each area, thereby improving the overall storage efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0042] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0043] Figure 1 A flowchart of the first embodiment of the error correction coding method of the present application is provided;
[0044] Figure 2 A detailed flow chart of the first embodiment of the error correction coding method of the present application is provided;
[0045] Figure 3 A flowchart of the second embodiment of the error correction coding method of the present application is provided;
[0046] Figure 4 A detailed flow chart of the second embodiment of the error correction coding method of the present application is provided;
[0047] Figure 5 A flowchart of the third embodiment of the error correction coding method of the present application is provided;
[0048] Figure 6 A schematic diagram of the device structure of the hardware operating environment involved in the error correction coding method in the embodiment of the present application. DETAILED DESCRIPTION
[0049] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0050] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0051] The main solution of the embodiment of the present application is: monitor the operating status of the storage device to obtain the operating status information of the storage device; when the monitored operating status information reaches a preset condition, obtain the target physical address of the storage area in the storage device that meets the preset condition; when a data storage request is received, determine the target error correction coding scheme according to the target physical address; and perform error correction coding according to the target error correction coding scheme.
[0052] At present, in order to reduce the error rate of 3D NAND flash during use, data is generally migrated from areas that are more prone to wear and errors to more stable areas through intelligent data allocation algorithms. However, in the process of data migration, additional storage and I / O resources are consumed, and the risk of data migration errors may be introduced. In addition, frequent data migration may also affect the overall performance and stability of the system.
[0053] In order to solve the above problems, the present application provides an error correction coding method, which continuously monitors the operating status of a storage device to obtain the operating status information of the storage device. When it is monitored that the operating status information meets the preset conditions, the target physical address of the storage area that meets the preset conditions in the storage device is obtained, and the storage area where the problem occurs or meets specific conditions can be accurately located through the physical address. When a data storage request is received, the target error correction coding scheme is determined according to the obtained target physical address, and the data to be stored is error-corrected according to the scheme.
[0054] In this application, by triggering error correction coding according to the usage level of the storage area, it is possible to avoid the current error correction coding scheme being unable to correctly correct errors due to excessive usage, resulting in damaged blocks and data migration. In addition, the target error correction coding scheme is determined based on the target physical address of the storage area, and different coding error correction schemes are configured for different storage areas, which can be optimized according to the specific needs of each area, thereby improving the overall storage efficiency.
[0055] It should be noted that the execution subject of this embodiment can be a computing service device with network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, Nandflash (non-flash memory) and its corresponding controller, SSD (solid state drive) and its corresponding controller, etc., or an electronic device or device capable of realizing the above functions. The following takes an error correction coding device as an example to illustrate this embodiment and the following embodiments.
[0056] Based on this, the embodiment of the present application provides an error correction coding method, referring to Figure 1 , Figure 1 This is a flowchart of the first embodiment of the error correction coding method of the present application.
[0057] In this embodiment, the error correction coding method includes steps S100 to S400:
[0058] Step S100, monitoring the operating status of the storage device to obtain operating status information of the storage device.
[0059] It should be noted that the storage device may be a flash memory device. A flash memory device, also known as a flash memory or a read-only memory, is a non-volatile electronic memory. It can maintain stored data without external power supply and has the characteristics of high information density, large-scale reading and writing, and short random access time. Flash memory changes the charge state in the floating gate by controlling the gate voltage, thereby realizing data reading and writing.
[0060] 3D NAND flash (three-dimensional non-volatile NAND flash) is a technology that uses three-dimensional space to stack storage cells to increase storage density. Compared with traditional 2D NAND flash, 3D NAND flash can store more data in the same volume by vertically stacking multiple storage layers. In 3D NAND flash, high-frequency write / erase areas are more likely to reach their life limit due to frequent data operations. In a multi-layer stacked 3D NAND structure, storage cells in edge or corner areas may have lower durability due to manufacturing process limitations. In areas of storage devices that are close to heat sources or have poor heat dissipation, storage cells in these areas are more prone to errors and wear in high temperature environments.
[0061] In this embodiment, the operating status information is obtained by checking the system log of the operating system or the SSD firmware (Solid State Drive) log. You can also use a performance monitor to add counters related to the flash memory or SSD you want to monitor. By counting the real-time data of these counters, you can evaluate the operating status and usage of the flash memory. The operating status information can read and write speed, response time, amount of data written, and number of programming / erasing times. The operating status information can help predict the remaining life information of the flash memory.
[0062] Step S200, when it is monitored that the running status information reaches a preset condition, a target physical address of a storage area in the storage device that meets the preset condition is obtained.
[0063] It should be noted that a physical address usually includes multiple parts, which are used to uniquely identify each storage unit in a flash memory chip. These parts may include a block address, a page address, a column address, etc.
[0064] In this embodiment, when a storage request is obtained, a logical address is allocated to the data to be stored based on the type and parameters of the storage request, the data is accessed through the logical address, and the flash memory controller maps the logical address to the physical address inside the flash memory, performs a storage operation, and writes the data to the specified physical address location in the flash memory. When the usage stage of the target storage area changes, it means that the storage area is used too much and has entered the late life stage, and its remaining service life is small.
[0065] Please refer to Figure 2 In a feasible implementation manner, the operation status information includes the number of erase and write times, and step S200 may include steps S210 to S220:
[0066] Step S210, obtaining the number of erase and write times of the storage device.
[0067] Step S220, when the number of erasure and writing times reaches a preset first threshold, obtaining a target physical address corresponding to each storage area in the storage device whose number of erasure and writing times reaches the first threshold.
[0068] It should be noted that in flash memory, the erase operation (P / E) refers to the programming operation and the erase operation. The programming operation (P, Program), also known as the write operation, refers to the process of writing data into a flash memory cell. The erase operation (E, Erase) operation refers to the process of clearing the data in the flash memory cell and restoring it to its initial state. The number of erase times refers to the fact that after each storage cell has undergone a certain number of programming, writing and erasing operations, its performance will be significantly reduced, and it may even be unable to reliably store data. This is due to the physical characteristics of flash memory, and each erase operation will cause certain wear to the storage cell. In flash memory, the smallest unit of the erase operation is usually a block, a subblock or a deck. The smallest unit of the write operation is usually a page. The complete process of a flash memory cell undergoing one programming and one erasing is called a P / E cycle. Due to the physical characteristics of the flash memory cell, each P / E cycle will cause certain wear to it. The number of P / E cycles is an important indicator for measuring the life of flash memory. As the number of P / E cycles increases, the performance of the flash memory cell will gradually degrade and eventually may no longer be able to reliably store data.
[0069] In addition, it should be noted that the physical address is a unique identifier for each storage unit in the flash memory. When you want to access a piece of data in the flash memory, you need to know the physical address at which the data is stored. Then, the flash memory controller will find the corresponding storage unit based on the physical address and read or write the data.
[0070] In this embodiment, the number of erase and write times of the storage device is monitored in real time or according to a preset cycle. The use stage and degree of wear of each storage area of the flash memory are evaluated by the number of erase and write times. The erase and write operation can be monitored, and the number of erase and write times can be obtained according to the interface for monitoring the erase and write operation of the flash memory. The interface may also include an interrupt service routine, an event callback of the file system layer, or a special flash memory management API. One or more callback functions are registered to respond to the erase and write event. These callback functions are called each time an erase and write operation occurs. Inside the callback function, the detailed information of the erase and write event is captured, including the address of the erased block or page. Based on the captured address information, the specific range to be erased is determined (which can be a single block, page, or part of a page). Finally, based on the erase and write range, the erase and write times counter of the corresponding block or page is updated. If the erase and write operation spans multiple blocks or pages, their counters are updated separately.
[0071] When the number of erasures of the flash memory reaches a preset first threshold, the target physical address information of the erasure event is obtained, and all target physical address information is recorded. For example, the expected life of a 3D NAND Flash device is 3000 erasures, the first threshold is 40% of the total erasures, that is, 1200 erasures; the second threshold is 80% of the total erasures, that is, 2400 erasures.
[0072] In this embodiment, the success rate of error correction is improved by setting a first threshold and dynamically adjusting the error correction coding strategy when it is reached. The number of jumps in the data stored in the storage device increases with the number of erase and write times. When the number of erase and write times reaches a certain threshold, the number of jumps in the data stored in the storage device will increase to a level that the error coding cannot correct, resulting in the inability to read the data. At this time, obtaining a new error correction coding scheme, such as upgrading from the BCH algorithm to the LDPC algorithm and increasing the length of the check area, can significantly enhance the error correction capability, thereby allowing the flash memory to withstand more erase and write times without losing data.
[0073] In another feasible implementation, the use stage of the storage device can also be evaluated according to the write volume of the flash memory unit. When the monitoring system detects that the write volume of a certain storage area reaches a preset condition, the target physical address of the storage area is recorded and obtained.
[0074] For example, suppose there is an SSD with a capacity of 1TB, which contains multiple storage areas, each with an expected size of 256MB. When the amount of writing to any storage area reaches 60% of its capacity, its physical address is recorded. When the amount of writing to any storage area reaches 80% of its capacity, its physical address is recorded.
[0075] Step S300: When a data storage request is received, a target error correction coding scheme is determined according to a target physical address.
[0076] Step S400: performing error correction coding according to a target error correction coding scheme.
[0077] It should be noted that the role of error correction coding (ECC) is to detect and correct errors in stored data. Since the flash memory may be affected by various factors during data storage, data errors may occur. Error correction coding can automatically detect and correct these errors when reading data to ensure the accuracy and reliability of the data. The length of the check area is closely related to the error correction capability of the error correction code. The stronger the error correction capability of the error correction code, the more check data it requires, so the length of the check area is also longer.
[0078] In this embodiment, the corresponding target error correction coding scheme can be determined by the mapping relationship between the physical address and the error correction scheme. An address mapping table is set to record the mapping relationship between each physical address and the corresponding error correction code. This mapping table can be adjusted as needed when the system is running. When it is necessary to obtain a target error correction scheme associated with a specific target physical address, the entry corresponding to the address is searched in the address mapping table. This search process can be a direct index (if the mapping table is arranged in order of physical addresses) or a more complex search algorithm (if the mapping table is unordered or compression technology is used).
[0079] For example, a flash memory system contains 100 physical blocks, each with a unique physical address. To optimize performance and reliability, different ECC check area lengths are set for these blocks: blocks with physical addresses 0-49 use a 16-byte ECC check area. Blocks with physical addresses 50-74 use a 24-byte ECC check area to improve the error correction capability of these blocks. Blocks with physical addresses 75-99 use a 32-byte ECC check area because these blocks may be more susceptible to interference or damage.
[0080] In this embodiment, by detecting the state of the target flash memory, it is determined whether the usage level of a storage area in the target flash memory has reached a preset condition. When a storage area reaches the preset condition, a target error correction coding scheme is obtained for error correction coding, which can trigger the error correction coding operation according to the usage level of the storage area, thereby avoiding the degradation of the flash memory device performance caused by excessive usage. The error correction scheme corresponding to the storage area is obtained according to the preset mapping relationship, and different coding error correction schemes are configured for different storage areas, while improving data reliability and storage device performance, the data storage capacity is maximized, and the impact of the reduction in data capacity caused by the increase in the length of the check area in the coding scheme is reduced.
[0081] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above-mentioned embodiment 1 can be referred to the above introduction, and will not be repeated in the following. Figure 3 Before step S100, the error correction coding method further includes steps A100 to A200:
[0082] Step A100, performing an erase test on the storage device and generating an error distribution matrix according to the test result.
[0083] In this embodiment, the storage device is subjected to erase and write test in a certain period. Optionally, this period can be a fixed time interval, such as monthly, quarterly or annual evaluation. It can also be triggered based on specific conditions, such as when the error rate of a certain area exceeds a preset threshold, an evaluation is performed immediately.
[0084] Step A200, calculating the weight corresponding to each error storage area where an error occurs according to the error distribution matrix.
[0085] In this embodiment, during the erase and write test of the storage device, the following data are obtained and recorded: (1) Error rate: the number of errors or error rate of each area in the most recent evaluation cycle. (2) Usage frequency: the frequency of data read and write in each area, which may also include the write mode (such as sequential write, random write, etc.) and data type (such as metadata, user data, etc.). (3) Other related factors: factors such as ambient temperature and system load that may affect the performance and error rate of the flash memory. After the error information is collected, an error distribution matrix is generated based on the collected data.
[0086] In this embodiment, the weight of each area is calculated. The weight calculation method can be based on a combination of error rate, frequency of use and other relevant factors. For example, a weighted sum method can be used to assign a weight coefficient to each factor, and then the weighted sum is calculated as the final weight of the area. In order to simplify the process and ensure the smooth progress of regular evaluation and adjustment, system management tools or scripts are used to regularly perform data collection, weight calculation, evaluation and adjustment tasks, and monitor the system status to detect any potential anomalies or problems. Exemplarily, in an automated script process, a cron job (Linux) or a task scheduler (Windows) is used to set up a script that is executed regularly. The script collects the necessary data from the flash controller or system log and calls a function or program to calculate the weight of each area.
[0087] In one possible implementation, please refer to Figure 4 , step A100 may further include the following steps A110 to A130:
[0088] A110, obtains the number of errors detected by the error correction coding algorithm and the logical address where the error occurred.
[0089] A120, maps the logical address to the physical structure of the storage device to obtain the physical address of the error storage area.
[0090] In this embodiment, the number of errors and location information can be confirmed in the statistical data generated by the ECC (Error Correction Code) mechanism. After the error information is obtained, the location information needs to be converted into a physical address in order to more accurately locate the specific location where the error occurred in the memory. The location information in the ECC error log may appear in various forms, such as memory slot number, memory channel number, memory row number (rank), memory bank number (bank), etc. Different servers and memory configurations may have different memory mapping methods and address encoding rules. Address conversion is performed according to different memory mapping methods and address encoding rules, and the location information (such as slot number, channel number, etc.) is combined with the base address, offset or index value to obtain the final physical address.
[0091] For example, the document is searched for the part about memory mapping to obtain the description of the organization of memory, which may include: (1) Slot number: the number and position of each memory slot. (2) Channel number: if multi-channel technology (such as DDR4 dual channel) is supported, how each channel is allocated. (3) Row (Rank) and Bank (Bank): the organization inside the memory module, how each row and bank affects the address space. (3) Base address and size: the starting address of each slot, channel, row and bank and its size in the physical address space. Assume that the document states: the base address of slot 1 is 0x00000000 and the size is 2GB (i.e. 0x80000000). Each slot supports a single channel, each channel contains two ranks, and each row has two banks. In the error location information obtained, it is pointed out that the error occurs at a certain position of "slot 1, channel 0, row 0, bank 1". It is known from the document that the base address of slot 1 is 0x00000000. Assuming that the size of each row is 1GB (i.e. 0x40000000), the offset between banks may be defined by the manufacturer. Here, it is assumed that the offset of each bank is half the row size (i.e. 0x20000000). Therefore, the base address of row 0 is 0x00000000, and the offset of bank 1 is 0x20000000. Add the base address to the row and bank offsets to get the starting point of the physical address. In this example, the starting point of the physical address is 0x00000000+0x20000000=0x20000000. When the error location information provides a specific byte offset (such as 0x12345678), it is added to the starting address to obtain the final physical address 0x20000000+0x12345678=0x32345678.
[0092] In a possible implementation, some server manufacturers may provide dedicated management tools or diagnostic software, which can directly display the physical address of the ECC error.
[0093] A130 generates an error distribution matrix based on the number of errors and the physical address.
[0094] In this embodiment, an error distribution matrix based on the number of errors and the physical address where the error occurred. A hash table (or dictionary) can be used to store the key-value pair relationship between the number of errors and the physical address where the error occurred. After obtaining each physical address and its corresponding number of errors, traverse the data and update the hash table for each physical address and number of errors. If the address already exists, the number of errors is accumulated; if it does not exist, a new entry is created. When representing the error distribution in matrix form, first, determine the address range, and create a two-dimensional list based on this range, with all initial values being 0. Secondly, the number of errors in the hash table is assigned to the corresponding positions of the matrix, thereby obtaining an error distribution matrix.
[0095] In a feasible implementation manner, step A200 may further include steps A210 to A220:
[0096] Step A210: input the error distribution matrix into a preset mapping function.
[0097] Step A220, calculating the weight value corresponding to each element in the error distribution matrix.
[0098] In this implementation, the weight is calculated using a preset mapping function and a specific value of the number of errors. First, you need to determine the minimum and maximum number of errors in your data set. Assume that the minimum number of errors is min_error and the maximum number of errors is max_error. In the preset mapping function f(x), x is the number of errors and f(x) is the calculated weight. When x = min_error, f(x) = 0 (or a small value very close to 0 to avoid division by 0) When x = max_error, f(x) = 1.
[0099] Optionally, linear interpolation is used to calculate the weights. The function form is: f(x) = max_error - min_error + 2∈x - min_error + ∈. ε is a small positive number (such as 0.001) to avoid division by 0 and ensure that when x = min_error, f(x) is close to but not equal to 0.
[0100] Optionally, a Sigmoid function is used to calculate the weights. The function form is: f(x) = 1 / 1 + ek(x-mid), where k controls the steepness of the function (i.e., the slope), and mid is the center point of the function (which can be set to (min_error + max_error) / 2). The values of mid and k can be adjusted appropriately.
[0101] In a feasible implementation manner, determining a target error correction coding scheme according to a target physical address may include the following steps:
[0102] The error storage areas are grouped according to the preset classification criteria and weights.
[0103] According to the grouping result, the error correction coding algorithm and the length of the error correction coding check area corresponding to each group weight are determined.
[0104] An error correction coding scheme is generated according to the error correction coding algorithm and the length of the error correction coding check area, and there is a mapping relationship between the error correction coding scheme and the corresponding physical address.
[0105] In this embodiment, the groups can be grouped according to the regional weight size and the preset threshold, and different check area lengths can be set for each group. First, the weight matrix is obtained. The weight matrix is usually a two-dimensional array, in which each element represents a specific weight value. After obtaining the weight matrix, the entire weight matrix is traversed by programming to extract each weight value. Secondly, the corresponding error storage areas are grouped according to the preset range standard and weight size. The preset range standard can be dynamically adjusted according to actual needs. Finally, a different check area length is set for each group of error storage areas. The length of the check area may be determined based on factors such as the importance of the weight value, the frequency of data updates, or storage space limitations.
[0106] In this embodiment, by setting different check area lengths for error storage areas with different weights, more check area lengths are allocated to storage areas with weak retention capabilities to improve data reliability; while less check area lengths are allocated to storage areas with strong retention capabilities to save storage space. In addition, through reasonable partitioning and check area length settings, the physical layout of data can be optimized, the addressing time and delay when accessing data can be reduced, thereby improving data access speed.
[0107] In another feasible implementation, the grouping can be a simple division based on a threshold (for example, dividing areas with errors above a certain threshold into "high error areas"), or a more complex clustering method (such as K-means clustering) to identify areas with similar error characteristics. After the area division, the attention weight is calculated for each area. Features are extracted for each area, which may include the average number of errors in the area, the standard deviation of the error density, the size of the area, etc. A model is designed to calculate the attention weight based on the extracted features. This model can be a simple rule (such as the higher the number of errors, the greater the weight of the area), or it can be a more complex machine learning model (such as a neural network). Finally, the calculated attention weight is applied to the original data or the intermediate representation of the model, and areas with high weights will be considered more carefully or given higher priority.
[0108] In this embodiment, there is an address mapping table in the flash memory system, which records the mapping relationship between each physical address and the corresponding ECC configuration (including the length of the check area). This mapping table can be adjusted as needed when the system is running. When it is necessary to obtain the length of the ECC check area associated with a specific target physical address, the system will search for the entry corresponding to the address in the address mapping table. This search process can be a direct index (if the mapping table is arranged in order of physical addresses) or a more complex search algorithm (if the mapping table is unordered or compression technology is used). Once the entry for the target physical address is found in the mapping table, the length of the check area stored in the entry is read.
[0109] For example, a flash memory system contains 100 physical blocks, each with a unique physical address. To optimize performance and reliability, different ECC check area lengths are set for these blocks:
[0110] Blocks with physical addresses 0-49 use a 16-byte ECC check area. Blocks with physical addresses 50-74 use a 24-byte ECC check area to improve the error correction capability of these blocks. Blocks with physical addresses 75-99 use a 32-byte ECC check area because these blocks may be more susceptible to interference or damage.
[0111] In this embodiment, a mapping relationship table or function is established based on the positive correlation between the regional weight and the check bit length, which defines the check bit length corresponding to different weight ranges. After the mapping relationship is established, the error rate and weight of each region can be regularly evaluated, and the check bit length can be adjusted according to the evaluation results, and the mapping relationship can be updated. For example, a dynamic optimization method is used to adjust the check bit length. The historical data is trained using a machine learning algorithm to predict the trend of future error rate changes, and the check bit length is adjusted accordingly.
[0112] For example, the following mapping relationship can be set: for areas with weights lower than X1, the check bit length is N1; for areas with weights between X1 and X2, the check bit length is N2; for areas with weights higher than X2, the check bit length is N3, and N3>N2>N1.
[0113] In this embodiment, by selecting different check area length gears for different storage areas, different coding error correction schemes can be configured for different storage areas, and a higher check area length can be set for areas that are more prone to wear and errors. While improving data reliability and storage device performance, the data storage capacity is maximized and the impact of reduced data capacity due to increased check area length is reduced.
[0114] Based on the first embodiment of the present application, in the third embodiment of the present application, the same or similar contents as those in the above-mentioned embodiment 1 can be referred to the above introduction, and will not be repeated in the following. Figure 5 , after step S400, steps S500 to S800 are also included:
[0115] Step S500, after reading the test data corresponding to the data storage request, the test data is parsed according to the target error correction scheme to obtain comparison data.
[0116] Step S600: If the test data and the comparison data are inconsistent, a first error message is generated and an error log is updated.
[0117] In this embodiment, the test data includes a specific error pattern injected in advance. These data containing errors are written into the storage device. When the data is read and decoded, the performance of the error correction coding scheme can be evaluated by comparing the decoded data with the original data.
[0118] In this embodiment, when a test storage request for a specific target physical address is received, the request generally includes data to be written and related error correction coding scheme information. According to the test storage request, test data is extracted from the request. The target error correction coding scheme is used to encode the test data and write it to the target physical address. After writing the test data, the data just written is read from the target physical address. Using the same target error correction coding scheme, the read data is decoded to try to correct possible errors therein. The decoded data is called comparison data. The decoded comparison data is compared with the original test data to check whether the error correction coding scheme can correctly detect and correct errors. If the test data is found to be inconsistent with the comparison data during the comparison process, the storage device or the related test system should issue an error prompt. This prompt can be a simple status code, an error message, or more detailed diagnostic information to help locate the cause of the problem.
[0119] Step S700, if the test data and the comparison data are consistent, then determine whether the generated error cause matches the preset error mode;
[0120] Step S800: If there is no match, a second error message is generated and the error log is updated.
[0121] In this embodiment, the error log or status information related to the reading and decoding process is checked. The detected error cause is matched with the pre-injected error pattern. If the decoded data is consistent with the original data, but the error cause given does not match the preset error pattern, a second error prompt is generated. The second error information is added to the error log for subsequent analysis and processing. The error log should contain enough information to track the source, time of occurrence and context of the error.
[0122] In this embodiment, when error correction fails and considering updating the target error correction coding scheme, it can be verified whether the check bit length can meet the number of errors. During the verification process, the theoretical error correction capability of the error correction coding in the target error correction coding scheme can be calculated. For example, in the LDPC code, the LDPC code must be able to correct t errors, and its minimum code distance d should satisfy d≥2t+1. Compare the determined number of errors with the calculated theoretical error correction capability. If the number of errors is less than or equal to the error correction capability, theoretically, the errors should be successfully corrected. If the number of errors is greater than the error correction capability, the current check bit length is not enough to correct all errors. It is necessary to update the target error correction coding scheme and increase the check bit length to improve the error correction capability.
[0123] Through the above steps, it can be ensured that when error correction fails, the target error correction coding scheme can be appropriately updated, and the reliability of data and the stability of the system can be improved by verifying whether the check bit length meets the error number requirement.
[0124] The present application provides an error correction coding device, which includes: at least one processor; and a memory that is communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the error correction coding method in the above-mentioned embodiment one.
[0125] Reference below Figure 6 , which shows a schematic diagram of the structure of an error correction coding device suitable for implementing the embodiment of the present application. The error correction coding device in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The error correction coding device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0126] like Figure 6As shown, the error correction coding device may include a processing device 1001 (e.g., a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 to a random access memory (RAM: Random Access Memory) 1004. In RAM1004, various programs and data required for the operation of the error correction coding device are also stored. The processing device 1001, ROM1002, and RAM1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the error correction coding device to communicate with other devices wirelessly or by wire to exchange data. Although the error correction coding device with various systems is shown in the figure, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have alternatively.
[0127] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0128] The error correction coding device provided by the present application adopts the error correction coding method in the above embodiment, which can solve the technical problem of how to reduce the error rate of the storage device during use. Compared with the prior art, the beneficial effects of the error correction coding device provided by the present application are the same as the beneficial effects of the error correction coding method provided by the above embodiment, and the other technical features in the error correction coding device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.
[0129] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0130] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0131] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, and the computer-readable program instructions are used to execute the error correction coding method in the above-mentioned embodiment.
[0132] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0133] The computer-readable storage medium may be included in the error correction coding device; or may exist independently without being assembled into the error correction coding device.
[0134] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the error correction coding device, the error correction coding device: monitors the operating status of the storage device and obtains the operating status information of the storage device; when it is monitored that the operating status information meets the preset conditions, obtains the target physical address of the storage area in the storage device that meets the preset conditions; when a data storage request is received, determines the target error correction coding scheme according to the target physical address; and performs error correction coding according to the target error correction coding scheme.
[0135] Computer program code for performing the operations of the present application may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0136] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0137] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.
[0138] The readable storage medium provided by the present application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned error correction coding method, and can solve the technical problem of how to reduce the error rate of the storage device during use. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as the beneficial effects of the error correction coding method provided by the above-mentioned embodiment, and will not be repeated here.
[0139] The above descriptions are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect applications in other related technical fields are included in the patent protection scope of the present application.
Claims
1. An error correction coding method, characterized in that: The method includes: Monitor the operating status of the storage device and obtain the operating status information of the storage device; When it is monitored that the operation status information reaches a preset condition, obtaining a target physical address of a storage area in the storage device that meets the preset condition; When receiving a data storage request, determining a target error correction coding scheme according to the target physical address; Error correction coding is performed according to the target error correction coding scheme.
2. The error correction coding method according to claim 1, characterized in that: The operation status information includes the number of erase and write times, and when the operation status information is monitored to meet a preset condition, the step of acquiring a target physical address of a storage area in the storage device that meets the preset condition includes: Obtaining the number of erase and write times of the storage device; When the number of erasure and writing times reaches a preset first threshold, a target physical address corresponding to each storage area in the storage device whose number of erasure and writing times reaches the first threshold is obtained.
3. The error correction coding method according to claim 1, wherein: Before the step of monitoring the operating status of the storage device and obtaining the operating status information of the storage device, the step further includes: Performing an erase and write test on the storage device, and generating an error distribution matrix according to the test result; Calculating the weights corresponding to the error storage areas where errors occur according to the error distribution matrix; The step of determining a target error correction coding scheme according to the target physical address comprises: The target error correction coding scheme corresponding to the error storage area is determined according to the weight.
4. The error correction coding method according to claim 3, characterized in that: The step of performing an erase test on the storage device and generating an error distribution matrix according to the test result comprises: Obtain the number of errors detected by the error correction coding algorithm and the logical address where the error occurred; Mapping the logical address to the physical structure of the storage device to obtain the physical address of the error storage area; The error distribution matrix is generated according to the number of errors and the physical address.
5. The error correction coding method according to claim 3, characterized in that: The step of calculating the weights corresponding to the error storage areas where errors occur according to the error distribution matrix comprises: Inputting the error distribution matrix into a preset mapping function; The weight corresponding to each element in the error distribution matrix is calculated.
6. The error correction coding method according to claim 3, characterized in that: The step of determining the error correction coding scheme corresponding to the error storage area according to the weight comprises: Grouping the error storage areas according to a preset division standard and the weight; According to the grouping result, determining the error correction coding algorithm and the length of the error correction coding check area corresponding to the weight of each group; An error correction coding scheme is generated according to the error correction coding algorithm and the length of the error correction coding check area, and there is a mapping relationship between the error correction coding scheme and the corresponding physical address.
7. The error correction coding method according to claim 1, characterized in that: The step of performing error correction coding according to the target error correction coding scheme comprises: When the target physical address receives a test storage request, the test data corresponding to the test storage request is encoded and stored according to the target error correction scheme.
8. The error correction coding method according to any one of claims 1 to 7, characterized in that: After the step of performing error correction coding according to the target error correction coding scheme, the error correction coding method further includes: After reading the test data corresponding to the data storage request, parsing the test data according to the target error correction scheme to obtain comparison data; If the test data and the comparison data are inconsistent, a first error message is generated and an error log is updated; If the test data and the comparison data are consistent, determining whether the generated error cause matches a preset error pattern; If there is no match, a second error message is generated and the error log is updated.
9. An error correction coding device, characterized in that The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the error correction coding method according to any one of claims 1 to 8.
10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the error correction coding method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Flash memory management method, flash memory storage device, and computer-readable storage medium
CN109164978A
Data error correction method and device for Nand Flash, electronic equipment and storage medium
CN111813591A
Method of operating memory cell and memory device
CN113010102A
Data processing system and data processing method
US20020032891A1
Error correcting code predication system and method
US20090144598A1
Cited By
Hard disk data processing method and device, electronic equipment and storage medium
CN120832100A