A method for managing data blocks, an electronic device, a storage medium, and a product
By monitoring the growth rate of bad blocks of solid-state drives and comparing them with the simulation test results, potential risks are discovered in a timely manner, and the timeliness of bad block management in solid-state drives are solved to avoid data loss.
Patent Information
- Application Number
- CN202510389008.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-31
AI Technical Summary
The existing technology cannot detect bad blocks in solid-state drives in a timely manner, nor can it predict the potential data failure risk under the cover of the error correction mechanism, resulting in the loss of data stored in solid-state drives.
By obtaining the total amount of data blocks and marking information of the storage unit at the current time, counting the number of bad blocks, calculating the growth rate of bad blocks, and comparing them with the simulation test results, sending early warning information to prompt for abnormal states.
Timely discover that bad blocks in solid-state drives grow too fast, avoid data loss and ensure data security.
Smart Images

Figure CN119917029B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of storage technologies, and in particular, to a method for managing data blocks, an electronic device, a storage medium, and a product. Background Art
[0002] With the rapid development and wide application of solid state drives (SSDs), data security has become one of the important concerns of SSDs. An SSD includes flash memory chips, each flash memory chip includes at least one memory cell, each memory cell includes a plurality of data blocks, and each data block includes at least one storage block; during the operation of an SSD, there may be a small probability that due to problems with the flash memory chips of the SSD or other issues, many new bad blocks are added to the SSD, or the number of bad blocks increases too much in a short period of time. At this time, although the error correction mechanism of the SSD temporarily maintains the normal transmission of data input and output, if the SSD continues to be used, it is very likely that the SSD will fail or the data stored in the SSD will be lost. Therefore, it is necessary to manage bad blocks.
[0003] Traditional bad block management usually marks a bad block when the number of faulty storage blocks included in the data blocks in the SSD reaches a certain threshold, resulting in failures in reading, writing, and erasing the data blocks. However, this method, although it can mark bad blocks to indicate that there is a fault in the SSD to a certain extent, cannot detect bad blocks in a timely manner, nor can it predict the potential data failure risk hidden under the error correction mechanism, and still may cause the data stored in the SSD to be lost. Summary of the Invention
[0004] This application provides a method for managing data blocks, an electronic device, a storage medium, and a product, so as to at least solve the problem in the related art that bad blocks in a solid state drive cannot be detected in a timely manner, nor can the potential data failure risk hidden under the error correction mechanism be predicted, resulting in the loss of data stored in the solid state drive.
[0005] The present application provides a method for managing data blocks, the method including: obtaining the total amount of data blocks included in the storage unit at the current moment, and the marking information of each data block, where the storage unit includes at least one data block, and the marking information is used to indicate whether the current data block has failed; determining whether there are bad blocks among the at least one data block based on the marking information, and counting the first quantity of the determined bad blocks; when the first quantity is less than the first threshold, obtaining the number of newly added bad blocks in the storage unit during N write-erase operations before the current moment, and calculating the ratio between the number of newly added bad blocks and N to obtain the average bad block growth rate, where N≥1 and is a positive integer; calculating the ratio between the average bad block growth rate and the total amount of data blocks to obtain the first bad block growth rate of the storage unit at the current moment; if the first bad block growth rate is greater than or equal to the second bad block growth rate in the simulation test, sending a first warning message to the host, where the first warning message is used to report to the host that the data block status of the storage unit is abnormal at the current moment.
[0006] The present application further provides a data block management device, the device including: a transceiver module, configured to obtain the total amount of data blocks included in the storage unit at the current moment, and the marking information of each data block, where the storage unit includes at least one data block, and the marking information is used to indicate whether the current data block has failed; a processing module, configured to determine whether there are bad blocks among the at least one data block based on the marking information, and count the first quantity of the determined bad blocks; the processing module is further configured to, when the first quantity is less than the first threshold, obtain the number of newly added bad blocks in the storage unit during N write-erase operations before the current moment, and calculate the ratio between the number of newly added bad blocks and N to obtain the average bad block growth rate, where N≥1 and is a positive integer; the processing module is further configured to calculate the ratio between the average bad block growth rate and the total amount of data blocks to obtain the first bad block growth rate of the storage unit at the current moment; the transceiver module is further configured to, if the first bad block growth rate is greater than or equal to the second bad block growth rate in the simulation test, send a first warning message to the host, where the first warning message is used to report to the host that the data block status of the storage unit is abnormal at the current moment.
[0007] The present application further provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any one of the above data block management methods when executing the computer program.
[0008] The present application further provides a computer-readable storage medium, where a computer program is stored in the computer-readable storage medium, and wherein the computer program, when executed by a processor, implements the steps of any one of the above data block management methods.
[0009] The present application further provides a computer program product, including a computer program, where the computer program, when executed by a processor, implements the steps of any one of the above data block management methods.
[0010] Through this application, when the first quantity of bad blocks is less than the first threshold, the quantity of newly added bad blocks in the storage unit during the N write-erase operations before the current moment can be obtained, and the ratio to N can be calculated to obtain the average bad block growth rate, and then the first bad block growth rate of the storage unit at the current moment can be calculated, that is, the real-time bad block growth rate of the storage unit can be obtained.
[0011] After that, since the second bad block growth rate is the normal bad block growth rate corresponding to the situation where the storage unit can operate normally based on the simulation test results, and then the real-time bad block growth rate is compared with the second bad block growth rate in the simulation test, it can be determined whether the bad block growth rate of the storage unit at the current moment exceeds the expected index. If so, it indicates that the current bad block growth rate is too fast, and a first warning message is sent to the host to prompt the user to pay attention to the abnormal state of the storage unit in time, and then it can be determined in time that there is a risk in the solid-state drive and the data stored in the solid-state drive can be prevented from being lost. Description of the Drawings
[0012] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0013] Figure 1 It is a topological structure diagram of a data block management system provided by an embodiment of the present application;
[0014] Figure 2 It is a schematic flowchart of a data block management method provided by an embodiment of the present application;
[0015] Figure 3 It is a schematic diagram of the determination process of the second bad block growth rate in the simulation test provided by an embodiment of the present application;
[0016] Figure 4 It is a schematic flowchart of a bad block marking method provided by an embodiment of the present application;
[0017] Figure 5 It is a structural block diagram of a data block management device provided by an embodiment of the present application;
[0018] Figure 6 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application. Detailed Embodiments
[0019] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the protection scope of the present application.
[0020] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and not to describe a specific order or sequence.
[0021] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0022] In combination with a specific application environment architecture or a specific hardware architecture on which the execution of a method for managing data blocks depends, the specific application environment architecture or the specific hardware architecture will be described herein.
[0023] The embodiments of the present application are applied in the process of bad block management and warning scenarios of solid-state drives.
[0024] In the related art, although traditional bad block management can mark bad blocks to indicate that the SSD has a fault to a certain extent, it cannot detect bad blocks in time, nor can it predict the potential data failure risk masked by the error correction mechanism, which will still cause the data stored in the SSD to be lost.
[0025] To solve the above technical problems, an embodiment of the present application provides a method for managing data blocks. The method includes: obtaining the total number of data blocks contained in the storage unit at the current moment and the marking information of each data block; determining whether there are bad blocks in at least one data block based on the marking information, and counting the first number of determined bad blocks; when the first number is less than the first threshold, obtaining the number of newly added bad blocks in the storage unit during the previous N write-erase operations before the current moment, and calculating the ratio between the number of newly added bad blocks and N to obtain the average bad block growth rate; calculating the ratio between the average bad block growth rate and the total number of data blocks to obtain the first bad block growth rate of the storage unit at the current moment; if the first bad block growth rate is greater than or equal to the second bad block growth rate in the simulation test, sending a first warning message to the host. In this way, when the bad block growth rate is too large, it can be timely determined that the bad block growth in the solid-state drive is too fast, and then it can be timely determined that there is a risk in the solid-state drive, avoiding data loss stored in the solid-state drive.
[0026] The following takes Figure 1 the data block management system shown as an example to describe the method provided by the embodiment of the present application.
[0027] As Figure 1 shown, Figure 1 is the topology structure diagram of a data block management system provided by an embodiment of the present application. Figure 1 In, the data block management system 100 includes a data block management device 101, a host 102, and a solid-state drive 103.
[0028] The data block management device 101 in the embodiment of the present application can be any device with communication and computing functions. For example, the data block management device 101 can be a server, a virtual machine, or a cloud server.
[0029] The host 102 in the embodiment of the present application can be a central host (Central Processing Unit, CPU), which is the core component of a computer system.
[0030] The solid-state drive 103 in the embodiment of the present application can be a storage device based on flash memory particles. The solid-state drive includes multiple storage units, and each storage unit is also called a die (also called DIE). Each storage unit includes multiple data blocks. Each data block includes multiple storage blocks.
[0031] Figure 1 The data block management system shown is only for illustration and is not used to limit the technical solutions of the present application. Those skilled in the art should understand that in the specific implementation process, the management system at the data block can also include more devices, which are not limited.
[0032] An embodiment of the present application provides a data block processing method, which is applied to Figure 1 the management device of the data block shown in Figure 2 as shown in Figure 2 which is a schematic flowchart of a data block management method provided by an embodiment of the present application. The data block management method includes the following steps:
[0033] S201: Obtain the total amount of data blocks contained in the storage unit at the current moment, and the marking information of each data block.
[0034] Among them, the storage unit contains at least one data block. The storage unit can be a single DIE in a solid-state drive.
[0035] The marking information is used to indicate whether the current data block has failed.
[0036] It can be understood that if the marking information of a data block is marked, it indicates that the data block has failed, and this data block is called a bad block. If the marking information of a data block is empty, it indicates that the data block has not failed, and this data block is a normal data block.
[0037] S202: Determine whether there are bad blocks in at least one data block based on the marking information, and count the first quantity of the determined bad blocks.
[0038] In one example, the data block management device determines whether each data block in at least one data block has failed based on the marking information of each data block, and counts the data blocks that have failed to obtain the first quantity of bad blocks.
[0039] S203: When the first quantity is less than the first threshold, obtain the number of newly added bad blocks in the storage unit during the previous N write-erase operations before the current moment, and calculate the ratio between the number of newly added bad blocks and N to obtain the average bad block growth.
[0040] Among them, N≥1 and is a positive integer.
[0041] In an embodiment of the present application, the first threshold is the maximum number of bad blocks in the entire life cycle of the storage unit recorded in the user manual of the solid-state drive, and this maximum number of bad blocks is also called MAX_BAD_CNT.
[0042] In an embodiment of the present application, N can be determined based on the total amount of data blocks in the storage unit, the maximum number of bad blocks, and the total number of write-erase operations that the storage unit has performed at the current moment. Among them, the number of write-erase operations is also called the number of Program / Eras (P / E) times.
[0043] In one example, the management device of the data block determines the value of N according to the first expression, the total amount of data blocks in the storage unit, the maximum number of bad blocks, and the total number of write / erase operations performed on the storage unit.
[0044] Among them, the first expression is:
[0045] Among them, is the total amount of data blocks; is the maximum number of bad blocks; is the total number of write / erase operations performed at the current moment.
[0046] It can be understood that the management device of the data block comprehensively considers the influence of the total amount of data blocks, the maximum number of bad blocks, and the P / E times on the life of the storage unit and the generation of bad blocks. The larger the total amount of data blocks, the corresponding increase in the possible number of bad blocks; the maximum number of bad blocks limits the upper limit of bad blocks that the storage unit can tolerate; and the P / E times at the current moment reflect the number of erase / write cycles experienced by the storage unit. The more times, the easier it is to generate bad blocks. The management device of the data block dynamically determines the value of N according to the specific storage unit parameters, so as to more reasonably obtain the number of new bad blocks in the storage unit during a certain number of write / erase operations before the current moment.
[0047] Optionally, when the first quantity is greater than or equal to the first threshold, the management device of the data block sends a second warning message to the host.
[0048] Among them, the second warning message indicates that the storage unit cannot be used continuously.
[0049] It can be understood that if the first quantity is greater than or equal to the first threshold, it indicates that the storage unit is damaged. At this time, the management device of the data block activates a red warning, sets the position of the red warning flag for this storage unit, reduces the write data concurrency, performs active speed limiting, and notifies the user through the SMART log, so that the user can handle this abnormal disk in time.
[0050] S204: Calculate the ratio between the average growth of bad blocks and the total amount of data blocks to obtain the first bad block growth rate of the storage unit at the current moment.
[0051] S205: If the first bad block growth rate is greater than or equal to the second bad block growth rate in the simulation test, send a first warning message to the host.
[0052] Among them, the first warning message is used to report to the host that the data block status of the storage unit is abnormal at the current moment.
[0053] It can be understood that if the growth rate of the first bad block is greater than or equal to the growth rate of the second bad block in the simulation test, it indicates that the growth rate of bad blocks is too fast. At this time, a yellow warning is issued, and the data block management device can send a first warning message to the host through the SMART log to notify the user, so that the user can pay attention to the subsequent log information of this solid-state drive.
[0054] Specifically, the process for the data block management device to determine the growth rate of the second bad block in the simulation test is as follows Figure 3 shown Figure 3 is a schematic diagram of the process for determining the growth rate of the second bad block in a simulation test provided by an embodiment of the present application. The data block management device performs the following steps:
[0055] S301: Obtain the current temperature of the storage unit at the current moment and the total number of write / erase operations that have been performed.
[0056] S302: Based on a preset number of write / erase operation counts and a plurality of temperatures, perform a bad block test on the storage unit to obtain a mapping table.
[0057] Among them, the mapping table includes at least one mapping relationship, and each mapping relationship is a mapping relationship between a write / erase operation count ladder, a temperature ladder, and a basic bad block growth rate.
[0058] In an example, based on a preset number of write / erase operation counts and a plurality of temperatures, the data block management device uses a test fixture to perform a bad block test on the storage unit at a plurality of P / E count ladders ( ) and a plurality of temperatures ( ), and then calculates the growth rate of each ladder ( ) respectively. As shown in Table 1, Table 1 is a schematic table of the mapping table. In Table 1, the average growth rate of bad blocks under different P / E count ladders and temperature T ladders can be recorded.
[0059] Table 1: Mapping Table
[0060]
[0061] It can be understood that the maximum P / E count and the maximum tolerable temperature of different storage units will vary, and the gradient of the ladder can also be made more refined for testing, without limitation.
[0062] S303: Based on at least one mapping relationship, perform a simulation on the storage unit to obtain an aging influence coefficient and a temperature influence coefficient.
[0063] Among them, the aging influence coefficient reflects the relationship between the average growth rate of bad blocks and the P / E count. The temperature influence coefficient reflects the relationship between the average growth rate of bad blocks and the temperature.
[0064] In one example, the management device of the data block simulates the storage unit through a linear additive mathematical model based on at least one mapping relationship, multiple write / erase operation count ladders, multiple temperature ladders, and multiple basic bad block growth rates, and obtains an aging influence coefficient and a temperature influence coefficient.
[0065] S304: Determine a second bad block growth rate based on the current temperature, the total number of times, the mapping table, the aging influence coefficient, and the temperature influence coefficient.
[0066] In some alternative embodiments, the management device of the data block searches in the mapping table for the target temperature ladder where the current temperature is located, the target write / erase operation count ladder where the total number of times is located, and the target basic bad block growth rate corresponding to the target temperature ladder and the target write / erase operation count ladder; and calculates the second bad block growth rate according to the basic temperature value corresponding to the target temperature ladder, the basic write / erase operation count corresponding to the target write / erase operation count ladder, the target basic bad block growth rate, the aging influence coefficient, the temperature influence coefficient, the current temperature, and the total number of times, according to a preset expression.
[0067] Wherein, the preset expression is: 。
[0068] is the second bad block growth rate; is the target basic bad block growth rate; is the aging influence coefficient; is the temperature influence coefficient; is the total number of write / erase operations that have been performed; is the basic write / erase operation count; is the current temperature; is the basic temperature value.
[0069] Based on the above Figure 3 The second bad block growth rate can be calculated by the method shown.
[0070] Before obtaining the marking information of each data block, the embodiments of the present application further provide a method for marking bad blocks, as Figure 4 shown, Figure 4 is a schematic flowchart of a method for marking bad blocks provided by the embodiments of the present application; the management device of the data block may further perform the following steps:
[0071] S401: At a first moment, receive a read request sent from a host.
[0072] The read request carries address information of at least one storage block included in the target data block. The target data block is a data block in the storage unit.
[0073] The first moment can be any moment before the current moment.
[0074] S402: Read the stored data in at least one storage block from the target data block according to the address information.
[0075] In one example, the management device of the data block reads the stored data corresponding to the address information from the target data block according to the address information.
[0076] S403: Perform hard decoding on the stored data and detect whether the hard decoding is successful.
[0077] S404: If it is detected that the hard decoding is successful, obtain the target data after hard decoding and record the first bit error rate of the stored data during the hard decoding process.
[0078] Among them, the bit error rate (Bit Error Rate, BER) is the ratio of the number of error bits in the bit stream received by the host to the total number of bits during the hard decoding process.
[0079] The target data is the data after the stored data is successfully decoded.
[0080] S405: Determine whether the first bit rate is greater than or equal to the second threshold.
[0081] In one example, the management device of the data block obtains multiple historical bit error rates within a historical time period; calculates the product of the preset ratio and the maximum historical bit error rate to obtain the second threshold.
[0082] Among them, each historical bit error rate is the bit error rate corresponding to each successful hard decoding. The maximum historical bit error rate is the maximum historical bit error rate among multiple historical bit error rates, or the maximum BER of successful hard decoding during normal reading of the host.
[0083] The historical time period can also be the time period when the management device of the data block performs simulation tests.
[0084] The preset ratio can be set according to actual needs. For example, the preset ratio can be 80%.
[0085] S406: If the first bit error rate is greater than or equal to the second threshold, write all the stored data in the target data block into another data block.
[0086] Among them, the other data block is any free data block in the storage unit except the target data block.
[0087] In one example, if the first bit error rate is greater than or equal to the second threshold, the management device of the data block performs a force Garbage Collection (force GC) process to write all the stored data in the target data block into other data blocks.
[0088] It can be understood that since the first bit error rate is greater than or equal to the second threshold, it indicates that the bit error rate corresponding to the target data block is relatively large, that is, it means that the current data in the target data block is not very stable, which may trigger error correction and requires early relocation processing. It is necessary to back up all the stored data in the data block to other data blocks.
[0089] S407: If the first bit error rate is less than the second threshold, determine that at least one storage block is normal.
[0090] It can be understood that if the first bit error rate is less than the second threshold, it indicates that the BER corresponding to the target data block is relatively low, that is, it means that the current data in the target data block is relatively healthy and no additional processing is required.
[0091] S408: If a hard decoding failure is detected, perform soft decoding on the stored data and detect whether the soft decoding is successful.
[0092] It can be understood that if a hard decoding failure is detected, it means that the normal reading of the data cannot be read out, and soft decoding error correction is required.
[0093] S409: If a soft decoding failure is detected, mark the target data block as a bad block.
[0094] It can be understood that if a soft decoding failure is also detected, it means that the stored data corresponding to at least one current storage block has failed, and bad block marking and data relocation processing are required (that is, writing all the stored data in the target data block into other data blocks).
[0095] Optionally, the management device of the data block can also perform error correction through other error correction means, which will not be elaborated here.
[0096] S410: If the soft decoding is successful, obtain the target data after soft decoding and the first iteration count required for error correction by the error correction algorithm during the soft decoding process.
[0097] The error correction algorithm is the Error Checking and Correction (ECC) algorithm, which is used to determine the location of the error and correct it according to the preset rules when an error occurs during data transmission or storage.
[0098] S411: Determine whether the first iteration count is less than the third threshold.
[0099] In one example, the management device of the data block obtains multiple historical iteration counts within a historical time period; calculates the product of a preset ratio and the maximum historical iteration count to obtain a third threshold.
[0100] Among them, each historical iteration count is the number of iterations required for the error correction algorithm to perform error correction after successful soft decoding each time. The maximum historical iteration count is the maximum iteration count among the multiple historical iteration counts, or the maximum historical iteration count of the error correction algorithm.
[0101] S412: If the first iteration count is less than the third threshold, write all the stored data in the target data block to other data blocks.
[0102] It can be understood that when the first iteration count is less than the third threshold, that is, when the number of iterations required for the error correction algorithm to perform error correction during the soft decoding process is low, it means that the stored data is easy to correct back. Therefore, only a relocation process is needed, that is, write all the stored data in the target data block to other data blocks.
[0103] S413: If the first iteration count is greater than or equal to the third threshold, record the current error correction address and the first number of write / erase operations that the target storage unit has performed at the first moment.
[0104] It can be understood that the target storage unit is the above-mentioned storage unit. When the first iteration count is greater than or equal to the third threshold, that is, when the number of iterations required for the error correction algorithm to perform error correction during the soft decoding process is high, it means that although the stored data can be corrected back through soft decoding, the BER is very high. In addition to performing a data relocation process on the current stored data, the current position also needs to be recorded.
[0105] S414: During the M write / erase operations after the first number, when the soft decoding of the data stored at the current error correction address is successful again and the second iteration count is greater than or equal to the third threshold, mark the target data block as a bad block.
[0106] Among them, the second iteration count is the number of iterations required for the error correction algorithm to perform error correction during the process of performing soft decoding on the data stored at the current error correction address again.
[0107] M ≥ 1 and is a positive integer. M can be set according to actual needs. For example, M is 5.
[0108] It can be understood that within the specified number of additional P / E cycles, when reading this position again, if the second iteration count is still greater than or equal to the third threshold, it indicates that the current target data block is probably a weak block. Mark it as a bad block in advance to prevent the risk of data loss, and at the same time, it can also reduce data error correction, thereby reducing the impact on data input / output performance.
[0109] S415: During the M write / erase operations after the first count, when the hard decoding of the data stored at the current error correction address is successful again, or when the soft decoding of the data stored at the current error correction address is successful again and the second iteration count is less than the third threshold, determine that at least one storage block is normal.
[0110] It can be understood that within the specified number of additional P / E cycles, when reading this location again, if the second iteration count is less than the third threshold, it indicates that the BER of the current target data block is relatively low, and there is no need to mark a bad block. Determine that at least one storage block is normal.
[0111] Based on the above Figure 2 As shown in the method, when the first quantity of bad blocks is less than the first threshold, the data block management device can obtain the quantity of newly added bad blocks in the storage unit during the N write / erase operations before the current moment, and the ratio to N to obtain the average bad block growth rate, and then calculate the first bad block growth rate of the storage unit at the current moment, that is, the real-time bad block growth rate of the storage unit can be obtained.
[0112] After that, since the second bad block growth rate is the normal bad block growth rate corresponding to the situation where the storage unit can operate normally determined based on the simulation test results, and then compare the real-time bad block growth rate with the second bad block growth rate in the simulation test, it can be judged whether the bad block growth rate of the storage unit at the current moment exceeds the expected index. If so, it indicates that the current bad block growth rate is too fast, and send the first warning message to the host to prompt the user to pay attention to the abnormal state of the storage unit in a timely manner, and then determine that there is a risk in the solid-state drive in a timely manner to avoid data loss in the solid-state drive.
[0113] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0114] The embodiment of the present application also provides a data block management device, as Figure 5 shown, Figure 5 is the structural block diagram of a data block management device provided by the embodiment of the present application; the device includes:
[0115] A transceiver module 501, configured to obtain the total quantity of data blocks included in the storage unit at the current moment, and the marking information of each data block. The storage unit includes at least one data block, and the marking information is used to indicate whether the current data block has a fault.
[0116] A processing module 502, configured to determine whether there are bad blocks in at least one data block based on the marking information, and count the first quantity of the determined bad blocks.
[0117] The processing module 502 is further configured to, when the first quantity is less than the first threshold, obtain the number of newly added bad blocks in the storage unit during the N write-erase operations before the current moment, and calculate the ratio between the number of newly added bad blocks and N to obtain the average bad block growth rate, where N≥1 and is a positive integer.
[0118] The processing module 502 is further configured to calculate the ratio between the average bad block growth rate and the total amount of data blocks to obtain the first bad block growth rate of the storage unit at the current moment.
[0119] The transceiver module 501 is further configured to, if the first bad block growth rate is greater than or equal to the second bad block growth rate in the simulation test, send a first warning message to the host, and the first warning message is used to report to the host that the data block state of the storage unit is abnormal at the current moment.
[0120] In some alternative embodiments, the transceiver module 501 is further configured to obtain the current temperature of the storage unit at the current moment and the total number of write-erase operations that have been performed; the processing module 502 is further configured to perform a bad block test on the storage unit based on a plurality of preset write-erase operation times and a plurality of temperatures to obtain a mapping table, where the mapping table includes at least one mapping relationship, and each mapping relationship is a mapping relationship between a write-erase operation times ladder, a temperature ladder, and a basic bad block growth rate; based on the at least one mapping relationship, simulate the storage unit to obtain an aging influence coefficient and a temperature influence coefficient; and determine the second bad block growth rate based on the current temperature, the total number, the mapping table, the aging influence coefficient, and the temperature influence coefficient.
[0121] In some alternative embodiments, the processing module 502 is specifically configured to search in the mapping table for a target temperature ladder where the current temperature is located, a target write-erase operation times ladder where the total number is located, and a target basic bad block growth rate corresponding to the target temperature ladder and the target write-erase operation times ladder; and calculate the second bad block growth rate according to the basic temperature value corresponding to the target temperature ladder, the basic write-erase operation times corresponding to the target write-erase operation times ladder, the target basic bad block growth rate, the aging influence coefficient, the temperature influence coefficient, the current temperature, and the total number according to a preset expression, where the preset expression is: 。
[0122] Wherein, is the second bad block growth rate; is the target basic bad block growth rate; is the aging influence coefficient; is the temperature influence coefficient; is the total number of write-erase operations that have been performed; is the basic write-erase operation times; is the current temperature; is the basic temperature value.
[0123] In some alternative embodiments, the transceiver module 501 is further configured to receive a read request sent from a host at a first moment, where the read request carries address information of at least one storage block included in a target data block, and the target data block is a data block in a storage unit; the processing module 502 is further configured to read stored data in at least one storage block from the target data block according to the address information; perform hard decoding on the stored data; if it is detected that the hard decoding is successful, obtain the target data after hard decoding, and record a first bit error rate of the stored data during the hard decoding process; if the first bit error rate is greater than or equal to a second threshold, write all the stored data in the target data block into another data block, and the other data block is any free data block in the storage unit other than the target data block.
[0124] In some alternative embodiments, the processing module 502 is further configured to determine that at least one storage block is normal if the first bit error rate is less than the second threshold.
[0125] In some alternative embodiments, the processing module 502 is further configured to perform soft decoding on the stored data if it is detected that the hard decoding fails; if it is detected that the soft decoding fails, mark the target data block as a bad block.
[0126] In some alternative embodiments, the processing module 502 is further configured to obtain the target data after soft decoding and a first number of iterations required for error correction by an error correction algorithm during the soft decoding process if it is detected that the soft decoding is successful; if the first number of iterations is less than a third threshold, write all the stored data in the target data block into another data block.
[0127] In some alternative embodiments, the processing module 502 is further configured to record a current error correction address and a first number of write / erase operations that the storage unit has performed at the first moment if the first number of iterations is greater than or equal to the third threshold; during M write / erase operations after the first number of operations, when soft decoding of the data stored at the current error correction address is successful again and a second number of iterations is greater than or equal to the third threshold, mark the target data block to which the stored data belongs as a bad block, where the second number of iterations is the number of iterations required for error correction by the error correction algorithm during the process of performing soft decoding on the data stored at the current error correction address again; M≥1 and is a positive integer.
[0128] In some alternative embodiments, the processing module 502 is further configured to determine that at least one storage block is normal during M write / erase operations after the first number of operations when hard decoding of the data stored at the current error correction address is successful again, or when soft decoding of the data stored at the current error correction address is successful again and the second number of iterations is less than the third threshold.
[0129] In some alternative embodiments, the transceiver module 501 is further configured to send a second warning message to the host when the first quantity is greater than or equal to the first threshold, where the second warning message indicates that the storage unit cannot be used continuously.
[0130] In some alternative embodiments, the transceiver module 501 is further configured to obtain a plurality of historical bit error rates within a historical time period, where each historical bit error rate is the bit error rate corresponding to each successful hard decoding; the processing module 502 is further configured to calculate the product of a preset ratio and the maximum historical bit error rate to obtain a second threshold, where the maximum historical bit error rate is the largest historical bit error rate among the plurality of historical bit error rates.
[0131] In some alternative embodiments, the transceiver module 501 is further configured to obtain a plurality of historical iteration times within a historical time period, where each historical iteration time is the number of iterations required for the error correction algorithm to correct errors after each successful soft decoding; the processing module 502 is further configured to calculate the product of a preset ratio and the maximum historical iteration time to obtain a third threshold, where the maximum historical iteration time is the largest iteration time among the plurality of historical iteration times.
[0132] For the description of the features in the corresponding embodiments of the data block management device, reference may be made to the relevant description in the corresponding embodiments of the data block management method, which will not be elaborated here one by one.
[0133] Embodiments of the present application further provide an electronic device, such as Figure 6 shown Figure 6 is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application; the electronic device includes a processor 10 and a memory 20, where a computer program is stored in the memory 20, and the processor 10 is configured to run the computer program to execute the steps in any of the above embodiments of the data block management method.
[0134] Embodiments of the present application further provide a computer-readable storage medium, where a computer program is stored in the computer-readable storage medium, and the computer program is configured to execute the steps in any of the above embodiments of the data block management method when running.
[0135] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as a USB flash drive, a read-only memory (ROM for short), a random access memory (RAM for short), a mobile hard disk, a magnetic disk, or an optical disc that can store a computer program.
[0136] An embodiment of the present application also provides a computer program product. The computer program product includes a computer program which, when executed by a processor, implements the steps in any of the above-described embodiments of the data block management method.
[0137] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program which, when executed by a processor, implements the steps in any of the above-described embodiments of the data block management method.
[0138] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.
[0139] The above has introduced in detail a data block management method, an electronic device, a storage medium, and a product provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A method for managing data blocks, characterized in that, The method includes: Obtaining the total number of data blocks contained in the storage unit at the current moment and the marking information of each data block, where the storage unit contains at least one data block, and the marking information is used to indicate whether the current data block has failed; Based on the marking information, determining whether there are bad blocks in the at least one data block, and counting the number of bad blocks determined; When the number of bad blocks is less than the first threshold, obtaining the number of newly added bad blocks in the storage unit during the previous N write-erase operations before the current moment, and calculating the ratio between the number of newly added bad blocks and N to obtain the average bad block growth rate, where N≥1 and is a positive integer; Calculating the ratio between the average bad block growth rate and the total number of data blocks to obtain the first bad block growth rate of the storage unit at the current moment; If the first bad block growth rate is greater than or equal to the second bad block growth rate in the simulation test, sending a first warning message to the host, where the first warning message is used to report to the host that the data block status of the storage unit is abnormal at the current moment; Among them, the determination process of the second bad block growth rate in the simulation test includes: Obtaining the current temperature of the storage unit at the current moment and the total number of write-erase operations that have been performed; Based on a preset plurality of write-erase operation times and a plurality of temperatures, performing a bad block test on the storage unit to obtain a mapping table, where the mapping table includes at least one mapping relationship, and each mapping relationship is a mapping relationship between a write-erase operation time step, a temperature step, and a basic bad block growth rate; Based on the at least one mapping relationship, simulating the storage unit to obtain an aging influence coefficient and a temperature influence coefficient; Based on the current temperature, the total number of times, the mapping table, the aging influence coefficient, and the temperature influence coefficient, determining the second bad block growth rate; The determining the second bad block growth rate based on the current temperature, the total number of times, the mapping table, the aging influence coefficient, and the temperature influence coefficient includes: Searching in the mapping table for the target temperature step where the current temperature is located, the target write-erase operation time step where the total number of times is located, and the target basic bad block growth rate corresponding to the target temperature step and the target write-erase operation time step; According to the basic temperature value corresponding to the target temperature step, the basic write-erase operation times corresponding to the target write-erase operation time step, the target basic bad block growth rate, the aging influence coefficient, the temperature influence coefficient, the current temperature, and the total number of times, calculating the second bad block growth rate according to a preset expression; The preset expression is as follows: Among them, is the second bad block growth rate; is the target basic bad block growth rate; is the aging influence coefficient; is the temperature influence coefficient; is the total number of write-erase operations that have been performed; is the basic write-erase operation times; is the current temperature; is the basic temperature value.
2. The method according to claim 1, wherein The method further includes: At a first moment, receiving a read request sent by the host, where the read request carries the address information of at least one storage block included in a target data block, and the target data block is a data block in the storage unit; According to the address information, reading the stored data in the at least one storage block from the target data block; Performing hard decoding on the stored data; If it is detected that the hard decoding is successful, the target data after the hard decoding is obtained, and the first bit error rate of the stored data during the hard decoding is recorded; If the first bit error rate is greater than or equal to a second threshold, all the stored data in the target data block is written into another data block, and the other data block is any free data block in the storage unit except the target data block.
3. The method according to claim 2, wherein The method further includes: If the first bit error rate is less than the second threshold, it is determined that the at least one storage block is normal.
4. The method according to claim 2, wherein The method further includes: If it is detected that the hard decoding fails, the stored data is soft decoded; If it is detected that the soft decoding fails, the target data block to which the stored data belongs is marked as the bad block.
5. The method according to claim 4, characterized in that, The method further includes: If it is detected that the soft decoding is successful, the target data after the soft decoding is obtained, and the first number of iterations required for the error correction algorithm to perform error correction during the soft decoding is obtained; If the first number of iterations is less than a third threshold, all the stored data in the target data block is written into the other data block.
6. The method according to claim 5, wherein The method further includes: If the first number of iterations is greater than or equal to the third threshold, the current error correction address and the first number of write / erase operations that the storage unit has performed at the first moment are recorded; During the M write / erase operations after the first number, when the data stored at the current error correction address is soft decoded successfully again and the second number of iterations is greater than or equal to the third threshold, the target data block is marked as the bad block, where the second number of iterations is the number of iterations required for the error correction algorithm to perform error correction during the process of soft decoding the data stored at the current error correction address again; M≥1 and is a positive integer.
7. The method according to claim 6, wherein The method further includes: During the M write / erase operations after the first number, when the data stored at the current error correction address is hard decoded successfully again, or when the data stored at the current error correction address is soft decoded successfully again and the second number of iterations is less than the third threshold, it is determined that the at least one storage block is normal.
8. The method according to claim 7, characterized in that, The method further includes: When the number of bad blocks is greater than or equal to the first threshold, a second warning message is sent to the host, and the second warning message indicates that the storage unit cannot be used continuously.
9. The method according to claim 8, wherein The method further includes: A plurality of historical bit error rates within a historical time period are obtained, and each historical bit error rate is the bit error rate corresponding to each successful hard decoding; The product of a preset ratio and the maximum historical bit error rate is calculated to obtain the second threshold, where the maximum historical bit error rate is the largest historical bit error rate among the plurality of historical bit error rates.
10. The method according to claim 9, wherein The method further includes: A plurality of historical numbers of iterations within the historical time period are obtained, and each historical number of iterations is the number of iterations required for the error correction algorithm to perform error correction during each successful soft decoding; The product of the preset ratio and the maximum historical number of iterations is calculated to obtain the third threshold, where the maximum historical number of iterations is the largest number of iterations among the plurality of historical numbers of iterations.
11. An electronic device, characterized in that, Includes: A memory for storing a computer program; A processor for implementing the steps of the method for managing data blocks as recited in any one of claims 1 to 10 when executing the computer program.
12. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the method for managing data blocks as recited in any one of claims 1 to 10 when executed by a processor.
13. A computer program product, comprising a computer program, characterized in that, The computer program implements the steps of the method for managing data blocks as recited in any one of claims 1 to 10 when executed by a processor.
Citation Information
Patent Citations
Bad block detection method and system for solid state disk
CN118394609A
Memory block replacement method and device based on solid state disk, equipment and medium
CN118585128A