Risk block screening method, device, storage device and storage medium

By marking problem blocks and dynamically adjusting the window range, and combining the risk block distribution rules, the risk blocks in the flash memory media are accurately screened, which solves the problems of insufficient analysis in the existing technology and improves the reliability and performance of the storage media.

CN119724306BActive Publication Date: 2025-07-25BIWIN STORAGE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510229058.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-07-25
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

In the prior art, the characteristics analysis of flash memory media is insufficient, and it is impossible to efficiently screen out potential risk weak blocks, resulting in the impact of the reliability and performance of the storage media.

Method used

By marking the memory blocks that meet the aging test failure conditions as problem blocks, setting the window range is centered on the problem block, traversing and detecting neighborhood memory blocks, dynamically adjusting the window range based on the distribution rules of risk blocks, and repeating the traversal and detection steps until all risk blocks are filtered out.

Benefits of technology

Accurately screening out risk blocks in storage media improves screening efficiency, reduces the impact of potential risk blocks on the performance of storage media, and improves the reliability and performance of storage media.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119724306B_ABST
    Figure CN119724306B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of memory, and discloses a method, device, storage device and storage medium for screening risk blocks. The risk block screening method of the present application marks the storage blocks that meet the aging test failure conditions in the storage medium as problem blocks. According to the set window range, with the current problem block as the center, it traverses the storage blocks within the window range to detect whether there are risk blocks within the window range. When there is at least one risk block, the window range is adjusted, and for the newly added storage blocks within the adjusted window range, the steps of traversal and risk block detection are repeatedly executed until all the risk blocks in the storage medium are screened out. The risk block screening method of the present application can accurately screen out potential risk blocks, improve the screening efficiency, reduce the impact of potential risk blocks on the performance of the storage medium, and significantly improve the reliability and performance of the storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of memory, and in particular, to a method, device, storage device, and storage medium for screening risk blocks. Background Art

[0002] With the continuous development of storage technology, flash memory-based storage products such as eMMC (Embedded Multimedia Card) and UFS (Universal Flash Storage) are widely used in fields such as smartphones, tablets, and automotive electronics. These storage products usually use flash memory as the storage medium. To ensure product quality, current mainstream storage products perform aging tests and good product screening on the production line during the production process, specifically by performing read, write, and erase operations on the memory to detect the reliability of the product. In these tests, the blocks or pages of the storage medium are the core test objects. For blocks or pages that fail during read, write, and erase operations, they are usually marked on the production line for subsequent further shielding or replacement.

[0003] However, although the current mainstream good product screening process can mark the blocks that fail during read, write, and erase operations, it lacks in-depth analysis of the characteristics of the flash memory medium. For example, the current detection process usually only marks the problem blocks without considering the impact brought by the characteristics of the flash memory medium, and does not further analyze and explore the neighboring storage blocks of the problem area.

[0004] Therefore, how to more effectively explore and analyze the blocks that fail in read, write, and erase operations and their neighboring areas based on the characteristics of the medium, and improve the efficiency of screening out potentially weak blocks with hidden dangers, has become an important problem that needs to be solved urgently at present. Summary of the Invention

[0005] In view of this, the embodiments of this application provide a method, device, storage device, and storage medium for screening risk blocks, which can effectively solve the problems of insufficient analysis of medium characteristics and inability to efficiently screen out potentially risky weak blocks in the prior art.

[0006] In a first aspect, the embodiments of this application provide a method for screening risk blocks, including:

[0007] Obtain the storage blocks in the storage medium that meet the aging test failure conditions and mark them as problem blocks;

[0008] According to the set window range, with the current problem block as the window center, traverse the storage blocks within the window range to detect whether there are risk blocks within the window range based on a preset judgment condition;

[0009] When there is at least one of the risk blocks, adjust the current window range, and repeat the steps of traversing and detecting risk blocks for the newly added storage blocks within the adjusted window range until all the risk blocks adjacent to each of the problem blocks in the storage medium are screened out.

[0010] In some embodiments, obtaining the storage blocks in the storage medium that meet the aging test failure condition and marking them as problem blocks includes:

[0011] Performing a read operation, a write operation, or an erase operation on the storage block;

[0012] When the storage block fails during any one of the read operation, the write operation, or the erase operation, or when it is detected that the error correction code value of the storage block exceeds a preset threshold, mark the storage block as a problem block.

[0013] In some embodiments, the setting of the window range includes:

[0014] Taking the position of the current problem block as the center of the window, setting the initial size of the window range based on preset parameters; the window range includes the problem block and several neighboring storage blocks close to the problem block.

[0015] In some embodiments, traversing the storage blocks within the window range to detect whether there are risk blocks within the window range based on a preset judgment condition includes:

[0016] Starting from the starting position of the window range, traverse all the neighboring storage blocks within the window range except the current problem block;

[0017] For the traversed neighboring storage blocks, read the error correction code values of the neighboring storage blocks;

[0018] When the error correction code value exceeds a preset threshold, mark the neighboring storage block as the risk block.

[0019] In some embodiments, after detecting whether there are risk blocks within the window range based on a preset judgment condition, it further includes:

[0020] Dynamically adjusting the size and direction of the window range based on the distribution law of the risk blocks in the storage medium.

[0021] In some embodiments, dynamically adjusting the size and direction of the window range based on the distribution law of the risk blocks in the storage medium further includes:

[0022] Statistical average number of times the neighboring storage blocks near the problem block are marked as risk blocks;

[0023] Set the size and expansion direction of the window range according to the statistically obtained average number of times.

[0024] In some embodiments, for the newly added problem blocks within the adjusted window range, repeat the steps of traversing and detecting risk blocks until all the risk blocks adjacent to each of the problem blocks in the storage medium are screened out, including:

[0025] Perform the read operation, write operation, erase operation or read the error correction code value on the newly added storage block within the adjusted window range to determine the attribute of the newly added storage block;

[0026] When the newly added storage block is marked as a newly added problem block, with the current newly added problem block as the center of the window, adjust the window range, traverse the newly added storage blocks within the adjusted window range, and detect whether there are risk blocks within the adjusted window range based on the preset judgment conditions;

[0027] Repeat the steps of traversing and detecting risk blocks until all the risk blocks adjacent to each of the problem blocks in the storage medium are screened out.

[0028] In a second aspect, an embodiment of the present application provides a risk block screening device, including:

[0029] A storage block detection module, configured to obtain the storage blocks in the storage medium that meet the aging test failure condition and mark them as problem blocks;

[0030] A risk block judgment module, configured to traverse the problem blocks within the window range with the current problem block as the center of the window according to the set window range, and detect whether there are risk blocks within the window range based on preset judgment conditions;

[0031] A risk block screening module, configured to, when there is at least one risk block, adjust the current window range, and for the newly added problem blocks within the adjusted window range, repeat the steps of traversing and detecting risk blocks until all the risk blocks adjacent to each of the problem blocks in the storage medium are screened out.

[0032] In a third aspect, an embodiment of the present application provides a storage device, the storage device includes a processor and a memory, the memory stores a computer program, and the processor is configured to execute the computer program to implement the risk block screening method in the first aspect above.

[0033] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, when the computer program is executed on a processor, implementing the risk block screening method in the first aspect above.

[0034] The embodiments of the present application have the following beneficial effects:

[0035] In the risk block screening method of the present application, storage blocks that meet the aging test failure conditions in the storage medium are marked as problem blocks. According to the set window range, starting from the current problem block, the storage blocks within the window range are traversed to determine their attributes, and all potential risk blocks within the window range are screened out. By combining the means of adjusting the window range, the detection area can be flexibly expanded according to the distribution law of risk blocks, improving the coverage rate of risk blocks. For the newly added storage blocks within the adjusted window range, the steps of traversal and risk block detection are repeated until all risk blocks in the storage medium are screened out, ensuring a comprehensive screening of all potential risk blocks in the storage medium and avoiding missed detections. Through the strategy of adjusting the window range and iterative loop, the present application can quickly and accurately identify the blocks with potential hazards in the storage medium, significantly improving the risk block screening efficiency. At the same time, through in-depth detection of problem blocks and their neighboring areas, the existence of risk blocks in the storage medium is reduced, thereby effectively reducing the failure rate of the storage medium in subsequent use and improving the overall performance and reliability of the storage medium. The risk block screening method of the present application can not only improve the accuracy and efficiency of risk block screening, but also provide an efficient and reliable technical means for the quality control and performance optimization of the storage medium. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] To more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0037] Figure 1 Shows a flowchart of a risk block screening method according to an embodiment of the present application;

[0038] Figure 2 Shows a schematic diagram of risk judgment in a risk block screening method according to an embodiment of the present application;

[0039] Figure 3 Shows a schematic diagram of setting the window range in a risk block screening method according to an embodiment of the present application;

[0040] Figure 4 Shows a schematic diagram of expanding the window range in a risk block screening method according to an embodiment of the present application;

[0041] Figure 5 Shows another schematic diagram of expanding the window range in a risk block screening method according to an embodiment of the present application;

[0042] Figure 6The structural schematic diagram of a risk block screening device according to an embodiment of the present application is shown. Detailed implementation manners

[0043] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments.

[0044] Generally, the components of the embodiments of the present application described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the present application claimed, but merely represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.

[0045] Hereinafter, the terms "including", "having" and their cognates that can be used in various embodiments of the present application are only intended to represent specific features, numbers, steps, operations, elements, components or combinations of the foregoing items, and should not be construed as first excluding the existence of one or more other features, numbers, steps, operations, elements, components or combinations of the foregoing items or increasing the possibility of one or more features, numbers, steps, operations, elements, components or combinations of the foregoing items. In addition, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0046] Unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as commonly understood by those of ordinary skill in the art to which the various embodiments of the present application belong. The terms (such as those defined in a general-use dictionary) will be interpreted as having the same meaning as the contextual meaning in the relevant technical field and will not be interpreted as having an idealized meaning or being overly formal, unless clearly defined in the various embodiments of the present application.

[0047] Next, some implementation manners of the present application will be described in detail with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0048] Based on the analysis of a large amount of actual production line data, it is shown that in the storage medium, there are often potential hazards in the areas near the blocks or pages where read / write / erase failures occur. These hazards are mainly manifested as significantly higher ECC (Error Correction Code) values, indicating that these neighboring blocks may have risks of unstable use, and these hazard blocks are more likely to evolve into bad blocks during subsequent use. Especially near the weak blocks (blocks with ECC values exceeding the set threshold but not completely failed), this phenomenon is particularly obvious. Considering that the existing screening processes in the prior art cannot effectively identify these potential problem blocks, it may lead to unstable storage blocks flowing into subsequent usage links, thereby affecting the overall reliability of the storage medium. Therefore, a method for screening risk blocks is proposed. By detecting whether the storage blocks meet the aging test conditions, the storage blocks that meet the aging test conditions are marked as problem blocks; taking the problem blocks as the window centers, based on the set window range, traverse the problem blocks within the window range, judge their attributes and screen out all risk blocks; at the same time, combined with the distribution law of the risk blocks, adjust the window range, and repeat the steps of traversing and detecting risk blocks for newly added storage blocks until all the risk blocks in the storage medium are screened out. The method of this application can accurately screen out risk blocks, improve the screening efficiency, reduce the impact of risk blocks on the performance of the storage medium, and significantly enhance the reliability and performance of the storage medium.

[0049] Figure 1 FIG. shows a flowchart of a method for screening risk blocks according to an embodiment of the present application. Exemplarily, the method for screening risk blocks includes the following steps:

[0050] Step S100, obtain the storage blocks in the storage medium that meet the aging test failure conditions and mark them as problem blocks.

[0051] Exemplarily, during the production process of eMMC (Embedded Multimedia Card), by simulating the workload of the storage medium in actual use, the reliability and potential problems of the storage medium are detected. That is, during the aging test process, each storage block in the storage medium is detected, and the detection methods can include read operations, write operations, erase operations, etc., to obtain the operating status and related parameters of the storage blocks. During the detection process, according to the pre-set aging test failure conditions (for example, excessive number of read / write / erase operation failures, error correction code (ECC) value exceeding a certain threshold, etc.), the storage blocks are detected to determine whether they meet these conditions. If the storage blocks in the storage medium meet the aging test failure conditions, the storage blocks are marked as "problem blocks (also known as bad blocks)".

[0052] By classifying and marking the abnormal conditions of the storage blocks through the production line aging test, the storage blocks in the storage medium that may cause data errors or equipment failures can be effectively identified and processed in advance, thereby improving the reliability and data security of the storage medium.

[0053] For example, in an eMMC storage medium, when it is detected that the error correction code (ECC) value of a certain storage block exceeds the set safety threshold during a read operation, the storage block is marked as a "problem block (also known as a bad block)".

[0054] In an alternative embodiment, in step S100, obtaining the storage blocks in the storage medium that meet the aging test failure condition and marking them as problem blocks includes:

[0055] Step S110, performing a read operation, a write operation, or an erase operation on the storage block.

[0056] Exemplarily, a read operation: is a process of reading the stored data from a certain storage block of the storage medium. By sending a read instruction, data is obtained from the specified storage block for user access, running, or other tasks. A write operation: refers to a process of writing new data into a certain storage block of the storage medium. The write operation will overwrite the original data to ensure that the storage medium stores the latest information. For the success or failure of the write operation, the integrity of the data written is usually confirmed through a verification mechanism (e.g., error correction code ECC). An erase operation: The erase operation is to clear the data in the storage block so that it is in a state where it can be written again. In NAND or eMMC, the storage block usually needs to be erased before writing data, which is part of the working mechanism of the storage medium.

[0057] In the detection or management of the storage medium, these operations can also be used to check the health status of the storage block. Among them, the read operation ensures data access, the write operation realizes data storage, and the erase operation prepares space for the subsequent use of the storage medium.

[0058] Step S120, when any one of the read operation, the write operation, or the erase operation on the storage block fails, or if it is detected that the error correction code value of the storage block exceeds the preset threshold, the storage block is marked as a problem block.

[0059] Exemplarily, a read operation failure means that the data in the storage block cannot be read correctly, which may be due to storage block damage or data error. A write operation failure means that the data cannot be written correctly into the storage block, which may be due to storage block wear or hardware defects. An erase operation failure means that the storage block cannot be cleared, which may be because the storage block has reached the end of its service life or is hardware damaged.

[0060] When a read operation is performed on a certain storage block, an attempt is made to read the stored data from that storage block, and at the same time, the error correction code (ECC) value of the data is detected. The ECC value is an important indicator reflecting the data error situation in the storage block. The higher the value, the more errors occur in that storage block. Among them, the threshold of the ECC value is preset as a criterion for judging the health status of the storage block. If the ECC value is lower than the threshold, it means that the storage block is still in a healthy or repairable range; if the ECC value exceeds the threshold, it indicates that the error rate of the storage block is relatively high and there is a potential failure risk. When it is detected that the ECC value of a certain storage block exceeds the preset threshold, that storage block is marked as a "problem block". By detecting the ECC value, potential problems of the storage block can be identified in advance before serious data errors or losses occur. This method improves the reliability of the storage medium and data security, and reduces the risk of device failures caused by problem blocks.

[0061] For example, in a NAND storage medium, when attempting to read the data of a certain storage block, it is detected that the ECC value of this block is 50, while the safety threshold is 40. Since the ECC value exceeds the threshold, this storage block is marked as a "problem block".

[0062] That is, as Figure 2 shown, when a storage block fails during any of the operations of read operation, write operation, erase operation, or when the error correction code value exceeds the preset threshold, the storage block is marked as a "problem block" to trigger window detection in subsequent steps to screen out all risk blocks.

[0063] Step S200, according to the set window range, with the current problem block as the window center, traverse the storage blocks within the window range to detect whether there are risk blocks within the window range based on the preset judgment conditions.

[0064] Among them, the window range is used to define the range of storage blocks to be detected. A risk block refers to a block when the ECC value exceeds a certain set threshold, and is sometimes also called a weak block.

[0065] Exemplarily, with the current problem block as the center, according to the set window range, determine the area that needs to be further detected. After determining the area, traverse all the storage blocks within the window range one by one, and analyze and judge the attributes of each storage block. According to the preset conditions (such as the ECC value exceeding the threshold, etc.), judge whether each problem block has potential problems. Mark all the problem blocks within the window range that meet the judgment conditions as risk blocks, so as to realize the screening and comprehensive identification of risk blocks.

[0066] It can be understood that through the window screening mechanism centered on the problem block, the detection range can be expanded to discover risk blocks that may have regional consistency problems, thereby avoiding missing potential risk blocks and further improving the reliability of the storage medium.

[0067] For example, in an eMMC, it is detected that abnormal operations (such as read failure, write failure, erase failure) occur in storage block N. Storage block N is marked as an Error Block (problem block). Centered on the problem block N, a window range is set to further detect whether there are potential risk blocks, i.e., weak blocks, among the storage blocks around the problem block. In this example, as Figure 3 shown, the window size is set to 5 to cover the storage block range from N - 2 to N + 2. It should be noted that since block N has been marked as a problem block, during the traversal within the window range, block N is skipped and does not need to be detected again. Starting from the leftmost block N - 2 of the window, the risk detection is performed on the storage blocks within the window range in sequence, and the ECC value is read: if the ECC value does not exceed the preset threshold, the next block is continued to be detected; if the ECC value exceeds the threshold, the block is marked as a risk weak block, and the window range is dynamically adjusted to further expand the detection (for example, extended to N - 3 or N + 3), and the potential risk weak blocks that may exist after expanding the window are continued to be screened. The termination condition of the detection is that all storage blocks (except the already marked problem blocks) within the window range meet the health condition, the ECC value does not exceed the preset threshold, or the window range no longer needs to be expanded.

[0068] In an alternative embodiment, in step S200, the setting of the window range includes: taking the position of the current problem block as the center of the window, and setting the initial size of the window range based on preset parameters; the window range includes the problem block and several neighboring storage blocks close to the problem block.

[0069] Exemplarily, the problem block is a storage block in the storage medium that meets the failure condition of the aging test. Taking the position of the current problem block as the center of the window, which is the starting point of the detection range, an initial window range is set according to preset parameters (such as window size, expansion direction, detection strategy, etc.) to limit the area of storage blocks to be detected. Among them, the size of the initial window range can be a fixed value (for example, covering several blocks on both sides of the problem block) or dynamically adjusted (according to the storage block distribution or device characteristics). The window range not only includes the currently marked problem block but also several adjacent storage blocks close to the problem block. So as to traverse all the neighboring storage blocks within the window range in subsequent steps to determine whether there are potential problems with these neighboring storage blocks - whether they are risk weak blocks. By setting the initial window range centered on the problem block, the storage blocks with a high correlation with the problem block can be quickly traversed, unnecessary full - disk scans can be reduced, and the detection efficiency can be improved.

[0070] For example, in an eMMC, Block N is detected as a problem block. Taking Block N as the center, if the preset parameter is 2, the determined window range includes Block N and two neighboring storage blocks adjacent to the left and right ends respectively, that is, the area covering from Block N - 2 to Block N + 2, and the initial size of the window range is 5 storage blocks in total.

[0071] In an alternative embodiment, in step S200, when traversing the storage blocks within the window range to detect whether there are risk blocks based on a preset judgment condition, it includes:

[0072] Step S210, starting from the starting position of the window range, traverse all neighboring storage blocks within the window range except the current problem block.

[0073] Exemplarily, starting from the starting position (usually the leftmost storage block in the window range), the detection is carried out sequentially in order. This sequential traversal method ensures that no storage block within the window range is missed. For each neighboring storage block within the window range, the error correction code value ECC is read one by one to determine whether there are potential risk problems in this storage block, so as to mark and take measures in time. At the same time, this method of traversing one by one ensures the comprehensiveness and accuracy of the detection.

[0074] Step S220, for the neighboring storage blocks traversed, read the error correction code value of the neighboring storage blocks.

[0075] Exemplarily, the error correction code (ECC) is a technology used to detect and repair data errors in storage blocks. The ECC value reflects the health status of the storage block: a lower ECC value usually indicates higher data reliability; a higher ECC value indicates that there are more data errors or potential risks in this storage block. During the entire traversal process, the reading of the ECC value is completed block by block to ensure that every neighboring storage block within the window range is detected and no possible potential problem block is missed.

[0076] Step S230, when the error correction code value exceeds the preset threshold, mark the neighboring storage block as a risk block, that is, a weak block.

[0077] For example, assume that the set window range is from Block N-2 to Block N+2, and the starting position of the window is Block N-2. Starting from Block N-2, successively detect Block N-2, Block N-1, Block N, Block N+1, and Block N+2. Each time, detect one storage block, read its attributes (e.g., the error correction code ECC value), and determine whether it needs to be marked as a weakly risky block. If the ECC value of a certain block is 40 and the preset safety threshold is 50, then this block is considered healthy; if the ECC value of a certain block is 55, which exceeds the preset safety threshold, this storage block will be marked as a weakly risky block.

[0078] In an alternative implementation, after step S210, it further includes:

[0079] Based on the distribution law of risky blocks in the storage medium, dynamically adjust the size and direction of the window range.

[0080] Among them, the distribution law refers to the distribution characteristics of risky blocks in the storage area recorded during the production or operation of the storage medium. These laws are obtained by statistically analyzing a large amount of data. For example, whether risky blocks are concentrated in certain specific areas, whether there is a trend that adjacent storage blocks are affected, the average number of risky blocks near problem blocks, etc.

[0081] Exemplarily, according to the above distribution law, the window range can be dynamically adjusted, that is, the size and direction of the detection area. For example: If the data shows that multiple consecutive risky blocks usually appear near a problem block, the window range needs to be expanded to cover these adjacent areas. If the distribution of problem blocks is relatively scattered, the window range can be adjusted more flexibly in terms of direction and size.

[0082] It can be understood that by reasonably expanding the window range, it is ensured to cover all areas where risky blocks may exist. By adjusting the size of the window range according to the distribution law, the ineffective detection of irrelevant areas is reduced, saving resources. That is, the dynamic adjustment of the window range depends on the statistical analysis of data.

[0083] For example, in the production line data of the storage medium, through analysis, it is found that: usually 2 weakly risky blocks are continuously distributed around a problem block. According to this law, when a problem block is detected, the window range will be set to cover the storage blocks on both sides of the problem block (for example, expand 2 storage blocks to the left with the problem block as the center). If a new weakly risky block is further detected, the window range will continue to be adjusted according to the actual detection results (for example, expand 2 storage blocks to the left again).

[0084] By combining the distribution law of risky blocks and the production line statistical data, determine the size and direction of the window range to achieve a comprehensive screening and marking of potential weakly risky blocks in the storage medium.

[0085] For example, as Figure 4 shown, a problem block (Error Block, Block N) is detected in the storage medium, and a window range is set centered on this. The initial size of the window is set to 5 (i.e., covering Block N - 2 to Block N + 2). Within this range, the error correction code (ECC) values of each storage block are sequentially read to determine whether there are weak blocks. First, starting from the left end of the window range (Block N - 2), Block N - 2, N - 1, N + 1, and N + 2 are sequentially detected (it should be noted that Block N is skipped from detection because it has already been marked as a problem block). When Block N - 1 is detected as a weak block, according to the characteristic of medium area consistency (i.e., the probability of adjacent blocks being weak blocks is relatively high), it is necessary to expand the detection area to the left of the window range, expanding the window range from the original size N = 5 to N = 6, covering from Block N - 3 to Block N + 2. Continue to detect the ECC value of the newly added left - end block (Block N - 3). If Block N - 3 is also determined to be a weak block, the window range is expanded again. The expansion detection process continues until the newly added storage block on the left side of the window range (e.g., Block N - 4) is no longer determined to be a weak block, at which point the expansion stops. At this time, the final size of the window range is determined to be N = 7 (covering from Block N - 3 to Block N + 3), and the screening is completed, as Figure 5 shown.

[0086] That is, by dynamically adjusting the size and direction of the window range, potential high - risk storage blocks in the neighborhood can be covered, avoiding the accumulation of data errors caused by read fail, write fail, Erase fail, and Weak block to surrounding storage blocks, thereby reducing the risk of uncorrectable (UNC).

[0087] In an alternative embodiment, the average number of times that the neighborhood storage blocks near the problem block are marked as risk blocks is statistically counted.

[0088] Exemplarily, by statistically counting the number of times that the neighborhood storage blocks near the problem block are marked as risk blocks, the distribution law of risk blocks in the storage medium and the health status of adjacent areas can be understood. The average number refers to the frequency of neighborhood storage blocks being marked as problem blocks after statistically counting multiple risk - block data. For example: If it is found during the statistical process that on average, 2 storage blocks in the neighborhood range of the problem block are also marked as risk blocks, this value can be used as a basis for subsequent analysis and detection. By calculating this average value, it can help the system more scientifically predict the distribution of potential problem blocks and optimize the detection range.

[0089] For example, assume that in a storage medium, the data of 100 problem blocks are counted, and it is found that the number of times the neighborhood storage blocks of these problem blocks are marked as risk blocks is as follows:

[0090] There are 2 storage blocks in the neighborhood of problem block 1 marked as risk blocks;

[0091] There are 4 storage blocks in the neighborhood of problem block 2 marked as risk blocks;

[0092] There are 3 storage blocks in the neighborhood of problem block 3 marked as risk blocks;

[0093] By counting the data of all problem blocks, the average number of times the neighborhood storage blocks of problem blocks are marked as risk blocks is calculated to be 3.

[0094] According to the average number obtained from the statistics, set the size and expansion direction of the window range.

[0095] Exemplarily, by counting the average number of times the neighborhood storage blocks of problem blocks in the storage medium on the production line are marked as risk blocks, the distribution law of risk blocks can be quantified, so as to scientifically set the initial size of the detection window, enabling the window range to cover the high-risk areas around the problem blocks.

[0096] For example, if the statistical average number is N, the size of the window range can be set to 2N to ensure that the neighborhood storage blocks on both sides of the problem block are covered, so as to comprehensively screen potential weak blocks. Further, according to the distribution law, the window expansion direction can be dynamically adjusted: one-way expansion (e.g., left, right), two-way expansion, etc.

[0097] For example, assume that through statistical analysis on the production line, it is found that the average number of times the neighborhood storage blocks of problem blocks are marked as risk blocks is 3 (N = 3): initially set the window range to 2N, that is, covering Block N - 3 to Block N + 3, a total of 7 storage blocks. If it is detected that Block N - 4 on the left side of the window range is also marked as a risk block, then according to the statistical law, the window range is expanded to the left to Block N - 4. If it is detected that Block N + 4 on the right side is also a risk block, then the window range is expanded to the right at the same time.

[0098] Step S300, when there is at least one risk block, adjust the current window range, and repeat the steps of traversing and risk block detection for the new storage blocks within the adjusted window range until all risk blocks adjacent to each problem block in the storage medium are screened out.

[0099] Exemplarily, after detecting at least one risk block, the window range is dynamically adjusted according to the distribution characteristics to cover more blocks. In the adjusted window range, each newly added storage block is detected one by one, its relevant attributes (e.g., the error correction code ECC value) are read, and it is determined whether these storage blocks meet the determination conditions of the risk blocks. If there are risk blocks among the newly added storage blocks, the window range is further adjusted and the steps of traversing and detecting risk blocks are repeated. This operation is continuously iterated to ensure that all newly added storage blocks can be detected after each adjustment of the window range. When all storage blocks in the adjusted window range have been detected and there are no new problem blocks, the detection stops. At this time, it can be confirmed that all potential risk blocks in the storage medium have been marked and the detection process ends.

[0100] In an alternative embodiment, in step S300, for the newly added storage blocks within the adjusted window range, the steps of traversing and detecting risk blocks are repeatedly executed until all risk blocks adjacent to each problem block in the storage medium are screened out, including:

[0101] Step S310, within the adjusted window range, perform a read operation, a write operation, an erase operation, or read the error correction code value of the newly added storage block to determine the attributes of the newly added storage block.

[0102] Exemplarily, the newly added storage block refers to the storage block newly included in the detection within the window range after adjusting the window range. The following operations are performed on the newly added storage block to determine whether its attributes are normal: Read operation: Read the data in the storage block and check whether the data can be correctly read. Write operation: Try to write data to the storage block and verify whether the write operation is successful and correct. Erase operation: Try to clear the data in the storage block to ensure that the erase operation can be successfully completed. Read the error correction code (ECC) value: When performing the read operation, obtain the error correction code (ECC) value of the storage block. The ECC value is an important indicator to measure the health status of the storage block. If the ECC value exceeds a certain preset threshold, it may indicate that the storage block is abnormal.

[0103] Step S320, when the newly added storage block is marked as a newly added problem block, take the storage location of the current newly added problem block as the center of the window, adjust the window range, and traverse the newly added storage blocks within the adjusted window range to detect whether there are risk blocks within the adjusted window range based on a preset judgment condition.

[0104] Exemplarily, the appearance of newly added problem blocks indicates that there may be more potential risk blocks within the current detection range. After detecting newly added problem blocks, the window range will be readjusted with the storage location of the problem block as the center. This means that the detection focus of the window range will shift from the previous center position (e.g., the previous problem block) to the position of the currently newly added problem block. According to the appearance location of the newly added storage block, the size and direction of the window range are dynamically adjusted, which may specifically include: Expanding the window range: increasing the number of storage blocks within the window range, for example, expanding from Block N-1 to Block N+2 to Block N-2 to Block N+3. Changing the detection direction: if problem blocks are concentrated on one side (e.g., the left side, the right side, or both sides), the window range can be preferentially expanded in that direction. The purpose of dynamic adjustment is to ensure that the window covers the storage blocks in the neighborhood of the newly added problem blocks, thereby detecting more potential risk blocks.

[0105] Step S330, repeat the steps of traversing and risk block detection until all risk blocks adjacent to each problem block in the storage medium are screened out.

[0106] Exemplarily, for the newly added storage blocks within the window range, read their error correction code values (ECC) to determine the status of the newly added storage blocks. If the newly added storage blocks meet the determination conditions of risk blocks (e.g., read failure, ECC value exceeding the standard), then mark them as risk blocks. Then continue to dynamically adjust the window range, expanding the detection area to cover more possible potential risk blocks. Among them, the adjustment range may include increasing the size of the detection window or changing the detection direction (e.g., expanding to the left, expanding to the right, or performing bilateral expansion). After each adjustment of the window range, repeat the above steps of "traversing" and "risk block detection" for the storage blocks newly included in the window range.

[0107] It should be noted that this process is a recursive loop, and it is necessary to continuously detect and expand the window to ensure that all potential risk blocks are covered. When no more problem blocks and risk blocks are found after judging the newly added storage blocks within the window range, and no more problem blocks and risk blocks are found throughout the storage medium, the loop operation ends, and the window size for this storage medium is determined.

[0108] The risk block screening method according to the embodiments of the present application marks the storage blocks that meet the aging test failure conditions as problem blocks. According to the set window range, with the problem blocks as the center, it determines whether there are risk blocks, combines the distribution rules of the risk blocks, gradually covers the potential risk areas, and screens out all potential risk blocks within the window range, realizing the efficient screening of risk blocks in the storage medium. The present application can optimize the window range during the detection process to ensure the comprehensiveness of the screening, while reducing the ineffective detection of normal areas. By iteratively judging the attributes of the storage blocks and adjusting the range multiple times, it can effectively identify and mark the problem blocks and risk blocks, thereby improving the stability and reliability of the storage medium and reducing the risk of storage failures caused by risk blocks.

[0109] Figure 6 FIG. shows a schematic structural diagram of a risk block screening device according to an embodiment of the present application. Exemplarily, the device includes:

[0110] A storage block detection module 61, configured to obtain the storage blocks in the storage medium that meet the aging test failure conditions and mark them as problem blocks;

[0111] A risk block judgment module 62, configured to traverse the problem blocks within the window range with the current problem block as the window center according to the set window range, and detect whether there are risk blocks within the window range based on a preset judgment condition;

[0112] A risk block screening module 63, configured to, when there is at least one of the risk blocks, adjust the current window range, and repeat the steps of traversing and risk block detection for the newly added problem blocks within the adjusted window range until all risk blocks adjacent to each of the problem blocks in the storage medium are screened out.

[0113] It can be understood that the device in this embodiment corresponds to the risk block screening method in the above embodiment, and the optional items in the above embodiment are also applicable to this embodiment, so they will not be repeated here.

[0114] The present application also provides a storage device. Exemplarily, the storage device includes a processor and a memory. Among them, the memory stores a computer program, and the processor runs the computer program to enable the storage device to execute the functions of the above risk block screening method or each module in the above risk block screening device.

[0115] Among them, the processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, including at least one of a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc., and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application.

[0116] The memory can be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc. Among them, the memory is used to store a computer program, and after receiving an execution instruction, the processor can execute the computer program accordingly.

[0117] The present application also provides a computer-readable storage medium for storing the computer program used in the above storage device. For example, the computer-readable storage medium can include, but is not limited to: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.

[0118] In several embodiments provided by this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and structure diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in an alternative implementation, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the structure diagram and / or flowchart, as well as the combination of blocks in the structure diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0119] In addition, in each embodiment of this application, the various functional modules or units can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0120] If the described functions are implemented in the form of software functional modules and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a storage device (which can be a smart phone, a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application.

[0121] As described above, the above are only the specific implementation manners of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application.

Claims

1. A risk block screening method, characterized in that, The method includes: Obtaining storage blocks in a storage medium that meet the aging test failure condition and marking them as problem blocks; classifying and marking abnormal conditions during the operation of the storage blocks through the aging test to identify the problem blocks; the obtaining of storage blocks in the storage medium that meet the aging test failure condition and marking them as problem blocks includes: Performing a read operation, a write operation, or an erase operation on the storage blocks; When the storage block fails during any one of the read operation, the write operation, or the erase operation, or if it is detected that the error correction code value of the storage block exceeds a preset threshold, marking the storage block as a problem block; According to a set window range, with the current problem block as the window center, traversing the storage blocks within the window range to detect whether there are risk blocks within the window range based on a preset judgment condition, including: Starting from the starting position of the window range, traversing all neighboring storage blocks within the window range except the current problem block; For the traversed neighboring storage blocks, reading the error correction code value of the neighboring storage blocks; When the error correction code value exceeds the preset threshold, marking the neighboring storage block as the risk block; The setting of the window range includes: taking the position of the current problem block as the window center and setting the initial size of the window range based on preset parameters; the window range includes the problem block and several neighboring storage blocks close to the problem block; Dynamically adjusting the size and direction of the window range based on the distribution law of the risk blocks in the storage medium; the distribution law refers to the distribution characteristics of the risk blocks in the storage area recorded during the production or operation of the storage medium; When there is at least one risk block, dynamically adjusting the current window range according to the distribution characteristics, and repeating the steps of traversing and detecting risk blocks for the newly added storage blocks within the adjusted window range until all risk blocks adjacent to each problem block in the storage medium are screened out.

2. The risk block screening method according to claim 1, wherein The dynamically adjusting the size and direction of the window range based on the distribution law of the risk blocks in the storage medium further includes: Counting the average number of times that the neighboring storage blocks near the problem block are marked as risk blocks; Setting the size and expansion direction of the window range according to the statistically obtained average number of times.

3. The risk block screening method according to claim 1, wherein The repeating the steps of traversing and detecting risk blocks for the newly added problem blocks within the adjusted window range until all risk blocks adjacent to each problem block in the storage medium are screened out includes: Performing a read operation, a write operation, an erase operation, or reading the error correction code value on the newly added storage blocks within the adjusted window range to determine the attributes of the newly added storage blocks; When the newly added storage block is marked as a newly added problem block, taking the current newly added problem block as the window center, adjusting the window range, and traversing the newly added storage blocks within the adjusted window range to detect whether there are risk blocks within the adjusted window range based on the preset judgment condition; Repeat the steps of the traversal and risk block detection until all the risk blocks adjacent to each of the problem blocks in the storage medium are screened out.

4. A risk block screening device, characterized in that, The device includes: A storage block detection module, configured to obtain storage blocks in the storage medium that meet the aging test failure condition and mark them as problem blocks; classify and mark abnormal conditions during the operation of the storage blocks through the aging test to identify the problem blocks; the obtaining of the storage blocks in the storage medium that meet the aging test failure condition and marking them as problem blocks includes: Performing a read operation, a write operation, or an erase operation on the storage block; When the storage block fails during any one of the read operation, the write operation, or the erase operation, or when it is detected that the error correction code value of the storage block exceeds a preset threshold, mark the storage block as a problem block; A risk block judgment module, configured to traverse the problem blocks within the window range with the current problem block as the window center according to a set window range, so as to detect whether there are risk blocks within the window range based on a preset judgment condition, including: Starting from the starting position of the window range, traverse all the neighborhood storage blocks within the window range except the current problem block; For the traversed neighborhood storage blocks, read the error correction code values of the neighborhood storage blocks; When the error correction code value exceeds the preset threshold, mark the neighborhood storage block as the risk block; The setting of the window range includes: Taking the position of the current problem block as the window center, setting the initial size of the window range based on preset parameters; the window range includes the problem block and several neighborhood storage blocks close to the problem block; Dynamically adjusting the size and direction of the window range based on the distribution law of the risk blocks in the storage medium; the distribution law refers to the distribution characteristics of the risk blocks in the storage area recorded during the production or operation of the storage medium; A risk block screening module, configured to, when there is at least one risk block, dynamically adjust the current window range according to the distribution characteristics, and repeat the steps of traversal and risk block detection for the newly added problem blocks within the adjusted window range until all the risk blocks adjacent to each of the problem blocks in the storage medium are screened out.

5. A storage device, characterized in that, The storage device includes a processor and a memory, the memory stores a computer program, and the processor is configured to execute the computer program to implement the risk block screening method according to any one of claims 1-3.

6. A computer-readable storage medium, characterized in that, It stores a computer program, and when the computer program is executed on a processor, it implements the risk block screening method according to any one of claims 1-3.

Citation Information

Patent Citations

  • Firmware-based SSD block failure prediction and avoidance scheme

    CN112711492A

  • Weak block screening method of storage device

    CN115116527A