Storage equipment repairing method and device, equipment and storage medium
By dividing functional areas and replacement areas in the storage device, accurately positioning and replacing faulty storage blocks, the problems of waste and low reliability of storage devices in the prior art are solved, and more efficient storage resource utilization and equipment reliability are achieved.
Patent Information
- Application Number
- CN202510914719.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-08-01
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, when dealing with storage devices failing storage units, the overall replacement method is usually adopted, resulting in waste of normal storage units and reduced resource utilization, making it difficult to meet reliability requirements.
By dividing functional areas and replacement areas in the storage device, accurately locate the faulty storage blocks and select alternate storage blocks from the replacement areas for replacement, replace the faulty storage blocks with replacement blocks, and update the logical address map to ensure data integrity and system availability.
It improves the resource utilization and reliability of storage devices, reduces waste of storage space, extends the service life of the device, and improves the availability and user experience of the system through a transparent address redirection mechanism.
Smart Images

Figure CN120407268A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data storage device management, and particularly relates to a storage device repair method, apparatus, device, and storage medium. Background Art
[0002] In the field of data storage, storage devices such as storage devices will inevitably have storage unit failure problems during use. Traditional technologies usually adopt a simple bad block management mechanism, and avoid data errors by marking faulty units and skipping their use. However, with the improvement of storage density and the complexity of storage structures, this extensive management method has been difficult to meet the reliability requirements.
[0003] In the process of implementing the present invention, it is found that the related technologies have at least the following problems: when dealing with faulty storage units, the related technologies often perform an overall replacement of the logical units of the storage device. This processing method not only causes a waste of a large number of normal storage units, but also significantly reduces the resource utilization rate of the storage system due to the overly extensive replacement granularity. Summary of the Invention
[0004] In view of the above problems, the present invention provides a storage device repair method, apparatus, device, and storage medium.
[0005] According to a first aspect of the present invention, there is provided a storage device repair method, including: in response to detecting a faulty storage block in a functional area of a storage device, obtaining a data plane of the faulty storage block in a target logical unit, wherein the storage device includes a plurality of logical units, each of the plurality of logical units includes a plurality of storage blocks and is divided into a functional area and a replacement area, the functional area is used for storing data, the replacement area includes spare storage blocks, and the target logical unit is the logical unit where the faulty storage block is located; according to the data plane, selecting a replacement storage block from the replacement area for replacing the faulty storage block, wherein the replacement area is used for replacing the faulty storage block, and the replacement area includes a plurality of logical units; and using the replacement storage block to replace the faulty storage block.
[0006] According to an embodiment of the present invention, selecting a replacement storage block from the replacement area for replacing the faulty storage block according to the data plane includes: in the replacement area, selecting a spare storage block in the data plane as the replacement storage block.
[0007] According to an embodiment of the present invention, selecting a spare storage block in the data plane as the replacement storage block in the replacement area includes: sequentially traversing the spare storage blocks in the data plane until an available storage block is determined, and determining the available storage block as the replacement storage block.
[0008] According to an embodiment of the present invention, replacing a faulty storage block with a replacement storage block includes: obtaining the original data stored in the faulty storage block and the logical address corresponding to the faulty storage block; storing the original data in the replacement storage block; modifying the logical address to point to the replacement storage block so that subsequent data storage requests and data read requests initiated through the logical address can be allocated to the replacement storage block; and marking the faulty storage block as an invalid storage block.
[0009] According to an embodiment of the present invention, the storage device repair method further includes: in response to the power-on of the storage device, traversing the bad block table of the storage device, and determining that a faulty storage block exists in the functional area of the storage device when it is determined that there is a newly added faulty storage block in the area corresponding to the functional area in the bad block table; the newly added faulty storage block is recorded in the bad block table in the following manner: performing a fault detection on the storage device to obtain a fault detection result; and when the fault detection result indicates that a storage block in the storage device has a fault, marking the storage block in the bad block table.
[0010] According to an embodiment of the present invention, when the fault detection result indicates that a storage block in the storage device is in one of the following situations, the fault detection result indicates that a storage block in the storage device has a fault: writing data to the storage block and the storage block returns a programming error; reading data from the storage block and the storage block feedbacks a data read failure or reads incorrect data.
[0011] According to an embodiment of the present invention, the storage device repair method further includes: after completing the replacement of the faulty storage block, recording the replacement storage block and the faulty storage block in an update record table so that in the case of a newly added faulty storage block, the usage situation of multiple storage blocks in the replacement area can be determined based on the update record table.
[0012] A second aspect of the present invention provides a storage device repair apparatus, including: a position acquisition module, configured to, in response to detecting that a faulty storage block exists in the functional area of the storage device, acquire the data plane of the faulty storage block in the target logical unit, where the storage device includes a plurality of logical units, each of the plurality of logical units includes a plurality of storage blocks and is divided into a functional area and a replacement area, the functional area is used to store data, the replacement area includes spare storage blocks, and the target logical unit is the logical unit where the faulty storage block is located; a storage block selection module, configured to select a replacement storage block for replacing the faulty storage block from the replacement area according to the data plane, where the replacement area is used to replace the faulty storage block and the replacement area includes a plurality of logical units; and a storage block replacement module, configured to replace the faulty storage block with the replacement storage block.
[0013] A third aspect of the present invention provides an electronic device, including: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.
[0014] A fourth aspect of the present invention further provides a computer-readable storage medium, on which a computer program or instruction is stored, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented.
[0015] A fifth aspect of the present invention further provides a computer program product, including a computer program or instruction, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented. Description of the Drawings
[0016] Through the following description of the embodiments of the present invention with reference to the drawings, the above content and other objects, features and advantages of the present invention will become clearer.
[0017] Figure 1 Shows a structural diagram of a storage device to which a storage device repair method, apparatus, device and storage medium according to an embodiment of the present invention are applicable.
[0018] Figure 2 Shows a flowchart of a storage device repair method according to an embodiment of the present invention.
[0019] Figure 3 Shows a flowchart of a storage device repair method according to another embodiment of the present invention.
[0020] Figure 4A Shows the health status of the storage blocks of the storage device before repair by the storage device repair method according to an embodiment of the present invention.
[0021] Figure 4B Shows the original health status and corresponding relationship of the storage blocks of the storage device after repair by the storage device repair method according to an embodiment of the present invention.
[0022] Figure 5 Shows a structural block diagram of a storage device repair apparatus according to an embodiment of the present invention.
[0023] Figure 6 Shows a block diagram of an electronic device suitable for implementing a storage device repair method according to an embodiment of the present invention. Detailed Embodiments
[0024] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, numerous specific details are set forth in order to provide a comprehensive understanding of the embodiments of the present invention. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.
[0025] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "comprising", "including" and the like used herein indicate the presence of the described features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.
[0026] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification, and should not be interpreted in an idealized or overly rigid manner.
[0027] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0028] An embodiment of the present invention provides a method for repairing a storage device, including: in response to detecting a faulty storage block in the functional area of the storage device, obtaining the data plane of the faulty storage block in the target logical unit, where the storage device includes a plurality of logical units, each of the plurality of logical units includes a plurality of storage blocks and is divided into a functional area and a replacement area, the functional area is used to store data, the replacement area includes spare storage blocks, and the target logical unit is the logical unit where the faulty storage block is located; according to the data plane, selecting a replacement storage block from the replacement area for replacing the faulty storage block, where the replacement area is used to replace the faulty storage block, and the replacement area includes a plurality of logical units; and using the replacement storage block to replace the faulty storage block.
[0029] Figure 1 The structural diagram of the storage device to which the storage device repair method, device, equipment and storage medium according to the embodiments of the present invention are applicable is shown.
[0030] As Figure 1As shown, the storage device 100 includes a first logical unit 101 and a second logical unit 102. Each of the first logical unit 101 and the second logical unit 102 includes multiple data planes, and each data plane includes multiple storage blocks.
[0031] The logical units in a storage device can be managed by a Logical Unit Number (LUN), which is used to identify and manage different storage devices or different logical units of the same storage device in a storage system.
[0032] A data plane in a logical unit is the basic unit for parallel data processing, which can improve the efficiency of data transmission.
[0033] A data plane includes multiple storage blocks, at least one data page register, and at least one flash register. Among them, a storage block is the smallest erasure unit in a storage device, and each storage block includes at least one data page for storing data.
[0034] The data page register is used to store the address information of the data page currently being processed. When the storage device needs to read or write data, the data page register can be used to record the address of the data page to ensure the correct reading and writing of data.
[0035] The flash register is used to cache the content of the most recently accessed data page, thereby reducing the number of accesses to the main memory during the data reading process and improving the access speed.
[0036] It should be understood that Figure 1 the numbers of logical units, data planes, storage blocks, data page registers, and flash registers in [[ ]] are merely illustrative. According to implementation requirements, there can be any number of logical units, data planes, storage blocks, data page registers, and flash registers.
[0037] The following will be based on Figure 1 the described scenario, and through Figures 2 to 3 , Figure 4A , Figure 4B describe in detail the storage device repair method of the inventive embodiments.
[0038] Figure 2 shows a flowchart of the storage device repair method according to an embodiment of the present invention.
[0039] As Figure 2 shown, the storage device repair method of this embodiment includes operations S210 to S230.
[0040] In operation S210, in response to detecting a faulty storage block in the functional area of the storage device, obtain the data plane of the faulty storage block in the target logical unit.
[0041] According to an embodiment of the present invention, a storage device may include a storage device based on a semiconductor storage medium, such as a solid state drive (SSD). The storage device includes a plurality of logical units, each of the plurality of logical units includes a plurality of storage blocks and is divided into a functional area and a replacement area. The functional area is used to store data, the replacement area includes spare storage blocks, and the target logical unit is the logical unit where the faulty storage block is located.
[0042] According to an embodiment of the present invention, a faulty storage block may include a storage block that has been faulty since the storage device left the factory, or may also include a storage block that newly appears faulty during the use of the storage device.
[0043] In operation S220, according to the data plane, a replacement storage block for replacing the faulty storage block is selected from the replacement area.
[0044] According to an embodiment of the present invention, the replacement storage block may be selected from the data plane where the faulty storage block is located.
[0045] In operation S230, the faulty storage block is replaced with the replacement storage block.
[0046] According to an embodiment of the present invention, after the replacement is completed, the data of the faulty storage block that originally needed to be stored in the functional area can be stored in the replacement storage block. Similarly, when reading data from the functional area, the relevant data will be read from the replacement storage block and the read result will be returned, thereby ensuring the performance and availability of the storage device.
[0047] According to an embodiment of the present invention, by detecting the faulty storage block in the functional area and accurately locating its data plane position in the logical unit, and replacing it with a storage block in the same data plane, a finer-grained fault management is achieved. Using the spare storage blocks in the replacement area for targeted replacement not only improves the repair efficiency, but also optimizes the utilization rate of storage resources through the replacement strategy at the storage block level within the same data plane, ensuring the reliability of data storage, reducing the waste of storage space caused by the traditional overall replacement idea, thereby improving the reliability and service life of the storage device and the resource utilization rate of the storage device.
[0048] According to an embodiment of the present invention, selecting a replacement storage block for replacing the faulty storage block from the replacement area according to the data plane includes: in the replacement area, selecting a spare storage block in the data plane as the replacement storage block.
[0049] According to an embodiment of the present invention, in the faulty area, spare storage blocks are sequentially selected from the data plane where the faulty storage block is located, and the usage situation of the spare storage blocks is judged.
[0050] According to an embodiment of the present invention, the usage status of a spare storage block can be used to indicate whether the spare storage block can be used to replace a faulty storage block in the current state. For example, in the case where the spare storage block itself is also faulty, etc., the usage status can indicate that the spare storage block is unavailable.
[0051] According to an embodiment of the present invention, by determining the specific position of a faulty block in a logical unit, a spare storage block is matched from a replacement area. This can reduce the selection granularity of the spare storage block and the replacement storage block, thereby improving the storage resource utilization rate of the storage device.
[0052] According to an embodiment of the present invention, in the replacement area, selecting a spare storage block in the data plane as a replacement storage block includes: sequentially traversing the spare storage blocks in the data plane until an available storage block is determined, and determining the available storage block as the replacement storage block.
[0053] According to an embodiment of the present invention, sequentially traverse the storage blocks in the data plane in the replacement area. When the first available storage block is determined, stop traversing and determine the available storage block as the replacement storage block.
[0054] According to an embodiment of the present invention, during the traversal process, for the current spare storage block, determine whether it is faulty or has been used (for example, has been used as a replacement storage block in the historical behavior of storage device repair). In the case where it is determined that it is faulty or has been used, determine it as an unavailable storage block and proceed to the next spare storage block until the current spare storage block is not faulty and has not been used. Determine the current spare storage block as an available storage block and end the traversal. After ending the traversal, the available storage block can be determined as the replacement storage block.
[0055] According to an embodiment of the present invention, by combining real-time detection of the health status and usage status of storage blocks, fast and reliable replacement block selection is achieved. This strategy ensures the quality of replacement blocks through strict availability verification, thereby ensuring the availability of the storage device after replacement.
[0056] According to an embodiment of the present invention, using a replacement storage block to replace a faulty storage block includes: obtaining the original data stored in the faulty storage block and the logical address corresponding to the faulty storage block; storing the original data in the replacement storage block; modifying the logical address to point to the replacement storage block so that subsequent data storage requests and data read requests initiated through the logical address can be allocated to the replacement storage block; and marking the faulty storage block as an invalid storage block.
[0057] According to an embodiment of the present invention, when reading a faulty storage block, if the reading is successful, the original data is obtained. If the reading fails, the storage log of the faulty storage block is queried to determine the latest storage log before the faulty storage block fails, and the original data is determined based on the latest storage log.
[0058] According to an embodiment of the present invention, after obtaining the original data, the original data is stored in a replacement storage block, so that during the subsequent data reading process from the replacement storage block, the data stored before the faulty storage block fails can be read, thereby avoiding data loss during the repair process of the storage device and ensuring data security.
[0059] According to an embodiment of the present invention, each storage block corresponds to a logical address, and the logical address can be used by upper-layer applications. The upper-layer applications do not need to know the specific physical address of the corresponding storage block. They only need to use the logical address to locate the storage block corresponding to the logical address and complete the data reading and writing operations on the storage block.
[0060] Therefore, when using a replacement storage block to replace a faulty storage block, only by modifying the logical address of the faulty storage block to point to the replacement storage block, the read and write operations triggered by the upper-layer applications during subsequent operations can be transferred to the replacement storage block.
[0061] According to an embodiment of the present invention, the physical address of the faulty storage block can be added to the invalid storage block table to mark the faulty storage block as an invalid storage block, so that during the subsequent use of the storage device, the faulty storage block will no longer be operated on, reducing the invalid operations on the storage block and improving the operation efficiency.
[0062] According to an embodiment of the present invention, by completely saving the original data and updating the logical address mapping, seamless replacement of the faulty block can be achieved, and data integrity is ensured. At the same time, through a transparent address redirection mechanism, upper-layer applications do not need to perceive the change of the underlying storage structure, improving the system availability and user experience.
[0063] According to an embodiment of the present invention, the storage device repair method further includes: in response to the power-on of the storage device, traversing the bad block table of the storage device. If it is determined that there is a newly added faulty storage block in the area corresponding to the functional area in the bad block table, it is determined that a faulty storage block exists in the functional area of the storage device; the newly added faulty storage block is recorded in the bad block table in the following way: performing a fault detection on the storage device to obtain a fault detection result; if the fault detection result indicates that a storage block in the storage device has a fault, the storage block is marked in the bad block table.
[0064] According to an embodiment of the present invention, the storage device can be fault-detected by regularly performing operations such as data reading and writing on the storage device. During the fault detection, if there is a storage block in which data cannot be written or read, it can be determined that the storage block in the storage device where data cannot be written or read has a fault, and the storage block is marked in the bad block table.
[0065] Through the above-mentioned fault detection means, during the power-on process of the storage device, the storage blocks can be periodically detected, and when a new faulty storage block is found, it is marked in the bad block table, so that when the storage device is powered on next time, the newly added faulty storage blocks can be determined through the bad block table.
[0066] Specifically, each time the storage device is powered on, the bad block table of the storage device can be traversed to obtain a traversal result, and based on the difference between the traversal result and the traversal result when the storage device was powered on last time, the newly added faulty storage blocks of the storage device during the period from the last power-on of the storage device to the current power-on of the storage device can be determined. Among them, the bad block table is a data table used to store the faulty storage blocks in the storage device, and the newly added faulty storage blocks are the storage blocks that did not have a fault when the storage device was powered on last time and have a fault when the storage device is powered on this time.
[0067] According to an embodiment of the present invention, by actively scanning the bad block table when the device is started to determine the newly added faulty storage blocks, it is possible to promptly discover the newly added faults, provide data support for the accurate replacement of subsequent storage blocks, and improve the preventive maintenance ability of the system.
[0068] According to an embodiment of the present invention, it is determined to determine a spare storage block from the same data plane in the replacement area according to the data plane where the faulty storage block is located in the target logical unit. In the case where all the spare storage blocks in this data plane are unavailable storage blocks, it can be determined that the faulty storage block cannot be replaced in the current storage device. Therefore, the faulty storage block can be marked as a permanent bad block, and the logical address corresponding to the permanent bad block is redirected to the spare storage block of the storage device. When a data write request or a data read request is received in the spare storage block, the spare storage block redirects the data write request or the data read request to the storage block that can be used for the load data write request or the data read request according to the usage conditions of multiple storage blocks in the functional area to perform data reading and writing.
[0069] According to an embodiment of the present invention, when the proportion of permanent bad blocks in the storage device to all storage blocks in the functional area reaches a preset bad block ratio, it can be determined that the storage device is scrapped, a scrap time is set, and a scrap reminder is sent to the user. When the storage device is powered on next time, the data stored in the storage device is read out and the data is saved in an electronic device connected to the storage device. After the scrap time is reached, it is determined that the storage device fails. By sending a scrap reminder and setting a scrap time, the health status of the storage device can be synchronized to the user in a timely manner. Automatically backing up data can avoid the situation of data loss caused by the user not backing up data because they are unaware that the storage device is about to be scrapped.
[0070] Figure 3 The flowchart of a storage device repair method according to another embodiment of the present invention is shown.
[0071] As Figure 3 shown, the process includes operation S301 to operation S312.
[0072] In operation S301, it is detected that the storage device is powered on, the bad block table is traversed, and a traversal result is obtained.
[0073] In operation S302, the traversal result is compared with the traversal result of the previous power-on of the storage device to obtain a comparison result.
[0074] In operation S303, it is determined whether the comparison result indicates that there are newly added faulty storage blocks in the storage device. If not, operation S304 is executed; if so, operation S305 is executed.
[0075] In operation S304, the storage device repair process is ended, and the storage device is waited for to be powered on next time.
[0076] In operation S305, the data plane where the newly added faulty storage block is located is determined.
[0077] In operation S306, according to the data plane, a spare storage block is selected from the replacement area.
[0078] In operation S307, it is determined whether the spare storage block has no faults and has not been used as a replacement storage block. If not, operation S308 is executed; if so, operation S311 is executed.
[0079] In operation S308, it is determined whether the currently selected spare storage block is the last spare storage block. If so, operation S309 is executed; if not, operation S310 is executed.
[0080] In operation S309, the faulty storage block is marked as a permanent bad block, and the storage device repair process is ended.
[0081] In operation S310, mark the currently selected spare storage block as an unavailable storage block, and perform operation S306.
[0082] In operation S311, determine the currently selected spare storage block as the replacement storage block.
[0083] In operation S312, replace the faulty storage block with the replacement storage block.
[0084] According to an embodiment of the present invention, when the fault detection result indicates that one of the following situations exists in the storage block of the storage device, the fault detection result indicates that there is a fault in the storage block of the storage device: when writing data to the storage block, the storage block returns a programming error; when reading data from the storage block, the storage block feedbacks that the data reading fails or incorrect data is read.
[0085] According to an embodiment of the present invention, during the process of fault detection, test data is written to multiple storage blocks of the storage device respectively, and then the test data is read out, so as to judge the health status of each of the multiple storage blocks.
[0086] According to an embodiment of the present invention, for each storage block, during the process of writing data to the storage block, if a programming error is returned, it can be determined that the storage block cannot write data, and it is judged that the storage block has a fault. If no programming error is returned, it can be determined that the data has been correctly written into the storage block, and the storage block has no writing fault.
[0087] According to an embodiment of the present invention, for each storage block, according to the previously written test data, each storage block is read respectively, and it is determined that the storage block that cannot correctly read the written test data has a fault, where the inability to query the written test data may include situations such as not being able to query the data or reading incorrect data.
[0088] According to an embodiment of the present invention, after the test of the storage device is completed, the test data written to each storage block during the test process can be deleted to ensure that the test process will not affect the actual stored data of the storage device.
[0089] Particularly, in the case where the storage capacity of the storage block has been fully occupied and cannot accommodate the test data to be written, the data also cannot be correctly written into the storage block. However, in this case, it cannot be determined whether the storage block has a writing fault. The above storage block can be used as a storage block to be tested for subsequent testing to determine whether the writing function of the storage block is normal. Since the test data cannot be written, the reading function of the storage block cannot be tested either.
[0090] For example, part of the original data can be read from the storage block to be tested as the data to be transferred. After saving the data to be transferred to the reserved space in the storage device, the data to be transferred is deleted from the storage block to be tested. It is necessary to ensure that the amount of data occupied by the data to be transferred is greater than the test data to be written during the test process.
[0091] After completing the above operations, test data can be normally written to the storage block to be tested, and the read function of the storage block to be tested can be tested based on the test data. After obtaining the test result of the storage block to be tested, the pre-transferred data to be transferred can be read from the reserved space, the data to be transferred can be stored back in the storage block to be tested, and the data to be transferred can be deleted from the reserved space.
[0092] According to the embodiments of the present invention, transferring the data in the storage block to be tested can test the storage block whose storage space has been fully occupied, thereby avoiding the omission of the storage block to be tested due to the lack of extra storage space in the storage block to be tested, resulting in the missed detection of the faulty storage block, improving the accuracy and comprehensiveness of the fault detection, and improving the real-time performance of the storage device repair.
[0093] According to the embodiments of the present invention, by using the read / write failure as the judgment condition for the faulty storage block, an objective fault identification mechanism is established. This standard covers the main operation scenarios of the storage device, can accurately distinguish real hardware faults and temporary errors, avoids invalid replacement operations caused by misjudgment, and improves the accuracy of fault diagnosis.
[0094] According to the embodiments of the present invention, the storage device repair method further includes: after replacing the faulty storage block, recording the replacement storage block and the faulty storage block in the update record table, so that in the case of a new faulty storage block, the usage of multiple storage blocks in the replacement area can be determined based on the update record table.
[0095] According to the embodiments of the present invention, the update record table can be used to record the faulty storage blocks involved and the replacement storage blocks for replacing the faulty storage blocks during each storage device repair process.
[0096] According to the embodiments of the present invention, updating and maintaining the update record table can perform update traceability based on the update record table, that is, determine whether there is a replacement situation for each storage block in the current storage device, and locate the storage blocks with replacement situations. It can also determine whether multiple storage blocks in the replacement area have been used as replacement storage blocks based on the update record table, and then determine the usage of each of the multiple storage blocks in the replacement area.
[0097] According to an embodiment of the present invention, by continuously tracking the replacement relationship, a complete repair history is constructed. This mechanism not only provides a decision-making basis for subsequent fault handling, but also dynamically optimizes the resource allocation in the replacement area, realizes the intelligent maintenance of the storage device, and significantly improves the long-term reliability and service life of the device.
[0098] Figure 4A Shows the health status of the storage blocks of the storage device before repair by the storage device repair method according to an embodiment of the present invention.
[0099] Among them, since the storage block is the smallest erasure unit in the storage device, the data pages in the storage block are no longer shown in the figure. In addition, the multiple storage blocks in each column in the figure can represent the storage blocks in the same data plane.
[0100] As Figure 4A shown, in the functional area, storage block 1 and storage block 3 of logical unit 2 are faulty.
[0101] Similarly, it can be determined that storage block 3 of logical unit 5 is faulty. Except for the above three faulty storage blocks, the other storage blocks are healthy storage blocks operating normally.
[0102] Figure 4B Shows the original health status and corresponding relationship of the storage blocks of the storage device after repair by the storage device repair method according to an embodiment of the present invention.
[0103] As Figure 4B shown, shows the health status of the storage blocks in the storage device and the corresponding relationship between the faulty storage blocks and the replacement storage blocks after the replacement is completed.
[0104] According to an embodiment of the present invention, during the process of replacing the storage blocks by applying the storage device repair method, the data plane and target location information of the three faulty storage blocks are as follows:
[0105] The first faulty storage block is storage block 1 of logical unit 2.
[0106] The second faulty storage block is storage block 3 of logical unit 2.
[0107] The third faulty storage block is storage block 3 of logical unit 5.
[0108] According to the above data plane and the target location, a spare storage block can be selected from the replacement area. For the first faulty storage block, the selected spare storage block is the first storage block (i.e., storage block 1') in the replacement area of the same data plane in this logical unit (i.e., logical unit 2). By judging the usage situation of this spare storage block, it can be determined that this spare storage block has no faults and has not been used as a replacement storage block. Therefore, this spare storage block is an available storage block, and this spare storage block can be used as a replacement storage block to replace the first faulty storage block.
[0109] For the second faulty storage block, the selected spare storage block is the first storage block (i.e., storage block 1') in the replacement area of the same data plane in this logical unit (i.e., logical unit 2). By judging the usage situation of this spare storage block, it can be determined that this spare storage block has been used as a replacement storage block. Therefore, this spare storage block is an unavailable storage block. Continue to select a spare storage block, and the next storage block (i.e., storage block 2') in this logical unit can be obtained. By judging the usage situation of this spare storage block, it can be determined that this spare storage block has faults. Therefore, this spare storage block is an unavailable storage block. Continue to traverse in logical unit 2, take storage block 3' as the spare storage block, and determine that the spare storage block has no faults and has not been used as a replacement storage block. Therefore, this spare storage block is an available storage block, and this spare storage block can be used as a replacement storage block to replace the second faulty storage block.
[0110] Similarly, for the third faulty storage block, the selected spare storage block is the first storage block (i.e., storage block 1') in the replacement area of the same data plane in this logical unit (i.e., logical unit 5). It is determined that this spare storage block has been used as a replacement storage block. Therefore, this spare storage block is an unavailable storage block. Continue to select a spare storage block, and the next storage block (i.e., storage block 2') in this logical unit can be obtained. By judging the usage situation of this spare storage block, it can be determined that this spare storage block has no faults and has not been used as a replacement storage block. Therefore, this spare storage block is an available storage block, and this spare storage block can be used as a replacement storage block to replace the third faulty storage block.
[0111] Based on the above storage device repair method, the present invention also provides a storage device repair device. The following will be combined with Figure 5 to describe this device in detail.
[0112] Figure 5 shows a structural block diagram of a storage device repair device according to an embodiment of the present invention.
[0113] As Figure 5 shown, the storage device repair device 500 of this embodiment includes a position acquisition module 510, a storage block selection module 520, and a storage block replacement module 530.
[0114] The location acquisition module 510 is configured to obtain the data plane of a faulty storage block in a target logical unit in response to detecting a faulty storage block in a functional area of a storage device. The storage device includes a plurality of logical units, each of the plurality of logical units includes a plurality of storage blocks and is divided into a functional area and a replacement area. The functional area is used to store data, the replacement area includes spare storage blocks, and the target logical unit is the logical unit where the faulty storage block is located. In one embodiment, the location acquisition module 510 may be configured to perform the operation S210 described above, which will not be elaborated here.
[0115] The storage block selection module 520 is configured to select a replacement storage block for replacing the faulty storage block from the replacement area according to the data plane. The replacement area is used to replace the faulty storage block, and the replacement area includes a plurality of logical units. In one embodiment, the storage block selection module 520 may be configured to perform the operation S220 described above, which will not be elaborated here.
[0116] The storage block replacement module 530 is configured to replace the faulty storage block with the replacement storage block. In one embodiment, the storage block replacement module 530 may be configured to perform the operation S230 described above, which will not be elaborated here.
[0117] According to an embodiment of the present invention, the storage block selection module 520 includes a replacement block selection sub-module.
[0118] The replacement block selection sub-module is configured to select a spare storage block in the data plane in the replacement area as the replacement storage block.
[0119] According to an embodiment of the present invention, the replacement block selection sub-module includes a storage block traversal block.
[0120] The storage block traversal block is configured to sequentially traverse the spare storage blocks in the data plane until an available storage block is determined, and determine the available storage block as the replacement storage block.
[0121] According to an embodiment of the present invention, the storage block replacement module 530 includes an address determination sub-module, a data storage sub-module, an address correction sub-module, and a storage block marking sub-module.
[0122] The address determination sub-module is configured to obtain the original data stored in the faulty storage block and the logical address corresponding to the faulty storage block.
[0123] The data storage sub-module is configured to store the original data in the replacement storage block.
[0124] The address correction sub-module is configured to correct the logical address to point to the replacement storage block, so that subsequent data storage requests and data read requests initiated through the logical address can be allocated to the replacement storage block.
[0125] A storage block marking sub-module, configured to mark a faulty storage block as an invalid storage block.
[0126] According to an embodiment of the present invention, the storage device repair apparatus 500 further includes a data table traversal module, a fault detection module, and a storage block marking module.
[0127] The data table traversal module is configured to traverse the bad block table of the storage device in response to the power-on of the storage device, and determine that a faulty storage block exists in the functional area of the storage device when it is determined that a new faulty storage block exists in the area corresponding to the functional area in the bad block table.
[0128] The fault detection module is configured to perform a fault detection on the storage device to obtain a fault detection result.
[0129] The storage block marking module is configured to mark the storage block in the bad block table when the fault detection result indicates that a storage block in the storage device is faulty.
[0130] According to an embodiment of the present invention, the storage device repair apparatus 500 further includes a fault detection module.
[0131] The fault detection module is configured to, when the fault detection result indicates that a storage block in the storage device is in one of the following situations, the fault detection result indicates that a storage block in the storage device is faulty: when writing data to the storage block, the storage block returns a programming error; when reading data from the storage block, the storage block feedbacks a data read failure or reads incorrect data.
[0132] According to an embodiment of the present invention, the storage device repair apparatus 500 further includes a record update module.
[0133] The record update module is configured to, after replacing the faulty storage block, record the replacement storage block and the faulty storage block in an update record table, so that in the case of a new faulty storage block, the usage of multiple storage blocks in the replacement area can be determined based on the update record table.
[0134] According to an embodiment of the present invention, any multiple of the position acquisition module 510, the storage block selection module 520, and the storage block replacement module 530 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the position acquisition module 510, the storage block selection module 520, and the storage block replacement module 530 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or may be implemented by any other reasonable means such as hardware or firmware for integrating or packaging circuits, or may be implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the position acquisition module 510, the storage block selection module 520, and the storage block replacement module 530 may be at least partially implemented as a computer program module, and when the computer program module is run, corresponding functions may be executed.
[0135] Figure 6 A block diagram of an electronic device suitable for implementing a storage device repair method according to an embodiment of the present invention is shown.
[0136] As Figure 6 shown, the electronic device 600 according to an embodiment of the present invention includes a processor 601, which may perform various appropriate actions and processes according to a program stored in a read only memory (ROM) 602 or a program loaded from a storage section 608 into a random access memory (RAM) 603. The processor 601 may include, for example, a general microprocessor (such as a CPU), an instruction set processor and / or a related chipset and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), and so on. The processor 601 may also include on-board memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0137] In the RAM 603, various programs and data required for the operation of the electronic device 600 are stored. The processor 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. The processor 601 performs various operations of the method flow according to the embodiments of the present invention by executing the programs in the ROM 602 and / or the RAM 603. It should be noted that the programs can also be stored in one or more memories other than the ROM 602 and the RAM 603. The processor 601 can also perform various operations of the method flow according to the embodiments of the present invention by executing the programs stored in the one or more memories.
[0138] According to an embodiment of the present invention, the electronic device 600 may further include an input / output (I / O) interface 605, and the input / output (I / O) interface 605 is also connected to the bus 604. The electronic device 600 may further include one or more of the following components connected to the input / output (I / O) interface 605: an input part 606 including a keyboard, a mouse, etc.; an output part 607 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage part 608 including a storage device, etc.; and a communication part 609 including a network interface card such as a LAN card, a modem, etc. The communication part 609 performs communication processing via a network such as the Internet. A driver 610 is also connected to the input / output (I / O) interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the driver 610 as needed so that a computer program read from it can be installed into the storage part 608 as needed.
[0139] The present invention also provides a computer-readable storage medium, which may be included in the device / device / system described in the above embodiments; or may exist separately without being assembled into the device / device / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present invention is implemented.
[0140] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, storage devices, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include one or more memories other than the ROM 602 and / or RAM 603 and / or ROM 602 and RAM 603 described above.
[0141] An embodiment of the present invention further includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the method provided by the embodiment of the present invention.
[0142] When the computer program is executed by the processor 601, it executes the above functions defined in the system / apparatus of the embodiment of the present invention. According to an embodiment of the present invention, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0143] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and is downloaded and installed through the communication part 609, and / or installed from the removable medium 611. The program code included in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0144] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or installed from the removable medium 611. When the computer program is executed by the processor 601, it executes the above functions defined in the system of the embodiment of the present invention. According to an embodiment of the present invention, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0145] According to embodiments of the present invention, program code for executing the computer programs provided by the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).
[0146] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the block may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0147] Those skilled in the art can understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.
[0148] The above describes the embodiments of the present invention. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments are described separately above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.
Claims
1. A method for repairing a storage device, characterized in that, The method includes: In response to detecting a faulty storage block in a functional area of a storage device, obtaining a data plane in a target logical unit where the faulty storage block is located, where the storage device includes a plurality of logical units, each of the plurality of logical units includes a plurality of storage blocks and is divided into a functional area and a replacement area, the functional area is used to store data, the replacement area includes spare storage blocks, and the target logical unit is the logical unit where the faulty storage block is located; Selecting, according to the data plane, a replacement storage block from the replacement area for replacing the faulty storage block; and Replacing the faulty storage block with the replacement storage block.
2. The method according to claim 1, wherein The selecting, according to the data plane, a replacement storage block from the replacement area for replacing the faulty storage block includes: In the replacement area, selecting a spare storage block in the data plane as the replacement storage block.
3. The method according to claim 2, wherein The selecting a spare storage block in the data plane as the replacement storage block in the replacement area includes: Sequentially traversing the spare storage blocks in the data plane until an available storage block is determined, and determining the available storage block as the replacement storage block.
4. The method according to claim 1, characterized in that The replacing the faulty storage block with the replacement storage block includes: Obtaining original data stored in the faulty storage block and a logical address corresponding to the faulty storage block; Storing the original data into the replacement storage block; Modifying the logical address to point to the replacement storage block, so that subsequent data storage requests and data read requests initiated through the logical address can be allocated to the replacement storage block; and Marking the faulty storage block as an invalid storage block.
5. The method according to claim 1, wherein The method further includes: In response to the storage device being powered on, traversing a bad block table of the storage device, and determining that a faulty storage block is detected in the storage device when it is determined that there is a newly added faulty storage block in an area corresponding to the functional area in the bad block table; The newly added faulty storage block is recorded in the bad block table in the following manner: Performing a fault detection on the storage device to obtain a fault detection result; When the fault detection result indicates that a storage block in the storage device is faulty, marking the storage block in the bad block table.
6. The method according to claim 5, characterized in that, When the fault detection result indicates one of the following situations for a storage block in the storage device, the fault detection result indicates that the storage block in the storage device is faulty: Writing data to the storage block, and the storage block returns a programming error; Reading data from the storage block, and the storage block feedbacks a data read failure or reads incorrect data.
7. The method according to any one of claims 1 to 6, characterized in that The method further includes: After completing the replacement of the faulty storage block, recording the replacement storage block and the faulty storage block in an update record table, so that in the case of a newly added faulty storage block, the usage conditions of multiple storage blocks in the replacement area can be determined based on the update record table.
8. A storage device repair apparatus, characterized in that The apparatus includes: A location acquisition module, configured to acquire a data plane of a faulty storage block in a target logical unit in response to detecting that there is a faulty storage block in a functional area of a storage device, where the storage device includes a plurality of logical units, each of the plurality of logical units includes a plurality of storage blocks and is divided into a functional area and a replacement area, the functional area is used to store data, the replacement area includes spare storage blocks, and the target logical unit is the logical unit where the faulty storage block is located; A storage block selection module, configured to select a replacement storage block for replacing the faulty storage block from the replacement area according to the data plane, where the replacement area is used to replace the faulty storage block, and the replacement area includes a plurality of the logical units; and A storage block replacement module, configured to replace the faulty storage block with the replacement storage block.
9. An electronic device, comprising: One or more processors; A memory, configured to store one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, The computer program or instruction, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Wear leveling method and device for storage equipment and related equipment
CN111913647A
Storage block replacement method and device based on solid state disk
CN118098320A
Memory block replacement method and device based on solid state disk, equipment and medium
CN118585128A