Data repair method and device, electronic equipment, storage medium and program product
Patent Information
- Application Number
- CN202611142836.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-30
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-07-30
AI Technical Summary
[0004]本申请提供了数据修复方法、装置、电子设备、存储介质及程序产品,以至少解决相关技术中快照回滚和写前拷贝方式修复坏块时,需暂停业务、修复效率低、数据一致性难以保障等问题
[0010]通过本申请,采用快照卷映射与双写修复机制,在检测到坏块后建立源逻辑地址与目标逻辑地址之间的映射关系,将写请求数据同时存储于备用物理块和目标物理块,并在备用物理块可读时进行数据一致性比较,仅在比较结果不一致时执行数据修复。因此,可以解决相关技术中坏块修复需暂停业务、全量复制耗时长、数据一致性难以保障的技术问题,达到业务零中断、增量修复、数据一致性保障的技术效果。
Smart Images

Figure CN122653915B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data recovery technology, and in particular to data recovery methods, apparatus, electronic devices, storage media and program products. Background Technology
[0002] In the field of data recovery, snapshot technology is commonly used for data protection. When bad blocks appear on the original volume, snapshot rollback or copy-before-write methods are typically used for repair.
[0003] However, all of the above repair methods require suspending business operations for data replication, resulting in service interruption during the repair process. Furthermore, full replication is time-consuming and consumes significant storage space. The copy-before-write method also faces the problem of read / write amplification due to excessively long snapshot chains. Therefore, how to efficiently repair bad blocks while ensuring business continuity is a pressing technical problem that needs to be solved. Summary of the Invention
[0004] This application provides data repair methods, apparatus, electronic devices, storage media, and program products to at least solve the problems in related technologies, such as the need to suspend business operations, low repair efficiency, and difficulty in ensuring data consistency when repairing bad blocks using snapshot rollback and copy-before-write methods.
[0005] This application provides a data repair method, comprising: in response to detecting at least one bad block in a source data volume, determining the source logical address range of the at least one bad block, and establishing a first mapping relationship between the source logical address range and a target logical address range in a target volume, wherein the target volume is a snapshot volume of the source data volume before the bad block was detected; in response to receiving a write request for any source logical address within the source logical address range, storing the data to be written carried by the write request in a spare physical block in the source data volume and a target physical block corresponding to the target logical address range, respectively, according to the first mapping relationship; and pointing the source logical address to the spare physical block; if it is determined that the spare physical block is readable, comparing the data in the target physical block with the data in the spare physical block to obtain a comparison result; if the comparison result is inconsistent, storing the data in the target physical block in the spare physical block and releasing the first mapping relationship.
[0006] This application also provides a data repair apparatus, comprising: a bad block detection module, configured to, in response to detecting at least one bad block in a source data volume, determine the source logical address range of the at least one bad block, and establish a first mapping relationship between the source logical address range and a target logical address range in a target volume, wherein the target volume is a snapshot volume of the source data volume before the bad block was detected; a write request processing module, configured to, in response to receiving a write request for any source logical address within the source logical address range, store the data to be written carried by the write request in a spare physical block in the source data volume and a target physical block corresponding to the target logical address range, respectively, according to the first mapping relationship; and a data repair module, configured to, if it is determined that the spare physical block is readable, compare the data in the target physical block with the data in the spare physical block to obtain a comparison result; if the comparison result is inconsistent, store the data in the target physical block in the spare physical block and release the first mapping relationship.
[0007] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of any of the above methods.
[0008] This application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above methods.
[0009] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above methods.
[0010] This application employs a snapshot volume mapping and dual-write repair mechanism. Upon detecting a bad block, a mapping relationship is established between the source logical address and the target logical address. Write request data is simultaneously stored in both the backup physical block and the target physical block. Data consistency is compared only when the backup physical block is readable, and data repair is performed only if the comparison results are inconsistent. Therefore, this addresses the technical problems of related technologies, such as the need to pause business operations for bad block repair, the time-consuming nature of full replication, and the difficulty in guaranteeing data consistency. It achieves the technical effects of zero business interruption, incremental repair, and guaranteed data consistency. Attached Figure Description
[0011] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This application provides an illustration of an application scenario for a data repair method.
[0013] Figure 2 A flowchart illustrating the data repair method provided in this application embodiment;
[0014] Figure 3 This is a schematic diagram of the write request dual-write processing flow provided in an embodiment of this application;
[0015] Figure 4 This is a schematic diagram of the backup physical block processing flow provided in the embodiments of this application;
[0016] Figure 5 This is a schematic diagram of the overall data repair process provided in the embodiments of this application;
[0017] Figure 6 This is a structural block diagram of a data repair device according to an embodiment of this application. Detailed Implementation
[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0019] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0020] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0021] Embodiments of this application provide a data recovery method, apparatus, electronic device, storage medium, and program product.
[0022] Figure 1 This is an application scenario diagram of a data repair method provided in an embodiment of this application.
[0023] like Figure 1As shown, the application scenarios according to this embodiment include interactive scenarios for data repair, specifically including a host 101, a network 102, and a storage system 103. The network 102 is the communication medium between the host 101 and the storage system 103, and may include connection types such as wired, wireless communication links, or fiber optic cables.
[0024] In this application scenario, the storage system 103 is configured with a source data volume and a target volume. The target volume is a snapshot of the source data volume before bad blocks are detected, used to save the data state of the source data volume at the snapshot time. When a physical block of the source data volume is damaged, the storage system 103 detects at least one bad block through the storage controller, determines the source logical address range corresponding to the bad block, and establishes a first mapping relationship between the source logical address range and the target logical address range in the target volume. This first mapping relationship is used to redirect read and write requests for bad blocks to the target volume.
[0025] When host 101 issues a write request for any source logical address within the source logical address range, storage system 103 receives the write request through the storage controller, obtains the data to be written carried in the write request, and stores the data to be written in a spare physical block in the source data volume and a target physical block corresponding to the target logical address range, respectively, according to the first mapping relationship. The spare physical block is a pre-configured healthy physical block in the source data volume, used to temporarily take over business writes during bad block repair; the target physical block is a physical block in the target volume corresponding to the source logical address, used to save the correct data at the snapshot time.
[0026] After storing the data to be written in the spare physical block, the storage system 103 determines whether the spare physical block is readable. If the spare physical block is determined to be readable, the storage system 103 reads data from the target physical block and compares the read data with the data stored in the spare physical block to obtain a comparison result. This comparison operation is used to determine whether the data in the spare physical block is consistent with the data in the target physical block, in order to determine whether data repair needs to be performed. If the comparison result is inconsistent, the storage system 103 stores the data from the target physical block in the spare physical block to correct the data in the spare physical block so that it is consistent with the data in the target physical block, and removes the first mapping relationship, thereby stopping the redirection of requests for the source logical address to the target volume and restoring normal read and write operations to the source data volume.
[0027] For example, when host 101 issues a read request for any source logical address within the source logical address range, storage system 103 directs the read request to a target logical address within the target logical address range according to the first mapping relationship, reads data from the target physical block corresponding to the target logical address, and returns it to host 101. In this way, storage system 103 continuously processes read and write requests issued by host 101 during bad block repair, achieving zero service interruption during the repair process.
[0028] It should be noted that the storage system 103 can be a standalone storage server, a distributed storage cluster, a cloud storage platform, a storage controller or storage chip integrated into a computing device, or a combination of the above-mentioned devices. This application embodiment does not limit the specific implementation of the storage system, as long as it can interact with the host and execute data repair methods. The host 101 can be a single physical server, a virtual machine or container instance, or a computing cluster composed of multiple physical devices; this application embodiment does not limit this.
[0029] The following will be based on Figure 1 The described scene, through Figures 2-5 The data repair method of the disclosed embodiments will be described in detail.
[0030] Figure 2 A flowchart of the data repair method provided in the embodiments of this application.
[0031] like Figure 2 As shown, the data repair method of this embodiment may include operations S210 to S230.
[0032] In operation S210, in response to detecting at least one bad block in the source data volume, the source logical address range of the at least one bad block is determined, and a first mapping relationship is established between the source logical address range and the target logical address range in the target volume.
[0033] In this embodiment, the target volume is a snapshot of the source data volume before bad blocks are detected. The source data volume is a logical storage unit in the storage system that carries business data and provides data read / write services to the host. Bad blocks are physical storage units in the storage medium that cannot be read or written normally due to physical damage, data errors, or other reasons.
[0034] The source logical address range can be an interval of one or more consecutive logical addresses corresponding to the bad block in the source data volume. The target volume can be a snapshot volume created for the source data volume using snapshot technology before the bad block is detected, used to save the complete data state of the source data volume at the time of the snapshot. The first mapping relationship can be a one-to-one correspondence between the source logical address range and the target logical address range, used to guide read and write requests for bad blocks to the corresponding locations in the target volume.
[0035] In operation S220, in response to receiving a write request for any source logical address within the source logical address range, the data to be written carried by the write request is stored in the spare physical block in the source data volume and the target physical block corresponding to the target logical address range, respectively, according to the first mapping relationship, and the source logical address is pointed to the spare physical block.
[0036] In this embodiment, the write request is initiated by the host, carrying the data to be written and the target logical address. The spare physical block can be a healthy physical block dynamically allocated by the storage system from the spare block pool reserved in the source data volume after detecting a bad block. This block is used to temporarily take over the write data for the bad block, ensuring uninterrupted service writing. The target physical block can be the physical storage location in the target volume corresponding to the source logical address, used to store the correct data at the time of the snapshot. The storage system determines the target logical address based on the first mapping relationship, and then locates the target physical block.
[0037] By storing the data to be written in both a standby physical block and a target physical block simultaneously, it ensures that the data written by the business can be immediately read subsequently in the standby block, while a correct copy of the data is maintained in the target block, providing a data foundation for subsequent consistency comparisons. The entire dual-write process is completely transparent to the business.
[0038] It's important to note that this involves pointing the source logical address to the spare physical block. Specifically, this operation includes the storage system updating the metadata mapping of the source data volume, switching the physical address corresponding to the source logical address from the original bad block to the physical address of the spare physical block. Through this pointing operation, subsequent read requests targeting that source logical address can directly read data from the spare physical block without needing to query the first mapping or access the target volume again.
[0039] This pointer operation ensures that during the window period between the completion of a write request and the completion of data repair, business read operations can normally retrieve the written data, avoiding read request failures due to bad blocks and ensuring uninterrupted business read operations during bad block repair. Simultaneously, this pointer operation allows the spare physical block to officially take over the data service of the source logical address, providing the correct data source for subsequent data consistency comparisons and repairs.
[0040] In operation S230, if the backup physical block is determined to be readable, the data in the target physical block is compared with the data in the backup physical block to obtain a comparison result; if the comparison result is inconsistent, the data in the target physical block is stored in the backup physical block and the first mapping relationship is released.
[0041] In this embodiment, the readability determination of the spare physical block is used to confirm whether the data written to the spare physical block can be read normally. If the spare physical block is readable, data is read from the target physical block and compared with the data in the spare physical block. This comparison operation is used to detect whether the data in the spare physical block is consistent with the data in the target physical block.
[0042] When the comparison result is consistent, it indicates that the data in the spare physical block is correct and no further repair is needed. When the comparison result is inconsistent, it indicates that the data in the spare physical block is deviated, and the data in the target physical block needs to be stored in the spare physical block to complete the data correction. After the repair is completed, the first mapping relationship is released, redirecting requests for that logical address to the target volume is stopped, and the normal read / write mode of the source data volume is restored.
[0043] In this embodiment, bad blocks in the source data volume are detected, and a first mapping relationship is established between the source logical address range and the target logical address range in the target volume. Write request data is simultaneously stored in both the spare physical block and the target physical block. When the spare physical block is readable, the data consistency between the target physical block and the spare physical block is compared. Data repair and demapping are only performed when the comparison results are inconsistent. This scheme, through mapping redirection and dual-write verification mechanisms, ensures uninterrupted business read / write operations during bad block repair, allowing write requests to be processed normally and read requests to retrieve correct data from the target volume. By repairing on demand instead of full replication, repair time is significantly shortened and storage resource consumption is reduced. The dual-write mechanism ensures that newly written data is not lost during repair, guaranteeing data consistency and improving the efficiency and reliability of bad block repair.
[0044] It should be noted that snapshot technology is a point-in-time data protection technique. When a snapshot operation is performed on a source data volume, the storage system creates a corresponding snapshot volume for the source data volume, which is the target volume in this application, and establishes a relationship between the source data volume and the target volume. The snapshot volume is used to save the data state of the source data volume at the time of snapshot creation. When the snapshot is first created, the target volume does not actually store a physical copy of the source data volume. Instead, it copies the metadata mapping relationship of the source data volume so that the metadata of the target volume points to the physical data blocks of the source data volume at the time of snapshot.
[0045] When a data block in the source data volume is modified for the first time after a snapshot is created, the snapshot mechanism first copies the original data of that data block to the corresponding storage location in the target volume before allowing modification of the data block in the source data volume. This mechanism is called copy-on-write. In this way, the target volume stores a complete view of the source data volume at the time of the snapshot. When bad blocks appear in the source data volume, the correct data of the bad blocks at the time of the snapshot can be obtained from the target volume, providing a reliable data source for data repair.
[0046] In this embodiment of the application, in response to receiving a read request for any source logical address within the source logical address range, the read request is directed to a target logical address within the target logical address range according to the mapping relationship, and data is read from the target physical block corresponding to the target logical address.
[0047] In this embodiment, the read request is a business read operation of the source data volume, carrying the starting logical address and data length information. Upon receiving the read request, the starting logical address is extracted, and it is determined whether this starting logical address falls within the range of source logical addresses recorded in the bad block mapping table. When the starting logical address falls within the range of source logical addresses, the target logical address corresponding to the source logical address in the target volume is determined according to the correspondence recorded in the first mapping relationship. The access target of the read request is switched from the source data volume to the target volume, using the target logical address as the new read location. Through the target volume's metadata mapping table, the target logical address is converted into the corresponding target physical block address, data is read from the target physical block, and the read data is returned as the response data for the read request.
[0048] It's important to note that a Logical Block Address (LBA) is a number used in a storage system to identify the location of data within the logical storage space; it serves as the logical location identifier for upper-layer applications accessing the data. A Physical Block Address (PBA) is the actual storage location of data within the physical storage medium; it identifies the specific storage unit of data on physical devices such as disks or solid-state drives. A Target Logical Block Address (TLBA) is the logical address within the target volume that corresponds to the source logical address; it's used to locate the storage location within the target volume corresponding to the location of the bad block's source logical address.
[0049] For example, the source logical address range is LBA 100 to LBA 199, where the target logical address corresponding to LBA 150 is TLBA 150. When a read request for LBA 150 is received, the target logical address TLBA 150 corresponding to LBA 150 is looked up in the bad block mapping table. The read request is forwarded to the target volume, which uses its own metadata mapping table to translate TLBA 150 into the physical block address PBA 500, and then reads the data from PBA 500.
[0050] In this embodiment, by directing read requests to the target logical address and converting them to the target physical block, services can read data normally during bad block repair. This process does not involve direct access to bad blocks in the source data volume, and read operations are no longer affected by bad blocks, thus achieving uninterrupted read operations.
[0051] In this embodiment of the application, the target logical address corresponding to the source logical address is found from the target logical address range according to the mapping relationship; the read request is forwarded to the target logical address.
[0052] In this embodiment, the target logical address corresponding to the source logical address is searched from the bad block mapping table. The bad block mapping table stores a one-to-one correspondence between source logical address ranges and target logical address ranges. Using the source logical address as an index, the entries of the source logical address range in the bad block mapping table are traversed to determine the source logical address range to which the source logical address belongs. After determining the range, the target logical address range corresponding to that range is obtained. Based on the position offset of the source logical address within its respective source logical address range, the target logical address at the same offset within the target logical address range is determined. After obtaining the target logical address, a redirection request is constructed, replacing the target address field in the original read request with the target logical address, and then the redirected read request is sent to the target volume.
[0053] For example, the bad block mapping table records the source logical address range LBA 100 to LBA 199 and the corresponding target logical address range TLBA 100 to TLBA 199. A read request is received for LBA 150, where LBA 150's offset within the LBA 100 to LBA 199 range is 50. Based on this offset, the target logical address TLBA 150 is determined within the TLBA 100 to TLBA 199 range, and the read request is forwarded to the target logical address TLBA 150.
[0054] In this embodiment, the read request is accurately forwarded to the target logical address by using the correspondence between the source logical address range and the target logical address range in the bad block mapping table. This process ensures that the read request correctly reaches the target location during bad block repair, guaranteeing the accuracy of data reading.
[0055] In this embodiment of the application, in response to receiving a snapshot creation instruction, a snapshot operation is performed on the source data volume to generate a target volume; the snapshot creation instruction is triggered according to a preset time strategy, or in response to manual operation of the object.
[0056] In this embodiment, the snapshot creation command is a control signal that triggers the snapshot operation. The preset time strategy is a configured timed snapshot strategy, such as a rule to execute snapshot creation every day at midnight, or a rule to execute snapshot creation once a week. When the clock reaches the preset time point, the snapshot creation command is automatically generated. Manual operation of the object refers to the snapshot creation operation triggered through the storage system's management interface, command-line interface, or application programming interface. When it is necessary to manually save the data state of the source data volume at a certain moment, a snapshot creation command is sent to the system in the above manner. After receiving the snapshot creation command, a snapshot operation is performed on the source data volume to generate the target volume. The snapshot operation uses a copy-on-write mechanism, copying the metadata mapping relationship of the source data volume as the metadata mapping relationship of the target volume during snapshot creation, so that the target volume points to the physical data blocks of the source data volume at the snapshot time.
[0057] For example, a preset time policy can be configured to create snapshots at 2:00 AM daily. When the clock strikes 2:00 AM, a snapshot creation command is automatically generated, and a snapshot operation is performed on the source data volume to generate the target volume. When a snapshot needs to be created manually, a snapshot creation command is sent through the management interface, and the snapshot operation is executed upon receiving the command to generate the target volume.
[0058] In this embodiment, snapshot creation is triggered through both a preset time strategy and manual operation, ensuring that the target volume existed before the bad blocks occurred. This target volume serves as the data source during the repair process, providing reliable historical data for bad block repair.
[0059] In this embodiment of the application, the response in operation S210 to detecting at least one bad block in the source data volume, determining the source logical address range of at least one bad block, may further include: performing read / write detection on each physical block of the source data volume; if an error occurs during read detection on any physical block, or if write failure occurs during write detection on any physical block, determining that the physical block is a bad block; obtaining the physical block address of at least one bad block, and determining the source logical address range corresponding to the physical block address based on preset metadata, wherein the preset metadata includes a second mapping relationship between each source logical address and physical address in the source data volume.
[0060] In this embodiment, the read / write status of each physical block in the source data volume is monitored in real time by the underlying driver of the storage controller. A read error occurs when a physical block returns an error status, such as error verification failure, inability to read, or read timeout, during a read operation. A write failure occurs when a physical block returns a write error, such as write error, programming failure, erase failure, or write timeout, during a write operation.
[0061] When the above situation occurs, the physical block is identified as a bad block. The physical block address of the bad block is recorded; this address represents the actual location number of the bad block in the physical storage medium. A reverse lookup is performed based on the second mapping relationship in the preset metadata. The second mapping relationship is the mapping relationship between each logical address and physical address in the source data volume, recording the physical block address currently pointed to by each logical address. Using the physical block address of the bad block as an index, all logical addresses pointing to that physical block address are searched in the second mapping relationship, and the found logical addresses are determined as the source logical address range corresponding to the bad block.
[0062] For example, if an error occurs during a read check of physical block address PBA 200 and cannot be recovered after multiple retries, PBA 200 is determined to be a bad block. PBA 200 is retrieved, and the logical address pointing to PBA 200 is searched in the second mapping relationship of the preset metadata. Multiple logical addresses in the range LBA 100 to LBA 199 are found to point to PBA 200; this range is then determined as the source logical address range corresponding to the bad block.
[0063] In this embodiment, bad blocks are identified in real time through read / write detection, and the logical address range corresponding to the bad blocks is determined by reverse lookup using the second mapping relationship. This process enables the precise location of bad blocks, providing accurate address information for subsequent establishment of the first mapping relationship and I / O redirection.
[0064] In this embodiment of the application, the establishment of a first mapping relationship between the source logical address range and the target logical address range in the target volume in operation S210 may further include: recording the source logical address range of at least one bad block and the physical block address corresponding to the source logical address range as source volume record information; recording the target logical address range in the target volume corresponding to the source logical address range as target volume record information; and establishing a first mapping relationship by associating the source volume record information and the target volume record information.
[0065] In this embodiment, the source volume record information is a source volume-side record entry in the bad block mapping table, used to store source data volume information related to bad blocks. The source logical address range of the bad block and the corresponding physical block address are written into the source volume record area of the bad block mapping table. The target volume record information is a target volume-side record entry in the bad block mapping table, used to store the location information of the target volume corresponding to the bad block.
[0066] Write the target logical address range corresponding to the source logical address range in the target volume into the target volume record area of the bad block mapping table. Establish an association between the source volume record area and the target volume record area in the bad block mapping table, and bind the source volume record information and the target volume record information through pointers or indexes to form a one-to-one correspondence between the source logical address range and the target logical address range. This correspondence is the first mapping relationship.
[0067] For example, the source logical address range LBA 100 to LBA 199 and the physical block address PBA 200 of the bad block are recorded as source volume record information, and the target logical address range TLBA 100 to TLBA199 corresponding to LBA 100 to LBA 199 in the target volume are recorded as target volume record information. The source volume record information and the target volume record information are then bound together in the bad block mapping table using an association identifier to establish the first mapping relationship.
[0068] It should be noted that the structure of the bad block mapping table is shown in Table 1:
[0069] Table 1
[0070]
[0071] The bad block mapping table records the source logical address ranges, repair statuses, and correspondences with target logical address ranges for multiple bad blocks. Each row corresponds to one bad block entry. The first row shows that the target logical address range for bad block LBA ranges LBA1 to LBA2 is TLBA1 to TLBA2, and the repair status of this bad block is "pending repair." The second row shows that the target logical address range for bad block LBA ranges LBA3 to LBA4 is TLBA3 to TLBA4, and the repair status of this bad block is "under repair," indicating that data synchronization is currently being performed on this bad block.
[0072] The source logical address ranges in Table 1 correspond to LBA 100 to LBA 199 in the examples above, and the target logical address ranges correspond to TLBA 100 to TLBA 199 in the examples above. The bad block mapping table records the current repair progress of each bad block through a status field, including "Pending Repair," "Repairing," and "Repaired." "Pending Repair" indicates that the bad block has been detected but repair has not yet begun; "Repairing" indicates that the bad block is undergoing data synchronization; and "Repaired" indicates that the bad block has completed data correction and the mapping relationship can be removed.
[0073] The structure of the aforementioned bad block mapping table enables the storage system to uniformly manage the repair status and mapping relationships of multiple bad blocks. Each bad block entry independently records its logical address range, target logical address range, and repair status, providing a unified data index and status query basis for I / O redirection, dual-write operations, and incremental repair. Based on the status fields of each entry in the bad block mapping table, the storage system determines the list of bad blocks that need to be processed and performs corresponding processing operations for bad blocks in different statuses.
[0074] In this embodiment, a precise correspondence between bad blocks and the target volume is established by associating source volume record information and target volume record information. This correspondence provides a clear mapping between the data source and target location for I / O redirection, serving as the basis for subsequent read / write request redirection operations.
[0075] In this embodiment of the application, the operation S220, which stores the data to be written carried by the write request in the spare physical block in the source data volume and the target physical block corresponding to the target logical address range according to the first mapping relationship, may further include: determining the spare physical block from the spare block pool in the source data volume and storing the data to be written in the spare physical block; searching for the target logical address corresponding to the source logical address from the target logical address range according to the first mapping relationship; and storing the data to be written in the target physical block indicated by the target logical address.
[0076] In this embodiment, the spare block pool is a physical block resource pool reserved when the source data volume is created, used to store spare healthy physical blocks. Upon receiving a write request, an idle healthy physical block is determined from the spare block pool as the spare physical block. The data to be written, carried by the write request, is written to this spare physical block, completing the write operation on the source data volume side. Simultaneously, based on the correspondence recorded in the first mapping relationship, the target logical address corresponding to the source logical address is searched from the target logical address range.
[0077] Using the source logical address as an index, the corresponding target logical address is searched in the bad block mapping table. After determining the target logical address, the target logical address is converted into a target physical block address through the target volume's metadata mapping table. The data to be written is then written to the target physical block indicated by the target physical block address, completing the write operation on the target volume side.
[0078] In the embodiments of this application, Figure 3 This is a schematic diagram of the write request dual-write processing flow provided in an embodiment of this application. Figure 3 The process illustrates the data flow after a write request arrives, in which a double write operation is performed according to the first mapping relationship.
[0079] When the storage system receives a write request for any source logical address within the source logical address range, the write request carries the data to be written and the specific source logical address. Taking source logical address LBA 150 as an example, this LBA is located within the source logical address range LBA 100 to LBA 199, belonging to the logical address interval corresponding to detected bad blocks. The storage system queries the bad block mapping table based on LBA 150 to determine its corresponding source logical address range, and then obtains the first mapping relationship corresponding to that source logical address range.
[0080] Based on the first mapping, the write request data is routed to two target locations. The first target location is a spare physical block in the source data volume. This spare physical block is a healthy physical block dynamically allocated from the spare block pool of the source data volume, used to temporarily take over write data for bad blocks. The second target location is a target physical block in the target volume. This target physical block is located by converting the target logical address TLBA 150 recorded in the first mapping to a physical address using the target volume's metadata mapping table.
[0081] During a dual-write operation, the data to be written is simultaneously written to both the backup physical block and the target physical block. This dual-write mechanism ensures that the data requested for the write request has a healthy physical storage location on the source data volume side, while maintaining a correct data copy on the target volume side. After the dual-write operation is complete, the source logical address LBA 150 points to the backup physical block, enabling subsequent read requests targeting LBA 150 to read the latest written data from the backup physical block, guaranteeing that business read operations can access the data that was just written.
[0082] This dual-write process achieves the technical effect of uninterrupted business writes during bad block repair. Write requests are received and processed normally, and the data to be written is securely stored in two locations: a backup physical block and a target physical block, providing a data foundation for subsequent consistency comparisons and data repair. The entire dual-write process is completely transparent to the business.
[0083] For example, an idle physical block PBA 300 is selected from the spare block pool as a spare physical block, and the data to be written is written to PBA 300. Based on the first mapping relationship, the target logical address TLBA 150 corresponding to the source logical address LBA 150 is found. TLBA 150 is converted to the target physical block PBA 500 through the target volume's metadata mapping table, and the data to be written is written to PBA 500.
[0084] In this embodiment, a double-write operation for write requests is achieved by allocating a spare physical block from the spare block pool and finding the target physical block using a first mapping relationship. This process ensures that the data written by the service can be immediately read subsequently in the spare block, while retaining a correct data copy in the target block, providing a data foundation for subsequent consistency comparisons.
[0085] In this embodiment, a free physical block is determined from the spare block pool in the source data volume as a spare physical block; if there is no free physical block in the spare block pool, a free physical block is allocated from the preset storage pool of the source data volume as a spare physical block.
[0086] In this embodiment, the spare block pool is a set of physical blocks pre-allocated when the source data volume is created. A linked list of free blocks is maintained within the spare block pool, recording the addresses of all free physical blocks. When determining a spare physical block, a free physical block is taken from the head of the free block linked list as the spare physical block; this operation has constant time complexity. When all free physical blocks in the spare block pool are allocated, the free block linked list becomes empty. At this point, a new free physical block is requested from the global free space of the preset storage pool to which the source data volume belongs. The preset storage pool provides physical storage space for the source data volume and contains multiple physical storage devices. A space request is sent to the preset storage pool, which allocates a physical block from the global free space and returns it as a spare physical block.
[0087] For example, a free physical block PBA 300 is retrieved from the free block list of the spare block pool and used as a spare physical block. Once all free physical blocks in the spare block pool have been allocated, a space request is sent to the preset storage pool. The preset storage pool allocates physical block PBA 800 from the global free space and returns it as a spare physical block.
[0088] In this embodiment, a two-tier allocation mechanism—preferential allocation from the spare block pool combined with on-demand allocation from the global storage pool—ensures the continuous availability of spare physical blocks. This mechanism avoids the risk of write failures due to spare block exhaustion, while the reservation pool approach guarantees a fast response time for allocation under normal circumstances.
[0089] In this embodiment of the application, if it is determined that the spare physical block is unreadable, the data in the target physical block is stored in the spare physical block of the source data volume.
[0090] In this embodiment, "standby physical block unreadable" means that when attempting to read data from the standby physical block, the physical block returns an abnormal state such as read error, data verification failure, or read timeout. This indicates that although the write operation on the standby physical block returned success, the written data cannot be read normally. In this case, it is impossible to perform a consistency comparison between the data in the standby physical block and the data in the target physical block. The comparison operation is skipped, and data is directly read from the target physical block and written to the standby physical block of the source data volume.
[0091] For example, after a double write operation is completed, an attempt is made to read data from the spare physical block PBA 300 for a consistency comparison, but PBA 300 returns a read error. After confirming that PBA 300 is unreadable, data is read from the target physical block PBA 500 and then written to the spare physical block PBA 300.
[0092] In this embodiment, by skipping comparisons and directly copying, data repair can still proceed normally even when the backup physical block is unreadable. This process avoids the repair process being blocked due to abnormal backup physical blocks, ensuring the robustness of the repair process.
[0093] In this embodiment of the application, the operation S230 of comparing the data in the target physical block with the data in the backup physical block to obtain the comparison result may further include: calculating the first hash value of the data in the backup physical block and the second hash value of the data in the target physical block respectively; comparing the first hash value and the second hash value to obtain the comparison result.
[0094] In this embodiment, data stored in a spare physical block is read, and a hash algorithm is applied to this data to calculate a first hash value. Data stored in a target physical block is read, and the same hash algorithm is applied to the data in the target physical block to calculate a second hash value. The hash algorithm maps data of arbitrary length to hash values of fixed length. The same data content will produce the same hash value, and different data content will produce different hash values with a very high probability. The first hash value and the second hash value are compared. If the first hash value and the second hash value are the same, the comparison result is consistent; if the first hash value and the second hash value are different, the comparison result is inconsistent.
[0095] For example, reading 1MB of data stored in the spare physical block PBA 300 and calculating the first hash value using the Cyclic Redundancy Check (CRC) algorithm yields 0x3A. Reading 1MB of data stored in the target physical block PBA 500 and calculating the second hash value using the same algorithm yields 0x7F. Comparing 0x3A and 0x7F, they are different, resulting in an inconsistency.
[0096] In this embodiment, hash value comparison is used instead of full data comparison, significantly reducing the amount of data involved in the comparison operation and improving the efficiency of data consistency detection. The comparison result provides a basis for determining whether data repair needs to be performed.
[0097] In this embodiment of the application, the operation S230 of storing data from the target physical block to the spare physical block may further include: obtaining the physical address of the spare physical block; and storing the data read from the target physical block to the spare physical block according to the physical address.
[0098] In this embodiment, the physical block address of the spare physical block is read from the source volume record information of the bad block mapping table. This physical block address is the actual location number of the spare physical block allocated during the double-write operation in the physical storage medium. A write instruction is constructed based on this physical block address, and the data to be written is read from the target physical block. The read data is used as the write content, and a data write operation is initiated with the physical block address as the target location, writing the data to the physical storage unit corresponding to the physical block address.
[0099] For example, read the physical block address PBA 300 of the spare physical block from the bad block mapping table, and read data from the target physical block PBA500. Construct a write instruction to write the data to the physical memory unit corresponding to PBA 300.
[0100] In this embodiment, by obtaining the physical address of the backup physical block and performing data writing, accurate disk write of the repaired data is achieved. This process ensures that the data in the backup physical block is correctly corrected to the data in the target physical block.
[0101] In the embodiments of this application, Figure 4 This is a schematic diagram of the backup physical block processing flow provided in the embodiments of this application, such as... Figure 4 The process begins by determining whether the spare physical block is readable, and then proceeds to different processing paths based on the readability determination result.
[0102] Upon detection of at least one bad block in the source data volume, the data to be written, carried by the write request, is stored in both the spare physical block and the target physical block. After the data to be written is stored in the spare physical block, the storage system performs a read attempt on the spare physical block to determine if it is readable. The readable status of the spare physical block determines the subsequent data repair method.
[0103] If the spare physical block is determined to be unreadable, the process enters the unreadable branch. An unreadable spare physical block indicates that although the write operation returned success, the data in the spare physical block cannot be read normally, and therefore cannot be compared with the data in the target physical block. In this case, the storage system skips the consistency comparison operation, directly reads data from the target physical block, and stores the read data in the spare physical block of the source data volume. This direct copy method ensures that data repair can still proceed normally when the spare physical block is unreadable, avoiding blockage of the repair process due to spare physical block anomalies. After the direct copy is completed, the repair is complete, and the data in the spare physical block is consistent with the data in the target physical block.
[0104] If the spare physical block is determined to be readable, the process enters the readable branch. A readable spare physical block means the data stored in it can be read normally. The storage system reads data from the spare physical block and then from the target physical block, comparing the two to obtain the comparison result. If the comparison result is consistent, it means the data stored in the spare physical block is the same as the data stored in the target physical block, and the data in the spare physical block is already correct, requiring no additional data repair operations. In this case, the process terminates directly, marking the bad block repair as complete.
[0105] If the comparison results are inconsistent, it indicates that the data stored in the spare physical block is different from the data stored in the target physical block, and there is a deviation in the data in the spare physical block, requiring a data repair operation. In this case, the storage system reads the data stored in the target physical block and stores it in the spare physical block, completing the data correction. After data storage is completed, the data in the spare physical block is consistent with the data in the target physical block, and the repair is complete.
[0106] This process uses the readability of the spare physical block to implement differentiated processing for two scenarios: when the spare physical block is unreadable, it skips the comparison and copies directly; when the spare physical block is readable, it compares first and then copies as needed. This approach effectively addresses potential read anomalies that may occur after writing to the physical storage medium, improving the reliability and adaptability of the data recovery process.
[0107] In the embodiments of this application, if the data in the spare physical block of any bad block is the same as the data in the target physical block, or if the data in the target physical block has been written into the spare physical block, the repair status of the bad block is marked as repaired.
[0108] In this embodiment, the repair progress of each bad block in the bad block mapping table is monitored. When the comparison result of any bad block is consistent, that is, the data in the spare physical block is the same as the data in the target physical block, it is determined that the bad block does not need to perform data repair, and the repair status of the bad block is marked as repaired. When the comparison result of any bad block is inconsistent and the data in the target physical block has been successfully written to the spare physical block, it is determined that the bad block has completed data repair, and the repair status of the bad block is marked as repaired.
[0109] For example, if the data in the spare physical block PBA 300, which detects bad blocks LBA 100 to LBA 199, is consistent with the data in the target physical block PBA 500, the repair status of the bad block is marked as repaired.
[0110] In this embodiment, marking the repair status as "repaired" allows for accurate identification of the repair progress. This status provides a clear basis for subsequent determination of whether the mapping relationship can be removed, avoiding misunderstandings about removing the mapping relationship when the repair is incomplete.
[0111] In this embodiment of the application, when there is only one bad block, the repair status of the bad block is obtained; when the repair status of the bad block is repaired, the correspondence between the source logical address range and the target logical address range in the first mapping relationship is deleted.
[0112] In this embodiment, when only one bad block is recorded in the bad block mapping table, the repair status of that bad block is obtained. The current status identifier of the bad block is obtained by reading the status field in the bad block mapping table. When the repair status of the bad block is "repaired," the correspondence between the source logical address range and the target logical address range in the first mapping relationship is deleted. After deleting the correspondence, read / write requests for that bad block are no longer redirected, and subsequent read / write requests directly access the repaired spare physical block in the source data volume.
[0113] For example, the bad block mapping table only records bad blocks LBA 100 to LBA 199, and their repair status is read as repaired. The mapping relationship between LBA 100 to LBA 199 and TLBA 100 to TLBA 199 in the first mapping relationship is deleted.
[0114] In this embodiment, after a single bad block is repaired, the corresponding relationship is directly deleted, allowing the source data volume to resume normal read and write operations. This process simplifies the repair process in single bad block scenarios and avoids unnecessary state maintenance overhead.
[0115] In this embodiment of the application, during the process of writing data from the target physical block to the backup physical block, the repair status of any bad block is marked as repairing.
[0116] In this embodiment, when the write operation from the target physical block data to the spare physical block begins, the repair status of the bad block is updated from "Pending Repair" to "Repairing". This update is achieved by modifying the status field in the bad block mapping table. The "Repairing" status indicates that the bad block is currently undergoing data synchronization and has not yet been repaired. This status is maintained during the data write process until the data write operation is completed, at which point the status is updated to "Repaired".
[0117] For example, when writing data from the target physical block PBA 500 to the spare physical block PBA 300, the repair status of bad blocks LBA 100 to LBA 199 is updated from pending repair to repairing when the write operation begins.
[0118] In this embodiment, by marking the repair status, bad blocks that are being repaired and those that have not yet started repair can be distinguished. This status avoids the same bad block being repeatedly repaired, improving the accuracy of repair task scheduling.
[0119] In this embodiment of the application, the removal of the first mapping relationship in operation S230 may further include: when there are multiple bad blocks, obtaining the repair status of each bad block; when the repair status of multiple bad blocks is all repaired, deleting the correspondence between the source logical address range and the target logical address range in the first mapping relationship, and stopping the redirection of read and write requests for the source logical address to the target logical address.
[0120] In this embodiment, the number of bad blocks recorded in the bad block mapping table is detected. When the number of bad blocks is greater than or equal to two, a multi-bad block processing mode is entered. All bad block records in the bad block mapping table are traversed, and the repair status of each bad block is obtained one by one. It is determined whether the repair status of all bad blocks is "repaired". If any bad block's repair status is not "repaired", the system continues to wait for repair completion. When the repair status of all bad blocks is "repaired", the correspondence between all source logical address ranges and target logical address ranges in the first mapping relationship is deleted. Simultaneously, redirection operations are stopped for read / write requests targeting these source logical addresses, and subsequent read / write requests directly access the corresponding spare physical blocks in the source data volume.
[0121] For example, the bad block mapping table records two bad blocks, bad block A and bad block B. The repair status of bad block A and bad block B is set to "repaired". The mapping between the source logical address range and the target logical address range for bad block A and bad block B in the first mapping relationship is deleted, and all requests targeting these source logical addresses are stopped from being redirected to the target volume.
[0122] In this embodiment, in a multi-bad-block scenario, the first mapping relationship is uniformly removed after all bad blocks have been repaired. This process avoids misdiagnosis when some bad blocks are not repaired, which could lead to confusion in the repair process and ensures the correctness and completeness of the repair process in a multi-bad-block scenario.
[0123] In this embodiment, when the hash value comparison result is inconsistent, the storage system reads data from the storage location corresponding to the target physical address and writes the read data to the storage location corresponding to the backup physical address. This process corrects the data in the backup physical block, making the data in the backup physical block consistent with the data in the target physical block.
[0124] During the data writing process described above, the storage system updates the repair progress information corresponding to the current source logical address range in the bad block mapping table in real time. The repair progress information is recorded as a percentage, gradually increasing from 0% to 100%. When the repair progress reaches 100%, it indicates that all data blocks within that source logical address range have been corrected.
[0125] By updating repair progress information in real time, the storage system can monitor and manage the repair status of each bad block, making the repair process traceable and controllable. In scenarios where multiple bad blocks are repaired in parallel, the repair progress information provides a basis for scheduling and prioritizing repair tasks, enabling the storage system to dynamically adjust resource allocation based on the repair progress of each bad block, ensuring the efficient progress of the overall repair process.
[0126] In this embodiment of the application, before deleting the mapping relationship between the source logical address range and the target logical address range, the method may further include: calculating the source hash value of the source data volume and the target hash value of the target volume respectively; and deleting the correspondence between the source logical address range and the target logical address range if the source hash value and the target hash value are the same.
[0127] In this embodiment, after all bad blocks have been repaired, a full hash verification is performed on the source and target volumes before deleting the first mapping relationship. The contents of all data blocks in the source and target volumes are read respectively, and the same hash algorithm is applied to both sets of data to calculate the source and target hash values. The source and target hash values are compared. If they are the same, it indicates that the data in the source and target volumes is completely consistent, and the correspondence between the source and target logical address ranges in the first mapping relationship is deleted. If they are different, it indicates that there are unrepaired data differences, the first mapping relationship remains unchanged, and the repair operation continues until the full hash verification passes.
[0128] For example, if the full hash value of the source data volume is 0xA5 and the full hash value of the target volume is also 0xA5, then delete the correspondence between the source logical address range and the target logical address range in the first mapping relationship.
[0129] In this embodiment, a full hash verification is used to finally confirm data consistency before deleting the mapping relationship. This verification effectively avoids data inconsistencies caused by omissions or errors in the repair process, enhancing the reliability of the repair completion.
[0130] In the embodiments of this application, Figure 5 This is a schematic diagram illustrating the overall data repair process provided in an embodiment of this application. Figure 5 As shown, this process illustrates the main steps from the arrival of a read / write request to the completion of data repair.
[0131] When the storage system receives a read / write request for the source data volume, the process begins. The storage system first determines the logical address accessed by the read / write request and then queries the bad block mapping table to determine if the logical address falls within the range of source logical addresses recorded in the bad block mapping table. If the logical address does not fall within the range of source logical addresses recorded in the bad block mapping table, it indicates that the physical storage location corresponding to the logical address is in a healthy state, and the storage system directly performs read / write operations on the source data volume according to the normal read / write process.
[0132] If the logical address falls within the source logical address range recorded in the bad block mapping table, it indicates that the physical storage location corresponding to the logical address is a bad block, and the storage system enters the bad block handling process. At this time, the storage system distributes read and write requests according to the mapping relationship recorded in the bad block mapping table. For read requests, the storage system directs the read request to the corresponding target logical address in the target volume, reads data from the physical address corresponding to the target logical address, and returns it. For write requests, the storage system stores the data to be written carried by the write request in a spare physical block in the source data volume and a target physical block corresponding to the target logical address range, respectively, achieving dual write.
[0133] After a write request completes double-write, the source logical address points to the spare physical block, allowing subsequent read requests targeting that logical address to retrieve data from the spare physical block. In the background, the storage system initiates an incremental repair process. Incremental repair is performed at the data granularity level, which is the smallest unit of data management in the storage system, i.e., the smallest data block size involved in each read / write operation, such as 4 kilobytes or 8 kilobytes. The incremental repair process processes the range of logical addresses to be repaired recorded in the bad block mapping table one by one according to the data granularity. For each data block corresponding to a data granularity, the storage system determines whether the spare physical block is readable. If the spare physical block is readable, the data in the target physical block is compared with the data in the spare physical block. If the comparison result is inconsistent, the data in the target physical block is stored in the spare physical block to complete the repair for that data granularity. After repair is complete, the storage system removes the first mapping relationship and restores normal read / write operations to the source data volume.
[0134] The overall process achieves precise routing of read and write requests through a bad block mapping table, enabling read requests to return data normally during bad block repair and write requests to write data normally. At the same time, the background incremental repair gradually completes the data correction of all bad blocks according to the data granularity, achieving zero business interruption during the bad block repair process.
[0135] In the embodiments of this application, the above process achieves multiple technical effects, including uninterrupted repair mechanism, efficient utilization of streamlined volumes, incremental synchronization algorithm, and data consistency guarantee.
[0136] Regarding the uninterrupted repair mechanism, real-time input / output redirection seamlessly switches read requests for bad blocks to the target volume, while write requests employ a dual-write mechanism to ensure continuous business operation during the repair process, achieving zero interruption. For efficient utilization of thin volumes, leveraging the thin-volume characteristics of the target volume, only the data blocks corresponding to bad blocks are synchronized, avoiding full copying and reducing storage space usage and repair time. Regarding the incremental synchronization algorithm, incremental synchronization based on data block hash value comparison only transmits changed data blocks, significantly improving repair efficiency, especially suitable for large-size original volumes. Regarding data consistency assurance, the dual-write mechanism for write operations ensures data consistency during the repair process; after repair, data integrity can be verified through global hash checking, reducing the risk of data loss to zero.
[0137] Furthermore, the steps in this embodiment collectively achieve the function of disaster recovery service data repair. For scenarios involving damaged blocks in the original volume, rapid repair without service interruption is achieved, ensuring the continuity of disaster recovery services and data reliability. Through input / output redirection and incremental repair, the problems of service interruption, low repair efficiency, and data inconsistency in traditional methods are solved.
[0138] At the design level, the input / output redirection design captures read / write requests for bad blocks through an input / output interception module. Read requests are redirected to the target volume, while write requests employ a dual-write mechanism. Business operations are uninterrupted, and data access latency increases by no more than 5%, ensuring business continuity. The incremental repair design uses hash value comparison for incremental synchronization, synchronizing only the differing data blocks. This reduces repair time by more than 80% compared to traditional full-copy methods and decreases storage space usage by 90%. The data consistency design achieves 100% data consistency during repair through a dual-write mechanism and global hash verification, resulting in zero data loss after repair.
[0139] Based on the above data repair method, this application also provides a data repair apparatus. The following will be combined with... Figure 6 The device is described in detail.
[0140] Figure 6 This is a structural block diagram of a data repair device according to an embodiment of this application.
[0141] like Figure 6 As shown, the data repair device 600 of this embodiment includes a bad block detection module 610, a write request processing module 620, and a data repair module 630.
[0142] The bad block detection module 610 is configured to, in response to detecting at least one bad block in the source data volume, determine the source logical address range of the at least one bad block, and establish a first mapping relationship between the source logical address range and a target logical address range in the target volume, wherein the target volume is a snapshot volume of the source data volume prior to the detection of the bad block. In one embodiment, the bad block detection module 610 may be used to perform the operation S210 described above, which will not be repeated here.
[0143] The write request processing module 620, in response to receiving a write request for any source logical address within the source logical address range, stores the data to be written carried in the write request in a spare physical block in the source data volume and a target physical block corresponding to the target logical address range, respectively, according to a first mapping relationship, and points the source logical address to the spare physical block. In one embodiment, the write request processing module 620 can be used to perform the operation S220 described above, which will not be repeated here.
[0144] The data repair module 630 is used to compare the data in the target physical block with the data in the backup physical block when the backup physical block is determined to be readable, and obtain a comparison result; if the comparison result is inconsistent, the data in the target physical block is stored in the backup physical block, and the first mapping relationship is released. In one embodiment, the data repair module 630 can be used to perform the operation S230 described above, which will not be repeated here.
[0145] It should be noted that the description of the features in the embodiment corresponding to the data repair device can be found in the relevant description of the embodiment corresponding to the data repair method, and will not be repeated here.
[0146] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above-described data restoration method embodiments.
[0147] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described data repair method embodiments at runtime.
[0148] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0149] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described data repair method embodiments.
[0150] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described data repair method embodiments.
[0151] Any of the components, modules, units, parts, methods, and operations described herein can be implemented using software, firmware, hardware (e.g., fixed logic circuitry), manual processing, or any combination thereof. Alternatively or additionally, any functionality described herein can be executed at least in part by one or more hardware logic components, such as, but not limited to, a central processing unit (CPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), an application-specific standard product (ASSP), a system-on-a-chip (SoC), a complex programmable logic device (CPLD), a microprocessor (MCU), etc. The terms "system," "computing device," or "apparatus" as used herein encompass various means, devices, and machines for processing data, including, for example, one or more programmable processors, computers, SoCs, or combinations thereof. The apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or one or more combinations thereof. The aforementioned computer program (also known as a program, software, software application, app, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as a standalone program or as a module, component, subroutine, object, or other unit suitable for a computing environment.
[0152] The units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.
[0153] The data recovery method and electronic device provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make several improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of this application.
Claims
1. A data repair method, characterized in that, The method includes: In response to detecting at least one bad block in a source data volume, the source logical address range of the at least one bad block is determined, and a first mapping relationship is established between the source logical address range and a target logical address range in a target volume, wherein the target volume is a snapshot volume of the source data volume before the bad block was detected; In response to receiving a write request for any source logical address within the source logical address range, the data to be written carried by the write request is stored in a spare physical block in the source data volume and a target physical block corresponding to the target logical address range, respectively, according to the first mapping relationship, and the source logical address is pointed to the spare physical block; If the backup physical block is determined to be readable, the data in the target physical block is compared with the data in the backup physical block to obtain a comparison result; if the comparison result is inconsistent, the data in the target physical block is stored in the backup physical block, and the first mapping relationship is released.
2. The method according to claim 1, characterized in that, The method further includes: In response to receiving a read request for any source logical address within the source logical address range, the read request is directed to a target logical address within the target logical address range according to the mapping relationship, and data is read from the target physical block corresponding to the target logical address.
3. The method according to claim 2, characterized in that, The step of directing the read request to a target logical address within the target logical address range according to the mapping relationship includes: Based on the mapping relationship, find the target logical address corresponding to the source logical address from the target logical address range; The read request is forwarded to the target logical address.
4. The method according to claim 1, characterized in that, The method further includes: In response to receiving a snapshot creation instruction, a snapshot operation is performed on the source data volume to generate the target volume; the snapshot creation instruction is triggered according to a preset time strategy or in response to manual operation of the object.
5. The method according to claim 1, characterized in that, In response to detecting at least one bad block in the source data volume, determining the source logical address range of the at least one bad block includes: Read and write tests are performed on each physical block of the source data volume. If an error occurs during the read test of any physical block, or if a write failure occurs during the write test of any physical block, the physical block is determined to be a bad block. Obtain the physical block address of at least one of the bad blocks, and determine the source logical address range corresponding to the physical block address based on preset metadata. The preset metadata includes a second mapping relationship between each source logical address and physical address in the source data volume.
6. The method according to claim 1, characterized in that, The step of establishing a first mapping relationship between the source logical address range and the target logical address range in the target volume includes: The source logical address range of the at least one bad block and the physical block address corresponding to the source logical address range are recorded as source volume record information; Record the target logical address range in the target volume that corresponds to the source logical address range as target volume record information; The first mapping relationship is established by associating the source volume record information and the target volume record information.
7. The method according to claim 1, characterized in that, The step of storing the data to be written carried by the write request in the spare physical block in the source data volume and the target physical block corresponding to the target logical address range according to the first mapping relationship includes: A spare physical block is determined from the spare block pool in the source data volume, and the data to be written is stored in the spare physical block; Based on the first mapping relationship, find the target logical address corresponding to the source logical address from the target logical address range; The data to be written is stored in the target physical block indicated by the target logical address.
8. The method according to claim 7, characterized in that, The step of determining spare physical blocks from the spare block pool in the source data volume includes: Idle physical blocks are determined from the spare block pool in the source data volume as the spare physical blocks; If there are no free physical blocks in the spare block pool, a free physical block is allocated from the preset storage pool of the source data volume as the spare physical block.
9. The method according to claim 1, characterized in that, The method further includes: If the spare physical block is determined to be unreadable, the data in the target physical block is stored in the spare physical block of the source data volume.
10. The method according to claim 1, characterized in that, The step of comparing the data in the target physical block with the data in the spare physical block to obtain the comparison result includes: Calculate the first hash value of the data in the spare physical block and the second hash value of the data in the target physical block, respectively; The first hash value and the second hash value are compared to obtain the comparison result.
11. The method according to claim 1, characterized in that, The step of storing the data in the target physical block in the spare physical block includes: Obtain the physical address of the spare physical block; The data read from the target physical block is stored in the spare physical block according to the physical address.
12. The method according to claim 1, characterized in that, The method further includes: If the data in the spare physical block of any bad block is the same as the data in the target physical block, or if the data in the target physical block has been written into the spare physical block, the repair status of the bad block is marked as repaired.
13. The method according to claim 12, characterized in that, The method further includes: If the bad block is a single block, obtain the repair status of the bad block; If the bad block is repaired, delete the correspondence between the source logical address range and the target logical address range in the first mapping relationship.
14. The method according to claim 12, characterized in that, The method further includes: During the process of writing data from the target physical block to the backup physical block, the repair status of any bad block is marked as repairing.
15. The method according to claim 12, characterized in that, The process of releasing the first mapping relationship includes: In the case where there are multiple bad blocks, the repair status of each bad block is obtained; If the repair status of multiple bad blocks is "repaired", delete the correspondence between the source logical address range and the target logical address range in the first mapping relationship, and stop redirecting read and write requests for the source logical address to the target logical address.
16. The method according to claim 15, characterized in that, Before deleting the mapping relationship between the source logical address range and the target logical address range, the method further includes: Calculate the source hash value of the source data volume and the target hash value of the target volume, respectively; If the source hash value and the target hash value are the same, delete the correspondence between the source logical address range and the target logical address range.
17. A data repair device, characterized in that, The device includes: A bad block detection module is configured to, in response to detecting at least one bad block in a source data volume, determine the source logical address range of the at least one bad block, and establish a first mapping relationship between the source logical address range and a target logical address range in a target volume, wherein the target volume is a snapshot volume of the source data volume before the bad block was detected; A write request processing module is used to respond to receiving a write request for any source logical address within the source logical address range, and according to the first mapping relationship, store the data to be written carried by the write request in a spare physical block in the source data volume and a target physical block corresponding to the target logical address range, respectively, and point the source logical address to the spare physical block; The data repair module is used to compare the data in the target physical block with the data in the backup physical block when the backup physical block is determined to be readable, and obtain a comparison result; if the comparison result is inconsistent, the data in the target physical block is stored in the backup physical block, and the first mapping relationship is released.
18. An electronic device comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 16.
19. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 16.
20. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 16.
Citation Information
Patent Citations
Data reading / writing method and apparatus for storage device
CN107122261A
Data storage method and system for embedded device
CN121879693A