Inspection and repair methods for mirrored volumes, electronic devices, storage media and products
By generating and distributing task requests to the home nodes of each domain, the problem of long processing paths for unowned mirror volume inspection and repair tasks is solved, improving execution efficiency and the ability to cope with node failures.
Patent Information
- Application Number
- CN202511419183.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-09-30
AI Technical Summary
In the existing technology, the inspection and repair task of unowned mirror volumes has a long processing path and low execution efficiency, and it is impossible to start the inspection and repair task on other ownership nodes besides the ownership node to which the preset starting logical block address belongs.
By generating task requests and distributing them to the home nodes of each domain, all home nodes of the unowned mirror volume can initiate inspection and repair tasks, shortening the processing path and improving efficiency.
This solves the problem that unowned mirror volumes cannot start inspection and repair tasks on nodes other than the node to which the preset starting logical block address belongs, shortens the processing path of inspection and repair tasks, and improves the ability to cope with node failures and execution efficiency.
Smart Images

Figure CN120892263B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of storage technology, and in particular to methods for inspecting and repairing mirrored volumes, electronic devices, storage media, and products. Background Technology
[0002] With the development of cloud computing and storage technologies, mirrored volume technology refers to a virtual disk corresponding to two copy volumes. When the host performs input / output writes, both copy volumes need to be written, and a synchronization bitmap of the mirrored volume is used to mark the data differences between the two copy volumes. However, if a data block failure causes data errors on one copy volume, resulting in data differences between the two copy volumes, the synchronization bitmap will not be bound to the copy volume. In this case, it is necessary to perform inspection and repair operations on the mirrored volume to restore data synchronization between the two copy volumes.
[0003] In related technologies, for mirrored volumes with ownership, inspection and repair tasks are typically initiated on the ownership node of its ownership domain to perform inspection and repair operations on the mirrored volume with ownership. For mirrored volumes without ownership, inspection and repair tasks are initiated from the ownership node of the specified starting logical address. However, when processing inspection and repair tasks for storage blocks corresponding to logical block addresses belonging to other nodes in a mirrored volume without ownership, it is necessary to forward the inspection and repair tasks from the ownership node of the specified starting logical address to the other node to which the currently processed logical block address belongs. This makes the processing path of the inspection and repair tasks for the mirrored volume longer, the time for performing inspection and repair operations longer, and the efficiency of performing inspection and repair operations low. Summary of the Invention
[0004] This application provides a method, electronic device, storage medium, and product for inspecting and repairing mirrored volumes, in order to at least solve the problems in related technologies such as long processing paths, long execution times, and low efficiency of inspection and repair operations for inspection and repair tasks corresponding to unowned mirrored volumes.
[0005] This application provides a method for inspecting and repairing mirrored volumes, applied to the control node of a first input / output group in a cluster; the first input / output group corresponds to a first number of domains; the method includes: receiving a volume repair command issued by a client, the volume repair command instructing an inspection and repair operation to be performed on a secondary copy volume in the mirrored volume corresponding to the target virtual disk, the volume repair command carrying a preset starting logical block address of the storage block to be repaired in the secondary copy volume, the secondary copy volume being divided into multiple storage units, and each storage unit being divided into multiple logical blocks; obtaining the storage unit size of a single storage unit, the logical block size of a single logical block, and the logical address range size corresponding to each domain; when the mirrored volume is an unowned mirrored volume, generating a task request based on the preset starting logical block address, logical block size, storage unit size, logical address range size, and the first number; distributing the task request to the owner node of each domain in the first input / output group, so that each owner node performs an inspection and repair operation on the storage block to be repaired in its respective domain based on the task request.
[0006] This application also provides a mirror volume inspection and repair device, applied to the control node of a first input / output group in a cluster; the first input / output group corresponds to a first number of domains; the device includes: a transceiver module, used to receive a volume repair command sent by a client, the volume repair command instructing an inspection and repair operation to be performed on a secondary copy volume in the mirror volume corresponding to the target virtual disk, the volume repair command carrying a preset starting logical block address of the storage block to be repaired in the secondary copy volume, the secondary copy volume being divided into multiple storage units, each storage unit being divided into multiple logical blocks; obtaining the storage unit size of a single storage unit, the logical block size of a single logical block, and the logical address range size corresponding to each domain; a processing module, used to generate a task request based on the preset starting logical block address, logical block size, storage unit size, logical address range size, and the first number when the mirror volume is an unowned mirror volume; the transceiver module is also used to distribute the task request to the owner node of each domain in the first input / output group, so that each owner node performs an inspection and repair operation on the storage block to be repaired in its respective domain based on the task request.
[0007] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of any of the above-described methods for inspecting and repairing mirrored volumes.
[0008] This application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of any of the above-described methods for inspecting and repairing mirrored volumes.
[0009] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described image volume inspection and repair methods.
[0010] This application enables all home nodes of a non-homed mirror volume to initiate inspection and repair tasks and perform inspection and repair operations, as task requests generated for non-homed mirror volumes can be sent to the home nodes of each domain. This solves the problem that non-homed mirror volumes cannot initiate inspection and repair tasks on home nodes other than the home node of the preset starting logical block address, shortens the processing path of inspection and repair tasks for mirror volumes, improves the ability to handle node failures, reduces the time for performing inspection and repair operations, and improves the efficiency of performing inspection and repair operations. Attached Figure Description
[0011] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 The flowchart for performing inspection and repair operations on a mirrored volume is provided in the embodiments of this application.
[0013] Figure 2 A topology diagram of a mirror volume inspection and repair system provided in this application embodiment;
[0014] Figure 3 A flowchart illustrating a method for inspecting and repairing a mirrored volume, provided in an embodiment of this application;
[0015] Figure 4 A flowchart illustrating another method for inspecting and repairing a mirrored volume provided in this application embodiment;
[0016] Figure 5 A flowchart illustrating another method for inspecting and repairing a mirrored volume provided in this application embodiment;
[0017] Figure 6 A structural block diagram of a mirror volume inspection and repair device provided in this application embodiment;
[0018] Figure 7 This is a structural block diagram of a mirrored volume inspection and repair device provided in an embodiment of this application. Detailed Implementation
[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0020] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0021] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0022] This application embodiment is applied to a scenario where, due to data errors on the secondary copy volume of a mirrored volume caused by a data block failure, data differences exist between the primary copy volume and the secondary copy volume of the mirrored volume. The scenario involves inspecting and repairing the secondary copy volume.
[0023] A mirrored volume consists of two copy volumes: one as the primary copy volume and the other as the secondary copy volume. When the data between the two copy volumes of the mirrored volume is not synchronized, the synchronization bitmap of the volume mirror is bound to the secondary copy volume, recording the positions on the secondary copy volume where data synchronization has not been completed. After the two copy volumes have completed data synchronization, the bitmap is unbound from the secondary copy volume.
[0024] However, if a data error on one copy volume is caused by a data block failure, resulting in data discrepancies between two copy volumes, the mirror volume's synchronization bitmap will not be bound to any copy volume and will not record the differences because this is not triggered by an input / output process. Therefore, in this case, data discrepancies between the mirror volumes require inspection of the mirror volume to repair, allowing the data on the two copy volumes to resynchronize and correct the data discrepancies caused by the data block corruption.
[0025] A mirrored volume with a home domain contains multiple logical blocks that belong to the same domain; a mirrored volume without a home domain contains multiple logical blocks that belong to multiple different domains.
[0026] In related technologies, such as Figure 1 As shown, Figure 1The flowchart for performing inspection and repair operations on a mirrored volume is provided in the embodiments of this application; Figure 1 In the cluster, the Control State Matching (CSM) deployed on the control node of the first input / output group receives the volume repair command (repairvdiskcopy command) issued by the client for the mirror volume corresponding to the virtual disk; it starts the state machine and sets the virtual disk to repair state; it mobilizes the agents of all nodes to issue volume repair tasks; each node agent determines whether it is the owner node of the preset starting logical block address. If it is the owner node of the preset starting logical block address, the agent of that node starts processing the repair task from the specified preset starting logical block address, and after repairing the logical block address belonging to its own node, it forwards the processing to the next node. After processing the last storage block to be repaired in the volume, it reports that the repair is complete; the control state machine changes the virtual disk to the normal state. If none of them are the owner nodes of the preset starting logical block address, the repair task is not executed, and the operation ends.
[0027] In the aforementioned related technologies, for mirrored volumes with ownership, the inspection and repair task for the mirrored volume with ownership is initiated on the owner node of its ownership domain; for mirrored volumes without ownership, the inspection and repair task for the mirrored volume without ownership is initiated from the owner node of the specified starting logical address. However, when processing the inspection and repair task for the storage blocks corresponding to the logical block addresses belonging to other nodes in the mirrored volume without ownership, it is necessary to forward the task from the owner node of the specified starting logical address to the other node to which the logical block address being processed belongs. This makes the processing path of the inspection and repair task corresponding to the mirrored volume longer, reduces the ability to cope with node failures, increases the time for performing inspection and repair operations, and reduces the efficiency of performing inspection and repair operations.
[0028] To address the aforementioned technical problems, embodiments of this application provide a method for inspecting and repairing mirrored volumes. This method determines the starting logical block addresses of unowned mirrored volumes in different domains and adds marker information to determine whether logical block addresses belonging to the current domain have been inspected and repaired. This enables all domain-owning nodes of the unowned mirrored volume to initiate inspection and repair tasks and individually notify the control node that all logical block addresses within the current domain have been inspected and repaired. This solves the problem that unowned mirrored volumes cannot initiate inspection and repair tasks on nodes other than the one to which the preset starting logical block address belongs, shortens the processing path of the inspection and repair tasks corresponding to the mirrored volume, improves the ability to handle node failures, reduces the time required to perform inspection and repair operations, and improves the efficiency of performing inspection and repair operations.
[0029] The following is based on Figure 2 The method provided in this application embodiment is described using the mirrored volume inspection and repair system shown as an example.
[0030] like Figure 2 As shown, Figure 2 This is a topology diagram of a mirror volume inspection and repair system provided in an embodiment of this application. Figure 2 In the system, the mirror volume inspection and repair system 200 includes a control node 201, a client 202, a first home node 203, a first backup node 204, a second home node 205, a second backup node 206, a third home node 207, a third backup node 208, a fourth home node 209, and a fourth backup node 210.
[0031] Control node 201 is the control node for the first input / output group in the cluster. The cluster includes multiple input / output groups. The first input / output group is any one of the multiple input / output groups.
[0032] The first input / output group includes a first number of domains. The first number can be set according to actual needs, and the first number can be 4.
[0033] Each domain consists of one Owner node and one Backup node. The Owner node serves as the primary processing node for transactions and I / O operations within the Domain. During mirror volume copy synchronization operations, the Domain's Owner node initiates the mirror volume copy synchronization process for all volumes belonging to that Domain.
[0034] The first home node 203 and the first backup node 204 belong to the first domain (Domain0); the second home node 205 and the second backup node 206 belong to the second domain (Domain1); the third home node 207 and the third backup node 208 belong to the third domain (Domain2); and the fourth home node 209 and the fourth backup node 210 belong to the fourth domain (Domain3).
[0035] Client 202 can be any device with display and communication capabilities. Client 202 can be a command-line client, used to communicate with control node 201 via command-line input instructions.
[0036] Figure 2 The mirrored volume inspection and repair system shown is for illustrative purposes only and is not intended to limit the technical solutions of this application. Those skilled in the art should understand that in specific implementations, the mirrored volume inspection and repair system may include more nodes, without limitation.
[0037] The embodiments of this application provide a method for inspecting and repairing mirrored volumes, applicable to... Figure 2 The control nodes shown are as follows: Figure 3 As shown, Figure 3 The flowchart illustrates a method for inspecting and repairing a mirrored volume, as provided in this embodiment of the application. The method includes the following steps:
[0038] S301 receives the volume repair command issued by the client.
[0039] The volume repair command instructs the execution of a check and repair operation on the copy volume of the mirrored volume corresponding to the target virtual disk. The volume repair command carries the preset starting logical block address of the storage blocks to be repaired in the copy volume. Optionally, the volume repair command may also carry the target virtual disk identifier.
[0040] The inspection and repair operations can compare the data of the primary and secondary copy volumes to check if the data in the secondary copy is completely consistent with the primary copy. If data misalignment, loss, or other deviations occur, the data content of the secondary copy volume will be restored to match the primary copy volume, ensuring the correctness of the data copy. It can also scan the metadata of the secondary copy volume to check for issues such as damaged, missing, or incorrectly labeled metadata records. Once metadata anomalies are detected, repair tools or relevant commands are used to correct the metadata information, restoring its integrity and accuracy. Furthermore, it can check the status of the underlying storage media on which the secondary copy volume depends, such as checking for bad sectors on the disk physical sectors and the stability of the storage device hardware connections. If a storage media failure risk or error is detected, it will attempt to reread and rewrite damaged sectors or migrate the secondary copy volume to a healthy storage location.
[0041] The secondary copy volume is divided into multiple storage units (Segments), and each storage unit is divided into multiple logical blocks.
[0042] The default starting logical block address (LBA) can be the starting LBA of the storage block to be repaired in the secondary copy volume.
[0043] For example, the client sends a volume repair command to the control node via a command-line tool. The control node receives the volume repair command from the client.
[0044] Optionally, the control node locates the stored data structure of the target virtual disk based on the target virtual disk identifier in the volume repair command. This data structure can be used as an array data member, divided by domain, to record the start and end LBA information for each domain of the target virtual disk during the repair task.
[0045] S302, obtain the storage cell size of a single storage unit, the logical block size of a single logical block, and the logical address range size corresponding to each field.
[0046] Each field corresponds to the same logical address range.
[0047] For example, the control node obtains the storage unit size of a single storage unit, the logical block size of a single logical block, and the logical address range size corresponding to each field from the data results by looking up the data structure of the target virtual disk.
[0048] For example, the size of a single storage unit can be 32MB. The size of a single logical block can be 512 bytes. The size of the logical address range corresponding to each field can be 0x10000.
[0049] S303, when the mirrored volume is an unowned mirrored volume, generate a task request based on the preset starting logical block address, logical block size, storage unit size, logical address range size and first quantity.
[0050] The task request is used to instruct each node to perform inspection and repair operations on the storage blocks to be repaired corresponding to the unowned mirror volume.
[0051] For example, when the mirrored volume is an unowned mirrored volume, the control node generates a task request based on the preset starting logical block address, logical block size, storage unit size, logical address range size, and a first quantity.
[0052] In some optional implementations, when the volume repair command also carries the preset end logical block address of the storage block to be repaired in the secondary copy volume, the control node can also generate a task request based on the start logical block address and the preset end logical block address of the corresponding storage block to be repaired in each domain.
[0053] Understandably, by using the start and end logical block addresses, the range of storage blocks to be repaired within each domain can be precisely defined, preventing repair operations from exceeding the target range. This precision reduces invalid scanning and repair actions, allowing each task request to focus on a specific area to be repaired, thereby improving overall repair efficiency and shortening repair time.
[0054] S304, Distribute the task request to the home node of each domain in the first input / output group, so that each home node can perform inspection and repair operations on the storage blocks to be repaired in its respective domain based on the task request.
[0055] For example, the CSM call function of the control node distributes the task request to the agent end of each node in each domain of the first input / output group; the backup node of each domain does not process the task request upon receiving it; the home node of each domain selects and sets its own starting LBA upon receiving the task request. The agent end of each home node, according to its own starting LBA set within the domain, starts and executes all background tasks belonging to its own domain corresponding to the inspection and repair operation.
[0056] Optionally, when a home node can be the home node of multiple domains, the home node can sequentially set itself as the starting LBA for performing inspection and recovery operations for the domain corresponding to the owner node.
[0057] Furthermore, after the Agent of each home node completes the inspection and repair operation on the storage blocks to be repaired in the domain based on the task request, the control node can also receive the response information returned from each home node in sequence; when the response information is received, the state machine is set to the preset state and the preset state is returned to the client.
[0058] Each response message is used to instruct each home node to perform inspection and repair operations on the storage blocks to be repaired within the domain based on the task request.
[0059] The default status is used to indicate that the secondary copy volume has completed the inspection and repair operation.
[0060] In some optional embodiments, when distributing task requests to the home nodes of each domain, the sending time of the task request for each home node is recorded; during the process of sequentially receiving response information returned from each home node, the time interval from the sending of the task request to the receiving of the response information is monitored in real time; if the time interval corresponding to the current home node is detected to be greater than or equal to a preset time, a task status query request is sent to the current home node; the query result generated by the current home node based on the task status query request is received; if the query result indicates that the current home node is in a first state, then the wait for the current home node to return the response information continues; or, if the query result indicates that the current home node is in a second state, then the task request is resent to the current home node until the current home node returns the response information.
[0061] The task status query request is used to obtain the node status of the currently owned node.
[0062] The first state indicates that the currently owned node is performing inspection and repair operations.
[0063] The second state is used to indicate that the currently owned node has interrupted the inspection and repair operation.
[0064] Understandably, by monitoring the response time interval in real time, tasks that may time out can be detected promptly. For nodes in the first state, continuing to wait avoids unnecessary task resending and ensures the continuity of the original task processing; for nodes in the second state, timely resending of the task can quickly restore the processing flow and reduce the probability of task interruption due to single point of failure.
[0065] based on Figure 3 The method shown allows the control node to receive a volume repair command from the client for the secondary copy volume in the mirror volume corresponding to the target virtual disk; obtain the storage unit size of a single storage unit, the logical block size of a single logical block, and the logical address range size corresponding to each domain of the secondary copy volume; when the mirror volume is an unowned mirror volume, generate a task request based on the preset starting logical block address, logical block size, storage unit size, logical address range size, and a first quantity; and distribute the task request to the ownership node of each domain in the first input / output group, so that each ownership node can perform inspection and repair operations on the storage blocks to be repaired in its respective domain based on the task request.
[0066] Because task requests generated for unowned mirrored volumes can be sent to the home nodes of each domain, all home nodes of the unowned mirrored volume in all domains can initiate inspection and repair tasks and perform inspection and repair operations. This solves the problem that unowned mirrored volumes cannot start inspection and repair tasks on home nodes other than the home node to which the preset starting logical block address belongs, shortens the processing path of the inspection and repair tasks corresponding to the mirrored volume, improves the ability to cope with node failures, reduces the time for executing inspection and repair operations, and improves the efficiency of executing inspection and repair operations.
[0067] In an optional example, based on the foregoing embodiments and as described above, when the mirrored volume is an unowned mirrored volume, a task request is generated according to the preset starting logical block address, logical block size, storage unit size, logical address range size, and a first quantity, as follows: See the detailed steps below. Figure 4 As shown, Figure 4 A flowchart illustrating another method for inspecting and repairing mirrored volumes provided in this application embodiment includes:
[0068] S401, when the mirror volume is an unowned mirror volume, determine the starting logical block address of the corresponding storage block to be repaired in each domain according to the preset starting logical block address, logical block size, storage unit size, logical address range size and first quantity.
[0069] In some optional implementations, the control node determines the sorting order of the target storage unit to which the preset starting logical block address belongs among multiple storage units based on the preset starting logical block address, logical block size, and storage unit size; determines the sorting order of the target domain to which the target storage unit belongs among a first number of domains based on the sorting order of the target storage unit among multiple storage units and a first number; and determines the starting logical block address of the corresponding storage block to be repaired in each domain based on the preset starting logical block address, the sorting order of the target domain among the first number of domains, the first number, and the logical address range size.
[0070] In one example, the control node determines the sorting order of the target storage unit to which the preset starting logical block address belongs among multiple storage units based on the preset starting logical block address, logical block size, and storage unit size, calculated using the following expression:
[0071]
[0072] in, Indicates the preset starting logical block address; Indicates the size of the logical block; D represents the size of the storage unit; D represents the sorting order of the target storage unit among multiple storage units.
[0073] In one example, the control node determines the sorting order of the target domain to which the target storage unit belongs within the first number of domains based on the sorting order of the target storage unit among multiple storage units and the first number, calculated using the following expression:
[0074]
[0075] in, Indicates the first quantity; This indicates the sorting order of the target domain among the first number of domains.
[0076] In some optional implementations, the control node determines the sorting order of each field based on a first quantity; determines the number of interval fields between each field and the target field based on the sorting order of each field and the sorting order of the target field in the first quantity of fields; and determines the starting logical block address of the corresponding storage block to be repaired in each field based on the preset starting logical block address, the number of fields, and the size of the logical address range.
[0077] In one example, the control node determines the starting logical block address of the corresponding storage block to be repaired in each field based on the preset starting logical block address, the number of fields, and the size of the logical address range, and calculates it using the following expression:
[0078]
[0079] in, This indicates the starting logical block address of the corresponding storage block to be repaired in each domain; For the number of fields; Indicates the size of the logical address range.
[0080] The following example uses a specific example to determine the starting logical block address of the corresponding storage block to be repaired in each domain. Taking a preset starting logical block address of 0x350000, a logical block size of 512 bytes, a storage unit size of 32MB, a logical address range size of 0x10000, and a first quantity of 4 as an example, the control node determines the sorting order of the target storage unit to which the preset starting logical block address belongs in the multiple storage units as 54 (Segment_id is 53) based on the preset starting logical block address, logical block size, and storage unit size; based on the sorting order of the target storage unit in the multiple storage units and the first quantity, the sorting order of the target domain to which the target storage unit belongs in the first quantity of domains is determined to be 2, that is, the target domain is the second domain in the first quantity of domains, namely Domain1.
[0081] Furthermore, based on the first quantity, the sorting order of each domain is determined; based on the sorting order of each domain and the sorting order of the target domain among the first quantity of domains, the number of interval domains between each domain and the target domain is determined, that is, the number of interval domains between Domain2 and Domain1 is determined to be 1, the number of interval domains between Domain3 and Domain1 is determined to be 2, and the number of interval domains between Domain0 and Domain1 is determined to be 3; based on the preset starting logical block address, the number of domains, and the size of the logical address interval, the starting logical block address of the corresponding storage block to be repaired in each domain is determined, that is, the starting logical block address of the storage block to be repaired corresponding to Domain1 is determined to be 0x350000, the starting logical block address of the storage block to be repaired corresponding to Domain2 is determined to be 0x360000, the starting logical block address of the storage block to be repaired corresponding to Domain3 is determined to be 0x370000, and the starting logical block address of the storage block to be repaired corresponding to Domain0 is determined to be 0x380000.
[0082] Optionally, the control node can also obtain the storage unit integrity coefficient of the repair requirement; calculate the number of storage blocks to be repaired corresponding to each storage unit in the multiple storage units included in each domain based on the starting logical block address, storage unit size, and logical block size of the corresponding storage blocks to be repaired in each domain; and determine the total number of logical blocks to be repaired in each domain based on the number of storage blocks to be repaired corresponding to each storage unit in the multiple storage units included in each domain and the storage unit integrity coefficient of the repair requirement.
[0083] The storage cell integrity coefficient for repair requests is a coefficient less than or equal to 1. A storage cell integrity coefficient of 1 indicates that all storage blocks within the domain need to be repaired.
[0084] In one example, the ratio between the size of a single storage cell and the size of a logical block is calculated, and this ratio is rounded up to obtain the number of storage blocks in a single storage cell. Based on the starting logical block address and the number of storage blocks, the area from the starting logical block address to the end of the storage cell within a single storage cell in each domain is determined as the repairable area range for each storage cell. Based on the ratio between the repairable area size and the logical block size, the number of repairable storage blocks corresponding to each storage cell in each domain is determined. The total number of repairable logical blocks in each domain within the year is determined by multiplying the number of repairable storage blocks corresponding to each storage cell in each domain with the storage cell integrity coefficient required for repair.
[0085] Understandably, by calculating the ratio of a single storage unit to a logical block (and rounding up), the total number of storage blocks contained in each storage unit is determined, providing a quantitative basis for the minimum management unit for subsequent division of the area to be repaired. This avoids under-repair or over-repair due to estimation errors, ensuring that repair resources (such as computing resources and bandwidth resources) are used only for necessary repair operations, thereby improving resource utilization.
[0086] S402, generate a task request based on the starting logical block address of the storage block to be repaired in each domain.
[0087] Understandably, for unowned mirrored volumes, by determining the starting logical address of the storage blocks to be repaired in each domain for inspection and repair operations, it is easier to subsequently distribute the operation to the home nodes of each domain. This allows each home node to execute inspection and repair operations within its own domain based on the corresponding starting logical address. This solves the problems of existing mirrored volume inspection and repair operations being initiated only on one home node and the need for inter-node forwarding of the logical block addresses of all storage blocks to be repaired corresponding to the inspection and repair operation. This improves the execution efficiency of inspection and repair operations for unowned mirrored volumes and enhances compatibility with node failures.
[0088] Optional, such as Figure 5 As shown, Figure 5 This is a flowchart illustrating another method for inspecting and repairing mirrored volumes provided in an embodiment of this application. Figure 5 In addition, the control node can also perform the following steps:
[0089] S501, when the mirrored volume is a mirrored volume with a home, the home domain of the storage unit to which the preset starting logical block address belongs is determined according to the preset starting logical block address, logical block size, storage unit size and first quantity.
[0090] In some optional implementations, when the mirrored volume is a mirrored volume with a home address, the sorting order of the storage units to which the preset starting logical block address belongs among multiple storage units is determined according to the preset starting logical block address, logical block size, and storage unit size; the home domain is determined according to the order and first quantity of the storage units to which the preset starting logical block address belongs among multiple storage units.
[0091] The methods for determining the sorting order of the storage unit to which the preset starting logical block address belongs among multiple storage units, and the methods for determining the domain to which the storage unit to which the preset starting logical block address belongs, are the same as those in S401 above, and will not be repeated here.
[0092] S502, generate a first task request according to the preset starting logical block address, and send the first task request to the target home node corresponding to the home domain, so that the target home node can perform inspection and repair operations on the storage blocks to be repaired in the domain based on the first task request.
[0093] Understandably, by presetting the starting logical block address, the first task request can directly locate the starting position of the storage block to be repaired within the domain, avoiding the repair operation from starting from the wrong position, thereby ensuring that the inspection and repair work is carried out in an orderly manner within the predetermined range, which is especially suitable for scenarios where a systematic repair of a continuous storage block area is required.
[0094] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.
[0095] Embodiments of this application also provide a mirror volume inspection and repair device, such as... Figure 6 As shown, Figure 6This application provides a device structure block diagram for a mirror volume inspection and repair apparatus; applied to a control node in a first input / output group in a cluster; the first input / output group corresponds to a first number of domains; the apparatus includes: a transceiver module 601, used to receive a volume repair command sent by a client, the volume repair command instructing an inspection and repair operation to be performed on a secondary copy volume in the mirror volume corresponding to the target virtual disk, the volume repair command carrying a preset starting logical block address of the storage block to be repaired in the secondary copy volume, the secondary copy volume being divided into multiple storage units, each storage unit being divided into multiple logical blocks; obtaining the storage unit size of a single storage unit, the logical block size of a single logical block, and the logical address range size corresponding to each domain; a processing module 602, used to generate a task request based on the preset starting logical block address, logical block size, storage unit size, logical address range size, and the first number when the mirror volume is an unowned mirror volume; the transceiver module 601 is also used to distribute the task request to the owner node of each domain in the first input / output group, so that each owner node performs an inspection and repair operation on the storage block to be repaired in its respective domain based on the task request.
[0096] In some optional implementations, the processing module 602 is specifically used to determine the starting logical block address of the corresponding storage block to be repaired in each domain based on the preset starting logical block address, logical block size, storage unit size, logical address range size and first quantity when the mirror volume is an unowned mirror volume; and to generate a task request based on the starting logical block address of the corresponding storage block to be repaired in each domain.
[0097] In some optional implementations, the processing module 602 is specifically used to determine the sorting order of the target storage unit to which the preset starting logical block address belongs among multiple storage units based on the preset starting logical block address, logical block size, and storage unit size; to determine the sorting order of the target domain to which the target storage unit belongs among a first number of domains based on the sorting order of the target storage unit among multiple storage units and a first number; and to determine the starting logical block address of the corresponding storage block to be repaired in each domain based on the preset starting logical block address, the sorting order of the target domain among the first number of domains, the first number, and the logical address range size.
[0098] In some optional implementations, the sorting order of the target domain to which the target storage unit belongs in the first number of domains is determined based on the sorting order of the target storage unit among multiple storage units and the first number, and is calculated by the following expression:
[0099]
[0100] in, Indicates the first quantity; This indicates the sorting order of the target domain among the first number of domains.
[0101] In some optional implementations, the processing module 602 is specifically used to determine the sorting order of each field according to the first quantity; determine the number of interval fields between each field and the target field according to the sorting order of each field and the sorting order of the target field in the first quantity of fields; and determine the starting logical block address of the corresponding storage block to be repaired in each field according to the preset starting logical block address, the number of fields and the size of the logical address interval.
[0102] In some optional implementations, the processing module 602 is further configured to, when the mirrored volume is a mirrored volume with a home address, determine the home domain of the storage unit to which the preset starting logical block address belongs based on the preset starting logical block address, logical block size, storage unit size and a first quantity; generate a first task request based on the preset starting logical block address, and send the first task request to the target home node corresponding to the home domain, so that the target home node can perform inspection and repair operations on the storage blocks to be repaired in the domain based on the first task request.
[0103] In some optional implementations, the processing module 602 is specifically used to determine the sorting order of the storage unit to which the preset starting logical block address belongs among multiple storage units based on the preset starting logical block address, logical block size, and storage unit size when the mirror volume is a mirror volume with an owner; and to determine the ownership domain based on the order and first quantity of the storage unit to which the preset starting logical block address belongs among multiple storage units.
[0104] In some optional implementations, the transceiver module 601 is further configured to receive response information returned from each home node in sequence, each response information being used to instruct each home node to complete the inspection and repair operation on the storage block to be repaired within the domain based on the task request; when each response information is received, the state of the state machine is set to a preset state, and the preset state is returned to the client, the preset state being used to instruct the secondary copy volume to complete the inspection and repair operation.
[0105] In some optional implementations, the transceiver module 601 is further configured to record the sending time of the task request for each home node when distributing the task request to the home node of each domain; monitor the time interval from the task request to the reception of the response information in real time during the process of receiving the response information returned from each home node in sequence; if the time interval corresponding to the current home node is detected to be greater than or equal to a preset time, send a task status query request to the current home node, the task status query request being used to obtain the node status of the current home node; receive the query result generated by the current home node based on the task status query request; the processing module 602 is further configured to continue waiting for the current home node to return the response information if the query result indicates that the current home node is in a first state, the first state being used to indicate that the current home node is performing an inspection and repair operation; or, the processing module 602 is further configured to resend the task request to the current home node if the query result indicates that the current home node is in a second state, until the current home node returns the response information, the second state being used to indicate that the current home node interrupts the execution of the inspection and repair operation.
[0106] For a description of the features in the embodiment corresponding to the mirrored volume inspection and repair device, please refer to the relevant description in the embodiment corresponding to the mirrored volume inspection and repair method, which will not be repeated here.
[0107] Embodiments of this application also provide an electronic device, such as... Figure 7 As shown, Figure 7 The present application provides a structural block diagram of a mirror volume inspection and repair device; the electronic device includes a processor 10 and a memory 20, the memory 20 storing a computer program, and the processor 10 is configured to run the computer program to execute the steps in any of the above-described mirror volume inspection and repair method embodiments.
[0108] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described embodiments of the inspection and repair method for mirrored volumes when it is run.
[0109] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0110] The embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described embodiments of the image volume inspection and repair method.
[0111] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described embodiments of the image volume inspection and repair method.
[0112] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0113] The foregoing has provided a detailed description of the inspection and repair method, electronic device, storage medium, and product for a mirrored volume provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only intended to help understand the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A method for inspecting and repairing mirrored volumes, characterized in that, Applied to the control node of the first input / output group in the cluster; The first input / output group corresponds to a first number of fields; the method includes: The system receives a volume repair command from the client. The volume repair command is used to instruct the system to perform a patrol and repair operation on the secondary copy volume in the mirror volume corresponding to the target virtual disk. The volume repair command carries the preset starting logical block address of the storage block to be repaired in the secondary copy volume. The secondary copy volume is divided into multiple storage units, and each storage unit is divided into multiple logical blocks. Obtain the storage unit size of a single storage unit, the logical block size of a single logical block, and the logical address range size corresponding to each field; When the mirror volume is an unowned mirror volume, a task request is generated based on the preset starting logical block address, the logical block size, the storage unit size, the logical address range size, and the first quantity. The task request is distributed to the home node of each domain in the first input / output group, so that each home node performs the inspection and repair operation on the storage blocks to be repaired in its respective domain based on the task request.
2. The method according to claim 1, characterized in that, When the mirrored volume is an unowned mirrored volume, a task request is generated based on the preset starting logical block address, the logical block size, the storage unit size, the logical address range size, and the first quantity, including: When the mirror volume is the unowned mirror volume, the starting logical block address of the corresponding storage block to be repaired in each domain is determined according to the preset starting logical block address, the logical block size, the storage unit size, the logical address range size and the first quantity; The task request is generated based on the starting logical block address of the storage block to be repaired corresponding to each domain.
3. The method according to claim 2, characterized in that, When the mirror volume is the unowned mirror volume, the starting logical block address of the corresponding storage block to be repaired in each domain is determined according to the preset starting logical block address, the logical block size, the storage unit size, the logical address range size, and the first quantity, including: Based on the preset starting logical block address, the logical block size, and the storage unit size, determine the sorting order of the target storage unit to which the preset starting logical block address belongs among the multiple storage units; Based on the sorting order of the target storage unit among the multiple storage units and the first quantity, determine the sorting order of the target domain to which the target storage unit belongs among the first number of domains; Based on the preset starting logical block address, the sorting order of the target domain in the first number of domains, the first number, and the size of the logical address range, the starting logical block address of the corresponding storage block to be repaired in each domain is determined.
4. The method according to claim 3, characterized in that, The step of determining the sorting order of the target domain to which the target storage unit belongs among the first number of domains, based on the sorting order of the target storage unit among the multiple storage units and the first number, is calculated using the following expression: in, This indicates the first quantity; This indicates the sorting order of the target domain among the first number of domains.
5. The method according to claim 3, characterized in that, The step of determining the starting logical block address of the corresponding storage block to be repaired in each of the domains based on the preset starting logical block address, the sorting order of the target domain in the first number of domains, the first number, and the size of the logical address interval includes: Based on the first quantity, determine the sorting order of each of the domains; The number of interval fields between each of the fields and the target field is determined based on the sorting order of each field and the sorting order of the target field among the first number of fields. Based on the preset starting logical block address, the number of fields, and the size of the logical address range, the starting logical block address of the corresponding storage block to be repaired in each field is determined.
6. The method according to claim 5, characterized in that, The method further includes: When the mirror volume is a mirror volume with a home, the home domain of the storage unit to which the preset starting logical block address belongs is determined based on the preset starting logical block address, the logical block size, the storage unit size, and the first quantity. A first task request is generated based on the preset starting logical block address, and the first task request is sent to the target home node corresponding to the home domain, so that the target home node can perform the inspection and repair operation on the storage block to be repaired in the domain based on the first task request.
7. The method according to claim 6, characterized in that, When the mirrored volume is a homeped mirrored volume, determining the home domain of the storage unit to which the preset starting logical block address belongs, based on the preset starting logical block address, the logical block size, the storage unit size, and the first quantity, includes: When the mirror volume is the mirror volume with a home, the sorting order of the storage unit to which the preset starting logical block address belongs among the multiple storage units is determined according to the preset starting logical block address, the logical block size, and the storage unit size; The domain is determined based on the order of the storage unit to which the preset starting logical block address belongs among the multiple storage units and the first quantity.
8. The method according to any one of claims 1-7, characterized in that, The method further includes: The system sequentially receives response information returned from each of the home nodes, and each response information is used to instruct each of the home nodes to perform the inspection and repair operation on the storage block to be repaired in the domain based on the task request. When all the response information is received, the state machine is set to a preset state and the preset state is returned to the client. The preset state is used to instruct the secondary copy volume to complete the inspection and repair operation.
9. The method according to claim 8, characterized in that, The method further includes: When distributing a task request to the home node of each domain, the sending time of the task request to each home node is recorded; During the process of sequentially receiving response information returned from each of the respective home nodes, the time interval from the issuance of the task request to the receipt of the response information is monitored in real time. If the time interval corresponding to the current home node is detected to be greater than or equal to a preset time, a task status query request is sent to the current home node. The task status query request is used to obtain the node status of the current home node. Receive the query results generated by the currently owned node based on the task status query request; If the query result indicates that the current home node is in the first state, then continue to wait for the current home node to return the response information. The first state is used to indicate that the current home node is performing the inspection and repair operation. Alternatively, if the query result indicates that the current home node is in the second state, the task request is resent to the current home node until the current home node returns a response. The second state is used to instruct the current home node to interrupt the inspection and repair operation.
10. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the inspection and repair method for a mirrored volume as described in any one of claims 1 to 9 when executing the computer program.
Citation Information
Patent Citations
Data protection method, device and system based on shared logical volume
CN110515778A
Non-attribution volume cloning method and device, electronic equipment and storage medium
CN120560595A