Mirror image volume inspection and repair method, electronic equipment, storage medium and product

By generating and distributing task requests to the home nodes of each domain, the problem of long processing paths for inspection and repair tasks of unowned mirror volumes is solved, improving execution efficiency and the ability to cope with node failures, and realizing efficient inspection and repair operations.

CN120892263AActive Publication Date: 2025-11-04INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511419183.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2025-11-04
Estimated Expiration
2045-09-30

AI Technical Summary

Technical Problem

In the existing technology, the inspection and repair task of unowned mirror volumes has a long processing path and low execution efficiency, and it is impossible to start the inspection and repair task on other ownership nodes besides the ownership node to which the preset starting logical block address belongs.

Method used

By generating task requests and distributing them to the home nodes of each domain, all home nodes of the unowned mirror volume in all domains can initiate inspection and repair tasks. The control node is used to determine the starting logical block address of the unowned mirror volume in different domains, and additional marker information is added to determine whether the logical block address has been inspected and repaired.

Benefits of technology

The processing path for inspection and repair tasks corresponding to mirror volumes has been shortened, the ability to cope with node failures has been improved, the time for performing inspection and repair operations has been reduced, and the execution efficiency has been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120892263A_ABST
    Figure CN120892263A_ABST
Patent Text Reader

Abstract

The invention discloses a mirror image volume inspection and repair method, electronic equipment, a storage medium and a product, and relates to the technical field of storage, the method comprises the steps that a volume repair command issued by a client is received, and the volume repair command is used for executing inspection and repair operation on an auxiliary copy volume in a mirror image volume corresponding to a target virtual disk; obtaining a storage unit size of a single storage unit of the auxiliary copy volume, a logic block size of a single logic block and a logic address interval size corresponding to each domain; when the mirror image volume is a non-attribution mirror image volume, generating a task request according to a preset initial logic block address, a logic block size, a storage unit size, a logic address interval size and a first number; and distributing the task request to the attribution node of each domain, so that each attribution node executes inspection repair operation on the to-be-repaired storage block in the respective domain based on the task request. The problem that the processing path of the inspection and repair task of the non-attribution mirror image volume is long is solved, and the efficiency of executing the inspection and repair operation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of storage, and in particular to a mirror volume inspection and repair method, an electronic device, a storage medium and a product. BACKGROUND

[0002] With the development of cloud computing and storage technology, the mirror volume technology refers to that a virtual disk corresponds to two copy volumes, and when a host end performs input and output writing, two copy volumes need to be written, and a synchronization bitmap of the mirror volume is used to mark the data difference between the two copy volumes. However, for the data difference between the two copy volumes caused by the data error on a copy volume due to the data block fault, the synchronization bitmap will not be bound with the copy volume. At this time, it is necessary to rely on the execution of the inspection and repair operation of the mirror volume to repair, so that the data of the two copy volumes is synchronized again.

[0003] In the related art, for a mirror volume with an attribute, a mirror volume with an attribute inspection and repair task is usually initiated on an attribute node of an attribute domain of the mirror volume with an attribute, so as to implement the execution of the inspection and repair operation of the mirror volume with an attribute; for a mirror volume without an attribute, a mirror volume without an attribute inspection and repair task is initiated from an attribute node of a specified starting logical address, and when processing the inspection and repair task of the storage block corresponding to the logical block address belonging to other nodes in the mirror volume without an attribute, it is necessary to forward the inspection and repair task from the attribute node of the specified starting logical address to the other nodes to which the logical block address belongs, which makes the processing path of the mirror volume corresponding to the inspection and repair task long, the execution time of the inspection and repair operation long, and the efficiency of the execution of the inspection and repair operation low. SUMMARY

[0004] The present application provides a mirror volume inspection and repair method, an electronic device, a storage medium and a product, to at least solve the problem of long processing path of the inspection and repair task corresponding to the mirror volume without an attribute, long execution time of the inspection and repair operation, and low efficiency of the execution of the inspection and repair operation in the related art.

[0005] The application provides a mirror volume inspection repair method, which is applied to a control node of a first input / output group in a cluster; the first input / output group corresponds to a first number of domains; the method comprises the following steps: receiving a volume repair command issued by a client, wherein the volume repair command is used for instructing to perform an inspection repair operation on a secondary copy volume in a mirror volume corresponding to a target virtual disk, and the volume repair command carries a preset starting logical block address of a storage block to be repaired in the secondary copy volume in the inspection repair operation; the secondary copy volume is divided into a plurality of storage units, and each storage unit is divided into a plurality of logical blocks; obtaining a storage unit size of a single storage unit, a logical block size of a single logical block, and a logical address interval size corresponding to each domain; when the mirror volume is a non-attribute mirror volume, generating a task request according to the preset starting logical block address, the logical block size, the storage unit size, the logical address interval size, and the first number; and distributing the task request to an attribute node of each domain in the first input / output group, so that each attribute node performs the inspection repair operation on the storage block to be repaired in the respective domain based on the task request.

[0006] The application further provides a mirror volume inspection repair device, which is applied to a control node of a first input / output group in a cluster; the first input / output group corresponds to a first number of domains; the device comprises the following steps: a transceiving module is used for receiving a volume repair command issued by a client, wherein the volume repair command is used for instructing to perform an inspection repair operation on a secondary copy volume in a mirror volume corresponding to a target virtual disk, and the volume repair command carries a preset starting logical block address of a storage block to be repaired in the secondary copy volume in the inspection repair operation; the secondary copy volume is divided into a plurality of storage units, and each storage unit is divided into a plurality of logical blocks; a storage unit size of a single storage unit, a logical block size of a single logical block, and a logical address interval size corresponding to each domain are obtained; a processing module is used for generating a task request according to the preset starting logical block address, the logical block size, the storage unit size, the logical address interval size, and the first number when the mirror volume is a non-attribute mirror volume; and the transceiving module is further used for distributing the task request to an attribute node of each domain in the first input / output group, so that each attribute node performs the inspection repair operation on the storage block to be repaired in the respective domain based on the task request.

[0007] The application further provides an electronic device, which comprises a memory for storing a computer program and a processor for executing the computer program to implement the steps of any one of the mirror volume inspection repair methods.

[0008] The application further provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of any one of the mirror volume inspection repair methods.

[0009] The application further provides a computer program product comprising a computer program which, when executed by a processor, implements the steps of any of the mirror volume patrol repair methods described above.

[0010] According to the application, the task request generated for the non-homing mirror volume can be sent to the homing nodes of each domain, so that the homing nodes of all domains of the non-homing mirror volume can initiate the patrol repair task and perform the patrol repair operation. In this way, the problem that the non-homing mirror volume cannot start the patrol repair task on the homing nodes other than the homing node to which the preset starting logical block address belongs is solved, the processing path of the patrol repair task corresponding to the mirror volume is shortened, the ability to cope with node failure is improved, the time for performing the patrol repair operation is reduced, and the efficiency of performing the patrol repair operation is improved. BRIEF DESCRIPTION OF DRAWINGS

[0011] In order to more clearly illustrate the embodiments of the application, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some of the embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.

[0012] Figure 1 The related technology provided for the embodiments of the application is a flowchart of performing the patrol repair operation on the mirror volume; Figure 2 The topological structure diagram of the mirror volume patrol repair system provided for the embodiments of the application; Figure 3 The flowchart of the mirror volume patrol repair method provided for the embodiments of the application; Figure 4 The flowchart of the mirror volume patrol repair method provided for the embodiments of the application; Figure 5 The flowchart of the mirror volume patrol repair method provided for the embodiments of the application; Figure 6 The device structure block diagram of the mirror volume patrol repair device provided for the embodiments of the application; Figure 7 The structure block diagram of the mirror volume patrol repair device provided for the embodiments of the application. DETAILED DESCRIPTION

[0013] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only some of the embodiments of the application, not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the application.

[0014] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0015] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0016] This application embodiment is applied to a scenario where, due to data errors on the secondary copy volume of a mirrored volume caused by a data block failure, data differences exist between the primary copy volume and the secondary copy volume of the mirrored volume. The scenario involves inspecting and repairing the secondary copy volume.

[0017] A mirrored volume consists of two copy volumes: one as the primary copy volume and the other as the secondary copy volume. When the data between the two copy volumes of the mirrored volume is not synchronized, the synchronization bitmap of the volume mirror is bound to the secondary copy volume, recording the positions on the secondary copy volume where data synchronization has not been completed. After the two copy volumes have completed data synchronization, the bitmap is unbound from the secondary copy volume.

[0018] However, if a data error on one copy volume is caused by a data block failure, resulting in data discrepancies between two copy volumes, the mirror volume's synchronization bitmap will not be bound to any copy volume and will not record the differences because this is not triggered by an input / output process. Therefore, in this case, data discrepancies between the mirror volumes require inspection of the mirror volume to repair, allowing the data on the two copy volumes to resynchronize and correct the data discrepancies caused by the data block corruption.

[0019] A mirrored volume with a home domain contains multiple logical blocks that belong to the same domain; a mirrored volume without a home domain contains multiple logical blocks that belong to multiple different domains.

[0020] In related technologies, such as Figure 1 As shown, Figure 1 The flowchart for performing inspection and repair operations on a mirrored volume is provided in the embodiments of this application; Figure 1In the cluster, the Control State Matching (CSM) deployed on the control node of the first input / output group receives the volume repair command (repairvdiskcopy command) issued by the client for the mirror volume corresponding to the virtual disk; it starts the state machine and sets the virtual disk to repair state; it mobilizes the agents of all nodes to issue volume repair tasks; each node agent determines whether it is the owner node of the preset starting logical block address. If it is the owner node of the preset starting logical block address, the agent of that node starts processing the repair task from the specified preset starting logical block address, and after repairing the logical block address belonging to its own node, it forwards the processing to the next node. After processing the last storage block to be repaired in the volume, it reports that the repair is complete; the control state machine changes the virtual disk to the normal state. If none of them are the owner nodes of the preset starting logical block address, the repair task is not executed, and the operation ends.

[0021] In the aforementioned related technologies, for mirrored volumes with ownership, the inspection and repair task for the mirrored volume with ownership is initiated on the owner node of its ownership domain; for mirrored volumes without ownership, the inspection and repair task for the mirrored volume without ownership is initiated from the owner node of the specified starting logical address. However, when processing the inspection and repair task for the storage blocks corresponding to the logical block addresses belonging to other nodes in the mirrored volume without ownership, it is necessary to forward the task from the owner node of the specified starting logical address to the other node to which the logical block address being processed belongs. This makes the processing path of the inspection and repair task corresponding to the mirrored volume longer, reduces the ability to cope with node failures, increases the time for performing inspection and repair operations, and reduces the efficiency of performing inspection and repair operations.

[0022] To address the aforementioned technical problems, embodiments of this application provide a method for inspecting and repairing mirrored volumes. This method determines the starting logical block addresses of unowned mirrored volumes in different domains and adds marker information to determine whether logical block addresses belonging to the current domain have been inspected and repaired. This enables all domain-owning nodes of the unowned mirrored volume to initiate inspection and repair tasks and individually notify the control node that all logical block addresses within the current domain have been inspected and repaired. This solves the problem that unowned mirrored volumes cannot initiate inspection and repair tasks on nodes other than the one to which the preset starting logical block address belongs, shortens the processing path of the inspection and repair tasks corresponding to the mirrored volume, improves the ability to handle node failures, reduces the time required to perform inspection and repair operations, and improves the efficiency of performing inspection and repair operations.

[0023] The following is based on Figure 2 The method provided in this application embodiment is described using the mirrored volume inspection and repair system shown as an example.

[0024] like Figure 2 As shown,Figure 2 A topology diagram of a mirror volume patrol repair system provided by an embodiment of the present application is shown. Figure 2 In the embodiment, the mirror volume patrol repair system 200 includes a control node 201, a client 202, a first home node 203, a first backup node 204, a second home node 205, a second backup node 206, a third home node 207, a third backup node 208, a fourth home node 209, and a fourth backup node 210.

[0025] The control node 201 is a control node of a first input / output group in a cluster. The cluster includes a plurality of input / output groups. The first input / output group is any one of the plurality of input / output groups.

[0026] The first input / output group includes a first number of domains (Domains). The first number can be set according to actual needs. The first number can be 4.

[0027] Each domain includes a home (Owner) node and a backup (Backup) node. The Owner node is a main processing node for transaction and input / output processing in the Domain. When performing a mirror volume copy synchronization operation, the Owner node of the Domain initiates a mirror volume copy synchronization processing for the Domain.

[0028] The first home node 203 and the first backup node 204 belong to a first Domain (Domain0); the second home node 205 and the second backup node 206 belong to a second Domain (Domain1); the third home node 207 and the third backup node 208 belong to a third Domain (Domain2); and the fourth home node 209 and the fourth backup node 210 belong to a fourth Domain (Domain3).

[0029] The client 202 can be any device with display and communication functions. The client 202 can be a command line client for communicating with the control node 201 through input instructions of a command line.

[0030] Figure 2 The mirror volume patrol repair system shown is only for example and is not intended to limit the technical solutions of the present application. Those skilled in the art should understand that, in the specific implementation process, the mirror volume patrol repair system can also include more nodes, which are not limited.

[0031] An embodiment of the present application provides a mirror volume patrol repair method, which is applied to Figure 2 The control node shown is, for example, Figure 3 As shown, Figure 3A flowchart of a mirror volume inspection repair method provided by an embodiment of the present application is shown in FIG. 1. The mirror volume inspection repair method comprises the following steps. S301, receiving a volume repair command issued by a client.

[0032] The volume repair command is used to instruct to perform an inspection repair operation on a secondary copy volume in a mirror volume corresponding to a target virtual disk. The volume repair command carries a preset starting logical block address of a storage block to be repaired in the secondary copy volume by the inspection repair operation. Optionally, the volume repair command can also carry a target virtual disk identifier.

[0033] The inspection repair operation can compare the data of the primary and secondary copy volumes, detect whether the data of the secondary copy volume is completely consistent with the primary copy volume, and if there is a deviation such as misplacement or loss of data, restore the data content of the secondary copy volume to be consistent with the primary copy volume to ensure the correctness of the data copy. The inspection repair operation can also scan the metadata of the secondary copy volume to check whether there is any problem such as damage, loss, or incorrect annotation in the metadata record. Once the metadata is found to be abnormal, a repair tool or related instruction is used to correct the metadata information to restore the integrity and accuracy of the metadata. The inspection repair operation can also detect the state of the underlying storage medium on which the secondary copy volume depends, such as whether there is a bad track in the physical sector of the disk, whether the hardware connection of the storage device is stable, etc. If it is found that the storage medium has a risk of failure or has already produced an error, attempts will be made to read and rewrite the damaged sector, migrate the secondary copy volume to a healthy storage location, etc.

[0034] The secondary copy volume is divided into multiple storage units (Segments), and each storage unit is divided into multiple logical blocks.

[0035] The preset starting logical block address (Logical Block Address, LBA) can be the starting LBA of the storage block to be repaired in the secondary copy volume.

[0036] For example, the client issues a volume repair command to the control node through a command line tool. The control node receives the volume repair command issued by the client.

[0037] Optionally, the control node finds the data structure of the stored target virtual disk based on the target virtual disk identifier in the volume repair command. The data structure can be used to record the starting LBA and the ending LBA information of each domain in the target virtual disk for performing a repair task.

[0038] S302, obtaining the storage unit size of a single storage unit, the logical block size of a single logical block, and the logical address interval size corresponding to each domain.

[0039] The logical address interval size corresponding to each domain is the same.

[0040] For example, the size of the single storage unit can be 32 MB. The size of the single logical block can be 512 bytes. The size of the logical address interval corresponding to each domain can be 0x10000.

[0041] For example, the size of the single storage unit can be 32 MB. The size of the single logical block can be 512 bytes. The size of the logical address interval corresponding to each domain can be 0x10000.

[0042] S303, when the mirror volume is a non-attribute mirror volume, generating a task request according to the preset starting logical block address, the logical block size, the storage unit size, the logical address interval size, and the first quantity.

[0043] The task request is used to instruct each node to perform a patrol repair operation on the to-be-repaired storage block corresponding to the non-attribute mirror volume.

[0044] For example, the control node generates a task request according to the preset starting logical block address, the logical block size, the storage unit size, the logical address interval size, and the first quantity when the mirror volume is a non-attribute mirror volume.

[0045] In some optional embodiments, when the volume repair command also carries a preset ending logical block address of the patrol repair operation on the to-be-repaired storage block in the secondary copy volume, the control node can also generate a task request according to the starting logical block address and the preset ending logical block address of the to-be-repaired storage block in each domain.

[0046] It can be understood that the range of the to-be-repaired storage block in each domain can be accurately delimited through the starting and ending logical block addresses, so as to avoid that the repair operation exceeds the target interval. Such accuracy reduces invalid scanning and repair actions, so that each task request focuses on a specific to-be-repaired area, thereby improving the overall repair efficiency and shortening the repair time consumption.

[0047] S304, distributing the task request to the attribute node of each domain in the first input / output group, so that each attribute node performs a patrol repair operation on the to-be-repaired storage block in the respective domain based on the task request.

[0048] For example, the CSM of the control node calls a function to distribute the task request to the Agent end of each node in each domain in the first input / output group; the backup node of each domain does not process the task request after receiving it; the attribute node of each domain selects to set its starting LBA after receiving the task request. The Agent end of each attribute node respectively starts to perform a patrol repair operation on all background tasks belonging to the domain according to the starting LBA set by itself in the domain.

[0049] Optionally, when a home node can also be a home node of multiple domains, the home node can in turn set itself as the start LBA of the execution of the patrol repair operation corresponding to the domain of the Owner node.

[0050] Further, when the Agent end of each home node completes the execution of the patrol repair operation on the storage blocks to be repaired in the domain based on the task request, the control node can in turn receive the response information returned from each home node; when the reception of each response information is completed, the state of the state machine is set to the preset state, and the preset state is returned to the client.

[0051] Each response information is used to indicate the completion of the execution of the patrol repair operation on the storage blocks to be repaired in the domain by each home node based on the task request.

[0052] The preset state is used to indicate the completion of the execution of the patrol repair operation by the secondary copy volume.

[0053] In some optional embodiments, when the task request is distributed to the home node of each domain, the sending time of the task request issued to each home node is recorded; in the process of receiving the response information returned from each home node in turn, the time interval from the issuance of the task request to the reception of the response information is monitored in real time; if it is monitored that the time interval corresponding to the current home node is greater than or equal to the preset time, a task state query request is sent to the current home node; the query result generated by the current home node based on the task state query request is received; if the query result indicates that the current home node is in the first state, the current home node is continued to be waited for the return of the response information; or, if the query result indicates that the current home node is in the second state, the task request is re-sent to the current home node until the current home node returns the response information.

[0054] The task state query request is used to obtain the node state of the current home node.

[0055] The first state is used to indicate that the current home node is executing the patrol repair operation.

[0056] The second state is used to indicate that the current home node interrupts the execution of the patrol repair operation.

[0057] It can be understood that by monitoring the response time interval in real time, the task that may appear to be timed out can be found in time. For the node in the first state, the unnecessary task retransmission is avoided, and the continuity of the original task processing is ensured; for the node in the second state, the timely retransmission of the task can quickly recover the processing flow, and the probability of task interruption caused by single-point abnormality is reduced.

[0058] Based on Figure 3The method shown, the control node can receive the client issued volume repair command for the copy volume in the mirror volume corresponding to the target virtual disk; obtain the storage unit size of the single storage unit corresponding to the copy volume, the logical block size of the single logical block, and the logical address interval size corresponding to each domain; when the mirror volume is a non-attribute mirror volume, generate a task request according to the preset starting logical block address, the logical block size, the storage unit size, the logical address interval size, and the first number; distribute the task request to the attribute node of each domain in the first input output group, so that each attribute node executes the inspection repair operation on the to-be-repaired storage block in the respective domain based on the task request.

[0059] Since the task request generated for the non-attribute mirror volume can be sent to the attribute node of each domain, the attribute node of all domains of the non-attribute mirror volume can initiate the inspection repair task and execute the inspection repair operation. Thus, the problem that the non-attribute mirror volume cannot start the inspection repair task on the attribute node other than the attribute node to which the preset starting logical block address belongs is solved, the processing path of the inspection repair task corresponding to the mirror volume is shortened, the ability to cope with node failure is improved, the time for executing the inspection repair operation is reduced, and the efficiency of executing the inspection repair operation is improved.

[0060] In an optional example, on the basis of the foregoing embodiments, when the mirror volume is a non-attribute mirror volume, a task request is generated according to the preset starting logical block address, the logical block size, the storage unit size, the logical address interval size, and the first number, as shown in the following method steps, for reference to Figure 4 as shown, Figure 4 Another flowchart of a mirror volume inspection repair method provided by the embodiments of the present application is shown in the figure, which includes: S401, when the mirror volume is a non-attribute mirror volume, the starting logical block address of the to-be-repaired storage block corresponding to each domain is determined according to the preset starting logical block address, the logical block size, the storage unit size, the logical address interval size, and the first number.

[0061] In some optional embodiments, the control node determines the sorting order of the target storage unit to which the preset starting logical block address belongs in the plurality of storage units according to the preset starting logical block address, the logical block size, and the storage unit size; determines the sorting order of the target domain to which the target storage unit belongs in the first number of domains according to the sorting order of the target storage unit in the plurality of storage units and the first number; and determines the starting logical block address of the to-be-repaired storage block corresponding to each domain according to the preset starting logical block address, the sorting order of the target domain in the first number of domains, the first number, and the logical address interval size.

[0062] In an example, the control node determines the order of the target storage unit in the plurality of storage units according to the preset starting logical block address, the logical block size, and the storage unit size, and calculates the order by the following expression:

[0063] wherein, represents the preset starting logical block address; represents the logical block size; represents the storage unit size; and D represents the order of the target storage unit in the plurality of storage units.

[0064] In an example, the control node determines the order of the target domain in the first number of domains according to the order of the target storage unit in the plurality of storage units and the first number, and calculates the order by the following expression:

[0065] wherein, represents the first number; represents the order of the target domain in the first number of domains.

[0066] In some optional embodiments, the control node determines the order of each domain according to the first number; determines the number of interval domains between each domain and the target domain according to the order of each domain and the order of the target domain in the first number of domains; and determines the starting logical block address of the corresponding storage block to be repaired in each domain according to the preset starting logical block address, the number of domains, and the logical address interval size.

[0067] In an example, the control node determines the starting logical block address of the corresponding storage block to be repaired in each domain according to the preset starting logical block address, the number of domains, and the logical address interval size, and calculates the starting logical block address by the following expression:

[0068] wherein, represents the starting logical block address of the corresponding storage block to be repaired in each domain; is the number of domains; represents the logical address interval size.

[0069] In the following, the starting logical block address of the corresponding to-be-repaired storage block in each domain is determined by taking a specific example. Taking the preset starting logical block address 0x350000, the logical block size 512 bytes, the storage unit size 32 MB, the logical address interval size 0x10000, and the first number 4 as examples, the control node determines the sorting order of the target storage unit to which the preset starting logical block address belongs in the plurality of storage units as 54 (Segment_id is 53) according to the preset starting logical block address, the logical block size, and the storage unit size; and determines the sorting order of the target domain to which the target storage unit belongs in the first number of domains as 2, i.e., the target domain is the second domain in the first number of domains, i.e., Domain1, according to the sorting order of the target storage unit in the plurality of storage units and the first number.

[0070] Further, the sorting order of each domain is determined according to the first number; and the domain number of the interval domain between each domain and the target domain is determined according to the sorting order of each domain and the sorting order of the target domain in the first number of domains, i.e., the domain number of the interval domain between Domain2 and Domain1 is 1, the domain number of the interval domain between Domain3 and Domain1 is 2, and the domain number of the interval domain between Domain0 and Domain1 is 3; and the starting logical block address of the corresponding to-be-repaired storage block in each domain is determined according to the preset starting logical block address, the domain number, and the logical address interval size, i.e., the starting logical block address of the corresponding to-be-repaired storage block in Domain1 is 0x350000, the starting logical block address of the corresponding to-be-repaired storage block in Domain2 is 0x360000, the starting logical block address of the corresponding to-be-repaired storage block in Domain3 is 0x370000, and the starting logical block address of the corresponding to-be-repaired storage block in Domain0 is 0x380000.

[0071] Optionally, the control node can also obtain a storage unit integrity coefficient of the repair requirement; calculate the number of to-be-repaired storage blocks corresponding to each storage unit in the plurality of storage units included in each domain according to the starting logical block address of the corresponding to-be-repaired storage block in each domain, the storage unit size, and the logical block size; and determine the total number of to-be-repaired logical blocks in each domain according to the number of to-be-repaired storage blocks corresponding to each storage unit in the plurality of storage units included in each domain and the storage unit integrity coefficient of the repair requirement.

[0072] The storage unit integrity coefficient of the repair requirement is a coefficient less than or equal to 1. When the storage unit integrity coefficient of the repair requirement is 1, it indicates that all to-be-repaired storage blocks in the domain need to be repaired.

[0073] In an example, a ratio between the single storage unit size and the logical block size is calculated, and the ratio is rounded up to obtain the number of storage blocks of the single storage unit. According to the starting logical block address and the number of storage blocks, a range of a to-be-repaired region in each domain in the single storage unit from the starting logical block address to the end of the storage unit is determined as the range of the to-be-repaired region of each storage unit. According to a ratio between the size of the to-be-repaired region and the logical block size, a number of to-be-repaired storage blocks corresponding to each storage unit in the plurality of storage units included in each domain is determined. A product between the number of to-be-repaired storage blocks corresponding to each storage unit in the plurality of storage units included in each domain and a storage unit integrity coefficient of the repair requirement is calculated to determine a total number of to-be-repaired logical blocks in each domain.

[0074] It can be understood that, by calculating the ratio between the single storage unit and the logical block (and rounding up), the total number of storage blocks included in each storage unit is determined, which provides a quantitative basis for the minimum management unit for subsequent division of the to-be-repaired region, avoids insufficient repair or excessive repair due to estimation deviation, ensures that repair resources (such as computing resources and bandwidth resources) are only used for necessary repair operations, and improves resource utilization.

[0075] S402, generating a task request according to the starting logical block address of the to-be-repaired storage block corresponding to each domain.

[0076] It can be understood that, for the image volume without a home, by respectively determining the starting logical address of the to-be-repaired storage block in each domain performing the patrol repair operation, the starting logical address corresponding to each home can be subsequently issued to the node of each domain, so that each home can perform the patrol repair operation in the corresponding domain based on the corresponding starting logical address, solve the problem that the existing image volume patrol repair operation only starts the patrol repair operation on one home node and the logical block address of all to-be-repaired storage blocks corresponding to the patrol repair operation needs to be forwarded between nodes, and improve the execution efficiency of the patrol repair operation on the image volume without a home and the compatibility to node faults.

[0077] Optionally, as shown in Figure 5 , Figure 5 FIG. 2 is a flow diagram of another image volume patrol repair method provided by the embodiments of the present application, in which Figure 5 the control node can further perform the following steps: S501, when the image volume is a home image volume, determining the home domain of the storage unit to which the preset starting logical block address belongs according to the preset starting logical block address, the logical block size, the storage unit size, and the first number.

[0078] In some optional embodiments, when the mirror volume is a home mirror volume, a preset starting logical block address belongs to a storage unit in the plurality of storage units according to a preset starting logical block address, a logical block size, and a storage unit size; and a home domain of the storage unit is determined according to the order of the storage unit in the plurality of storage units and the first quantity.

[0079] The determination method of the order of the storage unit to which the preset starting logical block address belongs in the plurality of storage units and the home domain of the storage unit to which the preset starting logical block address belongs is consistent with the determination method in S401, and is not described herein.

[0080] S502, a first task request is generated according to the preset starting logical block address, and the first task request is sent to a target home node corresponding to the home domain, so that the target home node performs a patrol repair operation on the storage block to be repaired in the domain based on the first task request.

[0081] It can be understood that, by the preset starting logical block address, the first task request can be directly positioned to the starting position of the storage block to be repaired in the domain, avoiding the repair operation from starting from an error position, and thus ensuring that the patrol repair work is orderly carried out according to a predetermined range, and is especially suitable for a scenario of systematic repair of a continuous storage block region.

[0082] Through the above description of the embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and a necessary general hardware platform, and of course can also be realized by hardware, but in many cases the former is a better embodiment.

[0083] Embodiments of the present application also provide a mirror volume patrol repair device, as shown in Figure 6 Figure 6 ​A device structure block diagram of an image volume inspection repair device is provided for an embodiment of the present application; a control node applied to a first input / output group in a cluster; the first input / output group corresponds to a first number of domains; the device comprises: a transceiver module 601, configured to receive a volume repair command issued by a client, the volume repair command being used to instruct to perform an inspection repair operation on a secondary copy volume in an image volume corresponding to a target virtual disk, the volume repair command carrying a preset starting logical block address of a storage block to be repaired in the secondary copy volume in the inspection repair operation, the secondary copy volume being divided into a plurality of storage units, and each storage unit being divided into a plurality of logical blocks; obtaining a storage unit size of a single storage unit, a logical block size of a single logical block, and a logical address interval size corresponding to each domain; a processing module 602, configured to, when the image volume is a non-attribute image volume, generate a task request according to the preset starting logical block address, the logical block size, the storage unit size, the logical address interval size, and the first number; and the transceiver module 601 is further configured to distribute the task request to an attribute node of each domain in the first input / output group, so as to instruct each attribute node to perform the inspection repair operation on the storage block to be repaired in the respective domain based on the task request.

[0084] In some optional embodiments, the processing module 602 is specifically configured to, when the image volume is a non-attribute image volume, determine a starting logical block address of the storage block to be repaired in each domain according to the preset starting logical block address, the logical block size, the storage unit size, the logical address interval size, and the first number; and generate the task request according to the starting logical block address of the storage block to be repaired in each domain.

[0085] In some optional embodiments, the processing module 602 is specifically configured to determine a sorting order of a target storage unit in the plurality of storage units according to the preset starting logical block address, the logical block size, and the storage unit size; determine a sorting order of a target domain to which the target storage unit belongs in the first number of domains according to the sorting order of the target storage unit in the plurality of storage units and the first number; and determine the starting logical block address of the storage block to be repaired in each domain according to the preset starting logical block address, the sorting order of the target domain in the first number of domains, the first number, and the logical address interval size.

[0086] In some optional embodiments, the sorting order of the target domain to which the target storage unit belongs in the first number of domains is determined according to the sorting order of the target storage unit in the plurality of storage units and the first number, and is calculated through the following expression:

[0087] wherein, represents the first number; represents the sorting order of the target domain in the first number of domains.

[0088] In some optional embodiments, the processing module 602 is specifically configured to determine the ranking order of each domain according to the first quantity; determine the domain quantity of the interval domain between each domain and the target domain according to the ranking order of each domain and the ranking order of the target domain in the first quantity of domains; and determine the starting logical block address of the corresponding to-be-repaired storage block in each domain according to the preset starting logical block address, the domain quantity, and the logical address interval size.

[0089] In some optional embodiments, the processing module 602 is further configured to, when the mirror volume is the owned mirror volume, determine the owned domain of the storage unit to which the preset starting logical block address belongs according to the preset starting logical block address, the logical block size, the storage unit size, and the first quantity; generate a first task request according to the preset starting logical block address, and send the first task request to the target owned node corresponding to the owned domain, so that the target owned node performs the inspection and repair operation on the to-be-repaired storage block in the domain based on the first task request.

[0090] In some optional embodiments, the processing module 602 is specifically configured to, when the mirror volume is the owned mirror volume, determine the ranking order of the storage unit to which the preset starting logical block address belongs in the plurality of storage units according to the preset starting logical block address, the logical block size, and the storage unit size; and determine the owned domain according to the ranking order of the storage unit to which the preset starting logical block address belongs in the plurality of storage units and the first quantity.

[0091] In some optional embodiments, the transceiver module 601 is further configured to sequentially receive response information returned from each owned node, each response information being used to indicate that the owned node completes the inspection and repair operation on the to-be-repaired storage block in the domain based on the task request; set the state of the state machine to a preset state and return the preset state to the client when the receiving of each response information is completed, the preset state being used to indicate that the secondary copy volume completes the inspection and repair operation.

[0092] In some optional embodiments, the transceiving module 601 is further configured to, when distributing the task request to the home node of each domain, record the sending time of the task request issued by each home node; in the process of sequentially receiving the response information returned from each home node, monitor the time interval from the issuance of the task request to the reception of the response information in real time; if the time interval corresponding to the current home node is greater than or equal to the preset time, send a task state query request to the current home node, the task state query request being used to obtain the node state of the current home node; receive the query result generated by the current home node based on the task state query request; the processing module 602 is further configured to, if the query result indicates that the current home node is in a first state, continue to wait for the response information returned by the current home node, the first state being used to indicate that the current home node is performing the patrol repair operation; or, the processing module 602 is further configured to, if the query result indicates that the current home node is in a second state, resend the task request to the current home node until the response information is returned by the current home node, the second state being used to indicate that the current home node interrupts the patrol repair operation.

[0093] The descriptions of the features in the embodiments of the mirror volume patrol repair device can be referred to the descriptions of the embodiments of the mirror volume patrol repair method, which will not be repeated here.

[0094] The embodiments of the present application also provide an electronic device, as shown in Figure 7 Figure 7 The embodiments of the present application provide a structural block diagram of a mirror volume patrol repair device; the electronic device includes a processor 10 and a memory 20, the memory 20 stores a computer program, and the processor 10 is configured to run the computer program to execute the steps in any of the above mirror volume patrol repair method embodiments.

[0095] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps in any of the above mirror volume patrol repair method embodiments when running.

[0096] In an exemplary embodiment, the above computer readable storage medium can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.

[0097] ​The embodiment of the present application further provides a computer program product, which comprises a computer program, and the computer program realizes the steps in any of the mirror volume patrol repair method embodiments when executed by a processor.

[0098] The embodiment of the present application further provides another computer program product, which comprises a nonvolatile computer readable storage medium, and the nonvolatile computer readable storage medium stores a computer program, and the computer program realizes the steps in any of the mirror volume patrol repair method embodiments when executed by a processor.

[0099] The skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in the present application can be realized by electronic hardware, computer software or combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in the above description in general. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. The skilled in the art can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0100] The above describes in detail the mirror volume patrol repair method, electronic device, storage medium and product provided by the present application. The principles and implementation manners of the present application are described by applying specific examples in the present application. The above description of the examples is only applicable to help understand the method of the present application and its core idea. It should be pointed out that, for the ordinary skilled in the art, without departing from the principles of the present application, the present application can be improved and modified in several ways. These improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A method for inspecting and repairing mirrored volumes, characterized in that, Applied to the control node of the first input / output group in the cluster; The first input / output group corresponds to a first number of fields; the method includes: The system receives a volume repair command from the client. The volume repair command is used to instruct the system to perform a patrol and repair operation on the secondary copy volume in the mirror volume corresponding to the target virtual disk. The volume repair command carries the preset starting logical block address of the storage block to be repaired in the secondary copy volume. The secondary copy volume is divided into multiple storage units, and each storage unit is divided into multiple logical blocks. Obtain the storage unit size of a single storage unit, the logical block size of a single logical block, and the logical address range size corresponding to each field; When the mirror volume is an unowned mirror volume, a task request is generated based on the preset starting logical block address, the logical block size, the storage unit size, the logical address range size, and the first quantity. The task request is distributed to the home node of each domain in the first input / output group, so that each home node performs the inspection and repair operation on the storage blocks to be repaired in its respective domain based on the task request.

2. The method according to claim 1, characterized in that, When the mirrored volume is an unowned mirrored volume, a task request is generated based on the preset starting logical block address, the logical block size, the storage unit size, the logical address range size, and the first quantity, including: When the mirror volume is the unowned mirror volume, the starting logical block address of the corresponding storage block to be repaired in each domain is determined according to the preset starting logical block address, the logical block size, the storage unit size, the logical address range size and the first quantity; The task request is generated based on the starting logical block address of the storage block to be repaired corresponding to each domain.

3. The method according to claim 2, characterized in that, When the mirror volume is the unowned mirror volume, the starting logical block address of the corresponding storage block to be repaired in each domain is determined according to the preset starting logical block address, the logical block size, the storage unit size, the logical address range size, and the first quantity, including: Based on the preset starting logical block address, the logical block size, and the storage unit size, determine the sorting order of the target storage unit to which the preset starting logical block address belongs among the multiple storage units; Based on the sorting order of the target storage unit among the multiple storage units and the first quantity, determine the sorting order of the target domain to which the target storage unit belongs among the first number of domains; Based on the preset starting logical block address, the sorting order of the target domain in the first number of domains, the first number, and the size of the logical address range, the starting logical block address of the corresponding storage block to be repaired in each domain is determined.

4. The method according to claim 3, characterized in that, The step of determining the sorting order of the target domain to which the target storage unit belongs among the first number of domains, based on the sorting order of the target storage unit among the multiple storage units and the first number, is calculated using the following expression: in, This represents the first quantity; This indicates the sorting order of the target domain among the first number of domains.

5. The method according to claim 3, characterized in that, The step of determining the starting logical block address of the corresponding storage block to be repaired in each of the domains based on the preset starting logical block address, the sorting order of the target domain in the first number of domains, the first number, and the size of the logical address interval includes: Based on the first quantity, determine the sorting order of each of the domains; The number of interval fields between each of the fields and the target field is determined based on the sorting order of each field and the sorting order of the target field among the first number of fields. Based on the preset starting logical block address, the number of fields, and the size of the logical address range, the starting logical block address of the corresponding storage block to be repaired in each field is determined.

6. The method according to claim 5, characterized in that, The method further includes: When the mirrored volume is a mirrored volume with a home, the home domain of the storage unit to which the preset starting logical block address belongs is determined based on the preset starting logical block address, the logical block size, the storage unit size, and the first quantity. A first task request is generated based on the preset starting logical block address, and the first task request is sent to the target home node corresponding to the home domain, so that the target home node can perform the inspection and repair operation on the storage block to be repaired in the domain based on the first task request.

7. The method according to claim 6, characterized in that, When the mirrored volume is a homeped mirrored volume, determining the home domain of the storage unit to which the preset starting logical block address belongs, based on the preset starting logical block address, the logical block size, the storage unit size, and the first quantity, includes: When the mirror volume is the mirror volume with a home, the sorting order of the storage unit to which the preset starting logical block address belongs among the multiple storage units is determined according to the preset starting logical block address, the logical block size, and the storage unit size; The domain is determined based on the order of the storage unit to which the preset starting logical block address belongs among the multiple storage units and the first quantity.

8. The method according to any one of claims 1-7, characterized in that, The method further includes: The system sequentially receives response information from each of the home nodes, and each response information is used to instruct each of the home nodes to perform the inspection and repair operation on the storage blocks to be repaired in the domain based on the task request. When all the response information is received, the state machine is set to a preset state and the preset state is returned to the client. The preset state is used to instruct the secondary copy volume to complete the inspection and repair operation.

9. The method according to claim 8, characterized in that, The method further includes: When distributing a task request to the home node of each domain, the sending time of the task request to each home node is recorded; During the process of sequentially receiving response information returned from each of the respective home nodes, the time interval from the issuance of the task request to the receipt of the response information is monitored in real time. If the time interval corresponding to the current home node is detected to be greater than or equal to a preset time, a task status query request is sent to the current home node. The task status query request is used to obtain the node status of the current home node. Receive the query results generated by the currently owned node based on the task status query request; If the query result indicates that the current home node is in the first state, then continue to wait for the current home node to return the response information. The first state is used to indicate that the current home node is performing the inspection and repair operation. Alternatively, if the query result indicates that the current home node is in the second state, the task request is resent to the current home node until the current home node returns a response. The second state is used to instruct the current home node to interrupt the inspection and repair operation.

10. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the inspection and repair method for a mirrored volume as described in any one of claims 1 to 9 when executing the computer program.

Citation Information

Patent Citations

  • Data protection method, device and system based on shared logical volume

    CN110515778A

  • Non-attribution mirror image volume data synchronization method and device, computer equipment and storage medium

    CN119248194A

  • Non-attribution volume cloning method and device, electronic equipment and storage medium

    CN120560595A

  • Asynchronous local and remote generation of consistent point-in-time snap copies in consistency groups

    US20190034286A1