Storage system and data difference management method in storage system
The storage system addresses data restoration challenges by forming redundancy groups and managing differential information across nodes, enabling effective data recovery even in the event of two node failures.
Patent Information
- Application Number
- JP2024109854
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-08
- Publication Date
- 2026-01-21
- Estimated Expiration
- 2044-07-08
AI Technical Summary
Conventional storage systems with multiple nodes face challenges in restoring data when two nodes fail, as existing differential rebuild methods are inadequate for such scenarios.
A storage system with four or more nodes, where each node has a processor, memory, and storage drive, forms redundancy groups with data blocks and parities, and manages differential information across nodes to enable data restoration even in the event of two node failures by converting and distributing parity information.
Enables data restoration through differential rebuild even if two nodes fail, ensuring efficient data recovery by maintaining and synchronizing differential information across nodes to handle multiple failures effectively.
Smart Images

Figure 2026009747000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a storage system and a data difference management method in a storage system. [Background technology]
[0002] Storage systems made up of multiple storage nodes are known. For example, the storage system is provided as a software-defined storage (SDS) by running specific software on each storage node (hereinafter referred to as a node).
[0003] Data protection methods for this type of storage system include erasure coding (EC), etc. If, for example, a failure occurs in a drive in a node and the node is blocked, a rebuild is performed to recover the data on the failed drive according to these data protection methods.
[0004] Here, rebuilds are divided into full rebuilds and differential rebuilds. In a full rebuild, all data on the failed drive is restored from data stored on drives in nodes other than the blocked node that has the failed drive. In contrast, a differential rebuild restores only the data that was updated by IO (Input Output) received from the host while the node was blocked. Compared to a full rebuild, a differential rebuild rebuilds only the differential data on the drive, allowing data to be restored in a shorter time.
[0005] Patent Document 1 discloses a differential rebuild method as described above. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] Japanese Patent Application Publication No. 2023-106886 Summary of the Invention [Problem to be solved by the invention]
[0007] However, the above-mentioned conventional technology has a problem in that, although data can be restored when there is one blocked node, data may not be restored when there are two blocked nodes.
[0008] The present invention has been made in consideration of the above-mentioned problems, and aims to enable data restoration through differential rebuild in a storage system having multiple nodes, even if drive failures occur in two nodes. [Means for solving the problem]
[0009] In order to achieve the above-mentioned object, one aspect of the present invention is a storage system comprising four or more nodes each having a processor, a memory, and a storage drive, wherein for user data stored in the storage drive, a redundancy group is formed including a plurality of data blocks of the user data and a first parity and a second parity based on the data blocks, and for each of the plurality of nodes, the storage drive of the node has a user area for storing the plurality of data blocks belonging to different redundancy groups and a parity area for storing the second parity, and the processor converts the second parity stored in the parity area of the node into the first parity generated based on the plurality of data blocks belonging to the different redundancy group that are stored in the user area of one of the nodes other than the node in question. and the plurality of data blocks belonging to the same redundancy group that are stored in a distributed manner in the user areas of the node and the plurality of nodes excluding the one node, and the data block stored in the user area of the node and the first parity and the second parity stored in the parity area are managed in association with difference information indicating the presence or absence of a difference related to an update of either or both of the data block, the first parity, and the second parity that belong to the same redundancy group as the data block, the first parity, and the second parity, and when an update is made to the data block stored in the user area of the node during a period in which the node is blocked, the difference information related to the update is managed in one of the two nodes other than the node that is not blocked and is operating normally. [Effects of the Invention]
[0010] According to the present invention, in a storage system having multiple nodes, even if drive failures occur in two nodes, data can be restored by differential rebuild. [Brief explanation of the drawings]
[0011] [Figure 1] FIG. 1 is a diagram showing an example of the physical configuration of a storage system according to a first embodiment. [Figure 2] FIG. 1 is a diagram showing an example of the logical configuration of a storage system according to the first embodiment. [Figure 3] FIG. 3 is a diagram showing an example of information in a memory according to the first embodiment. [Figure 4A] FIG. 4 is a diagram showing an example of a difference management table according to the first embodiment. [Figure 4B] FIG. 4 is a diagram showing an example of a difference management table according to the first embodiment. [Figure 5] FIG. 2 is a diagram showing an example of the configuration of stripes in a parity group according to the first embodiment. [Figure 6] FIG. 2 is a diagram showing an example of the relationship between data blocks, C1 parity blocks, and PQ parity blocks according to the first embodiment. [Figure 7] FIG. 10 is a diagram showing an example of an overview of a process related to a differential rebuild according to the first embodiment. [Figure 8] FIG. 10 is a diagram showing an example of an outline of a difference information recording process when one node is blocked according to the first embodiment. [Figure 9] FIG. 10 is a diagram showing an example of an outline of the difference information collection process and the difference information reflection process when two nodes are blocked according to the first embodiment. [Figure 10] FIG. 10 is a diagram showing an example of an outline of a differential information clearing process when a node recovers according to the first embodiment. [Figure 11A] 10 is a timing chart showing an example of a differential rebuild process when two nodes fail according to the first embodiment. [Figure 11B] 10 is a timing chart showing an example of a differential rebuild process when two nodes fail according to the first embodiment. [Figure 12A] 10 is a timing chart showing an example of a differential rebuild process when two nodes fail according to the second embodiment. [Figure 12B] 10 is a timing chart showing an example of a differential rebuild process when two nodes fail according to the second embodiment. [Figure 12C] 10 is a timing chart showing an example of a differential rebuild process when two nodes fail according to the second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0012] In the following description, an "interface device" may refer to one or more communication interface devices. The one or more communication interface devices may be one or more homogeneous communication interface devices (e.g., one or more NICs (Network Interface Cards)) or two or more heterogeneous communication interface devices (e.g., a NIC and an HBA (Host Bus Adapter)).
[0013] In the following description, "memory" refers to one or more memory devices, which are an example of one or more storage devices, and may typically be a primary storage device. At least one memory device in the memory may be a volatile memory device or a non-volatile memory device.
[0014] In the following description, a "persistent storage device" may refer to one or more persistent storage devices, which are an example of one or more storage devices. A persistent storage device may typically be a non-volatile storage device (e.g., an auxiliary storage device), and specifically may be, for example, a hard disk drive (HDD), a solid state drive (SSD), or a non-volatile memory express (NVMe) drive.
[0015] Furthermore, in the following description, a "processor" may refer to one or more processor devices. The at least one processor device may typically be a microprocessor device such as a CPU (Central Processing Unit), but may also be another type of processor device such as a GPU (Graphics Processing Unit). The at least one processor device may be a single-core or multi-core. The at least one processor device may also be a processor core. The at least one processor device may also be a processor device in a broader sense, such as a hardware circuit that performs part or all of the processing (for example, an FPGA (Field-Programmable Gate Array), a CPLD (Complex Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit)).
[0016] In the following description, information that provides an output for an input may be described using expressions such as "xxx table." However, the information may be data of any structure (for example, structured data or unstructured data), or may be a neural network that generates an output for an input, or a learning model such as a genetic algorithm or random forest. Therefore, the "xxx table" may be referred to as "xxx information." In the following description, the structure of each table is an example, and one table may be divided into two or more tables, or all or part of two or more tables may be one table.
[0017] In the following description, processing may be described using a "program" as the subject. However, a program is executed by a processor to perform a predetermined process using a storage device and / or an interface device, etc., as appropriate. Therefore, the subject of processing may be the processor (or a device such as a controller having the processor). A program may be installed in a device such as a computer from a program source. The program source may be, for example, a program distribution server or a computer-readable (e.g., non-transitory) recording medium. In the following description, two or more programs may be realized as one program, or one program may be realized as two or more programs.
[0018] In the following embodiment, the storage system is assumed to be configured to include six nodes. Furthermore, the data protection method adopted in the storage system is assumed to be MEC (Multi-stage Erasure Coding), a type of 4D+2P EC in which two parities (C1 parity and PQ parity) are provided for every four data blocks, corresponding to the six nodes in the storage system. However, the number of nodes is not limited to six, and the data protection method is not limited to 4D+2P. In other words, a storage system with (m+n) or more nodes may adopt mD+nP (m is an integer equal to or greater than 2, and n is an integer equal to or greater than 2) MEC.
[0019] [Embodiment 1] (Physical configuration of storage system 101) FIG. 1 shows an example of the physical configuration of a storage system 101.
[0020] The storage system 101 is configured to include a plurality of nodes 210. In this embodiment, the storage system 101 is configured to include six nodes 210. The storage system 101 has an interface (not shown) for connecting to a network 120, and is connected to a host 200 via the network 120 so as to be able to communicate with the host 200.
[0021] The node 210 may have the configuration of a typical server computer. The node 210 is configured to include, for example, one or more processor packages 213, one or more drives 214, and one or more ports 215. The components are connected via an internal bus 216. The processor package 213 includes a processor 211, a memory 212, etc. The drive 214 is an example of a storage drive and a persistent storage device. The port 215 is an example of an interface device.
[0022] The processor 211 is, for example, a CPU, and performs various processes. The processor 211 performs various processes in cooperation with other processors 211 of other nodes 210 other than the node 210 that has the processor 211.
[0023] The memory 212 stores control information and data necessary to realize the functions of the node 210. The memory 212 also stores, for example, a program executed by the processor 211. The memory 212 may be a volatile dynamic random access memory (DRAM), a non-volatile SCM, or another storage device.
[0024] The drive 214 stores various types of data, programs, etc. The drive 214 may be an HDD or SSD connected via SAS (Serial Attached SCSI) or SATA (Serial Advanced Technology Attachment), an NVMe-connected SSD, or an SCM, and is an example of a storage device.
[0025] The port 215 is connected to a network 220 and is communicably connected to other nodes 210 within the site 201. The network 220 is, for example, a LAN (Local Area Network), but is not limited to a LAN.
[0026] The physical configuration of the storage system 101 is not limited to the above. For example, the network 220 may be made redundant. Also, for example, the network 220 may be separated into a management network and a storage network, the connection standard may be Ethernet (registered trademark), InfiniBand (registered trademark), or wireless, and the connection topology is not limited to the configuration shown in FIG. 1.
[0027] (Logical configuration of storage system 101) FIG. 2 shows an example of the logical configuration of the storage system 101.
[0028] The drive 214 of each node 210 in the storage system 101 has multiple physical chunks 214a. The physical chunk 214a is an area carved out of the storage area for user data in each drive 214, and has multiple data blocks 214a1 and multiple parity blocks 214a3. In this embodiment, one parity block 214a3 corresponds to four data blocks 214a1. The storage area for the data blocks 214a1 in the physical chunk 214a is called the user area, and the storage area for the parity blocks 214a3 is called the parity area.
[0029] The data blocks 214a1 are storage areas that serve as storage units for user data. The parity blocks 214a3 are areas that store the P+Q parity of the corresponding four data blocks 214a1, but may store other parities other than the P+Q parity.
[0030] A chunk group 210a is a combination of one physical chunk 214a extracted from each node 210. A parity group 214d is configured by extracting and combining four data blocks 214a1 and one parity block 214a3 from each physical chunk 214a belonging to the chunk group 210a.
[0031] A differential information management table 414 is provided for each parity group 214d. The differential information management table 414 arranges differential information 414a, in which four data blocks 214a1 and one parity block 214a3 are arranged vertically for each node 210, horizontally for all nodes 210 belonging to the same chunk group 210a. In the example shown in FIG. 2, the column for node #1 in the differential information management table 414 stores the blocks D1S1, D2S6, D3S5, D4S4, C1S3, and PQS2 arranged in this order. The columns for nodes #2, #3, #4, #5, and #6 in the differential information management table 414 are also as shown in FIG. 2.
[0032] (Information in memory 212) FIG. 3 shows an example of information in memory 212.
[0033] The information in the memory 212 , including the control information table 410 and the storage program 360 , is read from a non-volatile storage area such as the drive 214 into the memory 212 and executed by the processor 211 .
[0034] The control information table 410 includes a cluster management table 411 , a storage pool management table 412 , a parity group management table 413 , and a difference information management table 414 .
[0035] The cluster management table 411 stores information for managing the configuration of the storage system 101, the nodes 210, and the drives 214. The storage pool management table 412 stores control information for the thin provisioning function provided by the storage pool. A storage pool is configured to include multiple logical chunks corresponding to the data blocks 214a1 of the physical chunks 214a, and virtualizes the capacity of the entire storage system 101.
[0036] The parity group management table 413 stores control information for managing the configuration of parity groups 414c (redundancy groups) formed by combining multiple physical chunks 214a. Details of the cluster management table 411, storage pool management table 412, and parity group management table 413 will be omitted.
[0037] The difference information management table 414 is a table for managing difference information indicating whether or not there is a difference related to the update of each block of data / parity stored in the node 210 of the storage system 101. Details of the difference information management table 414 will be described later.
[0038] The storage program 360 includes an IO processing program 421 and a rebuild processing program 422. The IO processing program 421 processes IOs from the host 200 (FIG. 1). The rebuild processing program 422 will be described in detail later.
[0039] (Differential information management table 414) 4A and 4B are diagrams showing the difference information management table 414 for explaining the difference management method in the storage system 101. Fig. 4A shows a portion of the difference information management table 414 for the nodes 210 (nodes #1 to #3), and Fig. 4B shows a portion of the difference information management table 414 for the nodes 210 (nodes #4 to #6).
[0040] The differential information management table 414 has columns of "data / parity" and "difference information" for each of nodes #1 to #6. The differential information management table 414 also has rows for data blocks #1 to #4, C1 parity blocks, and PQ parity blocks.
[0041] The differential information management table 414 has "data / parity" in the upper row and "difference information" in the lower row for each parity group 214d and for each node #1 to #6. In this embodiment, the differential information management table 414 stores, in association with the "data / parity" shown in the upper row that is stored in each node 210, differential information that indicates whether or not blocks of data / parity stored in other nodes 210 are subject to differential rebuild.
[0042] Here, stripe 0 in parity group 214d will be described. Figure 5 is a diagram showing an example of the configuration of stripe 414c in parity group 214d. For ease of explanation, identical differential information management tables 414L and 414R are shown side by side in Figure 5(a). The components below the diagonal line (the line connecting D1S1, D2S1, D3S1, D4S1, C1S1, and PQS1) of differential information management table 414L are displayed in differential information management table 414R.
[0043] As shown in Fig. 5(b), stripe 414c includes four data blocks 214a1, one C1 parity block 214a2, and one parity block 214a3. For example, as shown in Fig. 5(a), stripe 414c includes six stripes: stripes 414c1, 414c2, 414c3, 414c4, 414c5, and 414c6.
[0044] Stripe 414c1 (stripe S1) is composed of D1S1 of node #1, D2S1 of node #2, D3S1 of node #3, D4S1 of node #4, C1S1 of node #5, and PQS1 of node #6. DiS1 (i = 1, 2, ..., 4) is data block 214a1 of stripe S1. C1S1 is the C1 parity block and is the XOR of data blocks D1S1, D2S6, D3S5, and D4S4 in node #1. PQS1 is the PQ parity block of stripe S1. Stripes S2 to S6 are similar to stripe S1.
[0045] Hereafter, DiSj (i = 1, ..., 4, j = 1, ..., 6) will be written as Di of the Sj stripe. Also, C1Sj (j = 1, ..., 6) will be written as C1 of the Sj stripe. Also, PQSj (j = 1, ..., 6) will be written as PQ of the Sj stripe.
[0046] The stripes 414c2 (stripe S2), 414c3 (stripe S3), 414c4 (stripe S4), 414c5 (stripe S5), and 414c6 (stripe S6) are also as shown in FIG.
[0047] Returning to the description of Figures 4A and 4B, the information shown in Figures 4A and 4B is stored at the intersections of the rows and columns of the difference information management table 414.
[0048] For example, D1S1 (D1 of stripe S1) is stored in "data / parity" and PQS1(0,1) is stored in "difference information" at the intersection of the row of data block #1 and the column of node #1 in differential information management table 414. "PQS1(0,1)" stores "1" if there is a difference in PQS1 (PQ of stripe S1), and "0" if there is no difference.
[0049] Furthermore, at the intersection of the row for data block #2 and the column for node #1 in differential information management table 414, D2S6 (D2 of stripe S6) is stored in "data / parity," and D1S6(0,1) and PQS6(0,1) are stored in "differential information." "D1S6(0,1)" stores "1" if there is a difference in D1S6 (D1 of stripe S6), and "0" if there is no difference. "PQS6(0,1)" stores "1" if there is a difference in PQS6 (PQ of stripe S6), and "0" if there is no difference.
[0050] Furthermore, at the intersection of the row of the PQ parity block and the column of node #1 in the differential information management table 414, PQS2 (PQ of stripe S2) is stored in "data / parity." Furthermore, D1S2(0,1), D2S2(0,1), D3S2(0,1), and D4S2(0,1) are stored in "differential information." "PQS2(0,1)" stores "1" if there is a difference in PQS2 (PQ of stripe S2), and "0" if there is no difference. "DiS2(0,1)" (i=1,2,...,4) stores "1" if there is a difference in DiS2 (Di of stripe S2), and stores "0" if there is no difference.
[0051] The C1 parity block and C1 parity are an example of first parity, and the PQ parity block and PQ parity are an example of second parity.
[0052] In MEC, for example, when data block "D1S1" stored in node 210 (node #1) is updated, the following two blocks are updated simultaneously. That is, the PQ parity block "PQS1" in the same stripe S1 and the PQ parity block "PQS3" in the same stripe S3 as the C1 parity block "C1S3" in node 210 (node #1) are updated simultaneously. That is, when data block "D1S1" stored in node 210 (node #1) is updated, node 210 (node #2) and node 210 (node #6) are accessed simultaneously.
[0053] Therefore, it is preferable for the efficiency of data access that one piece of difference information for data block "D1S1" be stored in association with PQ parity block "PQS1" stored in node 210 (node #6).
[0054] Furthermore, in this embodiment, to accommodate failures in two nodes 210, difference information for one piece of data / parity is stored in two locations. Therefore, it is preferable for data access efficiency that the other piece of difference information is stored in association with the data / parity stored in node 210 (node #2). Therefore, the other piece of difference information is stored in association with data block "D2S1," which is the block of data / parity for stripe S1 stored in node 210 (node #2).
[0055] (Relationship between data and parity blocks) FIG. 6 is a diagram showing an example of the relationship between a data block 214a1, a C1 parity block 214a2, and a parity block 214a3.
[0056] 6, C1S3 of node #1 is generated by XORing D1S1, D2S6, D3S5, and D4S4 stored in node #1. The actual data of C1S3 is not stored in drive 214 but is held in memory 212. PQS3 stored in node #2 is PQ parity calculated based on D1S3 stored in node #3, D2S3 stored in node #4, D3S3 stored in node #5, D4S3 stored in node #6, and C1S3.
[0057] C1S4 and PQS4, C1S5 and PQS5, C1S6 and PQS6, and C1S1 and PQS1 are also calculated in the same manner as shown in FIG.
[0058] In this manner, in this embodiment, the data and PQ parity of the same stripe Sx (x=1, 2, ..., 4) are distributed across all the nodes 210 (nodes #1 to #6). As a result, when one of the nodes 210 fails, the data and PQ parity stored in the failed node can be restored by a full rebuild or a differential rebuild based on the data and PQ parity that are distributed across the other nodes 210.
[0059] Furthermore, if a two-node failure occurs in which another node 210 fails in addition to the one-node failure described above, the C1 parity of the same stripe Sx as the data block that cannot be acquired due to the node failure is used in place of the data block that cannot be acquired. As a result, in the event of two node failures in the node 210, the data and PQ parity stored in the failed node can be restored by a full rebuild or differential rebuild based on the data, C1 parity, and PQ parity that are distributed and allocated to the other nodes 210.
[0060] Furthermore, differential information of the data / parity of each node 210 is held in two other nodes 210 that are accessed simultaneously when updating data in the node 210 in which the data to be updated is stored. As a result, even in the event of a two-node failure in which failures occur in the node 210 in question and one other node 210, the differential information is not lost because it is held in the remaining other node 210, and so it is possible to restore the data and PQ parity by differential rebuild.
[0061] (Timing of each process related to differential rebuild) FIG. 7 is a diagram illustrating an example of an outline of processing related to differential rebuild.
[0062] In the differential rebuild of this embodiment, similar to the differential rebuild of the conventional technology, the differential information update process ST11 during recovery of the node 210 (recovery target node) blocked due to a drive failure is executed from the start of rebuild preparation ST1 to the IO stop ST2 before the start of the rebuild.
[0063] This is for the following reason: To perform a differential rebuild, it is necessary to collect differential information from a node 210 (surviving node) that is operating normally. At this time, it is necessary to prevent a mismatch between the update to the node 210 (surviving node) that holds the differential information and the collection of the differential information (difference information collection process ST12), and to ensure that no differential information is missed when it is collected. A surviving node is a node that is not blocked and is operating normally.
[0064] 8 is a diagram illustrating an example of an overview of the differential information recording process when one node is blocked according to embodiment 1. Blocking a node refers to putting the node 210 in question into a state where it cannot accept IO from the host 200 due to a failure or the like of the drive 214 of the node 210.
[0065] As shown in Figure 8, for example, if node #6 becomes a node blocker, update information for D1S6 of node #6 is stored in node #1 and node #6. Also, if node #6 becomes a node blocker, update information for D2S5 of node #6 is stored in node #1 and node #4. Also, if node #6 becomes a node blocker, update information for D3S4 of node #6 is stored in node #1 and node #3. Also, if node #6 becomes a node blocker, update information for D4S3 of node #6 is stored in node #1 and node #2.
[0066] IO is stopped in IO stop ST2, and the configuration of the configuration information of node 210 (recovery target node) is completed in IO available state transition ST3. When the configuration information configuration of node 210 (recovery target node) is completed, IO becomes available and IO resumes ST4, and thereafter, the difference information update process ST11 is not performed.
[0067] Next, the difference collection information collected by the difference information collection process ST12 is reflected in the node 210 (recovery target node). This is because the node 210 (recovery target node) also becomes a target for differential information collection after recovery, and the other nodes 210 (recovery target nodes) perform differential rebuilds based on the difference collection information collected from the node 210 (recovery target node) that recovered earlier.
[0068] In this embodiment, the "data / parity" difference information is held in two nodes 210 other than the node 210 itself where the "data / parity" is stored. For example, if a node 210 that is operating normally and from which difference information is to be collected is newly blocked while a certain node 210 is recovering, the difference information stored during the period when the recovered node 210 was blocked must be reflected in the recovered node 210, or this difference information will be lost.
[0069] Therefore, the difference information collection process ST12 is performed after the IO stop ST2 and before the rebuild start ST5 in a state where the difference update of the node being recovered has not been performed.
[0070] Because the differential information is required for the differential rebuild, it is collected before the differential rebuild is executed. This is because if the differential information is collected while the differential information update process ST11 of the recovering node 210 is being performed, there is a possibility that the differential information required for the differential rebuild will not be collected completely because the differential information update process ST11 and the collection of the differential information are not synchronized.
[0071] Furthermore, the differential information reflection process ST13 is performed before the recovering node 210 has fully recovered. This is for the following reason. After the rebuild is complete, redundancy has been restored, making it possible to deal with the blockage of another node 210. However, if the differential information reflection process ST13 is not performed, when another node 210 becomes blocked and recovery has begun, the differential information that should be reflected in this blocked node 210 will not exist in any of the nodes 210. If the node 210 that has begun recovery in this state performs the differential information collection process ST12, differential rebuild will not be possible because there will be differential information that has not been reflected.
[0072] 9 is a diagram illustrating an example of an overview of the difference information collection process ST12 and the difference information reflection process ST13 when two nodes are blocked according to embodiment 1. Fig. 9 illustrates the case of recovery processing of node #1 when two nodes are blocked due to failures in nodes #1 and #6.
[0073] Because the differential information for D1S6 of node #6 is managed by node #1 and node #5, when node #1 recovers, the differential information for D1D6 is collected from node #5 as differential collection information 414b and reflected in the differential information 414a of node #1. Also, because the differential information for D2S5 of node #6 is managed by node #1 and node #4, when node #1 recovers, the differential information for D2D5 is collected from node #4 as differential collection information 414b and reflected in the differential information 414a of node #1. Also, because the differential information for D3S4 of node #6 is managed by node #1 and node #3, when node #1 recovers, the differential information for D3D4 is collected from node #3 as differential collection information 414b and reflected in the differential information 414a of node #1. In addition, since the differential information for D4S3 of node #6 is managed by node #1 and node #2, when node #1 recovers, the differential information for D4D3 is collected from node #2 as differential collection information 414b and reflected in the differential information 414a of node #1.
[0074] Fig. 10 is a diagram showing an example of an outline of the differential information clearing process when a node recovers according to embodiment 1. When the differential information reflecting process ST13 is completed, as shown in Fig. 10, the recovered node #1 transmits a notification that node #1 has recovered to nodes #2 to #5, which are surviving nodes operating normally. Upon receiving the notification that node #1 has recovered, nodes #2 to #5 identify the differential information to be cleared at each node based on the stripe relationship (the portion enclosed by dashed lines in Fig. 10 identified from the block enclosed by solid lines), and only the identified differential information is cleared.
[0075] Furthermore, when a rebuild can be started, the node 210 to be restored also starts recording the difference (difference information recording process ST14). Because a differential rebuild is performed when a blocked node 210 is restored, it is necessary to record the differences resulting from updates to the blocked node even while the node is being restored. If the differential information recording process ST14 is not performed during node recovery and another node 210 that held the difference becomes newly blocked, the difference updated during node recovery will not be held. As a result, when an attempt is made to restore the originally blocked node 210, some areas will not be differentially rebuilt, and the differential rebuild will not be performed correctly.
[0076] Therefore, if there is an update in the blocked node after IO resumes ST4, the recovered node is accessed for parity update, and the differential information in the recovered node is retained.
[0077] (Differential rebuild process when two nodes fail) 11A and 11B are timing charts showing an example of differential rebuild processing when two nodes fail. The explanation of FIGS. 11A and 11B shows an example in which failures occur in two nodes 210 in the storage system 101, and the nodes 210 to be recovered (recovery target nodes) are sequentially recovered one by one. Only one node 210 to be sequentially recovered (recovery target node) is shown in FIGS. 11A and 11B. The series of processes shown in FIGS. 11A and 11B are executed sequentially for each node 210 to be recovered (recovery target node).
[0078] The differential rebuild process is executed by the rebuild process program 422 of the representative node 210 (representative node) in the storage system 101 in cooperation with the IO processing program 421 of the other nodes 210 (surviving nodes) and the node 210 (recovery target node). The node 210 (surviving node) is a node that has not experienced any failure and continues to operate normally.
[0079] 11A and 11B is the difference information stored in the difference information management table 414 managed by the node 210. The difference recovery information 414b held by the node 210 (recovery target node) is the difference information stored in the difference information management table 414 managed by a node 210 other than the node 210 itself.
[0080] 11A, the rebuild processing program 422 instructs the IO processing programs 421 of all nodes 210 (surviving nodes) to stop IO, which stops processing of IO from the host 200. Next, in step S102, the IO processing programs 421 of all nodes 210 (surviving nodes) stop IO in response to the instruction of step S101.
[0081] Next, in step S103, the rebuild processing program 422 instructs the IO processing program 421 of node 210 (recovery target node) to collect difference information. Next, in step S104, the IO processing program 421 of node 210 (recovery target node) instructs the IO processing program 421 of node 210 (surviving node) to collect difference information.
[0082] Next, in step S105, the IO processing program 421 of the node 210 (surviving node) acquires the difference information 414a and sends it to the IO processing program 421 of the node 210 (recovery target node).
[0083] Next, in step S106, the IO processing program 421 of the node 210 (recovery target node) aggregates the difference information received from the node 210 (surviving node) as difference collection information 414b, and stores it in the memory 212. Step S106 starts storing the difference information necessary for the rebuild.
[0084] Next, in step S107, the rebuild processing program 422 instructs the IO processing program 421 of the node 210 (recovery target node) to reflect the difference collection information. Next, in step S108, the IO processing program 421 of the node 210 (recovery target node) acquires the difference collection information 414b and reflects it in the difference information 414a. After step S108, if there are other blocked nodes in the storage system 101, recording of the difference information of the blocked node begins.
[0085] Next, in step S109, the rebuild processing program 422 instructs the node 210 (recovery target node) and all nodes 210 (surviving nodes) to resume IO. The node 210 (recovery target node) and all nodes 210 (surviving nodes) resume IO upon receiving the instruction. Because the configuration information of the node 210 (recovery target node) is reconstructed while IO is stopped, IO becomes possible at the same time as IO is resumed by the node 210 (surviving node).
[0086] Next, in step S110 of FIG. 11B, the rebuild processing program 422 instructs the IO processing program 421 of the node 210 (recovery target node) to start a differential rebuild in chunk group 210a units.
[0087] Next, in step S111, the IO processing program 421 of the node 210 (recovery target node) references the difference collection information 414b, and checks whether or not there is a difference in the data block 214a1 for the chunk group 210a specified in step S110. For data blocks 214a1 with no difference, the IO processing program 421 of the node 210 (recovery target node) can use the data stored in the drive 214 of the node 210 (recovery target node) as is, and therefore omits steps S113 to S115.
[0088] Next, in step S112, the IO processing program 421 of the node 210 (recovery target node) instructs the IO processing program 421 of the node 210 (surviving node) to perform a differential rebuild to restore the data block 214a1 determined to have a difference in step S111.
[0089] Next, in step S113, the IO processing program 421 of the node 210 (surviving node) references the drive 214, collects restoration data for performing a differential rebuild on the data block 214a1 instructed to be restored in step S112, and restores the data. The collection of restoration data is called a correction copy.
[0090] Next, in step S114, the IO processing program 421 of the node 210 (surviving node) sends the restored data restored in step S113 to the IO processing program 421 of the node 210 (recovery target node) that issued the data restoration instruction in step S112.
[0091] Next, in step S115, the IO processing program 421 of the node 210 (recovery target node) writes the restored data (data block 214a1) received from the IO processing program 421 of the node 210 (surviving node) to the drive 214. The restored data is stored in the drive 214 from step S115 onwards.
[0092] Steps S112 to S115 are repeated as long as there is a difference in the data block 214a1.
[0093] Next, in step S116, the IO processing program 421 of the node 210 (recovery target node) references the difference collection information 414b, and checks whether there is a difference in the parity block 214a3 for the chunk group 210a specified in step S110. The IO processing program 421 of the node 210 (recovery target node) omits steps S117 to S121 because for the parity block 214a3 with no difference, the PQ parity stored in the drive 214 of the node 210 (recovery target node) can be used as is.
[0094] Next, in step S117, the IO processing program 421 of the node 210 (recovery target node) instructs the IO processing program 421 of the node 210 (surviving node) to restore data by differential rebuilding of the parity block 214a3 determined to have a difference.
[0095] Next, in step S118, the IO processing program 421 of the node 210 (surviving node) references the drive 214 and collects (collection copies) data for restoration to restore the parity block 214a3 instructed to be restored in step S117. Next, in step S119, the IO processing program 421 of the node 210 (surviving node) transmits the data for restoration collected in step S118 to the IO processing program 421 of the node 210 (recovery target node) that issued the data restoration instruction in step S117.
[0096] Next, in step S120, the IO processing program 421 of the node 210 (recovery target node) restores the parity block 214a3 based on the restored data received from the IO processing program 421 of the node 210 (surviving node). Next, in step S121, the IO processing program 421 of the node 210 (recovery target node) writes the restored data restored in step S120 to the drive 214.
[0097] As long as there is a difference in the parity block 214a3, steps S117 to S121 are repeated.
[0098] Next, in step S122, the rebuild processing program 422, and the IO processing program 421 of the node 210 (recovery target node) and node 210 (surviving node) execute a difference clear process at the timing when the difference rebuild is completed. In the difference clear process, a notification of which node 210 has recovered is sent from the node 210 (recovery node) to each node 210. Each node 210 that receives the notification identifies the data block 214a1 for which difference information should be cleared based on the stripe relationship, and clears the associated difference information. In the difference clear process, only the difference information of the recovered node 210 is cleared, and the difference information of the blocked node 210 that has not yet recovered is maintained.
[0099] In node recovery processing after two blocked nodes are encountered, by clearing only the differential information related to the first node 210 that is recovered, differential rebuilds can also be performed in the recovery processing of the second node 210.
[0100] Next, in step S123, the rebuild process program 422 determines whether the processing of steps S110 to S122 has been completed for all chunk groups 210a. If the processing of steps S110 to S122 has been completed for all chunk groups 210a (YES in step S123), the rebuild process program 422 ends the differential rebuild process in the event of two-node failures. On the other hand, if there is a chunk group 210a for which the processing of steps S110 to S122 has not been completed (NO in step S123), the rebuild process program 422 selects a new chunk group 210a and returns the process to step S110.
[0101] [Embodiment 2] In the above-described first embodiment, when two nodes fail, the failed nodes are sequentially restored one by one. However, the two failed nodes are not limited to being sequentially restored, and may be restored all at once. In the second embodiment, an example in which the two failed nodes are restored all at once will be described.
[0102] In the second embodiment, explanations that overlap with those in the first embodiment will be omitted.
[0103] (Differential rebuild process when two nodes fail according to the second embodiment) 12A, 12B, and 12C are timing charts illustrating an example of the differential rebuild process when two nodes fail according to the second embodiment.
[0104] The differential rebuild process in the case of two-node failures in the second embodiment differs from the first embodiment in that the differential information collection instruction in step S103 is output simultaneously to two recovery nodes (node 210 (recovery target node 1) and node 210 (recovery target node 2)).
[0105] In step S103a following step S102, the rebuild process program 422 of node 210 (recovery target node 1) instructs node 210 (recovery target node 2) to perform differential recovery.
[0106] Next, in step S104a, the IO processing program 421 of the node 210 (recovery target node 1) that received the difference information collection instruction sends a difference information collection instruction to the node 210 (recovery target node 2) and the node 210 (surviving node).
[0107] Next, in step S105, the IO processing program 421 of the node 210 (surviving node) acquires the difference information 414a and sends it to the node 210 (recovery target node 1).
[0108] Also, in step S105a, the IO processing program 421 of node 210 (recovery target node 2) acquires the difference information 414a in node 210 (recovery target node 2) and sends it to node 210 (recovery target node 1). When step S105a is executed, node 210 (recovery target node 2) has recovered, and therefore becomes the target for recovery of difference information. However, at this time, node 210 (recovery target node 2) does not actually have difference information, and therefore acquires empty data (data with no difference) as the difference information 414a.
[0109] Similarly, in step S104b, the IO processing program 421 of the node 210 (recovery target node 2) that received the difference information collection instruction sends a difference information collection instruction to the node 210 (recovery target node 1) and the node 210 (surviving node).
[0110] Next, in step S105b, the IO processing program 421 of the node 210 (surviving node) acquires the difference information 414a and sends it to the node 210 (recovery target node 2).
[0111] Furthermore, in step S105c, the IO processing program 421 of node 210 (recovery target node 1) acquires the difference information 414a in node 210 (recovery target node 1) and sends it to node 210 (recovery target node 2). When step S105c is executed, node 210 (recovery target node 1) has been recovered, and therefore becomes the target for recovery of difference information. However, at this time, node 210 (recovery target node 1) does not actually have difference information, and therefore acquires empty data (data with no difference) as the difference information 414a.
[0112] Next, in step S106a, the IO processing program 421 of node 210 (recovery target node 1) aggregates the difference information received from node 210 (recovery target node 2) and node 210 (surviving node), and stores it as difference collection information 414b in memory 212. In step S106b, the IO processing program 421 of node 210 (recovery target node 2) aggregates the difference information received from node 210 (recovery target node 1) and node 210 (surviving node), and stores it as difference collection information 414b in memory 212. Steps S106a and S106b start storing the difference information necessary for the rebuild.
[0113] Next, in step S107a, the rebuild process program 422 instructs the node 210 (recovery target node 1) and the node 210 (recovery target node 2) to reflect the difference information.
[0114] Next, in step S108a, the IO processing program 421 of the node 210 (recovery target node 1) acquires the difference collection information 414b in the node 210 (recovery target node 1) and reflects it in the difference information 414a of the node 210 (recovery target node 1).
[0115] Furthermore, in step S108b, the IO processing program 421 of the node 210 (recovery target node 2) acquires the differential collection information 414b in the node 210 (recovery target node 2) and reflects it in the differential information 414a in the node 210 (recovery target node 2). After steps S108a and S108b, if there are other blocked nodes in the storage system 101, recording of the differential information of the blocked node begins.
[0116] In steps S111 to S115, the IO processing program 421 of the node 210 (recovery target node 1) executes the same data block recovery process as the IO processing program 421 of the node 210 (recovery target node) of embodiment 1. Also, in steps S111a to S115a, the IO processing program 421 of the node 210 (recovery target node 2) executes the same data block recovery process as the IO processing program 421 of the node 210 (recovery target node) of embodiment 1.
[0117] In steps S116 to S121, the IO processing program 421 of the node 210 (recovery target node 1) executes the same parity block restoration process as the IO processing program 421 of the node 210 (recovery target node) of embodiment 1. Also, in steps S116a to S121a, the IO processing program 421 of the node 210 (recovery target node 2) executes the same parity block restoration process as the IO processing program 421 of the node 210 (recovery target node) of embodiment 1.
[0118] (Effects of the embodiment) In the above-described embodiment, if an update is made to a data block stored in the user area of a node while that node is blocked, the differential information related to the update is managed by the other two nodes that are not blocked and are operating normally. Therefore, even if two nodes are blocked at the same time, a differential rebuild can be performed based on the differential information managed by either of the two nodes.
[0119] In the above-described embodiment, first, one of the two blocked nodes is recovered. After that, the differential information managed by the surviving nodes other than the two blocked nodes is collected as first differential collection information, and the differential information managed by one of the nodes is restored based on the first differential collection information. Therefore, even if another node is newly blocked after the differential information of one of the nodes is restored, two-node management of differential information can be maintained.
[0120] Furthermore, in the above-described embodiment, after the differential information managed in one node is restored, if the first differential recovery information indicates that there is a difference in the data block or second parity stored in one node, the data block or second parity is restored by differential rebuild. Therefore, in the embodiment, by immediately performing a differential rebuild on one node after a differential rebuild on this node, it is possible to speed up recovery from two-node blockage to one-node blockage and prevent a decrease in the fault tolerance of the storage system.
[0121] In the above-described embodiment, after the restoration of the data block or second parity stored in one node by differential rebuild is completed, only the differential information related to the data block and second parity is cleared. This allows the differential information to be maintained and allows for the restoration of the second of the two blocked nodes.
[0122] In the above-described embodiment, after one node is restored, the differential information of the other node is restored, a differential rebuild is performed, and the differential information is cleared. Therefore, when two nodes are blocked, the blocked nodes are sequentially restored, so node recovery can be performed even in an unstable storage system situation, such as when two nodes are blocked, then one node is blocked, and finally two nodes are blocked.
[0123] In the above-described embodiment, recovery of the two blocked nodes, differential rebuild, and clearing of differential information are performed simultaneously, thereby enabling rapid node restoration including data and differential information for the two blocked nodes.
[0124] In the above-described embodiment, each of the multiple nodes stores the first parity and differential information in memory. Therefore, by managing the data with two nodes, even if two nodes are blocked, the differential information managed in memory can be prevented from being lost.
[0125] In the above-described embodiment, the redundancy group is an mD+nP MEC stripe configured to include m (m is an integer of 2 or greater) data blocks and n (n is an integer of 2 or greater) parities, including first and second parities. Therefore, in an MEC that combines data read / write performance and fault tolerance, it is possible to perform differential rebuilds even when two nodes are blocked.
[0126] Although several embodiments have been described above, these are merely examples for explaining the present invention, and it is not intended that the scope of the present invention be limited to these embodiments. The present invention can be implemented in various other forms. [Explanation of symbols]
[0127] 101...storage system, 210...node.
Claims
1. A storage system comprising four or more nodes, each having a processor, a memory, and a storage drive, a redundancy group is configured for user data stored in the storage drive, the redundancy group including a plurality of data blocks of the user data, and a first parity and a second parity based on the plurality of data blocks; For each of the plurality of nodes, the storage drive of the node has a user area that stores the plurality of data blocks that belong to different redundancy groups, and a parity area that stores the second parity, The processor: generating the second parity to be stored in the parity area of the node based on the first parity generated based on the plurality of data blocks belonging to different redundancy groups stored in the user area of one of the nodes other than the node, and based on the plurality of data blocks belonging to the same redundancy group that are distributed and stored in the user areas of the plurality of nodes other than the node and the one node; managing, in association with the plurality of data blocks stored in the user area of the node, and the first parity and the second parity stored in the parity area, difference information indicating the presence or absence of a difference relating to an update of either or both of the data block and the second parity that belong to the same redundancy group as each of the data blocks, the first parity, and the second parity; When an update is made to the data block stored in the user area of the node during a period when the node is blocked, the difference information relating to the update is managed in the node that is not blocked and is operating normally, out of the two nodes other than the node. A storage system comprising:
2. 2. The storage system according to claim 1, recovering one of the two blocked nodes among the plurality of nodes; After the one node is recovered, the difference information managed in the plurality of nodes excluding the two blocked nodes is collected as first difference collection information; The difference information managed in the one node is restored based on the recovered first difference recovery information. A storage system comprising:
3. 3. The storage system according to claim 2, After the differential information managed in the one node is restored, if the first differential recovery information indicates that the difference exists in the data block stored in the user area of the one node or the second parity stored in the parity area of the one node, the data block or the second parity is restored by differential rebuild. A storage system comprising:
4. 4. The storage system according to claim 3, After completion of restoration by differential rebuild of the data block stored in the user area of the one node or the second parity stored in the parity area of the one node, only the differential information related to the data block and the second parity is cleared. A storage system comprising:
5. 4. The storage system according to claim 3, After clearing the difference information related to the data block and the second parity managed in the plurality of nodes other than the two blocked nodes, recovering the other node other than the one of the two blocked nodes, After the other node is recovered, the difference information managed in the plurality of nodes excluding the other node is collected as second difference collection information; restoring the difference information managed in the other node based on the recovered second difference recovery information; After the differential information managed in the other node is restored, if the second differential recovery information indicates that the difference exists in the data block stored in the user area of the other node or the second parity stored in the parity area of the other node, restore the data block or the second parity by differential rebuild; After completion of restoration by differential rebuild of the data block stored in the user area or the second parity stored in the parity area of the other node, clear the differential information related to the data block and the second parity managed in the plurality of nodes excluding the two blocked nodes. A storage system comprising:
6. 2. The storage system according to claim 1, simultaneously recovering a first node and a second node that are blocked among the plurality of nodes; In the first node, the difference information managed in the plurality of nodes excluding the first node is collected as first difference collection information; restoring the difference information managed in the first node based on the recovered first difference recovery information; after restoring the differential information managed in the first node, if the first differential recovery information indicates that the differential exists in the data block stored in the user area of the first node or the second parity stored in the parity area, recovering the data block or the second parity by differential rebuild; In the second node, the difference information managed in the plurality of nodes excluding the second node is collected as second difference collection information; restoring the difference information managed in the second node based on the recovered second difference recovery information; after restoring the differential information managed in the second node, if the second differential recovery information indicates that the differential exists in the data block stored in the user area of the second node or the second parity stored in the parity area, restoring the data block or the second parity by differential rebuild; After completion of restoration by differential rebuild of the data blocks stored in the user areas of the first node and the second node or the second parity stored in the parity areas, the differential information related to the data blocks and the second parity managed in the plurality of nodes excluding the first node and the second node is cleared. A storage system comprising:
7. 2. The storage system according to claim 1, Each of the plurality of nodes The first parity and the difference information are stored in the memory. A storage system comprising:
8. 2. The storage system according to claim 1, The redundancy group is an mD+nP erasure coding stripe configured to include m (m is an integer of 2 or more) of the data blocks and n (n is an integer of 2 or more) of parities including the first parity and the second parity. A storage system comprising:
9. A data difference management method in a storage system including four or more nodes, each having a processor, a memory, and a storage drive, comprising: a redundancy group is configured for user data stored in the storage drive, the redundancy group including a plurality of data blocks of the user data and a first parity and a second parity based on the data blocks; For each of the plurality of nodes, the storage drive of the node has a user area that stores the plurality of data blocks that belong to different redundancy groups, and a parity area that stores the second parity, the processor: generating the second parity to be stored in the parity area of the node based on the first parity generated based on the plurality of data blocks belonging to different redundancy groups stored in the user area of one of the nodes other than the node, and based on the plurality of data blocks belonging to the same redundancy group that are distributed and stored in the user areas of the plurality of nodes other than the node and the one node; managing, in association with the plurality of data blocks stored in the user area of the node, and the first parity and the second parity stored in the parity area, difference information indicating the presence or absence of a difference relating to an update of either or both of the data block and the second parity that belong to the same redundancy group as each of the data blocks, the first parity, and the second parity; When an update is made to the data block stored in the user area of the node during a period when the node is blocked, the difference information relating to the update is managed in the node that is not blocked and is operating normally, out of the two nodes other than the node. A data difference management method characterized by comprising each process.
Citation Information
Patent Citations
Storage system
JP2023106886A