Storage System and Storage Control Method
The storage system addresses the challenge of protecting control information and cache data by using non-volatile storage, logging, and data duplication across nodes, resulting in high-performance and reliable data storage with effective recovery mechanisms.
Patent Information
- Application Number
- JP2022169428
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-10-21
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-10-21
AI Technical Summary
Existing storage systems face challenges in protecting control information and cache data, particularly in ensuring non-volatility and high reliability amidst node failures.
A storage system comprising multiple storage nodes with non-volatile storage devices, storage controllers, and volatile memory, where data related to writing is stored in memory, and critical data is logged and duplicated across nodes for redundancy and recovery.
This configuration enables high-performance and reliable data storage by ensuring non-volatile protection of control information and cache data, with efficient recovery mechanisms in case of node failures.
Smart Images

Figure 0007684266000001 
Figure 0007684266000002 
Figure 0007684266000003
Abstract
Description
Technical Field
[0001] The present invention relates to a storage system and a storage control method.
Background Art
[0002] Conventionally, in a storage system, a redundant configuration has been adopted to improve availability and reliability. For example, Patent Document 1 proposes the following storage system. In a storage system having a plurality of storage nodes, each storage node is provided with one or more storage devices that provide a storage area, and one or more storage control units that read and write the requested data to the corresponding storage device in response to a request from a higher-level device. Each storage control unit holds predetermined configuration information necessary to read and write the requested data to the corresponding storage device in response to a request from a higher-level device, a plurality of control software are managed as a redundancy group, and the configuration information held by each control software belonging to the same redundancy group is updated synchronously, and the plurality of control software constituting the redundancy group are arranged in different said storage nodes so as to disperse the load of each storage node.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] According to Patent Document 1, a storage system can be constructed using software to build a storage system (Software Defined Storage: SDS) that can continue reading and writing even during node failures. In order to improve performance and reliability in such a storage system, it is required to protect various data by making it non-volatile. The present invention aims to propose a technique for protecting control information, cache data, etc. in a storage system.
Means for Solving the Problems
[0005] To achieve the above object, one representative storage system of the present invention is a storage system including a plurality of storage nodes each having a non-volatile storage device, a storage controller that processes reading and writing of data to and from the storage device, and a volatile memory. The storage controller stores data related to the writing of the data in the memory, stores data that needs to be made non-volatile among the data stored in the memory in the storage device as log data, duplicates the log data stored in the storage device among a plurality of storage nodes, and when a problem occurs in the log data stored in the storage device of any one of the storage nodes, performs recovery processing of the log data. Also, one representative storage control method of the present invention is a storage control method in a storage system including a plurality of storage nodes each having a non-volatile storage device, a storage controller that processes reading and writing of data to and from the storage device, and a volatile memory. The storage controller stores data related to the writing of the data in the memory, stores data that needs to be made non-volatile among the data stored in the memory in the storage device as log data, duplicates the log data stored in the storage device among a plurality of storage nodes, and when a problem occurs in the log data stored in the storage device of any one of the storage nodes, performs recovery processing of the log data.
Effects of the Invention
[0006] According to the present invention, a storage system with high performance and reliability can be realized.
Brief Description of the Drawings
[0007]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Embodiments for Carrying Out the Invention
[0008] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. The embodiments relate to, for example, a storage system including a plurality of storage nodes in which one or more SDSs are implemented. In the disclosed embodiments, the storage node stores control information and cache data in the memory. And the storage node includes a non-volatile device. When updating control information and data in response to a write request from the host, the storage node stores the updated data in this non-volatile device in log format. Thereby, the updated data can be made non-volatile. Then, it responds to the host. And, asynchronously with this, the data in the memory is destaged to the storage device. In destaging, a process of writing to the storage device is performed reflecting the data written to the storage system. In destaging, various storage functions such as thin provisioning, snapshots, and data redundancy are provided, and processes such as creating a logical conversion address are performed so that data can be searched and randomly accessed. On the other hand, the storage in the non-volatile device in log format is for the purpose of restoring the data in the memory when it is lost, so the storage process is light and fast. Therefore, when using volatile memory, the response performance can be improved by quickly storing the data in the non-volatile storage device in log format and sending a completion response to the host device. When storing in log format, control information and data are stored in append-write format. Since it is stored by append-writing, it is necessary to recover free space. Two types of methods, the base image evacuation method and the garbage collection method, are selectively used for free space recovery. The base image evacuation method writes out the entire certain target area of control information and cache data to a non-volatile device and discards (recovers as free space) all the update logs in between. The garbage collection method is a method of recovering the log area by identifying unnecessary logs that are not the latest among the update logs and moving the logs other than the unnecessary logs to another area. At the time of power failure, the control information and cache data can be restored to the memory using this base image evacuation and logs without being lost. By selectively using both methods for free space recovery, the management information for free space management can be reduced, and the overhead for free space recovery can be reduced, improving the performance of the storage. In addition, the data stored in the storage device in log format is made redundant in a plurality of storage nodes. Therefore, even when a failure occurs in the storage device of any one of the storage nodes, the control information and cache data can be recovered from the storage devices of other storage nodes. Furthermore, by synchronizing the states of the memories of each storage node, it becomes possible to recover the data of the storage device from the memory of its own node.
Example
[0009] (1) Example 1 (1-1) Configuration of the storage system of Example 1 In FIG. 1, the storage system of Example 1 is shown as a whole. The storage system 100 includes, for example, a plurality of host devices 101 (Host), a plurality of storage nodes 103 (Storage Node), and a management node 104 (Management Node). The host device 101, the storage node 103, and the management node 104 are interconnected via a network 102 composed of Fibre Channel, Ethernet (registered trademark), LAN (Local Area Network), etc.
[0010] The host device 101 is a general-purpose computer device that transmits a read request or a write request (hereinafter, these are collectively referred to as I / O (Input / Output) requests as appropriate) to the storage node 103 in response to a user operation, a request from an implemented application program, etc. Note that the host device 101 may be a virtual computer device such as a virtual machine.
[0011] The storage node 103 is a computer device that provides a storage area for reading and writing data to the host device 101. The storage node 103 is, for example, a general-purpose server device. The management node 104 is a computer device used by a system administrator to manage the entire storage system 100. The management node 104 manages a plurality of storage nodes 103 as a group called a cluster. Note that in FIG. 1, an example in which only one cluster is provided is shown, but a plurality of clusters may be provided in the storage system 100.
[0012] As described above, the storage system 100 is composed of two or more storage nodes 103, one or more host devices 101, and one management node 104. The illustrated configuration is an example, and the host device 101, the storage node 103, and the management node 104 may be the same node. Also, it may be realized by a virtual machine or a container, or may be configured to coexist as a process on one machine.
[0013] FIG. 2 is a diagram showing an example of the physical configuration of the storage node 103. The storage node 103 includes a CPU (Central Processing Unit) 1031, a memory 1032, a plurality of storage devices 1033 (Drive), and a communication device 1034 (NIC: Network Interface Card).
[0014] CPU 1031 is a processor that controls the operation of the entire storage node. Memory 1032 is composed of semiconductor memories such as SRAM (Static RAM (Random Access Memory)) and DRAM (Dynamic RAM). Memory 1032 is used to temporarily hold various programs and necessary data. By executing the programs stored in the volatile memory 1032 by CPU 1031, various processes as the entire storage node 103 as described later are executed.
[0015] The storage device 1033 is composed of one or more types of large-capacity non-volatile storage devices such as an SSD (Solid State Drive), a SAS (Serial Attached SCSI (Small Computer System Interface)) hard disk drive, and a SATA (Serial ATA (Advanced Technology Attachment)) hard disk drive. The storage device 1033 provides a physical storage area for reading or writing data in response to I / O requests from the host device 101.
[0016] The communication device 1034 is an interface for the storage node 103 to communicate with the host device 101, other storage nodes 103, or the management node 104 via the network 102. The communication device 1034 is composed of, for example, a NIC, an FC card, etc. The communication device 1034 performs protocol control during communication with the host device 101, other storage nodes 103, or the management node 104.
[0017] Figure 3 is a diagram showing an example of the logical configuration of the storage node 103. The storage node 103 includes a front-end driver 1081 (Front-end a front-end driver 1081, a back-end driver 1087, one or more storage controllers 1083, and a data protection controller 1086.
[0018] The front-end driver 1081 is software that controls the communication device 1034 and provides the CPU 1031 with an abstracted interface for communication between the host device 101, other storage nodes 103, or management nodes 104 with the storage controller 1083.
[0019] The back-end driver 1087 is software that controls each storage device 1033 within the self-storage node 103 and provides the CPU 1031 with an abstracted interface for communication with each storage device 1033.
[0020] The storage controller 1083 is software that functions as a controller for the SDS. The storage controller 1083 receives I / O requests from the host device 101 and issues I / O commands corresponding to the I / O requests to the data protection controller 1086. The storage controller 1083 also has a logical volume configuration function. The logical volume configuration function associates the logical chunks configured by the data protection controller with the logical volumes provided to the host. For example, a straight mapping method (associating the logical chunks and logical volumes one-to-one and making the addresses of the logical chunks and logical volumes the same) may be used, or a virtual volume function (Thin Provisioning) method (dividing the logical volumes and logical chunks into small-sized areas (pages) and associating the addresses of the logical volumes and logical chunks in page units) may be adopted.
[0021] In the case of Embodiment 1, each storage controller 1083 implemented in the storage node 103 is managed as a pair that forms a redundant configuration together with another storage controller 1083 arranged in another storage node 103. Hereinafter, this pair will be referred to as a storage controller group 1085. Note that FIG. 3 shows a case where one storage controller group 1085 is constituted by two storage controllers 1083. Hereinafter, the description will proceed on the assumption that one storage controller group 1085 is constituted by two storage controllers 1083, but one redundant configuration may be constituted by three or more storage controllers 1083.
[0022] In the storage controller group 1085, one storage controller 1083 is set to a state in which it can accept an I / O request from the host device 101 (the state of the active system, hereinafter referred to as the active mode). Also, in the storage controller group 1085, the other storage controller 1083 is set to a state in which it does not accept an I / O request from the host device 101 (the state of the standby system, hereinafter referred to as the standby mode). Note that the node in the active mode is called an active node, and the node in the standby mode is called a standby node. Then, in the storage controller group 1085, when a failure occurs in the storage node 103 where the storage controller 1083 set to the active mode (hereinafter referred to as the active storage controller) is arranged, etc., the state of the storage controller 1083 that has been set to the standby mode until then (hereinafter referred to as the standby storage controller) is switched to the active mode. Thereby, when the active storage controller becomes unable to operate, the I / O processing being executed by the active storage controller can be taken over by the standby storage controller.
[0023] The data protection control unit 1086 is software that allocates a physical storage area provided by a storage device 1033 within its own storage node 103 or within another storage node 103 to each storage controller group 1085, and reads or writes specified data to the corresponding storage device 1033 according to the above-described I / O commands given from the storage controller 1083. In this case, when the data protection control unit 1086 allocates a physical storage area provided by a storage device 1033 within another storage node 103 to the storage controller group 1085, it cooperates with the data protection control unit 1086 implemented in the other storage node 103, and exchanges data with the data protection control unit 1086 via the network 102, and reads or writes the data to the storage area according to the I / O commands given from the active storage controller of the storage controller group 1085.
[0024] FIG. 4 is a diagram for explaining an outline of the disclosed storage system and storage control method. For I / O processing from a host and various other processes, the storage controller updates control information and cache data. At this time, the control information / cache data on the memory is updated, and the log thereof is stored in the storage device to make it non-volatile. For this purpose, an update log is created in the control information log buffer or the cache log buffer. A log consists of the update data itself and a log header, and is information indicating how the control information and cache data on the memory are updated. As shown in FIG. 7, the log header includes information indicating the update position, the update size, and the order relationship between updates.
[0025] The update log on the log buffer is written to the log area on the storage device in an append write format. This writing may be performed immediately or asynchronously. In order to perform append writing, the free space in the log area on each device gradually decreases and eventually becomes full, making it impossible to write. To avoid this, it is necessary to reclaim the free space. Different methods are used for the log area for control information and the log area for cache data respectively.
[0026] For control information, the base image evacuation method is used. In the base image evacuation method, the entire control information is copied as a whole to the base image area on the storage device. After the copy is completed, all update logs before the start of the copy are invalidated (reclaimed as free space). On the other hand, the garbage collection method is used to reclaim the free space in the log area for cache data. Cache data logs become invalid when the cache data is overwritten (by the asynchronous stage described later) or deleted from the cache. The garbage collection method reclaims the log area as free space by copying valid old logs to the end of the log area as new logs, excluding the invalidated logs.
[0027] Figure 5 is an example of the memory configuration diagram. Inside the memory, storage control information 10321, cache data area 10323, cache data log header management table 10324, control information log buffer 10325, and cache data log buffer 10326 are stored.
[0028] Storage control information 10321 is an area where control information for realizing various storage functions is stored. As an example, there is a cache directory 10322. The cache data log header management table 10324 is a table that stores the log headers of all cache data logs on the disk. The control information log buffer 10325 temporarily holds the log of control information. The cache data log buffer 10326 temporarily holds the log of cache data.
[0029] FIG. 6 is an example of a configuration diagram of a memory device. On the memory device, there are a control information base image area 10332, a control information log area 10333, a cache data log area 10334, and a persistence area 10335. The control information base image area is an area for copying all control information in the base image evacuation process described later. The control information log area 10333 and the cache data log area 10334 are areas to which the logs are evacuated in the log evacuation process described later, respectively. The persistence area 10335 is an area for storing user data, which is managed by the data protection control unit 1086.
[0030] FIG. 7 shows the structure of a log header. The log header is a table included in each log stored in a log buffer area on the memory or a log area on the memory device. Each log header has fields for a log sequence number, an update address, an update size, an area type, and a valid flag.
[0031] The log sequence number field stores a log sequence number uniquely assigned to each log. The update address field stores the address of the control information or cache data to be updated for each log. The update size field stores the size to be updated. The area type field stores a value for identifying either control information or cache data. Here, it is assumed that the character strings "control information" or "cache data" are stored. The valid flag field is set with a value of "valid" or "invalid".
[0032] FIG. 8 is an explanatory diagram of log data generation and redundancy. The storage controller of the active node processes I / O and updates the control information and cache data in the memory according to its operation. Then, it stores the control information and cache data in the log buffer, and creates log data from the control information and cache data in the log buffer. Specifically, the storage controller of the active node stores the control information in the control information log buffer and the cache data in the cache data log buffer. Then, it generates log data of the control information from the control information in the control information log buffer and stores it in the control information log area of the storage device. Similarly, it generates log data of the cache data from the cache data in the cache data log buffer and stores it in the cache data log area of the storage device. In addition, the storage controller of the active node transfers the control information and cache data in the log buffer to the standby node.
[0033] The storage controller of the standby node does not perform I / O processing, but has replicas of the control information and cache data to take over the service when the storage controller of the active node stops. Therefore, the storage controller of the standby node stores the control information received from the active node in the control information log buffer and the cache data received from the active node in the cache data log buffer. Then, it uses the control information and cache data stored in the log buffer to make the state of the memory of the standby node consistent with the state of the memory of the active node. Furthermore, the storage controller of the standby node generates log data of the control information from the control information in the control information log buffer and stores it in the control information log area of the storage device. Similarly, it generates log data of the cache data from the cache data in the cache data log buffer and stores it in the cache data log area of the storage device.
[0034] Figure 9 is a flowchart of the control information update process. When the active storage controller receives an I / O, if a failure occurs in the active storage controller and I / O reception becomes impossible, the standby storage controller takes over the I / O. To achieve this, the control information updates of the active and standby are also reflected in the standby.
[0035] The control information update process is called when updating the control information in the memory. When called, the memory address and size for identifying the control information to be updated, the update value, and information indicating whether non-volatilization is required are passed.
[0036] First, the storage controller of the active node (active storage controller) updates the control information in the memory (step S101). Next, it refers to the non-volatilization requirement passed and determines whether it is required (step S102). If non-volatilization is not required (step S102; NO), the process ends as it is. If non-volatilization is required (step S102; YES), the log creation process is called (S103).
[0037] After the log creation process, the active storage controller sends the log to the storage controller of the standby node (standby storage controller) (step S104) and ends the process.
[0038] The standby storage controller receives the log from the active storage controller (step S201), reflects the control information update of the active storage controller in its own node's memory (step S202), stores the log data in its own node's storage device (step S203), and ends the process.
[0039] Figure 10 is a flowchart of the log creation process. In this process, the "log buffer" refers to the control information log buffer when the update target is control information, and the user data cache log buffer when it is the user data cache.
[0040] First, the storage controller determines a log sequence number (step S301). The log sequence number is assigned in the order of log creation, and there is always one log corresponding to one log sequence number. Next, it secures an area in the log buffer where the next log will be written (step S302).
[0041] The log creation process may be executed by a plurality of parallel processes. In that case, it is necessary to perform exclusive processing so that the same log sequence number is not obtained by different processes and the same log buffer area is not secured by different processes.
[0042] Next, a log header is created (step S303). The log sequence number described above is stored in the sequence number field of the log header, and the values of the update address and update size in the memory passed to this log creation process are stored in the update address field and update size field. In the information type field, "control information" is stored when the control information is updated, and "cache data" is stored when the cache data is updated.
[0043] Next, the log is stored in the log buffer (step S304). The log consists of the log header and the update target data itself. The log header is stored at the beginning of the secured area secured on the log buffer, and the update data itself is stored at the memory address obtained by adding the log header size to the secured area. Finally, the log valid flag in the log header is set to "valid" (step S305), and this process ends.
[0044] Figure 11 is a flowchart of the cache data update process. Steps S401 to S404 are the same as steps S101 to S104, except that the update target is cache data instead of control information. In the cache data update process, different from the control information update process, when it is determined in step S402 that non-volatile storage is required, the processes of steps S405 to S407 are added.
[0045] In step S405, the active storage controller determines whether it is an overwrite (step S405). That is, it refers to the cache data log header management table, searches for whether there is a log at the same address, and if so, determines it is an overwrite. If it is an overwrite (step S405; YES), the log at the same address in the log header management table is invalidated (step S406). By this invalidation, the valid flag is set to "invalid". After step S406, or if it is not an overwrite (step S405; NO), the active storage controller adds the log header of the log created in step S403 to the log header management table (step S407) and ends the process.
[0046] The standby storage controller receives the log from the active storage controller (step S501), and reflects the cache data update of the active storage controller in its own node's memory (step S502). Then, the standby storage controller determines whether it is an overwrite (step S503). That is, it refers to the cache data log header management table, searches for whether there is a log at the same address, and if so, determines it is an overwrite. If it is an overwrite (step S503; YES), the log at the same address in the log header management table is invalidated (step S504). By this invalidation, the valid flag is set to "invalid". After step S504, or if it is not an overwrite (step S503; NO), the standby storage controller adds the log header of the log received in step S501 to the log header management table (step S505) and ends the process.
[0047] Figure 12 shows the processing flow of the log backup process. First, the active storage controller sends a failover instruction to the standby storage controller (step S601). Then, the active storage controller refers to the log buffer and reads out the non-failed-over logs (step S602). Next, the active storage controller stores the non-failed-over logs in the log area on the storage device (step S603). The write position is immediately after the last written log. When the writing is completed, the active storage controller deletes the logs from the log buffer in the memory (step S604) and ends the process. The standby storage controller receives the failover instruction from the active storage controller (step S701). Then, the standby storage controller refers to the log buffer and reads out the non-failed-over logs (step S702). Next, the standby storage controller stores the non-failed-over logs in the log area on the storage device (step S703). The write position is immediately after the last written log. When the writing is completed, the standby storage controller deletes the logs from the log buffer in the memory (step S704) and ends the process.
[0048] FIG. 13 is an explanatory diagram of log recovery (pattern 1). In the log recovery of pattern 1, logs are recovered from the storage devices of the paired nodes with a redundant configuration. When a storage device with a log area fails, since the log areas of the storage controller group 1085 are synchronized, replication is possible from the log areas of the paired nodes.
[0049] When a failure occurs, the failed node or the management node issues a log area transfer instruction to the paired node. The paired node that receives the instruction reads the logs from the storage device and transfers them to the failed node. The failed node recovers the log area according to the log creation processes of the control information update process and the cache data update process.
[0050] FIG. 14 is a flowchart of the log recovery process (pattern 1). The failed node sends a log transfer request to the paired node (step S801). When the peer node receives a log transfer request (step S901), it refers to the cache data log header management table and acquires valid log information (step S902). The peer node reads the valid log from the log area in the device (step S903) and transmits it to the node where the failure has occurred (step S904). The node where the failure has occurred receives the log from the peer node (step S802), writes the log to the drive (storage device 1033) (step S803), and ends the process.
[0051] FIG. 15 is an explanatory diagram of log recovery (pattern 2). In the log recovery of pattern 2, the node where the failure has occurred that has detected a failure in the storage drive and the peer node recover the log from their respective node memories. When a failure occurs in the storage device 1033 having a log area, since the information stored in the log area is also stored in the control information and cache data in the memory, replication from the memory is possible.
[0052] When a failure occurs, the node where the failure has occurred or the management node performs log recovery processing. In this log recovery processing, the node where the failure has occurred and the peer node evacuate the entire control information to the control information base image area. Also, the node where the failure has occurred and the peer node refer to the cache data log header management table to acquire valid log information, read data from the cache data, generate the log again, and store it in the cache data log area. The active side also executes data writing in order to synchronize the base image of the control information.
[0053] FIG. 16 is a flowchart of log recovery processing (pattern 2-1: control information recovery). First, the node where the failure has occurred transmits a control information base image evacuation request to the peer node (step S1001). Thereafter, the node where the failure has occurred evacuates the base image of the control information to its own node's storage device (step S1002) and ends the process. The peer node receives a control information-based image evacuation request from the failed node (step S1101), evacuates the base image of the control information to the storage device of its own node (step S1102), and ends the process.
[0054] Figure 17 is a flowchart of the log recovery process (pattern 2-2: cache data log recovery). The storage controller of the node that recovers the cache data log scans the cache data log header management table (step S1201) and determines whether the log is valid (step S1202). If the log is invalid (step S1202; NO), since the log has been destaged, the process for that log ends.
[0055] If the log is valid (step S1202; YES), the storage controller reads data from the cache data (step S1203) and regenerates the log again using the read data (step S1204).
[0056] After that, the storage controller determines whether it is an overwrite (step S1205). That is, it refers to the cache data log header management table, searches for whether there is a log at the same address, and if so, determines it is an overwrite. If it is an overwrite (step S1205; YES), the log at the same address in the log header management table is invalidated (step S1206). By this invalidation, the valid flag is set to "invalid". After step S1206, or if it is not an overwrite (step S1205; NO), the storage controller adds the log header of the log created in step S1203 to the log header management table (step S1207) and ends the process.
[0057] Figure 18 is a flowchart at the time of node failure recovery or node addition / removal. The log recovery process can also handle node failure recovery and node addition / removal. The log recovery process can handle either pattern 1 or pattern 2.
[0058] First, a memory transfer is requested from a node that performs failure recovery or addition / removal to a pair node (step S1301). Upon receiving the memory transfer request (step S1401), the pair node reads control information, cache data, and the log header management table from the memory (step S1402) and transfers them to the requester (step S1403). The node that performs failure recovery or addition / removal reflects the received content in the memory (step S1302), regenerates the log by log recovery processing (step S1303), and ends the process.
[0059] Note that the same applies during drive addition / removal. When the number of drives changes, it is necessary to relocate the log, so the storage controller reconfigures the log area after the drive number change. Also, after reconfiguration, the log is regenerated and re-stored by log recovery processing (capable of handling either pattern 1 or pattern 2).
[0060] FIG. 19 is a flowchart of garbage collection for the cache data log area. In this case, recovery processing pattern 2-2 is applied. The data stored in the cache data log area becomes unnecessary (invalid) due to overwriting or drive reflection of the cache. Since the area storage order and the invalidation order do not match, fragmentation occurs. Therefore, garbage collection is performed when the continuous free area falls below a certain value.
[0061] Specifically, the storage controller determines whether the continuous free area size of the log area is less than or equal to a certain value (step S1501). If the free area size exceeds the certain value (step S1501; NO), the process ends as it is.
[0062] If the free area size is equal to or less than a certain value (step S1501; YES), the storage controller performs the log recovery process pattern 2-2 on an area equal to or greater than a predetermined value, reproduces valid logs identified from the log header management table, and stores them in the log buffer (step S1502). After that, the storage controller writes the logs stored in the log buffer to the disk by log evacuation processing (step S1503), releases the area where the old logs of the area were stored (step S1504), and ends the process.
[0063] As described above, the disclosed storage system 100 is a storage system including a plurality of storage nodes 103 each having a non-volatile storage device 1033, a storage controller 1083 that processes data read / write to / from the storage device, and a volatile memory 1032. The storage controller 1083 stores data related to writing of the data in the memory 1032, stores data that needs to be made non-volatile among the data stored in the memory 1032 in the storage device 1033 as log data, duplicates the log data stored in the storage device 1033 among a plurality of storage nodes, and performs recovery processing of the log data when a problem occurs in the log data stored in the storage device 1033 of any one of the storage nodes. According to this configuration and operation, since the log data can be expanded, a storage system having both high performance and reliability can be realized.
[0064] In addition, the plurality of storage nodes duplicate the data stored in the memory 1032, and when a failure occurs in the storage device 1033 of any one of the storage nodes, the storage controller 1083 of each storage node re-creates the log data from the data stored in its own node's memory. Therefore, log re-redundancy without drive access and network communication can be realized.
[0065] In addition, when a failure occurs in the storage device 1033 of any of the storage nodes, the storage controller 1083 of the storage node can also acquire log data stored in the storage devices 1033 of other storage nodes and redundantize the log data again. Therefore, the log of the storage device 1033 can be reliably protected.
[0066] In addition, the plurality of storage nodes include an active node in an active state and a standby node in a standby state. The storage controller 1083 of the active node stores control information and cache data in a log buffer, creates the log data from the control information and cache data in the log buffer, and transfers the control information and cache data in the log buffer to the standby node. The storage controller 1083 of the standby node makes the state of the memory coincide with that of the active node using the control information and cache data received from the active node, and generates the log data using the control information and cache data received from the active node. Therefore, the states of the memories of the plurality of storage nodes can be synchronized.
[0067] In addition, the storage controller 1083 of the active node determines whether non-volatilization of the control information and cache data is required, and stores the control information and cache data that require non-volatilization in the log buffer. Therefore, the control information and cache data can be efficiently made non-volatile.
[0068] In addition, the storage controller 1083 of the active node performs garbage collection to write the cache data to the storage device 1033 in log units and recover free space. Therefore, the cache data can be efficiently made non-volatile.
[0069] Also, when the log data is lost in any of the plurality of storage nodes belonging to the same group, all the storage nodes belonging to the group store the base image of the control information as at least a part of the log data in the storage device of their own nodes. Therefore, it is possible to synchronize the log data of the control information possessed by the plurality of storage nodes belonging to the same group.
[0070] When the number of the storage nodes changes, the storage controller 1083 executes non-volatile storage of the data stored in the memory 1032. Therefore, even when the number of the storage nodes changes, it is possible to synchronize the log data.
[0071] Note that the present invention is not limited to the above-described embodiments and includes various modifications. For example, the above-described embodiments have been described in detail for easy understanding of the present invention, and are not necessarily limited to those having all the configurations described. Also, not only deletion of such a configuration is possible, but also replacement and addition of a configuration are possible.
Explanation of Reference Numerals
[0072] 100: Storage system, 101: Host device, 102: Network, 103: Storage node, 104: Management node, 1031: CPU, 1032: Memory, 1033: Drive, 1083: Storage controller
Claims
1. A storage system including a plurality of storage nodes each having a non-volatile memory device, a storage controller that processes reading and writing of data to and from the memory device, and a volatile memory, wherein the storage controller stores data related to writing of the data in the memory, stores, as log data, data that needs to be made non-volatile among the data stored in the memory in the memory device, redundantly stores the log data stored in the memory device among a plurality of storage nodes, and performs recovery processing of the log data when a problem occurs in the log data stored in the memory device of any one of the storage nodes, an active controller and a standby controller form a group, the plurality of storage nodes include active nodes each having an active controller and standby nodes each having a standby controller, the storage controller of the active node stores control information and cache data in a log buffer, creates the log data from the control information and cache data in the log buffer, and transfers the control information and cache data in the log buffer to the standby node, the storage controller of the standby node matches the state of the memory with that of the active node using the control information and cache data received from the active node, and generates the log data using the control information and cache data received from the active node A storage system characterized by the above.
2. The storage system according to claim 1, wherein the plurality of storage nodes redundantly store the data stored in the memory, when a failure occurs in the memory device of any one of the storage nodes, the storage controller of each storage node re-creates the log data from the data stored in the memory of its own node.
3. The storage system according to claim 1, wherein A storage system, wherein when a failure occurs in the storage device of any one of the storage nodes, the storage controller of the storage node acquires log data stored in the storage devices of other storage nodes and re - redundantizes the log data.
4. The storage system according to claim 1, wherein the storage controller of the active node determines whether non - volatility is required for the control information and cache data, and stores the control information and cache data that require non - volatility in the log buffer.
5. The storage system according to claim 1, wherein the storage controller of the active node performs garbage collection for writing the cache data to the storage device in log units to recover free space.
6. The storage system according to claim 1, wherein when the log data is lost in any of a plurality of storage nodes belonging to the same group, all storage nodes belonging to the group store the base image of the control information as at least a part of the log data in the storage device of their own nodes.
7. The storage system according to claim 1, wherein the storage controller performs non - volatility of the data stored in the memory when the number of the storage nodes changes.
8. A storage control method in a storage system including a plurality of storage nodes each having a non - volatile storage device, a storage controller for processing reading and writing of data to and from the storage device, and a volatile memory, wherein the storage controller stores data related to the writing of the data in the memory, stores, as log data in the storage device, data that requires non - volatility among the data stored in the memory, redundantizes the log data stored in the storage device among a plurality of storage nodes, and performs recovery processing of the log data when a problem occurs in the log data stored in the storage device of any one of the storage nodes. An active controller and a standby controller form a group. The plurality of storage nodes include an active node having an active controller and a standby node having a standby controller. The storage controller of the active node stores control information and cache data in a log buffer, creates the log data from the control information and cache data in the log buffer, and transfers the control information and cache data in the log buffer to the standby node. The storage controller of the standby node makes the state of the memory match that of the active node using the control information and cache data received from the active node, and generates the log data using the control information and cache data received from the active node. A storage control method characterized by the above.
Citation Information
Patent Citations
Disaster recovery method and system
JP2006338064A
Storage system and control software arrangement method
JP2019101703A
Storage apparatus, data management method, and data management program
JP2019192004A
Storage system, and method for recovering storage system
JP2020135138A
Remote copy system and remote copy management method
JP2021124889A