Storage System and Storage Control Method

The storage system addresses the challenge of protecting control information and cache data by using non-volatile storage and volatile memory to store base images and logs, resulting in high performance and reliability with efficient data recovery.

JP7695224B2Active Publication Date: 2025-06-18HITACHI VANTARA LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022188843
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-11-28
Publication Date
2025-06-18
Estimated Expiration
2042-11-28

AI Technical Summary

Technical Problem

Existing storage systems face challenges in protecting control information and cache data, particularly in ensuring data reliability and performance during node failures and power outages.

Method used

A storage system with non-volatile storage devices and volatile memory, where control information and cache data are stored as base images in multiple storage areas, and logs are maintained for recovery purposes, enabling efficient data restoration and high performance.

Benefits of technology

The proposed solution achieves high performance and reliability by ensuring data integrity and quick recovery from failures, while maintaining efficient storage operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007695224000001
    Figure 0007695224000001
  • Figure 0007695224000002
    Figure 0007695224000002
  • Figure 0007695224000003
    Figure 0007695224000003
Patent Text Reader

Abstract

To realize a storage system including both of a high performance and high reliability.SOLUTION: A storage system 100 includes one or more storage nodes 103 including a non-volatile storage device 1033, a storage controller 1083, and a volatile memory 1032. The storage device 1033 includes a plurality of base image storing regions that includes at least a first base image storing region and a second base image storing region as a region to store entire predetermined information stored in the memory 1032 as a base image. When completing storage of the base image into the first base image storing region, the storage controller 1083 starts processing of storing a next base image in the second base image storing region, and when predetermined information is lost from the memory 1032, reads out the base image that has been stored and restores the same in the memory.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a storage system and a storage control method.

Background Art

[0002] Conventionally, in a storage system, a redundant configuration has been adopted to improve availability and reliability. For example, Patent Document 1 proposes the following storage system. In a storage system having a plurality of storage nodes, each storage node is provided with one or more storage devices that provide a storage area, and one or more storage control units that read and write the requested data to the corresponding storage device in response to a request from a higher-level device. Each storage control unit holds predetermined configuration information necessary to read and write the requested data to the corresponding storage device in response to a request from a higher-level device, a plurality of control software is managed as a redundancy group, and the configuration information held by each control software belonging to the same redundancy group is updated synchronously, and the plurality of control software constituting the redundancy group is arranged on different said storage nodes so as to disperse the load of each storage node.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] According to Patent Document 1, a storage system can be constructed using software to build a storage system (Software Defined Storage: SDS) that can continue reading and writing even during node failures. In order to improve performance and reliability in such a storage system, it is required to protect various data by making it non-volatile. The present invention aims to propose a technique for protecting control information, cache data, etc. in a storage system.

Means for Solving the Problem

[0005] To achieve the above object, one of the typical storage systems of the present invention is a storage system including one or more storage nodes having a non-volatile storage device, a storage controller that processes reading and writing of data to and from the storage device, and a volatile memory. The storage device includes a plurality of base image storage areas including at least a first base image storage area and a second base image storage area as areas for storing the entire predetermined information stored in the memory as a base image. The storage controller performs a process of storing the base image in the first base image storage area. When the storage of the base image in the first base image storage area is completed, the storage controller starts a process of storing the next base image in the second base image storage area. When the predetermined information is lost from the memory, a recovery process is performed to read the stored base image and restore it to the memory. Also, one of the typical storage control methods of the present invention is a storage control method in a storage system including one or more storage nodes having a non-volatile memory device, a storage controller that processes reading and writing of data to and from the memory device, and a volatile memory, wherein the storage controller stores the entire predetermined information stored in the memory as a base image in a first base image storage area provided in the memory device; when the storage of the base image in the first base image storage area is completed, stores the next base image in a second base image storage area provided in the memory device; and when the predetermined information is lost from the memory, performs a recovery process of reading the stored base image and restoring it to the memory.

Effect of the Invention

[0006] According to the present invention, a storage system having both high performance and reliability can be realized.

Brief Description of the Drawings

[0007]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Mode for Carrying Out the Invention

[0008] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. The embodiments relate to, for example, a storage system including a plurality of storage nodes in which one or more SDSs are mounted. In the disclosed embodiments, the storage node stores control information and cache data in the memory. And the storage node is provided with a non-volatile device. When the storage node updates control information and data in response to a write request from the host, it stores the updated data in this non-volatile device in a log format. This enables the updated data to be made non-volatile. Then, it responds to the host. And, asynchronously with this, it destages the data in the memory to the storage device. In destaging, a process of writing to the storage device is performed reflecting the data written to the storage system. In destaging, various storage functions such as thin provisioning, snapshots, and data redundancy are provided, and processes such as creating a logical conversion address are performed to enable data to be searched and randomly accessed. On the other hand, storing in the non-volatile device in log format is for the purpose of restoring the data in the memory when it is lost, so the storage process is light and fast. Therefore, when using volatile memory, by quickly storing the data in the non-volatile storage device in log format and responding with completion to the host device, the response performance can be improved. When storing in log format, control information and data are stored in an append-write format. Since it is stored by append-write, it is necessary to recover the free space. Two types of methods, the base image evacuation method and the garbage collection method, are selectively used for free space recovery. The base image evacuation method is a method of writing out the entire certain target area of control information and cache data to the non-volatile device and discarding (recovering as free space) all the update logs in between. The garbage collection method is a method of identifying unnecessary logs that are not the latest among the update logs and recovering the log area by writing the logs other than the unnecessary logs to another area. At the time of power failure, the control information and cache data can be restored to the memory using this base image evacuation and logs without being lost. By selectively using both methods for free space recovery, the management information for free space management can be reduced, and the overhead for free space recovery can be reduced, improving the performance of the storage. In addition, data stored in the storage device in log format is made redundant across multiple storage nodes. Therefore, even if a failure occurs in the storage device of any one of the storage nodes, control information and cache data can be recovered from the storage devices of other storage nodes. Furthermore, by synchronizing the states of the memories of each storage node, it becomes possible to recover data from the storage device from the memory of its own node.

Example

[0009] Figure 1 is an explanatory diagram of the backup of control information. First, the storage node includes a storage controller 1083, a volatile memory 1032, and a non-volatile storage device 1033. The memory 1032 stores control information and a log buffer for the control information. The storage device includes a base image storage area #0 which is a first base image storage area, a base image storage area #1 which is a second base image storage area, a control information log storage area #0 which is a first log storage area, and a control information log storage area #0 which is a second log storage area.

[0010] The base image storage area #0 and the base image storage area #1 can store the entire control information in the memory 1032 as a base image. The control information log storage area #0 and the control information log storage area #1 can store the updated content of the control information, that is, the content of the control information log buffer.

[0011] The storage controller 1083 prepares two or more storage areas in order to always leave a base image in a backup-completed state in the storage device 1033, and alternately backs up the base image there (1). Figure 1 shows an example in which two storage areas, a base image storage area and a log backup area, are prepared for each side and they are made to correspond to each other.

[0012] The storage controller 1083 performs log backup processing as a normal operation during operation. Specifically, among the two sides (the base image storage area #0 and the base image storage area #1), the base image backup is performed on one side (2-1).

[0013] Also, the storage controller 1083 stores the updated content of the control information as a log in the log storage area corresponding to the backup surface of the base image. At this time, the logs are stored in order from the beginning of the storage area (2-2).

[0014] When the backup of the base image is completed and the corresponding log storage area approaches full, the storage controller 1083 performs "surface switching" (3-1). In surface switching, the surface for which backup has been completed (in the figure, the base image storage area #0) is set as the confirmed surface. Also, the old base image (the base image stored in the base image storage area #1 in the figure) is discarded and initialized, and a new base image is backed up (3-2). When performing surface switching, the storage controller 1083 invalidates the entire log storage area before the start of the backup of the confirmed base image all at once for the logs, and backs up the new logs (4).

[0015] Figure 2 is an explanatory diagram of the recovery process after loss of control information. When a power loss or the like occurs, the control information in the volatile memory 1032 is lost. When the control information is lost from the memory 1032, the storage controller performs a recovery process of reading the confirmed base image and logs from the storage device 1033 and applying them to the memory 1032.

[0016] In Figure 2, when the backup of the base image to the base image storage area #0 is completed, and then the logs in the log backup area #0 become full and surface switching is performed, and a power failure occurs during the backup of the base image to the base image storage area #1, the recovery procedures Step.1 to Step.3 are shown. Step.1 Apply the base image (#0) that was the confirmed surface at the time of the power failure. Step.2 Apply the logs in the log storage area (#0) of the confirmed surface. Step.3 Apply the logs in the other log storage area (#1). By this procedure, the storage controller 1083 recovers the state of the control information area at the time of the power failure (without rolling back).

[0017] Figure 3 is a configuration diagram of the storage system of Embodiment 1. The storage system 100 includes, for example, a plurality of host devices 101 (Host), a plurality of storage nodes 103 (Storage Node), and a management node 104 (Management Node). The host device 101, the storage node 103, and the management node 104 are interconnected via a network 102 composed of, for example, Fibre Channel, Ethernet (registered trademark), LAN (Local Area Network), etc.

[0018] The host device 101 is a general-purpose computer device that transmits a read request or a write request (hereinafter, these are collectively referred to as an I / O (Input / Output) request as appropriate) to the storage node 103 in response to requests from user operations, implemented application programs, etc. Note that the host device 101 may be a virtual computer device such as a virtual machine.

[0019] The storage node 103 is a computer device that provides a storage area for reading and writing data to the host device 101. The storage node 103 is, for example, a general-purpose server device. The management node 104 is a computer device used by a system administrator to manage the entire storage system 100. The management node 104 manages a plurality of storage nodes 103 as a group called a cluster. Note that in FIG. 1, an example in which only one cluster is provided is shown, but a plurality of clusters may be provided in the storage system 100.

[0020] As described above, the storage system 100 is composed of one or more storage nodes 103, one or more host devices 101, and one management node 104. The illustrated configuration is an example, and the host device 101, the storage node 103, and the management node 104 may be the same node. Further, it may be realized by a virtual machine or a container, or may be configured to coexist as a process on one machine. Also, the network frame may be made redundant, or may be separated into a management network and a storage network.

[0021] FIG. 4 is a diagram showing an example of the physical configuration of the storage node 103. The storage node 103 includes a CPU (Central Processing Unit) 1031, a memory 1032, a plurality of storage devices 1033 (Drive), and a communication device 1034 (NIC: Network Interface Card).

[0022] The CPU 1031 is a processor that controls the operation of the entire storage node. The memory 1032 is composed of a semiconductor memory such as SRAM (Static RAM (Random Access Memory)) or DRAM (Dynamic RAM). The memory 1032 is used to temporarily hold various programs and necessary data. By the CPU 1031 executing the programs stored in the volatile memory 1032, various processes of the entire storage node 103 as described below are executed.

[0023] The storage device 1033 is composed of one or more types of large-capacity non-volatile storage devices such as an SSD (Solid State Drive), a SAS (Serial Attached SCSI (Small Computer System Interface)) hard disk drive, or a SATA (Serial ATA (Advanced Technology Attachment)) hard disk drive. The storage device 1033 provides a physical storage area for reading or writing data in response to an I / O request from the host device 101.

[0024] The communication device 1034 is an interface for the storage node 103 to communicate with the host device 101, other storage nodes 103, or the management node 104 via the network 102. The communication device 1034 is composed of, for example, a NIC, an FC card, etc. The communication device 1034 performs protocol control during communication with the host device 101, other storage nodes 103, or the management node 104.

[0025] FIG. 5 is a diagram showing an example of the logical configuration of the storage node 103. The storage node 103 includes a front-end driver 1081 (Front-end driver), a back-end driver 1087 (Back-end driver), one or more storage controllers 1083 (Storage Controller), and a data protection control unit 1086 (Data Protection Controller).

[0026] The front-end driver 1081 is software that controls the communication device 1034 and has a function of providing an abstracted interface to the CPU 1031 for communication between the storage controller 1083 and the host device 101, other storage nodes 103, or the management node 104.

[0027] The back-end driver 1087 is software that controls each storage device 1033 within the own storage node 103 and has a function of providing an abstracted interface to the CPU 1031 for communication with each storage device 1033.

[0028] The storage controller 1083 is software that functions as a controller for the SDS. The storage controller 1083 receives I / O requests from the host device 101 and issues I / O commands corresponding to the I / O requests to the data protection control unit 1086. Further, the storage controller 1083 has a logical volume configuration function. The logical volume configuration function associates the logical chunks configured by the data protection control unit with the logical volumes provided to the host. For example, a straight mapping method (associating logical chunks and logical volumes one-to-one and making the addresses of logical chunks and logical volumes the same) may be used, or a virtual volume function (Thin Provisioning) method (dividing logical volumes and logical chunks into small-sized areas (pages) and associating the addresses of logical volumes and logical chunks with each other in page units) may be adopted.

[0029] In the case of the first embodiment, each storage controller 1083 implemented in the storage node 103 is managed as a storage controller group 1085 that forms a redundant configuration together with other storage controllers 1083 arranged in another storage node 103. When there are two storage nodes 103 belonging to the storage controller group 1085, it can also be called a storage controller pair.

[0030] In the storage controller group 1085, one storage controller 1083 is set to a state in which it can receive I / O requests from the host device 101 (the state of the active system, hereinafter referred to as the active mode). Further, in the storage controller group 1085, the other storage controller 1083 is set to a state in which it does not receive I / O requests from the host device 101 (the state of the standby system, hereinafter referred to as the standby mode). Note that the node in the active mode is called the active node, and the node in the standby mode is called the standby node. In the storage controller group 1085, when a failure occurs in the storage node 103 where the storage controller 1083 set to the active mode (hereinafter referred to as the active storage controller) is arranged, or the like, the state of the storage controller 1083 that has been set to the standby mode (hereinafter referred to as the standby storage controller) is switched to the active mode. Thereby, when the active storage controller becomes inoperable, the I / O processing executed by the active storage controller can be taken over by the standby storage controller.

[0031] The data protection control unit 1086 is software that allocates a physical storage area provided by the storage device 1033 in its own storage node 103 or another storage node 103 to each storage controller group 1085, and reads or writes the specified data to the corresponding storage device 1033 according to the above-described I / O command given from the storage controller 1083. In this case, when the data protection control unit 1086 allocates a physical storage area provided by the storage device 1033 in another storage node 103 to the storage controller group 1085, it cooperates with the data protection control unit 1086 implemented in the other storage node 103, and exchanges data with the data protection control unit 1086 via the network 102, and reads or writes the data to the storage area according to the I / O command given from the active storage controller of the storage controller group 1085.

[0032] FIG. 6 is a diagram showing an example of the logical configuration of the storage system. The storage controller 1083 provides a virtual storage area (virtual volume) to the host using the storage area (pool volume) provided by the data protection controller (so-called Thin Provisioning method). When receiving a Write to a page of the virtual volume from the host in an active process, the storage controller 1083 allocates a page of the pool volume.

[0033] The storage controller 1083 uses the storage area provided by the Backend driver as a base image storage area for control information and a log storage area. Mapping information from the pool volume to the virtual volume, etc. is written to the control information area in memory, and then the updated content is logged and stored in the log storage area. Also, the updated content is applied to the control information area of the standby process on the other node forming a pair, and is also stored in the log storage area of that node. The entire control information is stored in the storage area as a base image. Similarly, in the standby process forming a pair, the base image is stored in the storage area.

[0034] Figure 7 is a diagram showing an example of the software module structure of the storage controller 1083. The storage controller 1083 executes a base image storage status monitoring process, a commit point switching process, a base image storage redundancy process, a base image storage process, a read process, a write process, a log creation storage process, and a power failure recovery process. Details of each process will be described later.

[0035] Figure 8 is a specific example of the data stored in the memory 1032. In the memory areas managed by the active process and the standby process of the storage controller 1083, a base image storage area management table, a log storage area management table, a control information area, and a log buffer area are arranged.

[0036] The control information area stores information such as page mapping information between the virtual volume and the pool volume of Thin Provisioning. The log created when updating the control information area is first accumulated in the log buffer and then written to the log storage area.

[0037] Figure 9 is an explanatory diagram of the base image storage area management table. The base image storage area management table stores a confirmed surface ID that uniquely identifies a surface (confirmed surface) where evacuation is complete, an evacuation surface ID that uniquely identifies a surface (evacuation surface) where base image evacuation is performed, and management information for each storage area. The management information for each storage area stores the progress of base image evacuation (evacuation rate). An evacuation rate of 100% indicates that the storage of the base image for that storage area is complete and the base image can be used for recovery. If the evacuation rate is less than 100%, it indicates that the process of storing the base image with that storage area as the storage destination is in progress.

[0038] Figure 10 is an explanatory diagram of the log storage area management table. The log storage area management table has management information for each surface of the storage area.

[0039] For log storage, it is possible to use a method in which the same number of log storage areas as the base image storage areas are prepared and paired with each other, such as #0 and #1. The log storage area management table in the case of adopting this method has items for the start position address (storage position of the first log), end position address (end position of the last log), and usage rate of the storage area for each of the multiple log storage areas. When the base image is being evacuated to the base image storage area #0, the control information update content is started to be written as a log from the beginning of the log storage area #0. After the log writing is completed, the log end address of this table is updated.

[0040] The storage of logs may adopt a method in which the inside of one log storage area is divided into areas for log storage areas #0 and #1 and associated with base image storage areas #0 and #1 respectively. When this method is adopted, the log storage area management table has items for the start address and end address of the stored logs for each of the multiple log storage areas, and also has an item for the usage rate of the entire log storage area. In this method, after the log writing is completed, the log end address on the evacuation side is updated. When switching the surface, the start address of the new storage area to be the evacuation side is set to the end address of the old evacuation surface side storage.

[0041] FIG. 11 is an explanatory diagram of information on the storage device. For each process of each storage controller, a base image storage area information table, a log storage area information table, a base image storage area, a control information log storage area, and a persistence area are prepared on the storage device of the node. FIG. 11 shows an example in which both the base image storage area and the log storage area are prepared on two surfaces, #0 and #1.

[0042] The base image storage area information table stores which of the two surfaces of the storage area is the confirmed surface. The log storage area information table stores information on the area of the log storage area in which logs are stored. Each surface of the base image storage area and the log storage area may be divided and arranged on a plurality of storage devices. FIG. 11 shows an example in which it is divided and arranged on three storage devices. User data is stored in the persistence area.

[0043] The division of the base image storage area may be performed on the Backend driver side or on the storage controller side. In the former case, the base image received by the Backend driver layer from the Storage Controller is distributed and arranged on multiple drives. In the latter case, the Storage Controller layer distributes the base image across multiple drives.

[0044] Figure 12 is an explanatory diagram of the base image storage area information table. The base image storage area information table stores the ID of the completed evacuation surface and the sequence number for each storage area for the base image to be written to the storage device. These pieces of information are read during the recovery process.

[0045] Figure 13 is an explanatory diagram of the log storage area information table to be written to the storage device. The log storage area information table stores the range in which the valid logs of the log storage area are evacuated and is read during the recovery process. The management information for each storage area surface stores the start position address (the storage position of the first log) and the end position address (the end position of the last log) of the stored logs.

[0046] Regardless of whether the method of preparing the same number of log storage areas as the base image storage areas or the method of dividing one log storage area and allocating it to the base image storage areas is adopted, there is no difference in storing the log start position address and the log end position address in association with the base image storage areas. However, in the method of dividing one log storage area, the log end position address corresponding to one base image becomes the log start position address corresponding to the next base image.

[0047] Figure 14 shows the structure of the log header. The log header is a table included in each log stored in the log buffer area in memory or the log area on the storage device. The log header is attached to the head of the log when creating the log from the update content during control information update. The log header has fields such as a log sequence number, an update address, and an update size.

[0048] The log sequence number field stores a log sequence number uniquely assigned to each log. The update address field stores the address of the control information or cache data to be updated for each log. The update size field stores the size to be updated.

[0049] FIG. 15 is a flowchart showing the processing procedure of the base image evacuation status monitoring process. The base image evacuation status monitoring process monitors the usage rate of the log storage area and the evacuation status on the evacuation surface of the base image. When the usage rate of the log storage area is high and the evacuation of the base image is completed, a definite surface switching process is performed to start the evacuation of the base image to another storage surface. In the base image evacuation status monitoring process, the evacuation status of the base image is monitored by all the storage controllers 1083 that form pairs, and a surface switch is performed when all evacuations are completed.

[0050] Specifically, first, the active storage controller 1083 acquires the evacuation surface ID (step S101) and acquires the usage rate of the log storage area (step S102). If the usage rate of the log storage area is less than the threshold (step S103; No), the storage controller 1083 repeats step S102.

[0051] If the usage rate of the log storage area is equal to or greater than the threshold (step S103; Yes), the storage controller 1083 acquires the evacuation rate of the base image evacuation surface of its own node (step S104). If the evacuation rate of the base image evacuation surface is less than 100% (step S105; No), the storage controller 1083 repeats step S104.

[0052] If the evacuation rate of the base image evacuation surface is 100% (step S105; Yes), the storage controller 1083 transmits a base image evacuation rate request to another node (here, the standby storage controller 1083) (step S106). When the storage controller 1083 of another node receives the base image evacuation request, it acquires the evacuation rate of the base image evacuation surface in that node (step S107) and transmits it to the requesting node (the node that transmitted the base image evacuation rate request) as a base image evacuation rate response.

[0053] When the storage controller 1083 of the node that transmitted the base image evacuation rate request receives the base image evacuation rate response (step S108), it determines whether the evacuation rate at the other node is 100% (step S109). If the evacuation rate at the other node is less than 100% (step S109; No), it returns to step S106. If the evacuation rate at the other node is 100% (step S109; Yes), the storage controller 1083 performs the confirmed surface switching process (step S110) and returns to step S101.

[0054] FIG. 16 is a flowchart showing the processing procedure of the confirmed surface switching process (two-surface method). The confirmed surface switching process in the two-surface method is a process of switching the storage area for evacuating the base image in the storage controller 1083 having two base image storage areas. In the confirmed surface switching process, the storage controller 1083 updates the information in the base image storage area management table on the memory and the base image storage area information table on the storage device. The confirmed surface switching process updates the information at the node where the active process and the standby process operate. Also, a flag is set during the confirmed surface switching to stop the log creation process described later.

[0055] Specifically, the storage controller 1083 first determines the storage area to be the new evacuation surface (step S201). In the two-surface method, the storage areas are selected alternately. Next, the storage controller 1083 assigns a log sequence number to be set for the newly migrated base image (step S202), and sets a face switching flag (step S204). While the face switching flag is being set, log creation is stopped.

[0056] After that, the storage controller 1083 sends a request to update the confirmed face information to other nodes (step S204). The storage controller 1083 of the node that sent the request to update the confirmed face information and the storage controller 1083 of the node that received the request to update the confirmed face information each execute a confirmed face information update process. In FIG. 16, the confirmed face information update process of the node that sent the request to update the confirmed face information is shown as step S205, and the confirmed face information update process of the node that received the request to update the confirmed face information is shown as step S206.

[0057] After the confirmed face information update process, the storage controller 1083 of the node that sent the request to update the confirmed face information receives an update completion response for the confirmed face information from other nodes (step S207), releases the face switching flag (step S208), and starts migrating the base image (step S209). Step S209 includes a request to start the base image redundancy process.

[0058] In the confirmed face information update process, the storage controller 1083 first updates the confirmed face ID (step S301). Specifically, it updates the confirmed face ID in the base image storage area management table on the memory and the base image storage area information table on the storage device. Also, the storage controller 1083 updates the migrated face ID in the base image storage area management table on the memory (step S302).

[0059] After that, the storage controller 1083 updates the sequence number of the migrated face storage area (step S303). Specifically, the storage controller 1083 updates the sequence number on the migrated face side in the base image storage area information table on the storage device to the value assigned.

[0060] After step S303, the storage controller 1083 resets the evacuation rate of the evacuation surface side storage area (step S304). Specifically, it resets the evacuation rate on the evacuation surface side in the base image storage area management table in the memory to 0.

[0061] After step S304, the storage controller 1083 invalidates the logs before the sequence number in the storage area on the confirmed surface side (step S305). In the method of preparing the same number of log storage areas and base image storage areas, the log end side address on the side that newly becomes the evacuation surface in the log storage area management table in the memory and the log storage area information table on the storage device is corrected to the head address. In the method of preparing only one log storage area, the head and end addresses on the side that newly becomes the evacuation surface in the log storage area management table in the memory and the log storage area information table on the storage device are aligned with the end address on the confirmed surface side.

[0062] FIG. 17 is a flowchart showing the processing procedure of the confirmed surface switching process (3-surface method). The confirmed surface switching process of the 3-surface method is a process of switching the storage area for evacuating the base image in the storage controller 1083 having three base image storage areas. In the confirmed surface switching process of the 3-surface method, it is not necessary to stop log creation even during the confirmed surface switching.

[0063] Specifically, the storage controller 1083 first determines the storage area to be the new evacuation surface (step S401). In the 3-surface method, for example, it is selected in the order of #0, #1, #2, #0 ···. Next, the storage controller 1083 assigns a log sequence number to be set for the newly migrated base image (step S402), and sends a request to update the committed surface information to other nodes (step S404). The storage controller 1083 of the node that sent the request to update the committed surface information and the storage controller 1083 of the node that received the request to update the committed surface information each execute the committed surface information update process. In FIG. 17, the committed surface information update process of the node that sent the request to update the committed surface information is shown as step S404, and the committed surface information update process of the node that received the request to update the committed surface information is shown as step S405.

[0064] After the committed surface information update process, the storage controller 1083 of the node that sent the request to update the committed surface information receives a response indicating the completion of the update of the committed surface information from other nodes (step S406), and starts migrating the base image (step S407). Step S407 includes a request to start the base image redundancy process.

[0065] In FIG. 17, the details of the committed surface information update process are shown as steps S501 to S505. Since these processes are the same as steps S301 to S305 in FIG. 16, the description thereof is omitted.

[0066] FIGS. 18 and 19 are flowcharts of the base image migration redundancy process. FIG. 18 shows a method in which each node autonomously migrates the base image, and FIG. 19 shows a method in which the base image acquired by the active node is transferred to other nodes for storage.

[0067] In the method of FIG. 18, each process of the storage controller 1083 that forms a pair autonomously performs base image migration. Specifically, for example, the storage controller 1083 of the active node requests the other node (standby node) that forms a pair to migrate the base image (step S601), and also executes the base image migration process on its own node (step S602). Further, the other node that received the base image migration request also executes the base image migration process (step S603).

[0068] In the method of FIG. 19, among the processes of the storage controller 1083 that forms pairs, the base image acquired in one process is transferred to and stored in another process. Specifically, for example, the storage controller 1083 of the active node executes the base image storage process (step S701). Thereafter, the base image is transmitted to another node (standby node) that forms a pair, and a storage request is made (step S702). The other node that has received the base image stores the received base image in a designated storage area of its own node (step S703).

[0069] FIGS. 20 and 21 are flowcharts of the base image storage process. FIG. 20 shows a method of reading data from memory and storing it in a drive, and FIG. 21 shows a method of creating a base image from existing storage information.

[0070] In the method of FIG. 20, the storage controller 1083 reads all the control information on the memory and stores it in the storage area as a base image. Specifically, the storage controller 1083 first reads the storage area information (step S801). In this step, the storage controller 1083 acquires the storage surface ID from the base image storage area management table on the memory. After step S801, the storage controller 1083 reads the control information from the memory of its own node and writes it to the storage area (step S802).

[0071] In the method of FIG. 21, the storage controller 1083 obtains a base image in the state of the control information at the time of surface switching by applying the log of the corresponding log storage area to the base image on the confirmed surface side where storage is completed. Specifically, the storage controller 1083 first reads the confirmed surface base image from the base image storage area (step S901). Thereby, the confirmed surface side base image on the storage device can be acquired.

[0072] Next, the storage controller 1083 reads valid logs from the log storage area (step S902). Thereby, the finalized surface side logs on the storage device can be acquired. Next, the storage controller 1083 applies the logs to the finalized surface base image (step S903). Thereby, the logs can be applied to the base image, and the base image of the control information at the time of surface switching can be obtained. Thereafter, the storage controller 1083 reads the storage area information (step S904), and writes the base image to which the logs have been applied to the storage area (step S905).

[0073] FIG. 22 is a flowchart of the read process. Here, the process when there is a read to the virtual volume provided by the storage controller 1083 is shown. First, the storage controller 1083 analyzes the command (step S1001), and determines whether the page has been allocated to the access destination (step S1002). If the page has not been allocated (step S1002; No), 0 data is set as the response value (step S1008), the host is responded to (step S1007), and the process ends.

[0074] If the page has been allocated to the access destination (step S1002; Yes), the storage controller 1083 acquires the allocation destination address (step S1003), and acquires exclusive access (step S1004). Then, data is read from the drive (step S1005), the exclusive access is released (step S1006), the host is responded to (step S1007), and the process ends.

[0075] Figure 23 is a flowchart of the write process. Here, it shows the process when there is a write to the virtual volume provided by the storage controller 1083. The process is carried out by the active process in the pair of storage controllers that is in the state of accepting I / O. Also, when the page of the pool volume is not allocated to the page of the virtual volume to be written, after performing the process of allocating the page of the pool volume, the mapping information is logged and saved to the drive.

[0076] First, the storage controller 1083 analyzes the command (step S1101) and determines whether the access destination has been page-allocated (step S1102). If it has not been page-allocated (step S1102; No), a physical page is allocated to the logical area (step S1103), and the log creation process is executed (step S1104).

[0077] After the log creation process, or if the access destination has been page-allocated (step S1102; Yes), the storage controller 1083 acquires the destination address (step S1105), acquires exclusive access (step S1106). Then, data is written to the drive (step S1107), exclusive access is released (step S1108), a response is sent to the host (step S1109), and the process ends.

[0078] Figure 24 is a flowchart of the log creation and save process. In the log creation process, after the storage controller 1083 creates a log by attaching a log header to the updated content of the control information, it saves the log to the log storage area. At the same time, the log is sent to the process of the other storage controller 1083 that forms a pair, and the updated content is reflected in the control information area on its memory, and the log is saved to the log storage area. After the log is saved, the end addresses of the storage areas on the memory and the storage device are updated.

[0079] Specifically, first, the storage controller 1083 determines whether the surface switching flag is valid (step S1201). If the surface switching flag is valid (step S1201; Yes), it waits by repeating step S1201. If the surface switching flag is not valid (step S1201; No), it proceeds to step S1202.

[0080] In step S1202, the storage controller 1083 obtains the sequence number. After that, the storage controller 1083 creates a log header (step S1203) and assigns the log header (step S1204).

[0081] After that, the storage controller 1083 obtains the evacuation surface ID from the base image storage area management table in the memory (step S1205). Also, from the log storage area management table in the memory, it obtains the log end side address of the evacuation surface side log storage area and determines that address as the log storage position (step S1206). Then, it secures a log buffer (step S1207) and stores the log at the secured position (step S1208).

[0082] After that, the storage controller 1083 transfers the log to the log buffer of other nodes and sends an evacuation request (step S1209). Then, it performs the log evacuation process for its own node (step S1210). The node that has received the log transfer reflects the log (step S1211) and performs the log evacuation process for its own node (step S1212).

[0083] The storage controller 1083 of the node that has sent the evacuation request receives the evacuation completion notification from other nodes (step S1213) and ends the process.

[0084] In the log reflection process, the storage controller 1083 reads the log from the log buffer (step S1301), applies the log to the control information area of the memory (step S1302), makes the control information redundant, and ends the log reflection process.

[0085] In the log backup process, the storage controller 1083 reads the log from the log buffer (step S1401), writes the log to a specified position in the log storage area (step S1402). Then, it deletes the log on the log buffer (step S1403) and updates the log storage position information (step S1404). In the update of the log storage position information, the end addresses of the backup-side log storage areas in the log storage area management table on the memory and the log storage area information table on the storage device are updated. After that, it updates the log storage area usage rate information (step S1405) and ends the log backup process.

[0086] Figure 25 is a flowchart of the recovery process. Here, as an example of the process when performing recovery after losing the control information on the memory, the operation led by the active process is shown. In this recovery process, the base image and log backed up before the power failure are applied to the memory to recover the control information on the memory. Generally speaking, the finalized surface information backed up on the storage device is read out, the base image is read out from the base image storage area it points to and applied to the memory. After that, the logs after the sequence number of the finalized surface are read out from the log storage area, sorted, and then applied to the memory in sequence number order. If the base image or log cannot be read due to a drive failure, the storage controller reads from another paired process (standby process). If the process that was active before the power failure cannot be started due to a node failure, the standby process paired with it operates as the active process to perform the recovery process.

[0087] Specifically, when starting the recovery process, the storage controller 1083 obtains the finalized surface ID by reading the finalized surface information from the base image storage area information table on the storage device (step S1501). Also, it obtains the status of the drives that make up the base image storage area (step S1502). Then, it selects the node from which to read the base image (step S1503). For example, if a drive has failed, it attempts to read from other nodes.

[0088] If the node to read from is the local node (step S1504; Yes), the storage controller 1083 reads the base image from the finalized surface side storage area of the local node (step S1505). If the node to read from is not the local node (step S1504; No), it reads the base image from the finalized surface side storage area at the other node and transfers it to the active node (step S1506). After the end of step S1505 or step S1506, the storage controller 1083 applies the base image to the control information area of the memory (step S1507) and obtains the status of the drives that make up the log storage area (step S1508).

[0089] After step S1508, the storage controller 1083 selects the node from which to read the log (step S1509). For example, if a drive has failed, it attempts to read from other nodes.

[0090] If the node to read from is the local node (step S1510; Yes), the storage controller 1083 reads the log from the log storage area of the local node (step S1511). Specifically, it reads the log start address and end address of each storage area from the log storage area information table on the storage device, and then reads the log from within that range of each storage area. If the node to read from is not the local node (step S1510; No), it reads the log from the log storage area at the other node and transfers it to the active node (step S1512). After the completion of step S1511 or step S1512, the storage controller 1083 sorts the read logs in sequence number order (step S1513), applies the logs after the sequence number assigned to the applied base image to the control information area of the memory (step S1514), and ends the recovery process.

[0091] As described above, the disclosed storage system 100 is a storage system including one or more storage nodes 103 having a non-volatile storage device 1033, a storage controller 1083 that processes reading and writing of data to and from the storage device 1033, and a volatile memory 1032. The storage device 1033 includes a plurality of base image storage areas including at least a first base image storage area and a second base image storage area as areas for storing the entire predetermined information stored in the memory 1032 as a base image. The storage controller 1083 performs a process of storing the base image in the first base image storage area. When the storage of the base image in the first base image storage area is completed, the storage controller 1083 starts a process of storing the next base image in the second base image storage area. When the predetermined information is lost from the memory 1032, a recovery process is performed to read the stored base image and restore it to the memory. Therefore, even if a power failure occurs during the backup of the base image, the information on the memory can be recovered, and a storage system with high performance and reliability can be realized.

[0092] As an example, when the storage of the base image in the second base image storage area is completed, the storage controller 1083 starts a process of storing the next base image in the first base image storage area. When switching the storage destination of the base image between the first base image storage area and the second base image storage area, the execution of the read / write process for the storage device is suppressed. In this way, by alternately using the storage areas on two sides, it is possible to realize power-off protection while suppressing the size of the storage area.

[0093] Also, as an example, when the storage of the base image in the second base image storage area is completed, the storage controller 1083 starts the process of storing the next base image in the third base image storage area. When the storage of the base image in the third base image storage area is completed, the storage controller 1083 starts the process of storing the next base image in the first base image storage area. Even when the area to which the base image is to be stored is being switched, the read / write process for the storage device is executed. In this way, by sequentially using the storage areas on three sides, it is possible to realize power-off protection without stopping the read / write process.

[0094] Also, as an example, the predetermined information is control information. The storage device 1033 includes a plurality of log storage areas associated with the plurality of base image storage areas for storing the updated content of the control information as a log. While the storage controller 1083 is storing the base image of the control information in any of the base image storage areas, the updated content of the control information is stored in the corresponding log storage area. When performing the recovery process, the storage controller reads out the base image whose storage is completed from the base image storage area and writes it to the memory, and then reads out the log obtained from the start of storing the base image until the control information is lost from the log storage area and writes it to the memory to restore the control information. In this way, by associating a plurality of log storage areas with a plurality of base image storage areas, power-off protection can be easily realized.

[0095] Also, as an example, the predetermined information is control information, the storage device 1033 includes a log storage area for storing the update content of the control information as a log, and the storage controller 1083 records the storage position of the log storage area in response to the switching of the base image storage area used as the storage destination of the base image. When performing the recovery process, the storage controller reads the base image for which storage has been completed from the base image storage area and writes it to the memory, and then reads the log acquired from the start of storage of the base image until the control information is lost from the log storage area and writes it to the memory to restore the control information. In this way, by dividing one log storage area and corresponding it to a plurality of base image storage areas, the size of the log storage area can be suppressed.

[0096] Also, the storage controller 1083 copies the predetermined information existing in the memory and stores it in the base image storage area as the base image. Alternatively, when there is a base image for which storage has been completed, the log regarding the predetermined information may be applied to the base image and stored in the base image storage area as a new base image. In this way, any method can be used to create the base image.

[0097] Also, the predetermined information is control information, the plurality of storage nodes redundantly store the control information by storing the control information in their respective memories, and the storage controller 1083 of each storage node stores the entire control information stored in the memory of its own node in the storage device as the base image. Alternatively, the predetermined information is control information. Among the plurality of storage nodes, the storage controller 1083 of any one of the storage nodes uses the entire control information stored in the memory of its own node as the base image, stores it in the storage device of its own node, and transmits the base image to other storage nodes. Among the plurality of storage nodes, the storage controller of the storage node that has received the base image from another storage node may store the received base image in the storage device of its own node. In this way, each node may create a base image and perform evacuation, or each node may evacuate the base image created by any one of the nodes.

[0098] Further, the storage controller 1083 can divide and store the base image in a plurality of the storage devices. In this way, by making the base image redundant, the fault tolerance can be improved.

[0099] Also, when a failure occurs in the storage device of its own node, the storage controller 1083 acquires the base image stored in the base image storage area of another storage node and writes it to the memory of its own node. In this way, the base image can be recovered even when a failure occurs in the storage device.

[0100] Also, when a failure occurs in any one of the plurality of storage nodes, the storage controller of another storage node acquires the base image stored in the base image storage area of the storage node and writes it to the memory of the storage node. In this way, the base image can be recovered even when a failure occurs in the storage node.

[0101] Note that the present invention is not limited to the above embodiments, and various modifications are included. For example, the above-described embodiments have been described in detail for easy understanding of the present invention, and are not necessarily limited to those having all the configurations described. Further, not only deletion of such a configuration is possible, but also replacement or addition of a configuration is possible. For example, in the embodiment, the case where control information is saved as a base image is exemplified, but it is also applicable to the case where other data such as cache data is saved as a base image.

Description of Reference Numerals

[0102] 100: Storage system, 101: Host device, 102: Network, 103: Storage node, 104: Management node, 1031: CPU, 1032: Memory, 1033: Drive, 1083: Storage controller

Claims

1. A storage system comprising one or more storage nodes having a non-volatile memory device, a storage controller that processes reading and writing of data to and from the memory device, and a volatile memory, The memory device includes a plurality of base image storage areas including at least a first base image storage area and a second base image storage area as areas for storing the entire predetermined information stored in the memory as a base image, The storage controller, performs a process of storing the base image in the first base image storage area, and when the storage of the base image in the first base image storage area is completed, starts a process of storing the next base image in the second base image storage area, When the predetermined information is lost from the memory, performs a recovery process of reading out the base image for which the storage has been completed and restoring it to the memory, The storage controller, When the storage of the base image in the second base image storage area is completed, starts a process of storing the next base image in the first base image storage area, When switching the storage destination of the base image between the first base image storage area and the second base image storage area, suppresses the execution of the writing process to the memory device A storage system characterized by the above.

2. A storage system comprising one or more storage nodes having a non-volatile memory device, a storage controller that processes reading and writing of data to and from the memory device, and a volatile memory, The memory device includes a plurality of base image storage areas including at least a first base image storage area and a second base image storage area as areas for storing the entire predetermined information stored in the memory as a base image, The storage controller, Perform a process of storing the base image in the first base image storage area. When the storage of the base image in the first base image storage area is completed, start a process of storing the next base image in the second base image storage area. When the predetermined information is lost from the memory, perform a recovery process of reading out the base image for which the storage is completed and restoring it to the memory. The storage controller When the storage of the base image in the second base image storage area is completed, start a process of storing the next base image in the third base image storage area. When the storage of the base image in the third base image storage area is completed, start a process of storing the next base image in the first base image storage area. A storage system characterized in that even when the area to which the base image is to be stored is being switched, a read / write process for the storage device is executed.

3. The storage system according to claim 1, The predetermined information is control information, The storage device includes a plurality of log storage areas for storing update contents of the control information as logs, corresponding to the plurality of base image storage areas, While the storage controller is performing a process of storing the base image of the control information in any of the base image storage areas, store the update content of the control information in the corresponding log storage area, When performing the recovery process, the storage controller reads out the base image for which the storage is completed from the base image storage area and writes it to the memory, and then reads out the log acquired from the start of storing the base image until the control information is lost from the log storage area and writes it to the memory to restore the control information. A storage system characterized by this.

4. The storage system according to claim 1, wherein the predetermined information is control information, the storage device includes a log storage area for storing the updated content of the control information as a log, when switching the base image storage area used as the storage destination of the base image, the storage controller records the log end position address corresponding to the base image for which storage has been completed as the log start position address corresponding to the next base image, when performing the recovery process, the storage controller reads the base image for which storage has been completed from the base image storage area and writes it to the memory, and then reads the log after the log start position address corresponding to the base image from the log storage area and writes it to the memory to restore the control information. A storage system characterized by that.

5. The storage system according to claim 1, wherein the storage controller duplicates the predetermined information existing in the memory and stores it in the base image storage area as the base image. A storage system characterized by that.

6. A storage system including one or more storage nodes having a non-volatile storage device, a storage controller that processes reading and writing of data to and from the storage device, and a volatile memory, the storage device includes a plurality of base image storage areas including at least a first base image storage area and a second base image storage area as areas for storing the entire predetermined information stored in the memory as a base image, the storage controller, performs a process of storing the base image in the first base image storage area, and when the storage of the base image in the first base image storage area is completed, starts a process of storing the next base image in the second base image storage area, When the predetermined information is lost from the memory, a recovery process is performed to read out the base image for which the storage is completed and restore it to the memory. A storage system, characterized in that when the base image for which the storage is completed exists, a log related to the predetermined information is applied to the base image and stored in the base image storage area as a new base image.

7. A storage system including one or more storage nodes having a non-volatile storage device, a storage controller that processes reading and writing of data to and from the storage device, and a volatile memory, The storage device includes a plurality of base image storage areas including at least a first base image storage area and a second base image storage area as areas for storing the entire predetermined information stored in the memory as a base image. The storage controller Performs a process of storing the base image in the first base image storage area. When the storage of the base image in the first base image storage area is completed, a process of storing the next base image in the second base image storage area is started. When the predetermined information is lost from the memory, a recovery process is performed to read out the base image for which the storage is completed and restore it to the memory. The predetermined information is control information. The plurality of storage nodes make the control information redundant by storing the control information in their respective memories. A storage system, characterized in that the storage controller of each storage node stores the entire control information stored in the memory of its own node as the base image in the storage device.

8. A storage system including one or more storage nodes having a non-volatile storage device, a storage controller that processes reading and writing of data to and from the storage device, and a volatile memory, The memory device includes a plurality of base image storage areas including at least a first base image storage area and a second base image storage area as areas for storing the entire predetermined information stored in the memory as a base image. The storage controller performs a process of storing the base image in the first base image storage area, and when the storage of the base image in the first base image storage area is completed, starts a process of storing the next base image in the second base image storage area. When the predetermined information is lost from the memory, a recovery process is performed to read out the stored base image and restore it to the memory. The predetermined information is control information. Among the plurality of storage nodes, the storage controller of any one storage node uses the entire control information stored in its own node's memory as the base image, stores it in the storage device of its own node, and transmits the base image to other storage nodes. A storage system, wherein the storage controller of a storage node that has received the base image from another storage node among the plurality of storage nodes stores the received base image in the storage device of its own node.

9. A storage system according to claim 1, wherein the storage controller stores the base image by dividing it among a plurality of the memory devices.

10. A storage system including one or more storage nodes having a non-volatile memory device, a storage controller that processes reading and writing of data to and from the memory device, and a volatile memory. The memory device includes a plurality of base image storage areas including at least a first base image storage area and a second base image storage area as areas for storing the entire predetermined information stored in the memory as a base image. The storage controller performs a process of storing the base image in the first base image storage area, and when the storage of the base image in the first base image storage area is completed, starts a process of storing the next base image in the second base image storage area. When the predetermined information is lost from the memory, a recovery process is performed to read out the stored base image and restore it to the memory. The storage controller is characterized in that when a failure occurs in the memory device of its own node, it acquires the base image stored in the base image storage area of another storage node and writes it to the memory of its own node. A storage system.

11. A storage system including one or more storage nodes having a non-volatile memory device, a storage controller for processing reading and writing of data to and from the memory device, and a volatile memory, The memory device includes a plurality of base image storage areas including at least a first base image storage area and a second base image storage area as areas for storing the entire predetermined information stored in the memory as a base image. The storage controller performs a process of storing the base image in the first base image storage area, and when the storage of the base image in the first base image storage area is completed, starts a process of storing the next base image in the second base image storage area. When the predetermined information is lost from the memory, a recovery process is performed to read out the stored base image and restore it to the memory. When a failure occurs in any of a plurality of storage nodes, a storage controller of another storage node acquires a base image stored in the base image storage area of the storage node and writes it to the memory of the storage node. A storage system characterized by this.

12. A storage control method in a storage system including one or more storage nodes having a non-volatile storage device, a storage controller that processes reading and writing of data to and from the storage device, and a volatile memory, comprising: The storage controller Storing, as a base image, the entire predetermined information stored in the memory in a first base image storage area provided in the storage device; When the storage of the base image in the first base image storage area is completed, storing the next base image in a second base image storage area provided in the storage device; When the predetermined information is lost from the memory, performing a recovery process of reading out the stored base image and restoring it to the memory Including The storage controller When the storage of the base image in the second base image storage area is completed, starting a process of storing the next base image in the first base image storage area, When switching the storage destination of the base image between the first base image storage area and the second base image storage area, suppressing the execution of a write process to the storage device A storage control method characterized by this.

Citation Information

Patent Citations

  • Data backup / restore means of expanded image processing system composed of image processor and expansion controller

    JP2007122485A

  • Storage systems and storage methods

    JP2014517412A

  • Storage system and control software arrangement method

    JP2019101703A

  • Storage system, and method for recovering storage system

    JP2020135138A

  • Blobstore system for the management of large data objects

    US20190179710A1