Storage system and data protection method
The storage system addresses performance degradation by employing memory replication and log evacuation methods to protect cache data, ensuring high reliability and performance through efficient data handling during controller failures.
Patent Information
- Application Number
- JP2025081757
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-03-02
AI Technical Summary
Conventional storage systems experience significant performance degradation when switching to the write-through method due to controller failures, especially with data protection methods like RAID6, leading to prolonged write response times.
A storage system with non-volatile memory and multiple controllers that employ memory replication and log evacuation methods to protect cache data, switching between these methods based on controller status, ensuring high reliability and performance by reducing unnecessary drive accesses.
The system maintains high performance and reliability by minimizing drive access delays during controller failures, using memory replication and log evacuation strategies to ensure data integrity and quick destaging.
Smart Images

Figure 2025107625000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a storage system and a data protection method.
Background Art
[0002] A storage system records write data received from a host to a drive via a cache memory (hereinafter referred to as a cache). That is, write data requested to be written from the host is once held in the cache and then written to a predetermined drive. The method of writing data from the cache to the drive can be roughly classified into two types.
[0003] One method is, for example, a write-through method, in which write data is written to a drive before returning a response to a write request to the host. Another method is, for example, a write-back method or a write-after method, in which a response to a write request is returned to the host when the write data is stored in the cache. In the case of the write-back method, the writing of the write data to the drive is performed at a predetermined timing after the write data is stored in the cache.
[0004] Therefore, in the case of the write-back method, a response can be returned to the host without waiting for the completion of writing to the drive, so the response time can be shortened compared to the write-through method. On the other hand, in the case of the write-back method, the data for which the write from the host has been completed temporarily exists only in the cache. Therefore, it is necessary to appropriately protect the write data in the cache. For example, the storage system has a redundant configuration with a plurality of controllers, and the write data received by one controller is copied to the cache of another controller to ensure the redundancy of the write data. Also, in order to prepare for power outages and power failures, for example, the cache is protected by a battery.
[0005] A storage system is required to achieve both high reliability and high performance. For this purpose, the storage system can selectively use the above write-through method or write-back method according to the situation. As described in Patent Document 1, when the cache can be appropriately protected, it operates in the write-back method, and when the cache cannot be protected, it switches to the write-through method. By doing so, it is possible to return a response at high speed by the write-back method during normal times, and even when the controller fails and the redundancy of the cache is lost, reliability can be ensured by the write-through method.
Prior Art Documents
Patent Documents
[0006]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0007] However, in the above conventional method, the problem is that the performance when operating in the write-through method due to a controller failure or the like is significantly reduced compared to the performance during normal operation in the write-back method. In particular, in recent storage systems, data protection methods such as the RAID6 method have become widespread. In order to store write data in the drive, a plurality of data (old data and a plurality of protection data such as parity data) are read from the drive, the parity data is updated, and then the write data and the plurality of parity data need to be written to the drive. By waiting for this multiple drive access, the write response time deteriorates significantly compared to the write-back method.
[0008] An object of the present invention is to improve the performance when the redundancy of the cache is lost due to a controller failure or the like compared to the write-through method while ensuring high reliability in a storage system.
Means for Solving the Problems
[0009] In order to solve the above problems, a storage system according to the present invention is a storage system including a non-volatile memory device and a plurality of storage controllers that control reading and writing to the memory device, wherein the memory device is a drive for storing user data, and each of the plurality of storage controllers has a processor and a memory, and the storage controller includes a first memory protection method of a memory replication method for replicating data on the memory onto the memory of the corresponding storage controller, and a second memory protection method of a log evacuation method for generating a log related to updating of data on the memory and writing it to a non-volatile medium. The storage controller stores a write request from a host to the memory device as cache data in the memory, protects the cache data by the first memory protection method or the second memory protection method, and then returns a write completion response to the host. After the write response is completed, the cache data is destaged to the memory drive. The storage controller is characterized in that it switches which of the first memory protection method and the second memory protection method to use according to the operating state of another storage controller. The data protection method according to the present invention is a data protection method for a storage system including a non-volatile storage device and a plurality of storage controllers that control reading and writing to the storage device. The storage device is a drive for storing user data, and each of the plurality of storage controllers has a processor and a memory. The storage controller includes a first memory protection method of the memory replication method that replicates the data on the memory onto the memory of the corresponding storage controller, and a second memory protection method of the log evacuation method that generates a log related to the update of the data on the memory and writes it to a non-volatile medium. The method includes steps of: the storage controller storing a write request from a host to the storage device as cache data in the memory; the storage controller protecting the cache data by the first memory protection method or the second memory protection method; the storage controller returning a write completion response to the host; and the storage controller destaging the cache data to the storage drive after the write response is completed. The method further includes a step of the storage controller switching which of the first memory protection method and the second memory protection method to use according to the operating state of another storage controller.
Advantages of the Invention
[0010] According to the present invention, a storage system and a data protection method having both high performance and high reliability can be realized.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Modes for Carrying Out the Invention
[0012] Hereinafter, embodiments of the present invention will be described with reference to the drawings. The embodiments relate to, for example, a storage system including a plurality of storage controllers.
Examples
[0013] FIG. 1 is a diagram showing a configuration example of a storage system according to the present embodiment of the present invention. The storage system 100 of the present embodiment includes a plurality of controllers 103 and a drive 110 which is a storage device. The controller 103 is a device having a function of providing a volume to be read from and written to with respect to a host computer (hereinafter referred to as a host). The controller 103 includes a CPU 106, a memory 105, a memory backup drive 107, a front-end interface (FE I / F) 104, and a back-end interface (BE I / F) 108. The drive is, for example, an SSD (Solid State Drive) using a flash memory as a storage medium, an HDD (Hard Disk Drive) using a magnetic disk as a storage medium, or the like. The memory is, for example, a semiconductor memory such as DRAM. The memory backup drive is, for example, a drive such as an SSD, and is used to save the content of the memory, for example, when an external power loss occurs. The FE I / F is, for example, a Fibre Channel HBA (Host Bus Adapter) or a NIC (Network Interface Controller). The BE I / F is, for example, a SAS HBA, a PCI Express (hereinafter referred to as PCIe) adapter, or a NIC. Each controller and the drive are connected, for example, by a switch (BE Switch) 109. Also, the CPUs of the plurality of controllers are connected by an interconnect such as PCIe. Note that the CPUs may be connected via, for example, a PCIe switch. The storage system is connected to a storage area network (SAN) 101 such as Fibre Channel or Ethernet, and the host 102 is also connected to the SAN 101. The SAN 101 may include a switch or the like. Also, a plurality of hosts may be connected to the SAN 101.
[0014] FIG. 2 is a diagram showing an overview of the write operation when both controllers are normal. The CPU 106 of the controller 103 in the present embodiment receives a write request from the host 102, receives data from the host, and writes the data 201 to the memory 105 in its own controller and the memory 105 of the other controller. The CPU 106 also updates the control information (Metadata) 200 in the memory. Then the CPU 106 returns a write completion response to the host 102. Also, although omitted in the figure, the CPU 106 writes the write data written on the memory to the drive at a predetermined timing. In this write operation, the write data and control information on the memory are duplicated between the controllers to prepare for controller failures.
[0015] FIG. 3 is a diagram showing an overview of the write operation when one controller of the storage system in the present embodiment is blocked. The CPU 106 of the controller 103 in the present embodiment receives a write request from the host 102, receives data from the host, and writes the data 201 to the memory 105 in its own controller. The CPU 106 also updates the control information (Metadata) 200 in the memory. Further, the CPU 106 writes the updated content of the data as a log to the memory backup drive 107, and also writes the updated content of the control information as a log to the memory backup drive 107 (log backup). Then the CPU 106 returns a write completion response to the host 102. Also, although omitted in the figure, the CPU 106 writes the write data written on the memory to the drive at a predetermined timing (destage). In the present embodiment, the logs of the data and control information are described as being stored in the memory backup drive 107, but these logs may be recorded in, for example, a drive for storing user data or in another drive for storing logs.
[0016] In this write operation, the write data and control information in the memory are written as a log to the memory backup drive as a precaution in case the remaining controllers fail. If the remaining controllers fail, the storage system will stop (system down) once. However, by restoring the write data and control information in the memory using the log written to the memory backup drive after replacing the controller with a new one, data loss can be prevented. To prevent confusion in the following explanations, the difference between destaging and log backup is clarified here. Destaging refers to writing the dirty data in the cache to the final storage area of the user data storage drive, which is the final storage medium. The user data storage drive stores data with enhanced data protection, capacity efficiency, I / O performance, etc. through the storage functions provided by the storage system (mainly the controller). For example, in terms of data protection, it is protected by a method such as RAID6. In this case, parity data is generated during the destaging process, and the parity data is also written to the drive. After destaging is completed, the data is in a state where the data in the memory and on the drive match (clean), so the data can be lost from the memory without any problem. On the other hand, log backup refers to temporarily writing the updated content of the data and control information in the memory to a non-volatile storage medium (drive) such as a memory backup drive as a precaution in case of a controller failure. As described above, the dirty data written to the drive as a log can be deleted from the drive when the destaging of the dirty data is completed because there is no problem even if the data is lost from the memory.
[0017] In addition, when the memory backup drive 107 is used for log storage in this embodiment, the area allocated for backing up the memory content when both controllers are normal may be used as the area for storing logs when one controller is blocked. By doing so, since no additional drive or additional storage capacity is required for log storage, it is advantageous in terms of cost compared to the case of storing logs in the user data drive or mounting a separate drive for log storage.
[0018] By the way, since the data written from the host is generally handled in block units such as 512 bytes or 4 KiB, the log (cache data log) including the update content of the cache data has a relatively large granularity, while the control information is in byte units, so the log (control information log) including the update content of the control information has a relatively small granularity. Also, the ratio of the cache data area in the entire memory is relatively large. Therefore, for the control information log, the entire memory area (base image) where the control information is stored is periodically written to the drive, all the previously written logs are discarded, and the area where the log was written is recovered as a free area. This method is called the base image evacuation method. On the other hand, for the cache data log, unnecessary logs that are not the latest are identified and discarded (invalidated). This creates scattered free areas in the log area, so by packing only the valid logs into another area at a predetermined timing, a continuous free area is recovered. This method is called the garbage collection method. By properly using these two methods, it is possible to suppress the capacity consumption for base image evacuation, reduce the management information for free area management, and reduce the overhead for free area recovery.
[0019] FIG. 4 is a diagram showing an overview of the distinction of dirty data in the storage system of this embodiment. When one of the controllers fails during the operation of the storage system, if the remaining controller also fails with the un-stored write data (dirty data) remaining in the memory, this dirty data will be lost. This means that the data that has been written to the storage system (for which the write completion response has been received) is lost, which is a serious problem in a storage system that requires high reliability. Therefore, when one of the controllers fails, it is important to write the dirty data in the memory to the storage device in the shortest possible time to enhance the reliability of the storage system. However, the performance of the storage device that stores the dirty data as a log may become a bottleneck, and it may take a long time to write the dirty data. This can be a particular issue in a storage system configured to store logs on a small number of storage devices. Therefore, in the storage system according to the present invention, the dirty data existing in the memory at the time when one of the controllers fails (existing dirty) is distinguished from the dirty data created after one of the controllers fails (new dirty). The new dirty is written as a log to the log storage device, while the existing dirty is written to the drive for user data. By doing so, the bottleneck of the log storage device can be avoided or alleviated, the time required for the evacuation of the dirty data can be shortened, the probability of data loss can be reduced, and the reliability can be enhanced. Specifically, in this embodiment, two types of dirty queues, a dirty queue 400 for log protection and a dirty queue 401 outside the log protection, are provided in the memory. The new dirty is connected to the dirty queue for log protection, and the existing dirty is connected to the dirty queue outside the log protection.
[0020] Since the dirty data connected to the dirty queue to be log-protected is stored as a log in the memory backup drive 107 before the host response, in the destaging process of writing the dirty data to the user data drive, the dirty data connected to the dirty queue not subject to log protection is preferentially selected as the destaging target and written to the drive. The content of this destaging process will be described later using a flowchart.
[0021] In this embodiment, a method of distinguishing each dirty data using two types of dirty queues is exemplified. However, the types of dirty queues may be more. Also, the method of distinguishing dirty data is not limited to the method using a plurality of dirty queues. For example, it may be managed using a data structure such as a list, or a method of distinguishing dirty data by having identification information such as a flag in the control information of the cache may be used.
[0022] FIG. 5 is a diagram showing the content of the memory of the storage system of this embodiment. The memory 105 includes a storage control program 500, control information 200, cache data 501, a control information log buffer 502, and a cache data log buffer 503.
[0023] The storage control program 500 is a program for controlling the storage system and is executed by the CPU 106. Each process such as the write process described later is included in the content of this storage control program.
[0024] The control information 200 is data used by the storage control program 500 to control the execution of the program. The control information 200 includes, for example, cache control information including the correspondence between the address of the cache data and the logical address (LBA) in the volume, the state of the cache data (dirty / clean), configuration information including the type / capacity of the drive and the type / configuration of the RAID group, and the state of each controller (normal / blocked, etc.). The dirty queue described above also belongs to the cache control information in this control information 200.
[0025] Incidentally, when updating control information and cache data in the memory during a single controller blockage, it is not always necessary to individually write the logs related to the content one by one to the drive (memory backup drive 107), and they may be written together in a continuous area on the drive (memory backup drive 107). However, for example, before returning a write completion response to the host, by ensuring that the cache data and control information updated by the write process are written to the drive (memory backup drive 107), it is possible to prevent the written data from being lost due to a controller failure. The control information log buffer 502 and the cache data log buffer 503 are buffers for temporarily storing the logs in this way, and the control information log and the cache data log are temporarily stored respectively.
[0026] Figure 6 is a flowchart of the write process in the storage system of this embodiment, and the CPU 106 executes this process. First, the CPU 106 performs cache allocation (600). Cache allocation means allocating a part of the area in the memory for storing cache data for I / O processing and the like. Here, in order to store the write data sent from the host, an area of a size sufficient to store the data is allocated.
[0027] Subsequently, the CPU 106 performs cache data update processing (601). The content of the cache data update processing will be described later, but briefly, it is a process of receiving data from the host and storing the data in the previously allocated cache area.
[0028] Next, the CPU 106 determines whether the other controller is blocked (602). If it is blocked (Yes), the cache data duplication process is skipped. If it is not blocked (No), that is, if both controllers are operating, the cache data is duplicated (603). Cache data duplication is a process of copying the data received from the host to the memory of the other controller. For example, using the DMA built into the CPU 106, the data is copied from the memory of the own controller to the memory of the other controller.
[0029] Next, the CPU 106 performs control information update processing (604). The content of the control information update processing will be described later.
[0030] Next, the CPU 106 determines whether it is in the log backup mode (605). If it is in the log backup mode (Yes), the log backup process is performed (606). If it is not in the log backup mode (No), the log backup process is skipped. The content of the log backup process will be described later. After completing the above processing, the CPU 106 responds to the host that the write process has been completed (607). Thus, the write process is completed.
[0031] FIG. 7 is a flowchart of the memory protection method switching process in the storage system of this embodiment when one controller fails. This process is executed to switch the memory protection method from duplication between controllers to protection by logs when it is detected that the other controller has been blocked due to a failure or the like.
[0032] First, the CPU 106 determines whether the number of normal controllers remaining in the system is one (i.e., only the own controller) (700). If the number of remaining controllers is one (Yes), it proceeds to step 701. If it is not one (No), it skips all the remaining steps and ends this process.
[0033] Next, the CPU 106 sets the emergency detaching flag to ON (701). As a result, the operation of the detaching process described later will be changed. Also, while the emergency detaching flag is ON, the CPU 106 increases the execution frequency of the detaching process in order to store dirty data in the drive as soon as possible.
[0034] Next, the CPU 106 sets the log backup mode flag to ON (702). As a result, the CPU 106 will create a log when updating the memory. Finally, the CPU 106 executes the base image backup process (703). The details will be described later. Thus, the memory protection mode switching process at the time of single controller failure is completed.
[0035] FIG. 8 is a flowchart of the memory protection mode switching process at the time of controller recovery in the storage system of this embodiment. This process is executed to switch (restore) the memory protection mode from protection by log to duplication between controllers when it is detected that another controller has recovered due to maintenance replacement or the like and has become operational normally.
[0036] First, the CPU 106 performs control information duplication processing (800). This is a process of copying the control information on the memory to the memory of the recovered controller. It is completed when all the control information has been copied.
[0037] Next, the CPU 106 performs dirty data duplication processing (801). This is a process of copying the dirty data on the memory to the memory of the recovered controller. Also, each time a dirty data is copied, the cache control information regarding the dirty data is updated. It is completed when all the dirty data has been copied. Note that instead of copying the dirty data to another controller as in this embodiment, a method of protecting the dirty data by detaching it to the drive may be adopted.
[0038] Next, the CPU 106 sets the log backup mode flag to OFF (802). By doing so, subsequent memory updates will not be backed up to the drive as logs.
[0039] Finally, the CPU 106 performs a log deletion process (803). This is a process of deleting all the logs written on the log storage drive (memory backup drive 107) and the logs on the log buffer. For example, all the logs stored in the drive or memory may be overwritten with invalid data such as all zeros, or all the logs may be invalidated by setting the valid flag of all the log headers to OFF. The above completes the memory protection mode switching process when the controller recovers.
[0040] FIG. 9 is a flowchart of the destage process in the storage system of this embodiment. This destage process is started at a predetermined timing when there is dirty data in the memory. The startup frequency of the destage process is adjusted according to the amount of dirty data and the state of the storage system. For example, the higher the amount of dirty data, the higher the startup frequency. Also, when one of the controllers is blocked and there is dirty data outside the log protection stored in the drive, the destage process is started particularly frequently.
[0041] In the destage process, the CPU 106 first performs destage target data selection 900. The details of this process will be described later. Once the destage target data is determined, the CPU 106 then determines whether it can execute all stripe writes (901). This is a determination of whether all data for one stripe in a data protection method such as RAID 5 or RAID 6 exists in the cache. If all data for one stripe is present in the cache, new parity data can be generated without reading old data or old parity data from the drive. Therefore, if it is not possible to execute all stripe writes (No), the CPU 106 reads the old data and old parity data required for parity update from the drive (902), and if it is possible to execute all stripe writes (Yes), the CPU 106 skips this process.
[0042] Next, the CPU 106 generates new parity data (903) and writes the data and parity data to the drive (904). Subsequently, the CPU 106 performs control information update processing to delete the cache (905). In this process, the cache control information is updated to release the memory allocation of the cache data for which the destage is completed. Alternatively, identification information such as a flag indicating the dirty state may be set to OFF and the cache data may be left in the memory as clean (data that matches the data on the drive). The details of the control information update processing will be described later.
[0043] Next, if there are no dirty queues outside the log protection target, the CPU 106 sets the emergency destage flag to OFF (906). Finally, the CPU 106 invalidates the user data cache log related to the destaged dirty data. Thus, the destage process is completed.
[0044] Figure 10 is a flowchart of the destage target dirty selection process in the storage system of this embodiment. First, the CPU 106 determines whether the system is in the log backup mode (1000). Information indicating the log backup mode is held, for example, as a flag in the control information in the memory, and the CPU 106 makes the mode determination by referring to this information. If it is in the log backup mode (Yes), it proceeds to step 1001; if not (No), it proceeds to step 1003. In step 1001, it determines whether it is in an emergency message. Information indicating being in an emergency message is also held, for example, as a flag in the control information in the memory. If it is in an emergency message (Yes), the CPU 106 proceeds to step 1003; if not (No), the CPU 106 proceeds to step 1002. In step 1002, the CPU 106 selects the dirty data to be destaged from the log protection target dirty queue. Specifically, for example, the dirty data is taken out from the head of the dirty queue (dequeue), and the dirty data is made the target for destaging. On the other hand, in step 1003, the CPU 106 selects the dirty data to be destaged from the non-log protection target dirty queue.
[0045] FIG. 11 is a flowchart of the control information update process in the storage system of this embodiment. First, the CPU 106 updates the control information in the memory (1100). Next, the CPU 106 determines the necessity of non-volatilization (1101). If non-volatilization is necessary (Yes), it performs the log creation process (1102); if not (No), it skips the process. The content of the log creation process will be described later. Thus, the control information update process is completed.
[0046] FIG. 12 is a flowchart of the cache data update process in the storage system of this embodiment. First, the CPU 106 updates the cache data in the memory (1200). Specifically, for example, the data received from the host is written into the cache area already allocated in the memory. Next, the CPU 106 determines whether non-volatilization is required (1201). If non-volatilization is necessary (Yes), it proceeds to step 1202. If it is not necessary (No), it skips the subsequent processing and ends the cache data update process. Step 1202 is the log creation process. This is a process of creating a log regarding the updated cache data, the details of which will be described later.
[0047] Next, the CPU 106 determines whether the current cache data update is an overwrite (1203). That is, it checks whether there is a log regarding the cache data update within the address range included in the range of the cache area updated this time (this is called the "log of the same address") in the existing logs. If it exists, it is determined to be an overwrite. In the case of an overwrite (Yes), it invalidates the log of the same address written in the log header table (1204). In the case where it is not an overwrite (No), it skips this process. Finally, the CPU 106 updates the log header table (1205). Thus, the cache data update process is completed.
[0048] FIG. 13 is a flowchart of the log creation process in the storage system of this embodiment. First, the CPU 106 secures a sequence number (1300). The sequence number is a number indicating the order in which the logs are created, and the value is incremented by one each time a new log is created.
[0049] Next, the CPU 106 secures a log buffer for temporarily storing the log (1301). Specifically, when the data to be stored in the log is control information, it allocates an area of the size required to store the log to be created from the control information log buffer. When it is cache data, it allocates an area of the size required to store the log to be created from the cache data log buffer.
[0050] Subsequently, the CPU 106 creates a log header (1302). The log header includes the sequence number, the address of the target data in the memory, the size of the target data, etc. Next, the CPU 106 stores the log data in the log buffer (1303).
[0051] Finally, the CPU 106 performs activation processing on the created log (1304). Specifically, for example, the log header contains a flag indicating the validity / invalidity of the log, and the log is activated by turning on this flag. Thus, the log creation process is completed.
[0052] FIG. 14 is a flowchart of the log evacuation process in the storage system of this embodiment. The log evacuation process is a process of writing the logs accumulated in the log buffer to the drive, and is called when there is a need to write the logs to the drive, as was called before the host response in the aforementioned write process flowchart.
[0053] First, the CPU 106 extracts the unevacuated logs, that is, the logs that have not yet been written to the log storage drive, from the log buffer in the memory 105 (1400). Next, the CPU 106 writes the log to the log storage drive (memory backup drive 107) (1401). When the writing is completed, the CPU 106 deletes the written log from the log buffer (1402). Thus, the log evacuation process is completed.
[0054] FIG. 15 is a flowchart of the base image evacuation process in the storage system of this embodiment. As described above, the base image process is a process of writing the entire memory area to be protected to the drive, and in this embodiment, it is used for protecting control information and is executed at a predetermined timing, such as when a certain amount or more of the control information log on the drive has accumulated.
[0055] First, the CPU 106 refers to the sequence number and stores the latest sequence number at the current time (1500). Next, the CPU 106 writes the entire base image of the memory to the drive (1501). When this process is completed, the old logs become unnecessary. Next, the CPU 106 invalidates all the logs before the sequence number secured (stored) in step 1500 (1502). Thus, the base image backup process is completed.
[0056] FIG. 16 is a flowchart of the log recovery process in the storage system of this embodiment. This process is executed during the process of starting up the system after maintenance replacement work such as that on the controller has been performed after a system down due to the closing of both controllers, thereby recovering the control information and dirty data stored in the memory before the system down. This process is executed by the CPU 106 of a predetermined controller in the system before resuming the acceptance of I / O.
[0057] First, the CPU 106 reads the base image from the base image area on the log storage drive and stores it in the control information area on the memory (1600). Next, the CPU 106 reads the control information log and the cache data log from the log storage drive, sorts them in ascending order according to the sequence number, and reflects the contents from the oldest log to the newest log in order in the respective areas of the control information and cache data on the memory according to the address information written in the log header (1602). Thus, the log recovery process is completed.
Embodiment
[0058] Next, Embodiment 2 will be described. FIG. 17 is a diagram showing a configuration example of the storage system according to this embodiment. The storage system 100 of this embodiment includes a plurality of controllers 103 and a drive 110 which is a storage device. Each controller and the drive are connected by, for example, a switch (BE Switch) 109. Also, each controller 103 is connected to an interconnect switch 1701 and can communicate with each other. The interconnect switch is, for example, a PCIe switch, an Ethernet switch, an Infiniband switch, or the like.
[0059] Note that the controller of this embodiment includes a CPU, a memory, a memory backup drive, a front-end interface, and a back-end interface, similar to the controller of Embodiment 1, but is omitted in the drawing.
[0060] Also, in this drawing, a configuration in which two controllers are mounted in one controller enclosure is taken as an example, but the configuration for implementing the present invention is not necessarily limited to this configuration. For example, each controller may be in an independent housing, or three or more controllers may be mounted in one controller enclosure.
[0061] Also, in this drawing, a configuration in which a storage system includes four controllers is illustrated, but in this embodiment, the number of controllers may be three or more and is not necessarily limited to a configuration of four.
[0062] The storage system is connected to a storage area network (SAN) 101 such as Fibre Channel or Ethernet, and a host computer (hereinafter referred to as a host) 102 is also connected to the SAN 101. The SAN 101 may include a switch or the like. Also, a plurality of hosts may be connected to the SAN 101.
[0063] As is clear from the figure, the main difference between Embodiment 1 and this embodiment lies in the number of controllers. In this embodiment, even if one of three or more controllers fails, data can be made redundant on the memories of the remaining two or more controllers. Therefore, in this embodiment, while there are two or more normal controllers, the system operates in the same manner as when both controllers are normal in Embodiment 1, and control information and user data are made redundant between the memories of the controllers. Then, when only one normal controller remains, the system switches to the log backup mode and operates in the same manner as when one controller is blocked in Embodiment 1.
[0064] As described above, the disclosed storage system is a storage system including a drive 110 which is a non-volatile storage device, and a plurality of storage controllers 103 that control reading from and writing to the storage device. The storage device is a drive for storing user data. Each of the plurality of storage controllers 103 has a processor (CPU 106) and a memory 105. The storage controller has a first memory protection method of the memory replication method for replicating data on the memory onto the memory of the corresponding storage controller, and a second memory protection method of the log backup method for generating a log related to the update of data on the memory and writing it to a non-volatile medium. The storage controller stores a write request from a host to the storage device as cache data in the memory, protects the cache data by the first memory protection method or the second memory protection method, and then returns a write completion response to the host. After the write response is completed, the cache data is destaged to the storage drive. The storage controller switches which of the first memory protection method and the second memory protection method to use according to the operating state of other storage controllers. In addition, the storage controller is associated with another storage controller to form a redundant configuration, is used as the first memory protection method, and is used as the second memory protection method when the corresponding storage controller is blocked. In this way, the storage controller recognizes the states of other controllers in the system and operates in a write-back mode when other controllers are normal. Also, when other controllers are in an abnormal state such as a failure, the storage controller generates a log regarding the update of the memory content during the read / write operation and writes the log to the storage device. By doing so, compared with the write-through mode, the number of drive accesses required before the host response can be reduced, and thus the response performance can be improved. Therefore, a storage system and a data protection method with high performance and high reliability can be realized.
[0065] The non-volatile medium that is the write destination of the log is, for example, the memory backup drive 107 provided inside the storage controller 103. In this way, by providing the memory backup drive 107 inside the storage controller 103, the time until the writing of the log is completed can be shortened, and high performance can be achieved.
[0066] Also, as the non-volatile medium that is the write destination of the log, a part of the non-volatile storage device may be used. With this configuration, the configuration of the storage controller 103 can be simplified, and cost reduction can be realized.
[0067] Further, when the storage controller 103 detects an occlusion of the corresponding storage controller, it switches the operation from the memory replication method to the log backup method, preferentially destages the pre-switching cache data (dirty queue 401 not subject to log protection), which is the cache data before the operation switch, to the storage device, and destages the post-switching cache data (dirty queue subject to log protection), which is the cache data after the operation switch, after all the pre-switching cache data has been destaged. Thereafter, when the storage controller 103 detects recovery from the occlusion of the corresponding storage controller, it replicates the data in the memory onto the memory of the corresponding storage controller, deletes the log, and switches the operation from the log backup method to the memory replication method. Therefore, it is possible to quickly destage the pre-switching cache data that is not protected by the memory replication method and reduce the risk of data loss.
[0068] Note that the present invention is not limited to the above-described embodiments and includes various modifications. The above-described embodiments have been described in detail for easy understanding of the present invention and are not necessarily limited to those having all the configurations described. Further, not only deletion of such a configuration but also replacement or addition of a configuration is possible. For example, although the case where a controller fails has been described, it may be applied to the case where one of the controllers is put into a standby state for the purpose of reducing power consumption.
Explanation of Reference Numerals
[0069] 100: Storage system, 102: Host, 103: Storage controller, 105: Memory, 107: Memory backup drive, 110: Drive, 400: Dirty queue subject to log protection, 401: Dirty queue not subject to log protection.
Claims
1. A storage system comprising a non-volatile memory device and a plurality of storage controllers for controlling reading from and writing to the memory device, wherein the memory device is a drive for storing user data, each of the plurality of storage controllers has a processor and a memory, the storage controller includes a first memory protection method of a memory replication method for replicating data on the memory onto the memory of the corresponding storage controller, and a second memory protection method of a log evacuation method for generating a log regarding update of the data on the memory and writing it to a non-volatile medium, the storage controller stores a write request from a host to the memory device as cache data in the memory, protects the cache data by the first memory protection method or the second memory protection method, and then returns a write completion response to the host, and destages the cache data to the memory drive after the write response is completed, the storage controller switches which of the first memory protection method and the second memory protection method to use according to the operating state of another storage controller A storage system characterized by the above.
2. The storage system according to claim 1, wherein the storage controller is associated with another storage controller to form a redundant configuration, is used as the first memory protection method, and is used as the second memory protection method when the corresponding storage controller is blocked. A storage system characterized by this.
3. The storage system according to claim 2, wherein the non-volatile medium, which is the write destination of the log, is provided inside the storage controller. A storage system characterized by this.
4. The storage system according to claim 2, wherein a part of the non-volatile memory device is used as the non-volatile medium, which is the write destination of the log. A storage system characterized by this.
5. The storage system according to claim 1, wherein in the destaging, the data is stored in the final storage area of the memory drive using the storage function provided by the storage system. A storage system characterized by this.
6. The storage system according to claim 2, The storage controller performs an operation switch from the memory replication method to the log evacuation method when detecting an occlusion of the corresponding storage controller, preferentially destages the pre-switch cache data, which is the cache data before the operation switch, to the storage device, and destages the post-switch cache data, which is the cache data after the operation switch, after all of the pre-switch cache data has been destaged. A storage system characterized in that when detecting a recovery from the occlusion of the corresponding storage controller, it replicates the data on the memory onto the memory of the corresponding storage controller, deletes the log, and performs an operation switch from the log evacuation method to the memory replication method.
7. A data protection method for a storage system including a non-volatile storage device and a plurality of storage controllers that control reading and writing to the storage device, wherein the storage device is a drive for storing user data, each of the plurality of storage controllers has a processor and a memory, the storage controller includes a first memory protection method of the memory replication method for replicating the data on the memory onto the memory of the corresponding storage controller and a second memory protection method of the log evacuation method for generating a log regarding the update of the data on the memory and writing it to a non-volatile medium, the step of the storage controller storing a write request from a host to the storage device as cache data in the memory, the step of the storage controller protecting the cache data by the first memory protection method or the second memory protection method, the step of the storage controller returning a write completion response to the host, the step of the storage controller destaging the cache data to the storage drive after the write response is completed, and the storage controller further includes the step of switching which of the first memory protection method and the second memory protection method to use according to the operating state of another storage controller. A data protection method characterized by the above.
Citation Information
Patent Citations
Storage control device, storage device, and storage control method
JP2011170589A
Storage apparatus, data management method, and data management program
JP2019192004A
Storage control device , a storage system, a storage control method and a program thereof
US20110202791A1
Flushes after storage array events
US20180276142A1
Control system for disk cache
JP1994309232A