Storage device and storage destination control method
By configuring storage devices with multiple drives and controllers that manage cache data logs based on performance needs, the solution addresses the challenge of balancing cost, performance, and reliability in storage devices, particularly under complex data protection methods.
Patent Information
- Application Number
- JP2023198384
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-22
- Publication Date
- 2025-06-03
AI Technical Summary
Existing storage devices face challenges in balancing cost and performance while ensuring high reliability, particularly when using complex data protection methods like RAID6, which can lead to increased response times and potential performance bottlenecks.
The solution involves a storage device configuration with multiple types of non-volatile drives and controllers that selectively generate and write cache data logs to appropriate drives based on required performance, using a memory data protection function to manage log storage destinations.
This approach allows for high reliability while achieving a balance between cost and performance, by optimizing log storage based on performance requirements and using existing drive configurations efficiently.
Smart Images

Figure 2025084464000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a storage device and a storage destination control method, and is suitably applied to a storage device related to a technique of writing data in a memory and its updated content as a log to a drive when performing input / output processing of data with a host, for example.
Background Art
[0002] A storage device writes data received from a host computer (hereinafter abbreviated as "host") to a non-volatile drive (hereinafter abbreviated as "drive") such as an SSD or HDD via a cache memory (hereinafter abbreviated as "cache"). That is, data requested to be written from the host is once held in the cache and then written to a predetermined drive. There are roughly two types of methods for writing data from the cache to the drive.
[0003] One method is, for example, a method called write-through, in which data is written to the drive before returning a response to a write request to the host. Another method is, for example, a method called write-back or write-after, in which a response to a write request is returned to the host when the data is stored in the cache. In the case of the write-back method, writing of data to the drive is performed at an arbitrary timing after the data is stored in the cache.
[0004] Therefore, in the case of the write-back method, a response can be returned to the host without waiting for the completion of writing to the drive, so the response time can be shortened compared to the write-through method. On the other hand, in the case of the write-back method, the data for which the write from the host has been completed temporarily exists only in the cache. Therefore, it is necessary to appropriately protect the data in the cache. The storage device, for example, adopts a redundant configuration with a plurality of controllers and ensures data redundancy by copying the data received by one controller to the cache of another controller. Also, to prepare for power outages and power failures, for example, the cache is protected by a battery.
[0005] High reliability and high performance are required of the storage device. Therefore, the storage device can selectively use the above-described write-through method and write-back method according to the situation. In recent years, a usage method is known in which it operates in the write-back method when the cache can be appropriately protected, and switches to the write-through method when the cache cannot be protected. By doing so, a response can be returned at high speed by the write-back method during normal times, and reliability can be ensured by the write-through method even when, for example, a controller fails and the redundancy of the cache is lost.
[0006] However, recent storage devices apply relatively complex data protection methods such as RAID (Redundant Array of Independent Disks) 6, and the response time when operating in the write-through method often becomes long. For example, in the case of RAID6, every time a write request is received from the host, it is necessary to read old data and two old parities (P parity and Q parity) from the drive, generate parity data, and then write the new data and two new parities to the drive and return a response to the host. Since it is necessary to perform multiple drive accesses in this way, the response to the host becomes slow.
[0007] In contrast, in the method disclosed in Patent Document 1, the updated content of the data on the memory is written to the drive in the form of a log to protect the data on the memory. In this method, since a response can be returned to the host after finishing writing the updated content of the memory associated with the write request from the host as a log to the drive, the number of drive accesses is less than that of the write-through method, and the response time can be shortened.
Prior Art Documents
Patent Documents
[0008]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0009] In such a method disclosed in Patent Document 1, since the updated content itself of the data on the memory is written to the drive as a log, the amount of writing to the drive increases compared to the write-back method, so the drive for storing the log may become a performance bottleneck, and there is a risk that high reliability cannot be ensured. As such a drive for storing such a log, for example, a configuration using a large number of high-speed SSDs can also be considered, but if such a configuration is adopted, the cost for preparing the drive may increase. On the other hand, the requirements for the storage device are various. Sometimes performance is emphasized, and sometimes cost is to be suppressed even at the sacrifice of some performance. Therefore, it is required to appropriately balance cost and performance according to the requirements of the user.
[0010] The present invention has been made in consideration of the above points, and aims to propose a storage device capable of ensuring high reliability while balancing cost and performance, and a log storage destination control method capable of ensuring high reliability while balancing cost and performance.
Means for Solving the Problems
[0011] In order to solve such problems, in the present invention, there are provided a plurality of types of drives for non-volatile storage of data, a plurality of controllers for controlling reading and writing of the data between the host and the drives, and a volatile memory in which the data is temporarily stored. The controller has a memory data protection function of generating a cache data log including the data in the memory and a header related to the update of the data and writing the cache data log to any one of the plurality of types of drives. The controller selects a drive to store the cache data log from among the plurality of types of drives using the memory data protection function according to required performance, and stores the cache data log in the selected destination drive.
[0012] Further, in the present invention, there is provided a storage destination control method for a storage device including a plurality of types of drives for non-volatile storage of data, a plurality of controllers for controlling reading and writing of the data between the host and the drives, and a volatile memory for data cache control information. The controller has a data protection step of executing a memory data protection function of generating a cache data log including the data in the memory and a header related to the update of the data and writing the cache data log to any one of the plurality of types of drives. In the data protection step, the controller selects a drive to store the cache data log from among the plurality of types of drives using the memory data protection function according to required performance, and stores the cache data log in the selected destination drive.
Advantages of the Invention
[0013] According to the present invention, it is possible to ensure high reliability while achieving a balance between cost and performance.
Brief Description of the Drawings
[0014]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Embodiments for Carrying Out the Invention
[0015] Hereinafter, based on the drawings, an embodiment of the present invention will be described in detail. FIG. 1 is a block diagram showing an example of a configuration example of a storage device 100 according to the present embodiment. The storage device 100 includes a plurality of controllers 103 and a drive 110 for user data (hereinafter also referred to as “UDD”). The controller 103 has a function of providing a volume to be read from and written to data to a host computer (hereinafter referred to as “host”) 102.
[0016] Each of the plurality of controllers 103 includes a CPU 106, a memory 105, at least one memory backup drive 107, a front-end interface (hereinafter abbreviated as “FE I / F”) 104, and a back-end interface (hereinafter abbreviated as “BE I / F”) 108.
[0017] The drive 110 is a drive that stores data non-volatilely, and is, for example, an SSD (Solid State Drive) using a non-volatile device such as a flash memory as a storage medium, an HDD (Hard Disk Drive) using a magnetic disk as a storage medium, or the like. A base image area for storing a base image described later is allocated to a part of the storage area of the drive 110.
[0018] The memory 105 is a drive in which data is temporarily stored, and is a semiconductor memory such as a DRAM (Dynamic Random Access Memory). A cache area capable of temporarily storing data is allocated to a part of the storage area of the memory 105. In the present embodiment, the cache area may be simply referred to as “cache”.
[0019] The memory backup drive 107 is a drive that stores data non-volatilely, and is a drive for memory backup. The memory backup drive 107 may be a drive of a non-volatile device such as an SSD, for example. The memory backup drive 107 is used to store the stored content of the memory 105, for example, when there is a loss of external power supply.
[0020] The FE I / F 104 is, for example, a Fibre Channel HBA (Host Bus Adapter) or a NIC (Network Interface Controller). The BE I / F 108 is, for example, a SAS (Serial Attached SCSI) HBA or a PCI-Express (Peripheral Component Interconnect-Express) (hereinafter abbreviated as "PCIe") adapter or a NIC.
[0021] Each controller 103 and the drive 110 are connected by, for example, a Backend switch (hereinafter abbreviated as "BE switch") 109. Also, the CPUs 106 of the plurality of controllers 103 are connected by an interconnect such as PCIe, for example. Note that the CPU 196 and the CPU 106 of both controllers 103 may be connected via a PCIe switch, for example.
[0022] The storage device 100 is connected to a storage area network (hereinafter abbreviated as "SAN") 101 such as Fibre Channel or Ethernet, for example. Also, the host 102 is connected to the SAN 101. The SAN 101 may include a switch or the like. Also, a plurality of hosts may be connected to the SAN 101.
[0023] The management node 111 is connected to the storage device 100. The management node 111 is a computer on which a program for the administrator to perform settings of the storage device 100 and the like operates. The management node 111 is connected to each controller 103 via, for example, a management local area network (LAN). Note that the management node 111 may be connected to the SAN 101 instead. Note that the management node 111 does not necessarily have to be an independent computer. For example, the function of the management node 111 may be provided inside the storage device 100 or in the host 102.
[0024] In addition to the above-described drive 110, the storage device 100 is equipped with a memory backup drive 107 that stores data non-volatilely, and it can be said that the storage device 100 is equipped with a plurality of types of drives that store data non-volatilely. The above-described plurality of controllers 103 control the reading and writing of data between the host 102.
[0025] Each controller 103 has a memory data protection function of generating a cache data log including the data in the memory 105 and a header regarding data update in the data protection step and writing it to any one of the plurality of types of drives. In the data protection step, the controller 103 selects a drive to store the cache data log from among the plurality of types of drives using the memory data protection function according to the required performance.
[0026] In the present embodiment, when one of the plurality of controllers 103 is blocked, the other unblocked controller executes the memory data protection function in the data protection step.
[0027] In this embodiment, the above-described multiple types of drives include a memory backup drive 107 as a storage area for backing up the stored content of the memory 105 of one of the controllers 103 when the one controller 103 is blocked, and a drive 110 for user data (hereinafter also referred to as "UDD") for reading and writing data regardless of whether the one controller 301 is blocked.
[0028] In this embodiment, the data input and output between the host and the above is composed of a plurality of block unit data obtained by dividing the data into blocks. The above-described cache data log includes each block unit data and a header related to the update of each block unit data. The other controller 103 selects, in the data protection step, a drive for storing each block unit data and each header as a cache data log from among the multiple types of drives.
[0029] In the following description, when there is no need to particularly distinguish between the block unit data 201 and the data composed of the block unit data 201, it may sometimes be simply referred to as "data".
[0030] In this embodiment, when the other controller 103 that is not blocked is set (hereinafter referred to as "performance priority") in which it is prioritized to satisfy a predetermined performance requirement regarding, for example, the data read and write speed as the above-described necessary performance, the drive 110 is selected from among the multiple types of drives. On the other hand, when the other controller 103 is set (hereinafter referred to as "non-performance priority") in which it is not prioritized to satisfy a predetermined performance requirement as the above-described necessary performance, the memory backup drive is selected from among the multiple types of drives. In this embodiment, as the above-described predetermined performance requirement, for example, it can be mentioned that even if a cache data log described later, which has a larger data volume than a control information log described later, is written to the drive, the drive does not become a performance bottleneck.
[0031] Among the multiple types of drives in this embodiment, at least one SSD (Solid State Drive) is included. When the above-mentioned non-blocked other controller 103 is set such that it is prioritized to meet the predetermined performance requirements as the necessary performance described above in the data protection step, an SSD is selected from among the multiple types of drives.
[0032] The storage device 100 according to this embodiment includes, for example, in the memory 105, as an example of a storage destination management table that manages information regarding at least the type and capacity of multiple types of drives, although details will be described later, a cache data log storage destination management table (see FIG. 6). The other controller 103 refers to the cache data log storage destination management table and selects a drive to store the cache data log. The memory 105 also has a cache data log header table that manages the header of the cache data log.
[0033] FIG. 2 is a diagram showing an example of an overview of the write operation when both controllers 103 are normal. The CPU 106 of the controller 103 receives a write request from the host 102, receives block unit data 201 from the host 102, and writes the block unit data 201 to the memory 105 in its own controller 103 and the memory 105 of the other controller 103. The CPU 106 updates the control information 200 in the memory 105. The CPU 106 returns a write completion response to the host 102.
[0034] In this write operation, each block unit data 201 and each control information 200 on the memory 105 are duplicated among the multiple controllers 103 to prepare for a failure of the controller 103.
[0035] As described above, the memory backup drive 107 is used as a destination for backing up the block unit data 201 and the control information 200 written to the memory 105 during a power failure. That is, when the external power supply is cut off due to a power outage or the like, the controller 103 backs up the block unit data 201 and the control information 200 in the memory 105 to the memory backup drive 107 before the battery that backs up the power runs out. After the backup is completed, the storage device 100 is stopped.
[0036] After the power supply is restored, the controller 103 reads the block unit data 201 and the control information 200 backed up in the memory backup drive 107 into the memory 105 and operates the storage device 100, so that the operation of the storage device 100 can be resumed without losing the data before the interruption.
[0037] By the way, the block unit data 201 in the memory 105 is written to the drive 110 at an arbitrary timing after the write completion response (destage). At the time of destage, for example, parity data is generated for high reliability, and this parity data is written to a drive different from the block unit data 201.
[0038] FIG. 3 is a diagram showing an example of an outline of a write operation when one controller 103 of the storage device 100 is blocked.
[0039] The CPU 106 of the controller 103 receives a write request from the host 102, receives the block unit data 201 from the host 102, and writes the block unit data 201 to the memory 105 in its own controller. Further, the CPU 106 updates the control information 200 in the memory 105.
[0040] Furthermore, the CPU 106 writes, as a cache data log, the header regarding the update content together with the block unit data 201 written from the host 102 to the selected memory backup drive 107 or drive 110 as described above. At the same time, the CPU 106 writes, as a control information log, the header of the update content together with the control information 200 to the memory backup drive 107 or drive 110. Then, the CPU 106 returns a write completion response to the host 102.
[0041] Note that in this embodiment, the "header" indicates information including, for example, the address where data is stored as the update content of the data and the update order of the data, the "log" includes, for example, the cache data log and the control information log, the "cache data log" indicates the cache data (each block unit data 201) and the header indicating the update content thereof, and the "control information log" indicates the control information (each control information 200) and the header indicating the update content thereof. Note that the block unit data 201 written to the memory 105 as described above is destaged at an arbitrary timing after the write completion response.
[0042] In this write operation, in case the remaining controller 103 fails, each block unit data 201 on the memory 105 and its update content are written to the drive 110 as a cache data log (log backup). In this log backup, as a log, together with the cache data log, each control information 200 and its update content may also be written to the drive 110. If the remaining controller 103 fails, the storage device 100 will stop (system down) once. However, after the controller 103 is repaired and replaced, data loss can be prevented by restoring each block unit data 201 and each control information 200 on the memory 105 using the cache data log and the control information log.
[0043] In the following description, to avoid confusion, the differences between "destage" and "log evacuation" are clarified here. First, "destage" means writing the dirty data on the cache to the drive 110, which is the final storage medium. The drive 110 is generally protected by a method such as RAID (Redundant Array of Independent Disks) 6. In that case, parity data is generated during destage, and the parity data is also written to the drive 110.
[0044] For the block unit data 201 for which destage is completed, since the block unit data 201 on the cache area of the memory 105 and the block unit data 201 on the drive 110 are in a consistent state (hereinafter referred to as "clean"), the block unit data 201 may be lost from the cache area without any problem.
[0045] On the other hand, "log evacuation" means temporarily writing the block unit data 201 (cache data) on the cache area of the memory 105 and its updated content, i.e., the cache data log, and the control information 200 and its updated content, i.e., the control information log, to a non-volatile storage medium (drive) such as the memory backup drive 107 in case of a failure of the controller 103. "Log evacuation" is included in the memory data protection function.
[0046] Here, in this embodiment, a state where the block unit data 201 on the cache area and the block unit data 201 on the drive 110 do not match is referred to as "dirty". Even after log evacuation is completed, the dirty data on the cache area is retained until destage is completed. When destage is completed, since the header of the block unit data 201 on the cache area becomes unnecessary, the log can be deleted from the drive when destage is completed.
[0047] In addition, when storing the cache data log in the memory backup drive 107 in this embodiment, an area already allocated for both controllers 103 to save the stored content of the memory 105 during normal operation may be used as an area for storing the cache data log by one of the controllers 103 when it is blocked. By doing so, since it is not necessary to add a drive for storing the cache data log or increase the storage capacity, it is advantageous in terms of cost compared to the case of storing the cache data log in the drive 110 or mounting a separate drive for storing the cache data log.
[0048] Incidentally, the block unit data 201 and the cache data log including its update content are generally in block units such as 512 bytes or 4 KB for the block unit data 201 written from the host 102, so the granularity is relatively large. On the other hand, since the control information 200 is in byte units, for example, the control information log including the update content of the control information 200 has a relatively small granularity. Also, the ratio of the cache area in the entire memory 105 is relatively large. Therefore, for the control information log, periodically, the entire memory area (hereinafter referred to as "base image") storing the control information 200 is written to the drive 110, and all the previously written control information logs are discarded, and the area where the control information log was written is recovered as a free area. This method is called the "base image backup method".
[0049] On the other hand, for the cache data log, unnecessary cache data logs that are not the latest among the cache data logs are identified and discarded (invalidated). By doing so, free areas will be scattered in the storage area of the cache data log (cache log buffer described later), so by packing only the valid cache data logs into another area in advance at a predetermined timing, continuous free areas can be recovered. This method is called the "garbage collection method".
[0050] In this embodiment, by selectively using both of these two methods, it is possible to suppress the consumption of storage capacity for backing up the base image, reduce the management information for free space management, and reduce the overhead for free space recovery.
[0051] FIG. 4 is a diagram showing an example of the stored content of the memory 105. The memory 105 has a storage control program 400, control information 200, cache data 401, a control information log buffer 402, and a cache data log buffer 403. In the following description, the control information log buffer 402 and the cache data log buffer 403 are also collectively referred to as "log buffers".
[0052] The storage control program 400 is a program that controls the entire storage device 100 under the control of the CPU 106. Each process such as the writing process described later is executed by this storage control program 400.
[0053] The control information 200 is data used by the storage control program 400 to control the execution of various programs. The control information 200 includes a control information log storage destination management table 404 and a cache data log storage destination management table 405. The contents of these control information log storage destination management table 404 and cache data log storage destination management table 405 will be described later.
[0054] The control information 200 includes cache control information, configuration information, and information regarding the state (normal / blocked, etc.) of each controller 103. The cache control information includes information such as the correspondence between the address where the cache data is stored and the logical address (hereinafter also abbreviated as "LBA") within the volume, and the state (dirty / clean) of the cache data. The configuration information includes information such as the type and capacity of the drive, and the type and configuration of the RAID group.
[0055] Incidentally, in the present embodiment, when updating the control information 200 or the block unit data 201 in the memory 105, it is not always necessary to individually write each header related to the update content. For example, they may be collectively written in a continuous storage area. However, for example, before returning a write completion response to the host, by ensuring that the block unit data 201 and the control information 200 updated by the write are written, it is possible to prevent the block unit data 201 for which the write has been completed from being lost due to a failure of the controller 103.
[0056] The control information log buffer 402 and the cache data log buffer 403 are partial storage areas of the memory 105 and are buffers for temporarily storing the cache data log on the memory 105. The control information log buffer 402 and the cache data log buffer 403 temporarily store the control information log and the cache data log, respectively.
[0057] FIG. 5 is a diagram showing an example of the content of the control information log storage destination management table 599. In this control information log storage destination management table 599, records are stored one row for each area of the control information log storage destination, and each row includes columns for an ID (Identifier) 500, a drive number 502, a logical block address (LBA) 503, and a capacity size 504.
[0058] When there is only one storage destination for the control information log in this control information log storage destination management table 599, for example, a value indicating an example of being invalid (the "N / A" shown in the figure) is stored in the remaining rows.
[0059] The ID 500 is, for example, a sequential number starting from 0. The type 501 indicates the types of a plurality of types of non-volatile drives, and for example, values indicating an example such as a memory backup drive (hereinafter also abbreviated as "MBD") 107 and a drive for user data (hereinafter also abbreviated as "UDD") 110 are stored.
[0060] Drive number 502 is the number of the drive that stores the control information log. The logical block address 503 is the logical block address (LBA) at the head of the area for storing the control information log. The capacity size 504 shows an example of the size of the area for storing the control information log. In the illustrated example, for simplicity of explanation of the capacity size 504, a value of the capacity size in GB units is used, but for example, a value indicating the number of blocks of the block unit data 201 is stored.
[0061] FIG. 6 is a diagram showing an example of the content of the cache data log storage destination management table 699. This cache data log storage destination management table 699 is a table having the same column structure as the above-described control information log storage destination management table 599, and the meaning of each column is also the same as that of the control information log storage destination management table 599.
[0062] In FIGS. 5 and 6 described above, the case where the storage destination of the control information log is assigned to the memory backup drive (MBD) 107 and the storage destination of the cache data log is mainly assigned to the drive (UDD) 110 is illustrated. The actual assignment of the log storage destination does not necessarily have to follow this example, but a suitable log storage destination assignment policy can be mentioned from the viewpoints of the characteristics of each log and cost and performance. The policy is shown below.
[0063] First, since the amount of data in the control information log is smaller than that in the cache data log, the input / output load on the drive for storing the control information log is relatively small compared to the cache data log. On the other hand, since the amount of data in the cache data log is larger than that in the control information log, its input / output load is large, and it can be said that the drive 110 may become a performance bottleneck.
[0064] Next, the memory backup drive 107 has an area already allocated for backing up the stored content of the memory 105 in case of a power failure as described above. This area has a capacity equal to or greater than the capacity of the memory 105. When one of the controllers 103 is blocked, the other controller 103 uses this area as a log storage area, thereby suppressing the capacity consumption of the drive 110 and reducing the cost accordingly. On the other hand, multiple drives 110 are installed in the storage device 100, and in particular, in cases where high performance is required, many high-speed SSDs (Solid State Drives) are often installed.
[0065] Therefore, it is desirable to allocate the storage destination of the control information log to the memory backup drive 107. Also, when the storage destination of the cache data log is set with a priority on meeting a predetermined performance requirement (hereinafter referred to as "performance priority"), it is allocated to the drive 110. On the other hand, when the setting does not prioritize meeting a predetermined performance requirement (hereinafter referred to as "non-performance priority") (for example, when cost is prioritized), it is desirable to allocate it to the memory backup drive 107. Alternatively, a part of the storage destination of the cache data log may be allocated to the drive 110 and the rest to the memory backup drive 107 so as to meet the predetermined performance requirement. Based on the above guidelines, a flowchart showing an example of the procedure for setting the specific log storage destination will be described.
[0066] FIG. 7 is a flowchart showing an example of the procedure for setting the storage destination of the control information log. The control information log storage destination setting process is executed by the storage control program 400 under the control of the CPU 106. In the illustrated example, it is described as if there are, for example, i memory backup drives (MBDs) 107. Also, although the main body of the following processes is the storage control program 400 operating under the control of the CPU 106, it will be described below as being executed by the CPU 106.
[0067] First, the CPU 106 sets the counter i to "0" (step S700). The CPU 106 allocates an area for storing the control information log to the i-th memory backup drive 107 (step S701). Next, the CPU 106 increments i (step S702), and if i is smaller than the number of memory backup drives 107, it returns to step 701 and repeats this process.
[0068] FIG. 8 is a flowchart showing an example of the procedure for setting the cache data log storage destination. First, the CPU 106 determines whether the setting of the storage device 100 is "performance priority".
[0069] The "performance priority" setting may be set, for example, by the administrator of the storage device 100, or may be automatically set by the storage control program 400 of the storage device 100 or the program of the management node 111. In the case of automatic setting in this embodiment, for example, the following determination is made based on the device configuration of the storage device 100.
[0070] First, the CPU 106 mounted on the controller 103 determines that its own performance is low and the memory backup drive 107 can follow the operation of the CPU 106, and it is predicted that the memory backup drive 107 will not become a bottleneck even if all cache data logs are stored in the memory backup drive 107. In this case, it is set to "non-performance priority (cost priority)".
[0071] On the other hand, the CPU 109 mounted on the controller 103 determines that its own performance is high and the memory backup drive 107 cannot follow the operation of the CPU 106, and if there is free space in a high-speed drive such as an SSD mounted as the drive 110, it is set to "performance priority".
[0072] As another method of automatic setting, for example, when the target performance specified at the time of provisioning the storage device 100 is higher than a predetermined threshold, it may be set to "performance priority", while when the target performance is equal to or lower than the predetermined threshold, it may be set to "non-performance priority".
[0073] When it is not set to "performance priority" (that is, when cost priority is set), the CPU 106 executes step S807. In step 807, the CPU 106 allocates an area for storing cache data to the memory backup drive 107 (MBD). Since the procedure is the same as the control information log storage destination setting process, the description thereof is omitted.
[0074] On the other hand, when it is set to "performance priority", the CPU 106 sets the counter i to 0 (step S801). The CPU 106 determines whether there is an empty area in the i-th drive 110 (here, for example, an SSD is assumed) where the cache data log can be stored (step S802).
[0075] If there is an empty area, the CPU 106 allocates an area for storing the cache data log to the i-th SSD (step S803). On the other hand, if there is no empty area, the CPU 106 skips this step S803 and executes step S804.
[0076] Next, the CPU 106 increments i (step S804), and if i is smaller than the number of SSDs, it returns to step 802 and repeats the process from there (step S805). When i becomes equal to or greater than the number of SSDs, the CPU 106 executes step S806.
[0077] In step S806, the CPU 106 determines whether the capacity required to store the cache data has been allocated. If it has been allocated, the CPU 106 ends the process, while if the allocation of the required capacity has not yet been completed, it executes step 807 described above.
[0078] In the above flowchart, the area is allocated to the SSD among the drives 110 for user data. For example, the reason for not allocating the area to an HDD (Hard Disk Drive) is that the performance of the HDD is low. In this embodiment, even if it is an SSD, a slow one may not be subject to area allocation. Also, if a high-speed drive other than an SSD can be made in the future, the area may be allocated to that drive as well.
[0079] FIG. 9 is a flowchart showing an example of the procedure of the writing process. First, the CPU 106 performs cache allocation (step S900). Cache allocation means, for example, allocating an area in the memory 105 for storing cache data for I / O processing or the like. Here, in order to store the block unit data 201 transmitted from the host 102, an area of a size sufficient to store the block unit data 201 is allocated.
[0080] Subsequently, the CPU 106 performs cache data update processing (step S901). The details of the cache data update processing will be described later. Briefly speaking, it is a process of receiving the block unit data 201 from the host and storing the block unit data 201 in the cache area allocated previously.
[0081] Next, the CPU 106 determines whether one of the controllers 103 is blocked (step S902). If one of the controllers 103 is blocked, the CPU 106 skips the cache data duplication process. On the other hand, if one of the controllers 103 is not blocked, that is, if both controllers 103 are operating, the CPU 106 duplicates the cache data (step S903). Duplication of cache data is a process of copying block unit data 201 received from the host 102 to the memory 105 of the other controller 103. Here, for example, using DMA (Direct Memory Access) built into the CPU 106, data is copied from the memory 105 of its own controller 103 to the memory 105 of the other controller 103.
[0082] Next, the CPU 106 performs control information update processing (step S904). The details of the control information update processing will be described later.
[0083] Next, the CPU 106 determines whether it is in the log evacuation mode (step S905). If it is in the log evacuation mode (Yes), the CPU 106 performs the kilolog evacuation process. On the other hand, if it is not in the log evacuation mode (No), the CPU 106 skips the log evacuation process (step S906). The details of the log evacuation process will be described later. After completing the above processing, the CPU 106 responds to the host that the write process has been completed (step S907). Thus, the write process is completed.
[0084] Figure 10 is a flowchart showing an example of the procedure of the destage process. This destage process is started at a predetermined timing when there is dirty data on the memory 105. The activation frequency of the destage process is adjusted according to the amount of dirty data and the state of the storage device 100. For example, the higher the amount of dirty data, the higher the activation frequency of the destage process. Also, when one of the controllers 103 is blocked and there is dirty data that is not subject to cache data log protection and is not stored in the drive 110, the destage process is started particularly frequently.
[0085] In this destage process, the CPU 106 first selects the data to be destaged (step S1000). Next, the CPU 106 determines whether all stripe writes can be executed (step S1001). Here, being able to execute all stripe writes means that, for example, when adopting a configuration of 3D + 1P (one parity data block for three data blocks) in RAID5, all the data of the three stripe blocks included in one stripe are present in the cache.
[0086] As such, step S1001 is, for example, a determination as to whether all the data for one stripe in a data protection method such as RAID5 or RAID6 exists in the cache. When all the data for one stripe is present in the cache, the CPU 106 can generate new parity data without reading old data or old parity data from the drive 110. Therefore, when not all stripe writes can be executed, the CPU 106 reads the old data and old parity data necessary for updating the parity data from the drive 110 (step S1002), and when all stripe writes can be executed, the CPU 106 skips this step S1002.
[0087] Next, the CPU 106 generates new parity data (step S1003) and writes the data and the new parity data to the drive 110 (step S1004). Subsequently, the CPU 106 performs control information update processing (step S1005). In this control information update processing, the CPU 106 updates the control information 200 and releases the allocation of the cache data for which the destage has been completed. Alternatively, the CPU 106 may turn off identification information such as a flag indicating an example of the dirty state and leave it on the memory 105 as clean state (a state where the content matches the data on the drive) cache data. The content of the control information update processing will be described later.
[0088] Finally, the CPU 106 invalidates the cache data log related to the dirtied dirty data (step S1006). With this, the destaging process is completed.
[0089] FIG. 11 is a flowchart showing an example of the procedure of the control information update process. First, the CPU 106 updates the control information in the memory 105 (step S1100). Next, the CPU 106 determines whether non-volatility is required (step S1101). When non-volatility is necessary, the CPU 106 performs a log creation process (step S1102), while when it is not necessary, the CPU 106 skips the log creation process. The content of the log creation process will be described later. With this, the control information update process is completed.
[0090] FIG. 12 is a flowchart showing an example of the procedure of the cache data update process. First, the CPU 106 updates the block unit data 201 in the memory 105 (step S1200). Specifically, the CPU 106 writes, for example, the block unit data 201 received from the host 102 into the cache area already allocated in the memory 105. The written block unit data 201 is also referred to as cache data.
[0091] Next, the CPU 106 determines whether non-volatility is required (step S1201). When non-volatility is necessary, the CPU 106 executes step 1202, while when it is not necessary, the CPU 106 skips the subsequent steps and ends the cache data update process.
[0092] Step 1202 is a log creation process. This is a process of creating a cache data log related to the updated cache data, and this process will be described later.
[0093] Next, the CPU 106 determines whether the update of the current cache data is an overwrite (step S1203). This is done by checking whether a cache data log regarding the update of the cache data within the address range included in the range of the currently updated cache area exists in the existing cache data log. If it exists, it is determined to be an overwrite.
[0094] Next, in the case of an overwrite, the CPU 106 invalidates the log of the same address written in the log header table that manages the log header for overwrite (step S1204). On the other hand, in the case of not being an overwrite, this step S1204 is skipped. Finally, the CPU 106 updates this log header table. Thus, the cache data update process is completed.
[0095] FIG. 13 is a flowchart showing an example of the procedure of the log creation process. First, the CPU 106 secures a sequence number (step S1300). The sequence number is a number indicating an example of the order in which the logs are created, and the value is incremented by one each time a new log is created and attached to the new log.
[0096] Next, the CPU 106 secures a log buffer for temporarily storing the log to be created (step S1301). Specifically, when the object to be stored in the log buffer is the control information 200, an area of the size required to store the log of the object is allocated from the control information log buffer 402. When the object is cache data, an area of the size required to store the log of the object is allocated from the cache data log buffer 403.
[0097] Subsequently, the CPU 106 creates a log header which is the header of the log (step S1302). The log header includes the sequence number, the address of the target data on the memory 105, and the size of the target data. Next, the CPU 106 stores the log in the log buffer (step S1303).
[0098] Finally, the CPU 106 performs an activation process on the created log (step S1304). Specifically, for example, the CPU 106 enables the log by including a flag indicating the validity / invalidity of the log in the log header and turning on this flag. Thus, the log creation process ends.
[0099] FIG. 14 is a flowchart showing an example of the procedure of the log evacuation process. The log evacuation process is a process of writing the logs accumulated in the log buffer to the drive. The log evacuation process is called when there is a need to write the logs to the drive, as it was also called before the response to the host in step S806 of the writing process shown in FIG. 9.
[0100] First, the CPU 106 extracts the unevacuated logs, that is, the logs that have not yet been written to the drive, from the log buffer (step S1400). The CPU 106 refers to the cache data log storage destination management table 699 described above (step S1401) and checks whether the type 601 of the ID for identifying each cache data is "MBD" (see FIG. 6).
[0101] When the type 601 is "MBD", the CPU 106 determines that the storage destination drive is the memory backup drive 107. On the other hand, when the type 601 is not "MBD", the CPU 106 determines that the storage destination drive (corresponding to the "drive number" shown in the figure) is "UDD (drive 110)". Next, the CPU 106 writes the cache data log to the storage destination drive specified by the type 601 and the drive numbers 502 and 602 (step S1405).
[0102] After the writing is completed, the CPU 106 deletes the written log from the log buffer (step S1406). Thus, the log evacuation process is completed.
[0103] FIG. 15 is a flowchart showing an example of the procedure of the base image backup process. The base image backup process is a process of writing the stored content of the memory 105 to the drive. In the present embodiment, the base image backup process is used for protecting the control information 200 and is executed at a predetermined timing, for example, when the control information log on the drive has accumulated a certain amount or more.
[0104] First, the CPU 106 refers to the sequence number and stores the latest sequence number at the current time (step S1500). Next, the CPU 106 writes the entire control information 200 as a base image to the drive (step S1501). Note that this writing to the drive may be executed in multiple times. When this process is completed, the old control information log becomes unnecessary, so the CPU 106 invalidates all the control information logs before the sequence number saved (stored) in step 1500 (step S1502). Thus, the base image backup process is completed.
[0105] FIG. 16 is a flowchart showing an example of the procedure of the log recovery process. This log recovery process is executed during the process of starting up the storage device 100 after maintenance replacement work such as the controller 103 is performed after a system down due to both controllers 103 being blocked. Thereby, the control information 200 and the dirty data stored in the memory 105 before the system down can be recovered. This log recovery process is executed by the CPU 106 of the controller 103 of the storage device 100 before resuming acceptance of input / output.
[0106] First, the CPU 106 reads the base image from the base image area on the drive 110 and stores it in the area of the control information 200 on the memory 105 (step S1600).
[0107] The CPU 106 reads the control information log and the cache data log (hereinafter also abbreviated as "log"), and sorts them in ascending order according to the sequence number (step S1601). Next, the CPU 106 reflects the contents from the oldest log to the newest log in order in the control information log buffer 402 and the cache data log buffer 403 on the memory 105 according to the addresses indicated by the respective log headers (step S1602). Thus, the log recovery process is completed.
[0108] FIG. 17 is a flowchart showing an example of the procedure of the memory protection method switching process when one of the controllers 103 is blocked. This process is executed when the other controller 103 detects that the one controller 103 has been blocked due to a failure or the like, in order to switch the memory protection method from duplication between the controllers 103 to protection by logging.
[0109] First, the CPU 106 determines whether there is only one normal controller 103 remaining in the storage device 100 (i.e., only its own controller). If the number of the remaining controllers 103 is one (Yes), step 1701 is executed. On the other hand, if it is not one (No), all the remaining steps are skipped and this process is terminated.
[0110] Next, the CPU 106 sets the emergency message in progress flag indicating whether it is in the emergency message in progress state to ON (step S1701). While the emergency message in progress flag is ON, the CPU 106 increases the execution frequency of the message processing in order to store the dirty data in the drive 110 as soon as possible.
[0111] Next, the CPU 106 sets the log evacuation mode flag indicating whether it is in the mode of evacuating the log to ON (step S1702). Finally, the CPU 106 executes the base image evacuation process.
[0112] FIG. 18 is a flowchart showing an example of the procedure for switching the memory protection method during recovery by the controller 103. This process is executed to switch the memory protection method from protection by logging back to duplication between the controllers 103 when it is detected that another controller 103 has recovered due to maintenance replacement or the like and has become operational normally.
[0113] First, the CPU 106 performs control information duplication processing (step S1800). This control information duplication processing is a process of copying the control information 200 on the memory 105 to the memory 105 of the recovered controller 103. When all the control information 200 has been copied, the control information duplication processing ends.
[0114] Next, the CPU 106 performs dirty data duplication processing (step S1801). This duplication processing is a process of copying the dirty data on the memory 105 to the memory 105 of the other recovered controller 103. Also, each time a dirty data is copied, the cache control information regarding the dirty data is updated. When all the dirty data has been copied, it is completed. Note that, as in this embodiment, instead of copying the dirty data to another controller 103, a method of protecting the dirty data by destaging it to the drive 110 may be adopted.
[0115] Next, the CPU 106 sets the log backup mode flag to OFF (step S1802).
[0116] Finally, the CPU 106 performs log deletion processing (step S1803). This log deletion processing is a process of deleting all the written logs and the logs on the log buffers (control information log buffer 402, cache data log buffer 403). For example, all the logs stored in the drive 110 or the memory 105 may be overwritten with invalid values such as all zeros, or alternatively, all the logs may be invalidated by, for example, setting the valid flags of all the log headers to OFF.
[0117] As described above, the storage device 100 according to the present embodiment includes a plurality of types of drives 107 and 110 that store data non-volatiley, a plurality of controllers 103 that control reading and writing of data to the plurality of types of drives 107 and 110 in response to data input / output between the host 102 and the drives, and a volatile memory 105 (cache memory) in which data is temporarily stored. The controller 103 has a data protection step of executing a memory data protection function that generates a log including the data in the memory 105 and a header related to data update and writes the log to one of the plurality of types of drives 107 and 110. In the data protection step, the controller 103 selects a drive to store the log from among the plurality of types of drives 107 and 110 using the memory data protection function according to required performance.
[0118] By doing so, in order to select a drive to store the log according to required performance, it is possible to make the specifications of the drive the minimum necessary while ensuring sufficient performance. Therefore, it is possible to ensure high reliability while achieving a balance between cost and performance.
[0119] In the present embodiment, when one of the plurality of controllers 103 is blocked, the other unblocked controller 103 executes the memory data protection function. By doing so, the other unblocked controller 103 can more reliably store the log in a drive with the minimum necessary specifications.
[0120] The storage device 100 according to this embodiment includes a memory backup drive 107 as a storage area for storing the stored content of the memory 105 of one controller 103 when one controller 103 is blocked, and a drive 110 for user data in which data reading and writing are performed regardless of whether one controller 103 is blocked. By doing so, as multiple types of non-volatile drives, a memory backup drive 107 as a storage area for storing the stored content of the memory 105 of the controller 103 and a drive 110 for user data in which data reading and writing are performed regardless of whether one controller 103 is blocked are used. Therefore, there is no need to newly provide a dedicated drive, and the cost can be suppressed.
[0121] In this embodiment, when the other controller 103 is set (corresponding to the above-mentioned "performance priority") where it is prioritized to satisfy, as the above-mentioned necessary performance, for example, predetermined performance requirements regarding data reading and writing speeds, in the data protection step, a drive for storing a log is selected from among the multiple types of drives 107 and 110. By doing so, a drive that is performance-necessary and sufficient can be selected according to the setting.
[0122] In this embodiment, the drive 110 for user data includes at least one SSD (Solid State Drive), and the other controller 103 selects an SSD (Solid State Drive) as the drive for storing the log in the data protection step. By doing so, the log can be stored at a higher speed.
[0123] In this embodiment, data is composed of a plurality of block unit data obtained by dividing the data in block units. The cache data log includes the block unit data and a header related to the update of the block unit data. The other controller 103 selects, in the data protection step, a drive from among a plurality of types of drives 107 and 110 to store each block unit data and the header as the cache data log. By doing so, although the cache data log includes block unit data and thus has a large data volume, it is stored in a suitable drive from among a plurality of types of non-volatile drives according to whether or not it is a setting in which satisfying a predetermined performance requirement is prioritized.
[0124] The storage device 100 according to this embodiment includes a cache data log storage destination management table 699 as an example of a storage destination management table that manages information related to at least the type and capacity size of a plurality of types of drives. The other controller 103 refers to the cache data log storage destination management table 699 in the data protection step and selects a drive in which to store the cache data log. By doing so, if the other controller 103 has previously set in the cache data log storage destination management table 699 a drive in which to store the cache data log, the other controller 103 can refer to the cache data log storage destination management table 699 and store the cache data log in a suitable drive.
[0125] Note that the present invention is not limited to the above-described embodiment, and includes various modifications and equivalent configurations within the scope of the appended claims. For example, the above-described embodiment has been described in detail for easy understanding of the present invention, and the present invention is not necessarily limited to an embodiment including all the configurations described. Also, each of the elements described in parallel in this embodiment may be in a mode in which at least one of the elements is connected in series to another element.
Industrial Applicability
[0126] The present invention can be applied to a storage device related to a technique of writing data in a memory and its updated content as a data log to a drive when performing input / output processing of data with a host.
Explanation of Signs
[0127] 100……Storage device, 102……Host, 103……Controller, 105……Memory, 106……CPU, 107……Memory backup drive, 110……Drive
Claims
1. A plurality of types of drives for non-volatile storage of data, a plurality of controllers for controlling reading and writing of the data between the host and the drives, and a volatile memory in which the data is temporarily stored, wherein the controller has a memory data protection function of generating a cache data log including data in the memory and a header related to update content of the data, and writing the cache data log to any one of the plurality of types of drives, and the controller selects a drive to store the cache data log from among the plurality of types of drives using the memory data protection function according to required performance, and stores the cache data log in the selected drive A storage device characterized by the above.
2. When one of the plurality of controllers is blocked, the other unblocked controller executes the memory data protection function The storage device according to claim 1, characterized by the above.
3. The plurality of types of drives include a memory backup drive as a storage area for backing up stored content in the memory of the one controller when the one controller is blocked, and a drive for user data in which reading and writing of the data are performed regardless of whether the one controller is blocked The storage device according to claim 2, characterized by the above.
4. The other controller selects the memory backup drive from among the plurality of types of drives when the required performance is set such that satisfying a predetermined performance requirement regarding data reading and writing speed is prioritized, while selects the drive for user data from among the plurality of types of drives when the required performance is set such that satisfying the predetermined performance requirement is not prioritized The storage device according to claim 3, characterized by the above.
5. The plurality of types of drives include at least one SSD (Solid State Drive), and the other controller selects the SSD from among the plurality of types of drives when the required performance is set such that satisfying a predetermined performance requirement is prioritized The storage device according to claim 2, characterized by the above.
6. The data It is composed of a plurality of block unit data obtained by dividing the data in block units, The cache data log is, The block unit data, And a header regarding the update content of the block unit data, Including, The other controller, Selects a drive that stores each of the block unit data and each of the headers as the cache data log from among the plurality of types of drives The storage device according to claim 2, characterized in that.
7. It is provided with a storage destination management table that manages information regarding at least the type and capacity size of the plurality of types of drives, The other controller, Refers to the storage destination management table and selects a drive to store the cache data log The storage device according to claim 2, characterized in that.
8. A storage destination control method for a storage device including a plurality of types of drives that store data non-volatiley, a plurality of controllers that control reading and writing of the data between the host, and a volatile memory in which the data is temporarily stored, The data protection step of executing a memory data protection function in which the controller generates a cache data log including the data in the memory and a header regarding the update content of the data and writes it to one of the plurality of types of drives, In the data protection step, The controller selects a drive to store the cache data log from among the plurality of types of drives using the memory data protection function according to the required performance, and stores the cache data log in the selected drive The storage destination control method characterized by that.
9. When one of the plurality of controllers is blocked, in the data protection step, the other controller that is not blocked executes the memory data protection function The storage destination control method according to claim 8, characterized in that.
10. The plurality of types of drives are, A memory backup drive as a storage area where the stored content of the memory of the one controller is backed up when the one controller is blocked, A drive for user data in which reading and writing of the data are performed regardless of whether the one controller is blocked or not, The storage destination control method according to claim 9, characterized in that it includes.
11. In the data protection step, the other controller selects the memory backup drive from among the plurality of types of drives when, as the required performance, a setting is prioritized that satisfies a predetermined performance requirement regarding the data read and write speeds, while selecting the drive for user data from among the plurality of types of drives when, as the required performance, a setting is not prioritized that satisfies the predetermined performance requirement The storage destination control method according to claim 10, characterized by the above.
12. The plurality of types of drives include at least one SSD (Solid State Drive), and in the data protection step, the other controller selects the SSD from among the plurality of types of drives when, as the required performance, a setting is prioritized that satisfies a predetermined performance requirement The storage destination control method according to claim 9, characterized by the above.
13. The data is composed of a plurality of block unit data obtained by dividing the data into blocks, and the cache data log includes each of the block unit data, and each header regarding the update content of each of the block unit data, and in the data protection step, the other controller selects a drive from among the plurality of types of drives to store each of the block unit data and each of the headers as the cache data log The storage destination control method according to claim 9, characterized by the above.
14. is provided with a storage destination management table that manages information regarding at least the type and capacity size of the plurality of types of drives, and in the data protection step, the other controller refers to the storage destination management table and selects a drive to store the cache data log The storage destination control method according to claim 9, characterized by the above.
Citation Information
Patent Citations
Data storage system and storage control method including storing a log related to the stored data
US11609698B1