Storage system and data processing method
By using a snapshot virtual device and compressing data within the storage system, the storage system addresses the throughput issues caused by the Redirect on Write method, achieving efficient resource management and improved performance.
Patent Information
- Application Number
- JP2023200544
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-28
- Publication Date
- 2025-06-09
- Estimated Expiration
- 2043-11-28
AI Technical Summary
In storage systems utilizing snapshot technology, the Redirect on Write (RoW) method increases the load on the storage controller and decreases throughput due to the need for frequent updates and references to mapping information during read/write operations.
The storage system employs a snapshot virtual device as a storage destination for primary and snapshot volumes, compresses data stored in this virtual device, and allocates new areas for small-size data writes, optimizing data storage and reducing mapping information updates.
This approach enhances throughput by efficiently managing virtual device resources, reducing the time required for address mapping changes, and improving garbage collection performance.
Smart Images

Figure 2025086518000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a storage system and a data processing method.
Background Art
[0002] In recent years, the need for data utilization has increased, and the opportunities for data replication have also increased. Along with this, in a storage system, the Snapshot function has become increasingly important. Conventionally, there is a Redirect on Write (RoW) method as a typical means of realizing Snapshot. The RoW method has an advantage that since there is no copy of data or meta information, the impact on I / O performance at the time of creating a Snapshot is small. The RoW method is widely adopted in All Flash Array (AFA) devices. The RoW method is a method of appending data. Appending data means that when writing data to a storage system, without overwriting the data stored before writing, storing the data to be written in a new area, and rewriting the meta information so as to refer to the data stored in the new area.
[0003] In such data management to which deduplication technology and snapshot technology are applied, when the address in the virtual device of data is changed due to garbage collection or other reasons, and the address is a reference destination of a plurality of different addresses in a plurality of snapshot families, it is necessary to change the reference destination address for each of the plurality of addresses. For this reason, it takes a long time to change the address mapping, and as a result, the time for the entire process involving the change of the address mapping becomes long.
[0004] In Patent Document 1, a method is presented in which data shared by a plurality of VOLs by deduplication and snapshot is allocated to a virtual space different from the storage destination of the single data referred to only from one VOL, thereby reducing the number of references to duplicate data and improving the efficiency of garbage collection processing.
Prior Art Documents
Patent Document
[0005]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] However, when adding a virtual device space as in Patent Document 1, there is a problem that the load on the storage controller increases and the throughput decreases due to an increase in the mapping information to be referenced and updated during read / write I / O processing. Therefore, an object of the present invention is to efficiently utilize the resources of virtual devices and achieve both garbage collection performance and high throughput.
Means for Solving the Problems
[0007] To achieve the above object, one of the typical storage systems of the present invention includes a storage device and a processor that accesses the storage device. The processor manages a primary volume that is the target of read / write by a host and a snapshot volume generated from the primary volume as a snapshot family. The processor uses a snapshot virtual device, which is a logical address space associated with the snapshot family, as a storage destination for the data of the primary volume and the snapshot volume, stores the data stored in the snapshot virtual device in a compressed form in a compressed virtual device, and stores the data stored in the compressed virtual device in the storage device. When the processor receives a write request from the host, it switches between an overwrite process of overwriting an area on the snapshot virtual device that has already been allocated to large-size data according to the size of the write destination address range, and a new allocation process of allocating a new area on the snapshot virtual device to the write destination address range for small-size data, and compresses and stores the plurality of small-size data stored in the new area in the compressed virtual device in a lump. Also, one of the representative data processing methods of the present invention is a data processing method for a storage system including a storage device and a processor that accesses the storage device, wherein the processor manages a primary volume that is a target of read / write by a host and a snapshot volume generated from the primary volume as a snapshot family, and the processor uses a snapshot virtual device, which is a logical address space associated with the snapshot family, as a storage destination for data of the primary volume and the snapshot volume, the processor compresses the data stored in the snapshot virtual device and stores it in a compressed virtual device, and the processor stores the data stored in the compressed virtual device in the storage device. When the processor receives a write request from the host, it switches between an overwrite process of overwriting an area on the snapshot virtual device that has already been allocated to large-sized data according to the size of the write destination address range, and a new allocation process of allocating a new area on the snapshot virtual device to the write destination address range for small-sized data, and compresses and stores the plurality of small-sized data stored in the new area together in the compressed virtual device.
Advantages of the Invention
[0008] According to the present invention, high throughput can be achieved by efficiently using the resources of the virtual device. Problems, configurations, and effects other than those described above will be clarified by the description of the following embodiments.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33
Figure 34
Figure 35
Figure 36
Modes for Carrying Out the Invention
[0010] Hereinafter, examples will be described with reference to the drawings.
Examples
[0011] Figure 1 shows the hardware configuration of a computer system. The computer system 100 includes a storage system 201, a server system 202, and a management system 203. The storage system 201 and the server system 202 are connected via a storage network 204 using FC (Fiber Channel) or the like. The storage system 201 and the management system 203 are connected via a management network 205 using IP (Internet Protocol) or the like. Note that the storage network 204 and the management network 205 may be the same communication network.
[0012] The storage system 201 includes a plurality of storage controllers 210 and a plurality of SSDs 220. A plurality of SSDs 220 are connected to the storage controller 210. The plurality of SSDs 220 are an example of a persistent storage device. A pool 13 is configured based on the plurality of SSDs 220. The data stored in page 14 of pool 13 is stored in one or more SSDs 220.
[0013] The storage controller 210 includes a CPU 211, a memory 212, a backend interface 213, a frontend interface 214, and a management interface 215.
[0014] The CPU 211 executes the program stored in the memory 212. The memory 212 stores the program executed by the CPU 211 and the data used by the CPU 211, etc. The memory may be duplicated by the combination of the memory 212 and the CPU 211.
[0015] The backend interface 213, the frontend interface 214, and the management interface 215 are an example of an interface device. The back-end interface 213 is a communication interface device that mediates the data exchange between the SSD 220 and the storage controller 210. A plurality of SSDs 220 are connected to the back-end interface 213. The front-end interface 214 is a communication interface device that mediates the data exchange between the server system 202 and the storage controller 210. The server system 202 is connected to the front-end interface 214 via the storage network 204. The management interface 215 is a communication interface device that mediates the data exchange between the management system 203 and the storage controller 210. The management system 203 is connected to the management interface 215 via the management network 205.
[0016] The server system 202 is configured to include one or more host devices. The server system 202 transmits an I / O request (write request or read request) specifying an I / O destination to the storage controller 210. The I / O destination is, for example, a logical volume number such as a LUN (Logical Unit Number), a logical address such as an LBA (Logical Block Address), and the like. The management system 203 is configured to include one or more management devices. The management system 203 manages the storage system 201.
[0017] Figure 2 shows an overview of the memory control of the storage system. In Figure 2, data denoted by capital letters (Data A, B, C, …) are block data, and data denoted by lowercase letters (data a, b, c, …) are sub-block data. The block data may be data in block units. The block may be a fixed-length logical storage area (logical address range). The sub-block data is compressed data of the block data, and the sub-block group (one or more sub-blocks) is the data stored at the storage destination. The sub-block may be a logical storage area smaller in size than the block. For example, an integer multiple of the sub-block may be the block.
[0018] A storage system having a storage device and a processor includes an SS-Family (Snapshot Family) 9, an SS-VDEV (Snapshot Virtual Device) 11S, a Dedup-VDEV (Deduplication Virtual Device) 11D, a CR-VDEV (Compression Append Virtual Device) 11C, and a pool 13.
[0019] The SS-Family 9 is a VOL group including a PVOL 10P and an SVOL 10S which is a snapshot of the PVOL 10P. The SS-VDEV 11S is a virtual device as a logical address space, and is used as the storage destination of data whose storage destination is any VOL 10 in the SS-Family 9. The Dedup-VDEV 11D is a virtual device as a logical address space different from the SS-VDEV 11S, and is used as the storage destination of duplicate data of two or more SS-VDEVs 11S.
[0020] The CR-VDEV 11C is a virtual device as a logical address space different from the SS-VDEV 11S and the Dedup-VDEV 11D, and is used as the storage destination of compressed data. Each of the plurality of CR-VDEV11C is associated with either SS-VDEV11S or Dedup-VDEV11D, and is not associated with both VDEV11S and 11D. That is, each CR-VDEV11C serves as the storage destination of data for which the VDEV (virtual device) corresponding to the CR-VDEV11C is the storage destination, and does not serve as the storage destination of data for which the VDEV not corresponding to the CR-VDEV11C is the storage destination. The compressed data for which CR-VDEV11C is the storage destination is stored in pool 13.
[0021] Pool 13 is a logical address space based on at least a part of a storage device (e.g., a persistent storage device) of the storage system. Pool 13 may be based on at least a part of an external storage device (e.g., a persistent storage device) of the storage system instead of or in addition to at least a part of the storage device of the storage system. Pool 13 has a plurality of pages 14 which are a plurality of logical areas. The compressed data for which CR-VDEV11C is the storage destination is stored in page 14 in pool 13. The mapping between the address in CR-VDEV11C and the address in pool 13 is 1:1. Pool 13 is composed of one or more pool VOLs.
[0022] According to the example shown in FIG. 2, the following storage control is performed. The processor creates SVOL10S0 as a snapshot of PVOL10P0, thereby creating SS-Family9-0 with PVOL10P0 as the root VOL. Further, the processor creates SVOL10S1 as a snapshot of PVOL10P1, thereby creating SS-Family9-1 with PVOL10P1 as the root VOL. According to FIG. 2, as examples of a plurality of SS-Family, there are SS-Family9-0 and 9-1.
[0023] The storage system has one or more SS-VDEV11S for each of the plurality of SS-Family9. For each SS-Family9, for the data whose destination is any VOL10 in the SS-Family9, among the plurality of SS-VDEV11S, the SS-VDEV11S corresponding to the SS-Family9 is set as the destination. Taking SS-Family9-0 as an example, specifically, for example, it is as follows. · For the data A whose destination is SVOL10S0 of SS-Family9-0, the processor sets SS-VDEV11S0 as the destination. The processor maps the address corresponding to data A among SVOL10S0 to the address corresponding to data A among the SS-VDEV11S0 corresponding to SS-Family9-0. · When there is the same data B in a plurality of VOLs (PVOL10P0 and SVOL10S0) of SS-Family9-0, the processor maps the plurality of addresses of the same data B (the address in PVOL10P0 and the address in SVOL10S0) among the plurality of VOL10 to the address of SS-VDEV11S0 of SS-Family9-0 (the address corresponding to data B).
[0024] For each of SS-Family9-0 and 9-1 (an example of two or more SS-Family9), the storage destination of non-duplicate data is the CR-VDEV11C corresponding to the SS-Family9, and the storage destination of duplicate data is the Dedup-VDEV11D. That is, since data C overlaps in SS-VDEV11S0 and 11S1 of SS-Family9-0 and 9-1 (an example of two or more SS-VDEVs), the processor maps the two addresses of the overlapping data C among SS-VDEV11S0 and 11S1 to the address corresponding to the overlapping data C in Dedup-VDEV11D. Then, the processor compresses the overlapping data C and sets the storage destination of the compressed data c to CR-VDEV11CC corresponding to Dedup-VDEV11D. That is, the processor maps the address (block address) of the overlapping data C in Dedup-VDEV11D to the address (sub-block address) of the compressed data c in CR-VDEV11CC. Also, the processor allocates page 14B to CR-VDEV11CC and stores the compressed data c in page 14B. The address of the compressed data c in CR-VDEV11CC is mapped to the address in page 14B of pool 13.
[0025] On the other hand, since data A in SS-VDEV11S0 does not overlap with the data in other SS-VDEV11S1, the processor compresses the non-overlapping data A and sets the compressed data a to CR-VDEV11C0 corresponding to SS-VDEV11S0. That is, the processor maps the address (block address) of the non-overlapping data A in SS-VDEV11S0 to the address (sub-block address) of the compressed data a in CR-VDEV11CC. Also, the processor allocates page 14A to CR-VDEV11C0 and stores the compressed data a in page 14A. The address of the compressed data a in CR-VDEV11C0 is mapped to the address in page 14A of pool 13.
[0026] CR-VDEV11C is an append-type VDEV. Therefore, the processor updates the address mapping when the CR-VDEV11C corresponding to SS-VDEV11S is set as the storage destination of the update data and when the CR-VDEV11C corresponding to Dedup-VDEV11D is set as the storage destination of the update data. Specifically, the processor performs, for example, the following storage control. · When the data to be stored for CR-VDEV11C0 is the updated data a' of the compressed data a, the processor sets the storage destination of the updated data a' as the free address in CR-VDEV11C0 and invalidates the address of the compressed data a before the update. The processor remaps the address mapped to the address of the compressed data a and being an address in SS-VDEV11S0 to the storage destination address of the updated data a' in CR-VDEV11C0 instead of the address of the compressed data a in CR-VDEV11C0. Also, the processor remaps the address mapped to the address of the compressed data a and being an address in Page 14A to the storage destination address of the updated data a' in CR-VDEV11C0 instead of the address of the compressed data a in CR-VDEV11C0. · When the data to be stored for CR-VDEV11CC is the updated data c' of the compressed data c, the processor sets the storage destination of the updated data c' as the free address in CR-VDEV11CC and invalidates the address of the compressed data c before the update. The processor remaps the address mapped to the address of the compressed data c and being an address in Dedup-VDEV11D to the storage destination address of the updated data c' in CR-VDEV11CC instead of the address of the compressed data c in CR-VDEV11CC. Also, the processor remaps the address mapped to the address of the compressed data c and being an address in Page 14B to the storage destination address of the updated data c' in CR-VDEV11CC instead of the address of the compressed data a in CR-VDEV11CC. Generally, the above-described mapping is managed in units of small sizes such as 4KB or 8KB in order to increase the data volume reduction effect by snapshot or deduplication.
[0027] CR-VDEV11C is the appendable VDEV as described above, and garbage collection is performed. That is, the processor can make the valid addresses (the addresses of the latest data) continuous and the addresses of the free areas continuous by performing garbage collection on CR-VDEV11C.
[0028] According to the example shown in FIG. 2, in addition to SS-VDEV11S which is the storage destination of data in SS-Family9, Dedup-VDEV11D which is the storage destination of duplicate data in two or more SS-Family9 is prepared. Therefore, even if the address of the duplicate data C in Dedup-VDEV11D changes, the address mapping to be changed only requires two mappings (the mappings for each of the two addresses in SS-VDEV11S0 and 11S1). On the other hand, in a comparative example, the storage destination of data in SS-Family9 and the storage destination of duplicate data in two or more SS-Family9 are the same VDEV. In this case, the address mapping to be changed for the duplicate data C requires four mappings (the mappings for each of the four addresses in VOL10P0, 10S0, 10P1, and 10S1). In this embodiment, it is expected that the address mapping can be changed in a shorter time than in the comparative example.
[0029] Separate from Dedup-VDEV11D, CR-VDEV11CC is prepared for Dedup-VDEV11D. Therefore, even if the address of the compressed data c in CR-VDEV11CC changes, the address mapping to be changed can be handled with a single mapping (the mapping for one address in Dedup-VDEV11D). On the other hand, in a comparative example, the storage destination of the compressed data of the data in SS-Family9 and the storage destinations of the compressed data of the duplicate data in two or more SS-Family9s are the same VDEV. In this case, the address mapping to be changed for the compressed data c of the duplicate data C is four mappings (the mappings for each of the four addresses in VOL10P0, 10S0, 10P1, and 10S1). In this embodiment, it is expected that the address mapping can be changed in a shorter time than in the comparative example.
[0030] The change of the address of the compressed data c is performed in the garbage collection of CR-VDEV11CC corresponding to Dedup-VDEV11D. For example, in the garbage collection of CR-VDEV11CC, the processor changes the address of the updated data c' (the updated data of the compressed data c) in CR-VDEV11CC, and maps the address that was mapped to the address before the change and is an address in Dedup-VDEV11D to the address after the change in CR-VDEV11CC. Since it is expected to change the address mapping for the compressed data of the duplicate data in a short time, it is expected to perform garbage collection in a short time. Note that the garbage collection of CR-VDEV11C corresponding to SS-VDEV11S includes, for example, the following processing. That is, the processor changes the address of the updated data a' (the updated data of the compressed data a) in CR-VDEV11C0, and maps the address that was mapped to the address before the change and is an address in SS-VDEV11S0 to the address after the change in CR-VDEV11C0.
[0031] For at least one CR-VDEV11C, an appendable VDEV in which uncompressed data is stored instead of CR-VDEV11C may be adopted. However, in this embodiment, CR-VDEV11C is adopted as the appendable VDEV. Therefore, the data finally stored in the storage device is compressed data, and thus the consumed storage capacity can be reduced.
[0032] FIG. 3 shows an overview of the management of the mapping between the addresses in SS-Family9 and the addresses in SS-VDEV11S. In the drawings, "GX" (X is an integer greater than or equal to 0) means generation X. Also, FIG. 3 takes SS-Family9-0 and SS-VDEV11S0 as examples.
[0033] The processor can manage the mapping between the addresses in VOL10 of SS-Family9-0 and the addresses in SS-VDEV11S0 using the meta information. The meta information includes Dir-Info (directory information) and SS-Mapping-Info (snapshot mapping information). The processor manages the data of PVOL10P0 and SVOL10S0 by associating Dir-Info with SS-Mapping-Info. For the data with VOL10 as the storage destination, Dir-Info has information representing the source address (the address in VOL10), and the corresponding SS-Mapping-Info for the data has information representing the destination address (the address in SS-VDEV11S0).
[0034] Furthermore, the processor manages the time series of PVOL10P0 and SVOL10S0 with the generation information associated with Dir-Info, and for each data with SS-VDEV11S0 as the storage destination, manages by associating the generation information indicating the generation in which the data was created with SS-Mapping-Info. In addition, the processor manages the latest generation information at that time as the latest generation.
[0035] Before the snapshot is taken, assume that there are data A0, B0, and C0 whose storage destination is PVOL10P0. Also assume that the latest generation is "0". The Dir-Info associated with PVOL10P0 has "0" associated with it as the generation # (number representing the generation), and includes reference information indicating the reference destinations of all data A0, B0, and C0 of PVOL10P0. Hereinafter, when the generation # associated with Dir-Info is "X", it can be stated that Dir-Info is the generation X.
[0036] SS-VDEV11S0 is the storage destination of data A0, B0, and C0, and SS-Mapping-Info is associated with each of data A0, B0, and C0. Also, "0" is associated with each SS-Mapping-Info as the generation #. When the generation # associated with SS-Mapping-Info represents "X", it can be stated that the data corresponding to SS-Mapping-Info is the data of generation X.
[0037] In the state before the snapshot is taken, for each of data A0, B0, and C0, the information in Dir-Info refers to the SS-Mapping-Info corresponding to that data. By associating Dir-Info and SS-Mapping-Info in this way, PVOL10P0 and SS-VDEV11S0 can be associated, and data processing for PVOL10P0 can be realized.
[0038] To take a snapshot, the processor sets the copy of Dir-Info as the read-only Dir-Info of SVOL10S0. Then, the processor increments the generation of the Dir-Info of PVOL10P0 and also increments the latest generation. As a result, for each of data A0, B0, and C0, SS-Mapping-Info is referred to from both the Dir-Info of generation 0 and the Dir-Info of generation 1.
[0039] In this way, a snapshot can be created by replicating Dir-Info, and the snapshot can be created without increasing the data on SS-VDEV11S0 or SS-Mapping-Info.
[0040] Here, when a snapshot is acquired, the snapshot (SVOL10S0) in which writing is prohibited and the data is fixed at the acquisition time becomes generation 0, and PVOL10P0 in which data can be written even after acquisition becomes generation 1. Generation 0 is "a generation one generation older in the direct line" with respect to generation 1, and is referred to as "parent" for convenience. Similarly, generation 1 is "a generation one generation newer in the direct line" with respect to generation 0, and is referred to as "child" for convenience. The storage system manages the parent-child relationship of generations as the Dir-Info generation management tree 70. Also, the generation # of Dir-Info is the same as the generation # of VOL10 corresponding to the Dir-Info. Also, the generation # of SS-Mapping-Info is the oldest generation # among the generation #s of one or more Dir-Info that refer to the SS-Mapping-Info.
[0041] FIG. 4 shows an overview of the management of the mapping between the address in SS-VDEV11S and the address in Dedup-VDEV11D, the management of the mapping between the address in SS-VDEV11S and the address in CR-VDEV11C, and the management of the mapping between the address in Dedup-VDEV11D and the address in CR-VDEV11C. FIG. 4 takes SS-VDEV11S0, CR-VDEV11C0, and 11CC as examples.
[0042] The processor can manage the mapping between the address in SS-VDEV11S0 and the address in Dedup-VDEV11D, the mapping between the address in SS-VDEV11S0 and the address in CR-VDEV11C0, and the mapping between the address in Dedup-VDEV11D and the address in CR-VDEV11CC by using the meta information. As described above, the meta information includes Dir-Info and CR-Mapping-Info. The processor manages the data of SS-VDEV11S0 and Dedup-VDEV11D by associating Dir-Info with CR-Mapping-Info. For the data with SS-VDEV11S0 as the storage destination, Dir-Info has information representing the source address (the address in SS-VDEV11S0), and the corresponding CR-Mapping-Info for the data has information representing the destination address (the address in CR-VDEV11C0 or the address in Dedup-VDEV11D). For the data with Dedup-VDEV11D as the storage destination, Dir-Info has information representing the source address (the address in Dedup-VDEV11D), and the corresponding CR-Mapping-Info for the data has information representing the destination address (the address in CR-VDEV11CC). By referring to the compression allocation information, the processor can identify the address in SS-VDEV11S or Dedup-VDEV11D from the address in CR-VDEV11C. Although omitted in the drawings, the storage system 201 holds the compression allocation management table 1011 on the memory 212 as information on the reverse mapping from the address in the CR-VDEV to the address in the SS-VDEV or Dedup-VDEV in order to move the valid data on the CR-VDEV and secure a continuous free area.
[0043] Figure 5 shows the configuration of the memory 212. The memory 212 has a control information section 901 in which control information (which may also be referred to as management information) is stored, a program section 902 in which programs are stored, and a cache section 903 in which data is temporarily stored.
[0044] FIG. 6 shows the information stored in the control information section 901. The control information section 901 stores an ownership management table 1001, a CR-VDEV management table 1002, a snapshot management table 1003, a VOL-Dir management table 1004, a latest generation table 1005, a recovery management table 1006, a generation management tree table 1007, a snapshot allocation management table 1008, a Dir management table 1009, an SS-Mapping management table 1010, a compression allocation management table 1011, a CR-Mapping management table 1012, a Dedup-Dir management table 1013, a Dedup allocation management table 1014, a Pool-Mapping management table 1015, and a Pool allocation management table 1016.
[0045] FIG. 7 shows the programs stored in the program section 902. The program section 902 stores a snapshot acquisition program 1101, a snapshot restore program 1102, a snapshot deletion program 1103, an asynchronous recovery program 1104, a read / write program 1105, a snapshot append program 1106, a Dedup append program 1107, a compression append program 1108, a destage program 1109, a GC (garbage collection) program 1110, a CPU determination program 1111, an ownership transfer program 1112, and a Snapshot allocation determination program 1113.
[0046] FIG. 8 shows the configuration of the ownership management table 1001. The owner right management table 1001 manages the owner rights of VOL10 or VDEV11. For example, the owner right management table 1001 has entries for each of VOL10 and VDEV11. The entry has information such as VOL# / VDEV#1201 and owner CPU#1202.
[0047] VOL# / VDEV#1201 represents the identification number of VOL10 or VDEV11. Owner CPU#1202 represents the identification number of the CPU that is the owner CPU of VOL10 or VDEV11 (the CPU that has the owner right of VOL10 or VDEV11). Note that the owner CPU may be assigned in units of CPU groups or in units of storage controllers 210 instead of being assigned in units of CPU211.
[0048] Figure 9 shows the configuration of the CR-VDEV management table 1002. The CR-VDEV management table 1002 represents the CR-VDEV11C associated with SS-VDEV11S or Dedup-VDEV11D. For example, the CR-VDEV management table 1002 has entries for each of SS-VDEV11S and Dedup-VDEV11D. The entry has information such as VDEV#1301 and CR-VDEV#1302. VDEV#1301 represents the identification number of SS-VDEV11S or Dedup-VDEV11D. CR-VDEV#1302 represents the identification number of CR-VDEV11C.
[0049] Figure 10 shows the configuration of the snapshot management table 1003. The snapshot management table 1003 exists for each PVOL10P (each SS-Family9). The snapshot management table 1003 represents the acquisition time of each snapshot (SVOL10S). For example, the snapshot management table 1003 has entries for each SVOL10S. The entry has information such as PVOL#1401, SVOL#1402, and acquisition time 1403. PVOL#1401 represents the identification number of PVOL10P. SVOL#1402 represents the identification number of SVOL10S. Acquisition time 1403 represents the acquisition time of SVOL10S.
[0050] Figure 11 shows the configuration of the VOL-Dir management table 1004. The VOL-Dir management table 1004 represents the correspondence between VOL and Dir-Info. For example, the VOL-Dir management table 1004 has an entry for each VOL10. The entry has information such as VOL#1501, Root-VOL#1502, and Dir-Info#1503.
[0051] VOL#1501 represents the identification number of PVOL10P or SVOL10S. Root-VOL#1502 represents the identification number of Root-VOL. If VOL10 is PVOL10P, then Root-VOL is the corresponding PVOL10P. If VOL10 is SVOL10S, then Root-VOL is the PVOL10P corresponding to the SVOL10S. Dir-Info#1503 represents the identification number of the Dir-Info corresponding to VOL10.
[0052] Figure 12 shows the configuration of the latest generation table 1005. The latest generation table 1005 exists for each PVOL10P (for each SS-Family9) and represents the generation (generation#) of the corresponding PVOL10P.
[0053] Figure 13 shows the configuration of the recovery management table 1006. The recovery management table 1006 can be, for example, a bitmap and exists for each PVOL10P (for each SS-Family9), in other words, it exists for each Dir-Info generation management tree 70. The recovery management table 1006 has an entry for each Dir-Info. The entry has information such as Dir-Info#1701 and recovery request 1702. Dir-Info#1701 represents the identification number of Dir-Info. The recovery request 1702 indicates whether to request the recovery of Dir-Info. "1" means to request recovery, and "0" means not to request recovery.
[0054] Figure 14 shows the configuration of the generation management tree table 1007. The generation management tree table 1007 exists for each PVOL10P (for each SS-Family9), in other words, for each Dir-Info generation management tree 70. The generation management tree table 1007 has an entry for each Dir-Info. The entry has information such as Dir-Info#1801, generation#1802, Prev1803, and Next1804. Dir-Info#1801 represents the identification number of Dir-Info. Generation#1802 represents the generation of VOL10 corresponding to Dir-Info. Prev1803 represents the Dir-Info of the parent (one level above) of Dir-Info. Next1804 represents the Dir-Info of the child (one level below) of Dir-Info. The number of Next1804 may be the same as the number of child Dir-Info. In Figure 14, since there are two child Dir-Info, there are two Next1804 (Next-A1804A and Next-B1804B).
[0055] Figure 15 shows the configuration of the snapshot allocation management table 1008. The snapshot allocation management table 1008 exists for each SS-VDEV11S and represents the mapping from the address in SS-VDEV11S to the address in VOL10. The snapshot allocation management table 1008 has an entry for each address in SS-VDEV11S. The entry has information such as block address 1901, status 1902, target VOL#1903, and target address 1904.
[0056] Block address 1901 represents the address of a block in SS-VDEV11S. Status 1902 indicates whether the block is assigned to the address of any VOL ("1" means assigned, "0" means free). The destination VOL #1903 represents the identification number of VOL10 (PVOL10P or SVOL10S) that has the destination address of the block ("n / a" means unassigned). The destination address 1904 represents the destination address (block address) of the block ("n / a" means unassigned).
[0057] Figure 16 shows the configuration of the Dir management table 1009. The Dir management table 1009 exists for each Dir-Info and represents the reference Mapping-Info for each data (each block data). For example, the Dir management table 1009 has an entry for each address (block address). The entry has information such as the VOL / VDEV address 2001 and the reference Mapping-Info #2002. The VOL / VDEV address 2001 represents the address (block address) in VOL10 (PVOL10P or SVOL10S), or the address in VDEV11 (SS-VDEV11S or Dedup-VDEV11D). The reference Mapping-Info #2002 represents the identification number of the reference Mapping-Info.
[0058] Figure 17 shows the configuration of the SS-Mapping management table 1010. The SS-Mapping management table 1010 exists for each Dir-Info of VOL10. The SS-Mapping management table 1010 has an entry for each SS-Mapping-Info corresponding to the Dir-Info of VOL10. The entry has information such as the Mapping-Info #2101, the destination address 2102, the destination SS-VDEV #2103, and the generation #2104.
[0059] Mapping-Info #2101 represents the identification number of the SS-Mapping-Info. The reference address 2102 represents the address (the address in SS-VDEV11S) that the SS-Mapping-Info refers to. The reference destination SS-VDEV #2103 represents the identification number of the SS-VDEV11S that has the address referred to by the SS-Mapping-Info. The generation #2104 represents the generation of the data corresponding to the SS-Mapping-Info.
[0060] Figure 18 shows the configuration of the compression allocation management table 1011. The compression allocation management table 1011 exists for each CR-VDEV11C and has compression allocation information for each sub-block in the CR-VDEV11C. The compression allocation management table 1011 has an entry corresponding to the compression allocation information for each sub-block in the CR-VDEV11C. The entry has information such as the sub-block address 2201, the data length 2202, the status 2203, the start sub-block address 2204, the destination VDEV #2205, and the destination address 2206.
[0061] The sub-block address 2201 represents the address of the sub-block. The data length 2202 represents the number of sub-blocks (one or more sub-blocks) that make up the sub-block group in which the compressed data is stored (for example, "2" means that the compressed data exists in two sub-blocks). The status 2203 represents the status of the sub-block ("0" means free, "1" means allocated, and "2" means a target for GC (garbage collection)). The start sub-block address 2204 represents the address of the start sub-block of one or more sub-blocks (one or more sub-blocks in which the compressed data is stored) that include the sub-block. The destination VDEV #2205 represents the identification number of the VDEV11 (SS-VDEV11S or Dedup-VDEV11D) that has the destination block of the sub-block. The destination address 2206 represents the address of the destination block of the sub-block (the block address in SS-VDEV11S or Dedup-VDEV11D).
[0062] FIG. 19 shows the configuration of the CR-Mapping management table 1012. The CR-Mapping management table 1012 exists for each Dir-Info of the Dedup-VDEV 11D and for each Dir-Info of the SS-VDEV 11S. The CR-Mapping management table 1012 has entries for each CR-Mapping-Info corresponding to the Dir-Info of the Dedup-VDEV 11D and for each CR-Mapping-Info corresponding to the Dir-Info of the SS-VDEV 11S. The entry has information such as Mapping-Info #2301, reference address 2302, reference destination CR-VDEV #2303, and data length 2304.
[0063] Mapping-Info #2301 represents the identification number of the CR-Mapping-Info. The reference address 2302 represents the address (the address of the first sub-block among the sub-block groups) referred to by the CR-Mapping-Info. The reference destination CR-VDEV #2303 represents the identification number of the CR-VDEV 11C having the sub-block address referred to by the CR-Mapping-Info. The data length 2304 represents the number of blocks (blocks in the Dedup-VDEV 11D) referred to by the CR-Mapping-Info, or the number of sub-blocks constituting the sub-block group referred to by the CR-Mapping-Info.
[0064] FIG. 20 shows the configuration of the Dedup-Dir management table 1013. The Dedup-Dir management table 1013 exists for each Dedup-VDEV 11D and corresponds to the Dedup-Dir-Info. The Dedup-Dir management table 1013 has entries for each address in the Dedup-VDEV 11D. The entry has information such as the Dedup-VDEV address 2401 and the reference destination allocation information #2402. The Dedup-VDEV address 2401 represents the address (block address) in Dedup-VDEV11D. The reference destination allocation information #2402 represents the identification number of the Dedup allocation information of the reference destination.
[0065] Figure 21 shows the configuration of the Dedup allocation management table 1014. The Dedup allocation management table 1014 exists for each Dedup-VDEV11D (for each Dedup-Dir-Info), and represents the reverse mapping from the Dedup allocation information corresponding to the address in Dedup-VDEV11D to the address in SS-VDEV11S. The Dedup allocation management table 1014 has an entry for each Dedup allocation information. The entry has information such as allocation information #2501, destination SS-VDEV #2502, destination address 2503, and linked allocation information #2504. The allocation information #2501 represents the identification number of the Dedup allocation information. The destination SS-VDEV #2502 represents the identification number of the SS-VDEV11S having the address referred to by the Dedup allocation information. The destination address 2503 represents the address (block address in SS-VDEV11S) referred to by the Dedup allocation information. The linked allocation information #2504 represents the identification number of the Dedup allocation information linked to the Dedup allocation information.
[0066] According to Figure 21, the Dedup allocation information #「3」is linked to the Dedup allocation information #「1」, and there is no Dedup allocation information linked to the Dedup allocation information #「3」. Therefore, it can be seen that the duplicate data at the Dedup-VDEV address corresponding to the Dedup allocation information #「1」is the duplicate data in the SS-VDEV11S referred to by the Dedup allocation information #「1」and the SS-VDEV11S referred to by the Dedup allocation information #「3」. Since the number of duplicate data is indefinite, the Dedup allocation information is linked according to the number of duplicate data. When there is duplicate data in N SS-VDEV11S, N sequential Dedup allocation information is prepared.
[0067] FIG. 22 shows the configuration of the Pool-Mapping management table 1015. The Pool-Mapping management table 1015 exists for each CR-VDEV11C. The Pool-Mapping management table 1015 has an entry for each area in units of page size in CR-VDEV11C. The entry has information such as a VDEV address 2601 and a page #2602. The VDEV address 2601 represents the start address of an area in units of page size (for example, a plurality of blocks). The page #2602 represents the identification number of the allocated page 14 (for example, the address of page 14 in pool 13). When there are a plurality of pools 13, the page #2602 may include the identification number of the pool 13 having page 14.
[0068] FIG. 23 shows the configuration of the Pool allocation management table 1016. The Pool allocation management table 1016 exists for each pool 13 when there are, for example, a plurality of pools 13. The Pool allocation management table 1016 represents the correspondence between page 14 and the area in CR-VDEV11C. The Pool allocation management table 1016 has an entry for each page 14. The entry has information such as a page #2701, an RG #2702, a start address 2703, a status 2704, an allocation destination VDEV #2705, and an allocation destination address 2706.
[0069] The page #2701 represents the identification number of page 14. The RG #2702 represents the identification number of the RAID group (in this embodiment, a RAID group composed of two or more SSDs 220) on which page 14 is based. The start address 2703 represents the start address of page 14. The status 2704 represents the status of page 14 (“1” means allocated, and “0” means free). The allocation destination VDEV #2705 represents the identification number of the CR-VDEV11C to which page 14 is allocated (“n / a” means unallocated). The allocation destination address 2706 represents the allocation destination address (the address in CR-VDEV11C) of page 14 (“n / a” means unallocated).
[0070] Figure 24 shows the flow of the snapshot acquisition process. The snapshot acquisition process is executed by the snapshot acquisition program 1101 in response to a snapshot acquisition instruction from the management system 203 (or another system such as the server system 202). In the snapshot acquisition instruction, for example, the target PVOL10P is specified.
[0071] First, the snapshot acquisition program 1101 allocates the Dir management table 1009 that will be the copy destination and updates the VOL-Dir management table 1004 (S2401). The snapshot acquisition program 1101 increments the latest generation # (S2402) and updates the generation management tree table 1007 (Dir-Info generation management tree 70) (S2403). At this time, the snapshot acquisition program 1101 sets the latest generation # for the copy source and the generation # before increment for the copy destination.
[0072] The snapshot acquisition program 1101 determines whether there is cache dirty data for the target PVOL10P (S2404). The "cache dirty data" may be data that has not yet been written to the pool 13 among the data stored in the cache unit 903.
[0073] If the determination result of S2404 is true (S2404: Yes), the snapshot acquisition program 1101 causes the snapshot append program 1106 to execute the snapshot append process (S2405). If the determination result of S2404 is false (S2404: No), or after S2405, the snapshot acquisition program 1101 copies the Dir management table 1009 of the target PVOL10P to the Dir management table 1009 of the copy destination (S2406).
[0074] After that, the snapshot acquisition program 1101 updates the snapshot management table 1003 (S2407) and ends the process. In S2407, an entry having a PVOL#1401 representing the identification number of the target PVOL10P, an SVOL#1402 representing the identification number of the acquired snapshot (SVOL10S), and an acquisition time 1403 representing the acquisition time is added.
[0075] Figure 25 shows the flow of the snapshot restore process. The snapshot restore process is executed by the snapshot restore program 1102 in response to a restore instruction from the management system 203 (or another system such as the server system 202). In the restore instruction, for example, the source SVOL for restoration and the destination PVOL for restoration are specified.
[0076] First, the snapshot restore program 1102 allocates the Dir management table 1009 to be the destination and updates the VOL-Dir management table 1004 (S2501). The snapshot restore program 1102 increments the latest generation # (S2502) and updates the generation management tree table 1007 (Dir-Info generation management tree 70) (S2503). At this time, the snapshot restore program 1102 sets the generation # before incrementing as the copy source and sets the latest generation # as the copy destination.
[0077] The snapshot restore program 1102 purges the cache area of the destination PVOL (the area in the cache unit 903) (S2504). The snapshot restore program 1102 copies the Dir management table 1009 of the source SVOL for restoration to the Dir management table 1009 of the destination PVOL for restoration (S2505).
[0078] After that, the snapshot restore program 1102 registers the Dir-Info# of the old Dir-Info at the restore destination in the recovery management table 1006 (S2506) and ends the process. In S2506, the recovery request 1702 corresponding to the Dir-Info# is set to "1".
[0079] Figure 26 shows the flow of the snapshot deletion process. The snapshot deletion process is executed by the snapshot deletion program 1103 in response to a snapshot deletion instruction from the management system 203 (or another system such as the server system 202). In the restore instruction, for example, the target SVOL is specified.
[0080] First, the snapshot deletion program 1103 refers to the VOL-Dir management table 1004 and invalidates the Dir-Info (Dir-Info#1503) of the target SVOL (S2601). Then, the snapshot deletion program 1103 updates the snapshot management table 1003 (S2602), registers the old Dir-Info# of the target SVOL in the recovery management table 1006 (S2603), and ends the process. In S2603, the recovery request 1702 corresponding to the Dir-Info# is set to "1".
[0081] Figure 27 shows the flow of the asynchronous recovery process. The asynchronous recovery process is executed, for example, periodically by the asynchronous recovery program 1104. First, the asynchronous recovery program 1104 identifies the Dir-Info# to be recovered from the recovery management table 1006. The "Dir-Info# to be recovered" is the Dir-Info# for which the recovery request 1702 is "1". The asynchronous recovery program 1104 refers to the generation management tree table 1007, checks the entry of the Dir-Info# to be recovered, and does not select a Dir-Info having two or more children.
[0082] Thereafter, the asynchronous recovery program 1104 determines whether there is an unprocessed entry (S2702). The "unprocessed entry" mentioned here refers to an entry in the recovery management table 1006 where the recovery request 1702 is "1" and the asynchronous recovery process is unprocessed.
[0083] If the determination result of S2702 is true (S2702: Yes), the asynchronous recovery program 1104 determines a processing target entry (an entry including the recovery request 1702 "1") from one or more unprocessed entries (S2703), and identifies the reference destination Mapping-Info#2002 from the Dir management table 1009 corresponding to the target Dir-Info (the Dir-Info identified from Dir-Info#1701 in the processing target entry) (S2704).
[0084] The asynchronous recovery program 1104 refers to the generation management tree table 1007 and determines whether there is a Dir-Info of a child generation of the target Dir-Info (S2705). If the determination result of S2705 is true (S2705: Yes), the asynchronous recovery program 1104 identifies the reference destination Mapping-Info#2002 from the Dir management table 1009 corresponding to the Dir-Info of the child generation, and determines whether the reference destination Mapping-Info#2002 of the target Dir-Info matches the reference destination Mapping-Info#2002 of the Dir-Info of the child generation (S2706). If the determination result of S2706 is true (S2706: Yes), the process returns to S2702.
[0085] If the determination result of S2706 is false (S2706: No), or if the determination result of S2705 is false (S2705: No), the asynchronous recovery program 1104 determines whether the generation # of the Dir-Info of the parent generation of the target Dir-Info is older than the generation #2104 of the reference destination Mapping-Info of the target Dir-Info (see Figure 21) (S2707). If the determination result of S2707 is false (S2707: No), the process returns to S2702.
[0086] When the determination result of S2707 is true (S2707: Yes), the asynchronous recovery program 1104 initializes the target entry in the SS-Mapping management table 1010 and releases the target entry in the snapshot allocation management table 1008 (S2708). Then, the process returns to S2702. The release in S2708 corresponds to the release of blocks in the SS-VDEV.
[0087] When the determination result of S2702 is false (S2702: No), the asynchronous recovery program 1104 updates the recovery management table 1006 (S2709) and updates the generation management tree table 1007 (Dir-Info generation management tree 70) (S2710), and ends the process.
[0088] Figure 28 shows the flow of the Write process (front end). The Write process (front end) is executed by the read / write program 1105 when a write request from the server system 202 is received.
[0089] First, the read / write program 1105 determines whether the target data of the write request results in a cache hit (S2801). "Cache hit" means that a cache area corresponding to the write destination VOL address (the VOL address specified in the write request) of the target data has been secured. When the determination result of S2801 is false (S2801: No), the read / write program 1105 secures a cache area corresponding to the write destination VOL address of the target data from the cache unit 903 (S2802). Then, the process proceeds to S2806.
[0090] When the determination result of S2801 is true (S2801: Yes), the read / write program 1105 determines whether the cached data (data in the secured cache area) is dirty data (data not reflected (not written) in the pool 13) (S2803). When the determination result of S2803 is false (S2803: No), the process proceeds to S2806.
[0091] When the determination result of S2803 is true (S2803: Yes), the read / write program 1105 determines whether the WR (Write) generation # of the dirty data matches the generation # of the data targeted by the write request (S2804). The "WR generation #" is stored, for example, in the management information (not shown) of the cache data. Also, the generation # of the data targeted by the write request is obtained from the latest generation #403. S2804 is a process to prevent overwriting the snapshot data by updating the dirty data with the data targeted by the write request before the append process for the target data (dirty data) of the immediately preceding snapshot is completed.
[0092] When the determination result of S2804 is false (S2804: No), the read / write program 1105 causes the snapshot append program 1106 to execute the snapshot append process (S2805).
[0093] After S2802, or when the determination result of S2804 is true (S2804: Yes), the read / write program 1105 writes the data targeted by the write request to the cache area secured in S2802 or to the cache area obtained through S2805 (S2806). Then, the read / write program 1105 sets the WR generation # of the data written in S2806 to the latest generation # compared in S2804 (S2807) and returns a normal response (Good response) to the server system 202 (S2808).
[0094] Figure 29 shows the flow of the Write process (backend). The Write process (backend) is a process of writing unreflected data (dirty data) to the pool 13 when the unreflected data is on the cache unit 903. The Write process (backend) is performed synchronously or asynchronously with the Write process (frontend). The Write process (backend) is executed by the read / write program 1105.
[0095] The read / write program 1105 determines whether there is dirty data on the cache unit 903 (S2901). If the determination result in S2901 is true (S2901: Yes), the read / write program 1105 causes the snapshot append program 1106 to execute the snapshot allocation determination process (S2902).
[0096] Figure 30 shows the flow of the Snapshot allocation determination process. In the Snapshot allocation determination process, it is determined whether to newly allocate an area on the SS-VDEV11S for the area of the PVOL10P targeted by the write request from the host, or to overwrite an area already allocated on the SS-VDEV11S. The Snapshot allocation determination process is executed by the Snapshot allocation determination program 1113 called from the read / write program 1105.
[0097] The Snapshot allocation determination program 1113 first determines whether the generation # of the Dir-Info of the target VOL (the VOL where data is to be written) matches the generation # of the SS-Mapping-Info before append (S3001). If the determination result in S3001 is false (S3001: No), the Snapshot allocation determination program 1113 executes the Snapshot append process (S3007).
[0098] If the determination result in S3001 is true (S3001: Yes), the process proceeds to S3002. If the determination result in S3001 is true, the data in the range for which the write request is received from the host is in a state where it is not being referenced from other snapshots (hereinafter referred to as the single state). Therefore, by performing an overwrite at the same address on the SS-VDEV, it is possible to eliminate the need to update the snapshot allocation management table and the SS-Mapping management table.
[0099] Next, the Snapshot allocation determination program 1113 acquires the transfer length of the Write data requested from the host computer and determines whether the length is less than the threshold value (S3002). If the transfer length is less than or equal to the threshold value (S3002: Yes), the process proceeds to S3003. If the transfer length is greater than the threshold value (S3002: No), the process proceeds to S3006. The threshold value here is set to switch between overwriting without allocating a new area on the Snapshot space and proceeding to the Dedup append process, and allocating a new area on the Snapshot space and collectively updating the mapping information (Snapshot append process) to determine which is more advantageous in terms of performance. The specific threshold value is set according to the actual program implementation and performance characteristics.
[0100] Next, the Snapshot allocation determination program 1113 determines whether it is possible to allocate a new area as a continuous area for the Write target on the SS-VDEV (S3003). Specifically, it refers to the snapshot allocation management table 1008 and determines whether a continuous unallocated area can be secured. If it is determined that allocation is possible (S3003: Yes), the Snapshot allocation determination program 1113 executes the Snapshot append process (S3007). If it is determined that allocation is impossible (S3003: No), the process proceeds to S3004.
[0101] Next, the Snapshot allocation determination program 1113 determines whether it is possible to allocate a new VDEV on the SS-VDEV (S3004). If it is determined that allocation is possible (S3004: Yes), the process proceeds to S3005, and a new VDEV is allocated to the SS-VDEV. Next, the Snapshot allocation determination program 1113 executes the Snapshot append process (S3007) and allocates a new area on the SS-VDEV. If it is determined that allocation of a new VDEV is impossible (S3004: No), the process proceeds to S3006, and overwriting is performed on the SS-VDEV.
[0102] When proceeding to S3006, the Snapshot allocation determination program 1113 executes the Dedup append process. In this case, unlike when proceeding to the snapshot append process of S3007, the snapshot append process is skipped and overwriting is performed at the same address on the SS-VDEV. This can eliminate the need to update the snapshot allocation management table and the SS-Mapping management table.
[0103] Figure 31 shows the flow of the snapshot append process. The snapshot append process is a process of newly allocating an area on the SS-VDEV11S for the area of the PVOL10P targeted by the write request from the host. The snapshot append process is executed by the snapshot append program 1106 called from the snapshot acquisition program 1101 or the read / write program 1105.
[0104] The snapshot append program 1106 updates the snapshot allocation management table 1008 to secure a new area (block address with status 1902 being "0") in the target SS-VDEV11S (the SS-VDEV11S corresponding to the SS-Family9 including the target VOL (for example, the SVOL to be acquired or the VOL where data is written)) (S3101). Then, the snapshot append program 1106 causes the Dedup append program 1107 to execute the Dedup append process (S3102).
[0105] After that, the snapshot append program 1106 updates the SS-Mapping management table 1010 (S3103). In S3103, for example, the snapshot append program 1106 sets the latest generation # (the generation # represented by the latest generation table 1005) to the generation #2104 corresponding to the Mapping-Info# of the target SS-Mapping-Info. The "target SS-Mapping-Info" here is the SS-Mapping-Info corresponding to the data in the target VOL.
[0106] The snapshot append program 1106 updates the directory management table 1009 corresponding to the Dir-Info of the target SS-VDEV 11S (S3104). In S3104, the SS-Mapping-Info (information representing the destination address in the SS-VDEV) for the data to be written is associated with the address in the VOL10 of the data.
[0107] The snapshot append program 1106 refers to the generation management tree table 1007 (Dir-Info generation management tree 70) (S3105) and determines whether the generation # of the Dir-Info of the target VOL (the write destination VOL of the data) matches the generation # of the SS-Mapping-Info before the append (S3106).
[0108] If the determination result in S3106 is true (S3106: Yes), it means that after the data is stored in the area on the SS-VDEV pointed to by the SS-Mapping-Info before the append, no snapshot that shares the area has been created. That is, it can be determined that the area on the SS-VDEV pointed to by the SS-Mapping-Info before the append has become garbage. Therefore, the snapshot append program 1106 releases the target entry in the snapshot allocation management table 1008 (S3107) and ends the process. At this time, the mapping information to the CR-VDEV or Dedup-VDEV space corresponding to the area on the SS-VDEV pointed to by the SS-Mapping-Info before the append remains valid. That is, the target entry in the SS-Mapping management table 1010 is not initialized.
[0109] When the determination result of S3106 is false (S3106: No), it means that after data is stored in the area on the SS-VDEV pointed to by the SS-Mapping-Info before appending, a snapshot that shares the area is created and the generation # of DIR-Info is incremented. In this case, since the area on the SS-VDEV before appending does not become garbage, the SS-Mapping management table 1010 before appending is left as it is and the process ends.
[0110] Figure 32 shows the flow of the Dedup append process. The Dedup append process is executed by the Dedup append program 1107 called from the snapshot append program 1106.
[0111] The Dedup append program 1107 determines whether there is duplicate data among the data stored in the Pool (S3201). Although not described in detail in the figure, if all data is checked for the existence of data with the same content, the computational complexity becomes extremely large. Therefore, for each data, an operation using a hash function is performed to calculate a representative value of the data such as a hash value, and a comparison process is performed only between data with matching representative values. When the determination result of S3201 is false (S3201: No), the Dedup append program 1107 causes the compression append program 1108 to execute the compression append process (S3207). As a result, the compressed data of the data whose storage destination is SS-VDEV11S is stored in CR-VDEV11C without passing through Dedup-VDEV11D and without changing the CPU 211 that is the processing subject.
[0112] When the determination result of S3201 is true (S3201: Yes), the Dedup append program 1107 updates the Dedup allocation management table 1014 (S3202). In S3202, an entry for the Dedup allocation information corresponding to the data whose storage destination is the target Dedup-VDEV11D is added to the Dedup allocation management table 1014.
[0113] The Dedup append program 1107 updates the CR-Mapping management table 1012 (S3203). In S3203, an entry for the CR-Mapping-Info corresponding to the data for which Dedup-VDEV11D is the storage destination is added to the CR-Mapping management table 1012.
[0114] The Dedup append program 1107 updates the Dir management table 1009 corresponding to the Dir-Info of the target Dedup-VDEV11D (S3204). In S3204, the CR-Mapping-Info (information representing the reference address in CR-VDEV11CC) for the duplicate data is associated with the address in the target Dedup-VDEV11D of the data. In this way, the duplicate data is associated with the already stored data to reduce the used capacity of the pool.
[0115] The Dedup append program 1107 invalidates the pre-update allocation information (S3205). In S3205, the Dedup append program 1107 updates the Dedup allocation management table 1014 and the compression allocation management table 1011. Also, in S3205, for the areas where the number of allocation destinations in the Dedup allocation management table 1014 becomes zero, the target entries of the compression allocation information are also garbage-collected.
[0116] Figure 33 shows the flow of the compression append process. The compression append process is executed by the compression append program 1108 called from the Dedup append program 1107. The compression append program 1108 determines whether the data length that can be compressed and appended is equal to or greater than a threshold value. This data length may be the data length requested from the aforementioned Dedup append program 1107, or may be the data length that is already stored in the cache memory and includes uncompressed data that has not been reflected on the drive. The threshold value is a predetermined data length, and it is assumed to be a size larger than the unit of mapping for managing snapshots and deduplication, such as 256 KB or 512 KB. When performing compression append processing on the data in the cache memory, it is possible to update the update processing (S3303) of the compression allocation management table and the update processing (S3305) of the CR-Mapping management table in one go by processing continuous data together as much as possible, which can reduce the processing overhead. If the determination result of S3301 is true (S3301: Yes), the compression append program 1108 proceeds to S3302. If the determination result of S3301 is false (S3301: No), the process ends. In this case, the data will be temporarily stored in an uncompressed state in the cache memory. As a result, when another data is written by the host computer, the data is continuously written in the cache memory by the Snapshot append process, which can reduce the update processing overhead of the mapping information. The compression append program 1108 compresses the data to be written (S3302). The compression append program 1108 updates the compression allocation management table 1011 (S3303). In S3303, for each of one or more sub-blocks that are the storage destinations of the compressed data in S3302, the entry corresponding to the sub-block is updated.
[0117] The compression append program 1108 causes the destage program 1109 to execute destage processing (S3304). The compression append program 1108 updates the CR-Mapping management table 1012 after appending (S3305). The compression append program 1108 updates the Dir management table 1009 (S3306).
[0118] The compression append program 1108 invalidates the pre-update allocation information (S3307). In S3307, the Dedup append program 1107 updates the Dedup allocation management table 1014 and the compression allocation management table 1011. Also in S3307, for the areas where the number of allocation destinations in the Dedup allocation management table 1014 becomes zero, the target entries of the compression allocation information are also garbage-collected.
[0119] Figure 34 shows the flow of the destaging process. The destaging process is executed by the destaging program 1109 called from the compression append program 1108. The destaging program 1109 determines whether the append data (one or more compressed data) for the RAID stripe is in the cache unit 903 (S3401). A "RAID stripe" is a stripe in a RAID group (a storage area spanning multiple SSDs 220 that make up the RAID group). When the RAID level of the RAID group requires parity, the size of the "append data for the RAID stripe" may be the size obtained by subtracting the parity size from the stripe size. If the determination result in S3401 is false (S3401: No), the process ends.
[0120] If the determination result in S3401 is true (S3401: Yes), the destaging program 1109 refers to the Pool-Mapping management table 1015 and determines whether page 14 has been allocated to the storage destination (the address in the CR-VDEV11C) of the append data for the RAID stripe (S3402). If the determination result in S3402 is false (S3402: No), the process proceeds to S3405.
[0121] If the determination result of S3402 is true (S3402: Yes), the destage program 1109 updates the Pool allocation management table 1016 (S3403). Specifically, the destage program 1109 allocates page 14. In S3403, among the Pool allocation management table 1016, the entries corresponding to the allocated page 14 (for example, status 2704, destination VDEV #2705, and destination address 2706) are updated.
[0122] The destage program 1109 registers the page #2602 of the allocated page in the entry corresponding to the storage destination of the additional data for the RAID stripe in the Pool allocation management table 1016 (S3404).
[0123] The destage program 1109 writes the additional data for the RAID stripe to the stripe that is the basis of the page. When the RAID level is a RAID level that requires parity, the destage program 1109 generates parity based on the additional data for the RAID stripe and also writes the parity to the stripe.
[0124] Figure 35 is a flowchart showing the processing procedure of the read process. The read process is executed by the Read / Write program 1105 in response to a read request from the host device. First, in S3500, the Read / Write program 1105 obtains the address of the data targeted by the read request from the server system 202 within the PVOL or Snapshot. Next, in S3501, the Read / Write program 1105 determines whether the data targeted by the read request results in a cache hit. If the data targeted by the read request results in a cache hit (S3501: Yes), the Read / Write program 1105 transfers the process to S3508. If there is no cache hit (S3501: No), the Read / Write program 1105 transfers the process to S3502.
[0125] In S3502, the Read / Write program 1105 refers to the Dir management table 1009 and the SS-Mapping management table 1010, and based on the PVOL / Snapshot internal address obtained in S3500, obtains the address on the referenced SS-VDEV. At this time, if the size of the target data of the read request is larger than the management unit of the SS-Mapping management table 1010, all entries are referenced to obtain the address on the SS-VDEV.
[0126] Next, in S3503, the Read / Write program 1105 refers to the Dir management table 1009 and the CR-Mapping management table 1012, and based on the SS-VDEV internal address obtained in S3502, obtains the address within the CR-VDEV or within the Dedup-VDEV.
[0127] Next, in S3504, it is determined whether the referenced address obtained in S3503 is on the Dedup-VDEV. Specifically, the referenced destination CR-VDEV#2303 in the CR-Mapping management table 1012 is obtained, and further referring to the CR-VDEV management table 1002, the VDEV#1301 where the CR-VDEV#2303 and the CR-VDEV#1302 match is identified. If the result of the determination in S3504 is false (S3504: No), the Read / Write program 1105 proceeds to S3506. On the other hand, if the result of the determination in S3504 is true (S3504: Yes), the Read / Write program 1105 refers to the Dir management table 1009 and the CR-Mapping management table 1012 based on the address within the Dedup-VDEV identified in S3503 to obtain the address within the CR-VDEV (S3505).
[0128] Next, in S3506, the Read / Write program 1105 stages the data stored at the CR-VDEV internal address specified in S3503 or S3505 into the cache memory while expanding the data.
[0129] Next, the Read / Write program 1105 determines whether all the data within the range requested by the host device has been read onto the cache (S3507). If the result of the determination is true, the Read / Write program 1105 proceeds to S3508, transfers the data that hit the cache in S3501 or the data staged in S3506 to the host device, and ends the process.
[0130] On the other hand, if the result of S3507 is false (S3507: No), the Read / Write program 1105 returns to S3502 and restages the data that is insufficient on the cache. In this way, when the data within the range requested by the host is stored in a continuous area on the CR-VDEV, the data can be made complete on the cache with just one staging from the drive. However, when it is stored dispersedly on the SS-VDEV, Dedup-VDEV, or CR-VDEV, since the reference of metadata or the staging operation from the drive occurs multiple times, the throughput performance deteriorates.
[0131] FIG. 36 shows the flow of the GC (Garbage Collection) process. The GC process is executed by the GC program 1110, for example, periodically (or in response to an instruction from the management system 203). The GC program 1110 refers to the Pool-Mapping management table 1015 and the compression allocation management table 1011 to identify the pages having sub-blocks in the garbage state (status 2203 “2”) (S3601). If there are no pages having sub-blocks in the garbage state, the GC process may end. Also, in S3601, the GC program 1110 may preferentially select the CR-VDEV11C with the least free area among the plurality of CR-VDEVs 11C. Also, in S3601, the GC program 1110 may preferentially identify the page having the most sub-blocks in the garbage state among the CR-VDEVs 11C. Also, the GC may be performed in units of areas different from page 14. The GC program 1110 determines (S3602) whether there is an unprocessed sub-block (not yet determined at S3603) in the page specified at S3601. If the determination result of S3602 is true (S3602: Yes), the GC program 1110 refers to the compression allocation management table 1011 to determine the sub-block to be processed (S3603). The GC program 1110 determines whether the status 2203 corresponding to the sub-block to be processed is "1" (allocated) (S3604). If the determination result of S3604 is false (S3604: No), the process returns to S3602. If the determination result of S3604 is true (S3604: Yes), the GC program 1110 determines whether the allocation of the sub-block to be processed is also valid on the SS-VDEV (S3605). Specifically, the GC program 1110 refers to the compression allocation management table 1011 to identify the allocated VDEV 2205 and the allocated address 2206 of the entry corresponding to the sub-block to be processed, thereby identifying the address on the SS-VDEV where the sub-block to be processed is allocated. Then, referring to the snapshot allocation management table 1008, it determines whether the Status 1902 of the entry corresponding to the address on the SS-VDEV is 1 (allocated). If Status 1902 is 1 (allocated), it is determined that the allocation of the sub-block to be processed is also valid on the SS-VDEV. If the determination result of S3605 is false (S3605: No), the process returns to S3602. If the determination result of S3605 is true (S3605: Yes), the GC program 1110 appends the processing target sub-block to another area (S3606). This "another area" may be a free sub-block (a sub-block with status 2203 being "0") in a CR-VDEV11C different from the CR-VDEV11C that is the target of the GC process (the CR-VDEV11C having the sub-block to which the page specified in S3601 is allocated). Also, this "different CR-VDEV11C" may be a CR-VDEV11C in which all sub-blocks are flow sub-blocks. Also, page 14 may be allocated to this "another area", and the compressed data in the processing target sub-block may be written to the page 14 (in other words, the compressed data may be moved from the page allocated to the processing target sub-block to the page allocated to another area). If the determination result of S3602 is false (S3602: No), the GC program 1110 updates all entries of the compression allocation management table 1011 corresponding to the CR-VDEV11C that is the target of the GC process (S3607). In S3607, for example, the status 2203 of all entries becomes "0". Also, the GC program 1110 updates the Pool-Mapping management table 1015 and the Pool allocation management table 1016 (S3608). In S3806, for example, the page #2602 of all entries of the Pool-Mapping management table 1015 corresponding to the CR-VDEV11C that is the target of the GC process is initialized, and the status 2704 corresponding to all pages allocated to the CR-VDEV11C that is the target of the GC process may be set to "0" (free). In this way, the GC process according to this embodiment may move valid compressed data (compressed data in the allocated sub-blocks) between CR-VDEV11Cs to make a plurality of allocated sub-blocks in a discontinuous state into a continuous state. Note that making a plurality of allocated sub-blocks in a discontinuous state into a continuous state may be performed without data movement between CR-VDEV11Cs.
[0132] As described above, the disclosed storage system 201 includes an SSD 220 as a storage device and a processor 211 that accesses the storage device. The processor 211 manages a primary volume 10P that is a target of read / write by a host and a snapshot volume 10S generated from the primary volume 10P as a snapshot family 9. The processor 211 uses a snapshot virtual device 11S, which is a logical address space associated with the snapshot family 9, as a storage destination for data of the primary volume 10P and the snapshot volume 10S. The data stored in the snapshot virtual device is compressed and stored in a compressed virtual device, and the data stored in the compressed virtual device is stored in the storage device. When the processor 211 receives a write request from the host, it switches between an overwrite process of overwriting an area on the snapshot virtual device 11S that has already been allocated to large-size data according to the size of the write destination address range, and a new allocation process of allocating a new area on the snapshot virtual device 11S to the write destination address range for small-size data, and compresses and stores the plurality of small-size data stored in the new area in the compressed virtual device in a lump. Therefore, in a storage system that provides snapshots, by optimizing the amount of updated mapping information and the amount of data transfer according to the length of data written from the host and the mapping state, and switching the allocation of virtual device areas, high throughput performance can be achieved. That is, in the case of random writes, the throughput performance is improved because the fragmented data can be moved together from the snapshot virtual device to the compressed virtual device.
[0133] Also, a continuous new area on the snapshot virtual device is allocated to the plurality of small-size data. Therefore, the storage of data from the snapshot virtual device to the compressed virtual device can be made efficient.
[0134] Further, the processor 211 uses a snapshot allocation management table 1008 indicating the mapping from the address in the snapshot virtual device 11S to the address in the primary volume and / or the snapshot volume, manages whether the address in the snapshot virtual device 11S is assigned to the address of any volume, and when the processor 211 receives the write request and performs the new allocation process, the area of the snapshot virtual device 11S that was assigned to the write destination address range before the new allocation process is updated in the snapshot allocation management table 1008 as an area not assigned to the address of any volume. When the processor 211 performs garbage collection processing, the processor 211 identifies the address of the snapshot virtual device 11S referenced by the storage area of the compression virtual device that is a candidate for recovery, refers to the snapshot allocation management table 1008 for the identified address, and sets as a condition for the recovery that the area is not assigned to the address of any volume. By using this process, it is not necessary to update the SS-Mapping management table 1010 etc. during writing, and the writing process can be speeded up. This is because the snapshot allocation management table 1008 is small in size and the update of the snapshot allocation management table 1008 is completed in a shorter time compared to the update of the SS-Mapping management table 1010 etc. The SS-Mapping management table 1010 is updated when reflecting the appended write data on the storage device collectively. Also, during garbage collection, in addition to the condition that the status 2203 of the compression allocation management table 1011 indicates "garbage collection required", the snapshot allocation management table 1008 is used to confirm whether the allocation on the snapshot virtual device is valid, so that the target of garbage collection can be accurately determined even if the CR-Mapping management table 1012 is not updated during writing.
[0135] Further, the processor 211 performs the overwriting process on the condition that there is no reference from the snapshot volume 10S to the write destination address range and the write destination address range has a transfer length equal to or greater than a threshold value on the snapshot virtual device 11S. In the case of a transfer length less than the threshold value, the new allocation process is performed.
[0136] Also, in the disclosed storage system, when a plurality of volumes belonging to the same snapshot family 9 have the same data, a predetermined area on the snapshot virtual device 11S is allocated to the same data, and the plurality of volumes refer to the predetermined area. Therefore, the data referred to by the snapshot can be efficiently managed.
[0137] Also, in the disclosed storage system, when a plurality of volumes belonging to different snapshot families 9 have the same data, a predetermined area on the deduplication virtual device 11D referred to from the snapshot virtual device 11S is allocated to the same data. Therefore, duplicate data can be efficiently managed.
[0138] Further, the processor 211 includes information regarding the generation of the snapshot in the mapping information that manages the correspondence between the address in the volume belonging to the snapshot family 9 and the area on the snapshot virtual device 11S. If the generation of the write destination address range does not match the latest generation, the new allocation process is performed. Therefore, it is possible to efficiently manage whether the write destination is independent data or not.
[0139] Note that the present invention is not limited to the above-described embodiments and includes various modifications. For example, the above-described embodiments have been described in detail for easy understanding of the present invention and are not necessarily limited to those having all the configurations described. Also, not only deletion of such a configuration but also replacement and addition of a configuration are possible.
Explanation of Symbols
[0140] 9: Snapshot Family, 10P: Primary Volume, 10S: Snapshot Volume, 11C: Compression Append Virtual Device, 11D: Deduplication Virtual Device, 11S: Snapshot Virtual Device, 13: Pool, 70: Dir-Info Generation Management Tree, 100: Computer System, 201: Storage System, 202: Server System, 203: Management System, 204: Storage Network, 205: Management Network, 210: Storage Controller, 211: CPU, 212: Memory, 213: Back-End Interface, 214: Front-End Interface, 215: Management Interface, 901: Control Information Section, 902: Program Section, 903: Cache Section
Claims
1. A storage system comprising: a memory device; and a processor that accesses the memory device, wherein the processor manages, as a snapshot family, a primary volume that is a target of read / write by a host and a snapshot volume generated from the primary volume, wherein the processor uses, as a storage destination of data of the primary volume and the snapshot volume, a snapshot virtual device that is a logical address space associated with the snapshot family, compresses data stored in the snapshot virtual device and stores the compressed data in a compressed virtual device, stores the data stored in the compressed virtual device in the memory device, wherein when the processor receives a write request from the host, the processor switches between an overwrite process of overwriting an area on the snapshot virtual device that has already been allocated to large-size data according to the size of the write destination address range, and a new allocation process of allocating a new area on the snapshot virtual device to the write destination address range for small-size data, compresses a plurality of small-size data stored in the new area and stores the compressed data together in the compressed virtual device. A storage system characterized by the above.
2. The storage system according to claim 1, wherein a continuous new area on the snapshot virtual device is allocated to a plurality of the small-size data. A storage system characterized by the above.
3. The storage system according to claim 1, wherein the processor uses a snapshot allocation management table indicating mapping from an address in the snapshot virtual device to an address in the primary volume and / or the snapshot volume, and manages whether an address in the snapshot virtual device is allocated to an address of any volume, wherein when the processor receives the write request and performs the new allocation process, the processor updates the snapshot allocation management table so that the area of the snapshot virtual device that has been allocated to the write destination address range before the new allocation process is an area not allocated to an address of any volume. When performing garbage collection processing, the processor identifies the address of the snapshot virtual device referred to by the storage area of the compressed virtual device that is a candidate for recovery, refers to the snapshot allocation management table for the identified address, and sets as a condition for the recovery that the area is not allocated to the address of any volume. A storage system characterized by this.
4. The storage system according to claim 1, wherein when there is no reference from the snapshot volume to the address range of the write destination and the data length of the write request is equal to or greater than a threshold value, the processor performs the overwrite process, and when it is less than the threshold value, the processor performs the new allocation process. A storage system characterized by this.
5. The storage system according to claim 1, wherein when a plurality of volumes belonging to the same snapshot family have the same data, a predetermined area on the snapshot virtual device is allocated to the same data, and the plurality of volumes refer to the predetermined area. A storage system characterized by this.
6. The storage system according to claim 1, wherein when a plurality of volumes belonging to different snapshot families have the same data, a predetermined area on the deduplication virtual device referred to from the snapshot virtual device is allocated to the same data. A storage system characterized by this.
7. The storage system according to claim 1, wherein the processor includes information regarding the generation of the snapshot in the mapping information that manages the correspondence between the address in the volume belonging to the snapshot family and the area on the snapshot virtual device, and when the generation of the address range of the write destination does not match the latest generation, the processor performs the new allocation process. A storage system characterized by this.
8. The storage system according to claim 3, wherein when there is no reference from the snapshot volume to the address range of the write destination and the data length of the write request is equal to or greater than a threshold value, the processor Determine whether the area required for the new allocation process exists on the snapshot virtual device. If the required area exists, perform the allocation. If the required area does not exist, expand the snapshot virtual device and then perform the new allocation process. A storage system characterized by this.
9. A data processing method for a storage system including a storage device and a processor that accesses the storage device, wherein the processor manages a primary volume that is the target of read / write by a host and a snapshot volume generated from the primary volume as a snapshot family, the processor uses a snapshot virtual device, which is a logical address space associated with the snapshot family, as a storage destination for the data of the primary volume and the snapshot volume, the processor compresses the data stored in the snapshot virtual device and stores it in a compressed virtual device, the processor stores the data stored in the compressed virtual device in the storage device, when the processor receives a write request from the host, depending on the size of the write destination address range, it switches between an overwrite process of overwriting an area on the snapshot virtual device that has already been allocated to large-sized data and a new allocation process of allocating a new area on the snapshot virtual device to the write destination address range for small-sized data, and compresses and stores the plurality of small-sized data stored in the new area together in the compressed virtual device. A data processing method characterized by this.
Citation Information
Patent Citations
Storage controller and storage control method
JP2020047036A
Storage system and method for controlling the same
JP2022083955A
Storage controller and storage control method
US20200097179A1
Storage system and control method for storage system
US20220164146A1
Storage controller and storage control method
US10817209B2