Storage System and Memory Control Method

By utilizing separate virtual devices for snapshot families and deduplication in storage systems, the storage system efficiently manages address mappings for duplicate data, addressing the time-consuming issue of updating multiple references in existing technologies.

JP7699571B2Active Publication Date: 2025-06-27HITACHI VANTARA LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022187521
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-11-24
Publication Date
2025-06-27
Estimated Expiration
2042-11-24

AI Technical Summary

Technical Problem

In storage systems using the RoW method for snapshot management, changing the address mapping for duplicate data across multiple snapshot families is time-consuming due to the need to update multiple references.

Method used

The storage system employs separate virtual devices for snapshot families and deduplication, allowing for efficient mapping of duplicate data addresses across these virtual devices, thereby reducing the time required for address mapping changes.

Benefits of technology

This approach enables rapid address mapping changes for duplicate data, significantly reducing the overall processing time involved in managing snapshot families and deduplication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007699571000001
    Figure 0007699571000001
  • Figure 0007699571000002
    Figure 0007699571000002
  • Figure 0007699571000003
    Figure 0007699571000003
Patent Text Reader

Abstract

To change address mapping in a short time even when addresses are changed for duplicated data of a plurality of snapshot families each including a PVOL and an SVOL which is a snapshot of the PVOL.SOLUTION: A snapshot virtual device (SS-VDEV) is prepared for each snapshot family (SS-Family) and a deduplication virtual device is prepared apart from the SS-VDEV. When the same data exists in a plurality of VOLs of the SS-Family, a storage system maps a plurality of addresses of the same data among the plurality of VOLs to the address of the SS-VDEVs of the SS-Family. When duplicated data exists in two or more SS-VDEVs, the storage system maps two or more addresses of the duplicated data of the two or more SS-VDEVs to addresses corresponding to the duplicated data among the deduplication virtual devices.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to memory control of a storage system.

Background Art

[0002] As one of the functions of a storage system, a snapshot function is known. Regarding the snapshot function, for example, the technology disclosed in Patent Document 1 is known. Patent Document 1 discloses a technology related to a snapshot function of the RoW (Redirect on Write) method. The RoW method is a method of appending data. Appending data means that when writing data to a storage system, without overwriting the data stored before writing, storing the data to be written in a new area, and rewriting the meta-information so as to refer to the data stored in the new area.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Regarding the snapshot obtained by the RoW method, the following data management can be adopted. That is, for a snapshot family which is a VOL group including a PVOL (Primary Volume) and an SVOL (Secondary Volume) which is a snapshot of the PVOL, a virtual device is provided as a logical address space. The address of the same data in the virtual device becomes the reference destination of different addresses of a plurality of different VOLs (Volumes) in the snapshot family.

[0005] In such data management, a plurality of snapshot families can be configured. In this case, a virtual device is shared by a plurality of snapshot families. For each snapshot family, a plurality of addresses of a plurality of VOLs in the snapshot family can be the source of reference to the same address in the virtual device.

[0006] When duplicate data is to be stored among snapshot families, the duplicate data is stored in the virtual device for each snapshot family.

[0007] In order to avoid storing such duplicate data in the virtual device, a deduplication technique can be applied. When the deduplication technique is applied, a plurality of addresses in a plurality of snapshot families can be the source of reference to the same address in the virtual device.

[0008] In such data management where the deduplication technique is applied, when the address of data in the virtual device is changed due to garbage collection or other reasons, if the address is the destination of a plurality of different addresses in a plurality of snapshot families, it is necessary to change the destination address for each of the plurality of addresses. For this reason, it takes a long time to change the address mapping, and as a result, the time for the entire process involving the change of the address mapping becomes long.

Means for Solving the Problem

[0009] The storage system has one or more snapshot virtual devices for each of a plurality of snapshot families. Also, the storage system has a deduplication virtual device as a virtual device different from the snapshot virtual device. Each snapshot virtual device is a logical address space that is the storage destination of the data of the VOL in the snapshot family corresponding to the snapshot virtual device.

[0010] When there is the same data in multiple VOLs of a snapshot family, the storage system maps the multiple addresses of the same data among the multiple VOLs to the address of the snapshot virtual device of the snapshot family. When there is duplicate data in two or more snapshot virtual devices of two or more snapshot families, the storage system maps the two or more addresses of the duplicate data among the two or more snapshot virtual devices to the address corresponding to the duplicate data among the deduplication virtual devices.

Effect of the Invention

[0011] According to the present invention, even if the address is changed for duplicate data of multiple snapshot families, the address mapping can be changed in a short time.

Brief Description of the Drawings

[0012]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

Figure 35

Figure 36

Figure 37

Figure 38

Figure 39

Figure 40

Figure 41

Figure 42

Figure 43

Mode for Carrying Out the Invention

[0013] In the following description, the "interface device" may be one or more interface devices. The one or more interface devices may be at least one of the following. · One or more I / O (Input / Output) interface devices. The I / O (Input / Output) interface device is an interface device for at least one of an I / O device and a remote display computer. The I / O interface device for the display computer may be a communication interface device. At least one I / O device may be either an input device such as a user interface device, for example, a keyboard and a pointing device, or an output device such as a display device. · One or more communication interface devices. The one or more communication interface devices may be one or more homogeneous communication interface devices (for example, one or more NICs (Network Interface Cards)) or two or more heterogeneous communication interface devices (for example, a NIC and an HBA (Host Bus Adapter)).

[0014] Also, in the following description, "memory" is one or more memory devices which are an example of one or more storage devices, and may typically be a main memory device. At least one memory device in the memory may be a volatile memory device or a non-volatile memory device.

[0015] Also, in the following description, "persistent storage device" may be one or more persistent storage devices which are an example of one or more storage devices. The persistent storage device may typically be a non-volatile storage device (for example, an auxiliary storage device), and specifically, for example, an HDD (Hard Disk Drive), an SSD (Solid State Drive), an NVME (Non-Volatile Memory Express) drive, or an SCM (Storage Class Memory).

[0016] Also, in the following description, "storage device" may be at least the memory of the memory and the persistent storage device.

[0017] Also, in the following description, the "processor" may be one or more processor devices. At least one processor device may typically be a microprocessor device such as a CPU (Central Processing Unit), but may also be other types of processor devices such as a GPU (Graphics Processing Unit). At least one processor device may be single-core or multi-core. At least one processor device may be a processor core. At least one processor device may also be a circuit that is an aggregate of gate arrays (e.g., an FPGA (Field-Programmable Gate Array), a CPLD (Complex Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit)) described by a hardware description language that performs part or all of the processing, which is a processor device in a broad sense.

[0018] Also, in the following description, expressions such as "xxx table" may be used to describe information from which an output can be obtained for an input, but the information may be data of any structure (e.g., structured data or unstructured data), or may also be a learning model such as a neural network, a genetic algorithm, or a random forest that generates an output for an input. Therefore, "xxx table" can be referred to as "xxx information". Also, in the following description, the configuration of each table is an example, and one table may be divided into two or more tables, or all or part of two or more tables may be one table.

[0019] Also, in the following description, the "program" may be used as the subject to explain the processing. However, since the program performs the defined processing by being executed by a processor while appropriately using a storage device and / or an interface device, the subject of the processing may be the processor (or a device or system having the processor). The program may be installed from a program source into a device such as a computer. The program source may be, for example, a program distribution server or a computer-readable recording medium (e.g., a non-transitory recording medium). Also, in the following description, two or more programs may be realized as one program, or one program may be realized as two or more programs.

[0020] Also, in the following description, "VOL" is an abbreviation for logical volume, which may be a logical storage area. The VOL may be a physical VOL (RVOL) or a virtual VOL (VVOL). "RVOL" may be a VOL based on physical storage resources (e.g., one or more RAID groups) possessed by the storage system that provides the RVOL (where "RAID" is an abbreviation for Redundant Array of Independent (or Inexpensive) Disks). "VVOL" may be any of an external connected VOL (EVOL), a thin provisioning VOL (TPVOL), and a snapshot VOL (SSVOL). The EVOL may be a VOL based on the storage space (e.g., a VOL) of an external storage system and conforming to storage virtualization technology. The TPVOL may be a VOL composed of a plurality of virtual areas (virtual storage areas) and conforming to capacity virtualization technology (typically Thin Provisioning). The SSVOL may be a VOL provided as a snapshot of an original VOL. The SSVOL may be an RVOL. Typically, the SSVOL is positioned as a secondary VOL with the original VOL as the primary VOL (PVOL). A "pool" is a logical storage area (e.g., a collection of a plurality of pool VOLs) and may be prepared for each use. For example, there may be at least one of a TP pool and a snapshot pool as the pool. The TP pool may be a storage area composed of a plurality of real areas (physical storage areas). When a real area is not allocated to the virtual area (the virtual area of the TPVOL) to which the address specified by the write request received from the host system belongs, a real area is allocated from the TP pool to that virtual area (the write destination virtual area) (even if another real area has already been allocated to the write destination virtual area or a new real area is allocated to the write destination virtual area). The storage system may write the write target data associated with the write request to the allocated real area. The snapshot pool may be a storage area in which data evacuated from the PVOL is stored. One pool may be used as both a TP pool and a snapshot pool."Pool VOL" may be a VOL that is a component of a pool. The pool VOL may be an RVOL or an EVOL.

[0021] Also, the "storage system" may be a system including a controller that performs I / O of data with respect to a plurality of persistent storage devices (or a device having a plurality of persistent storage devices), or may be a system including one or more physical computers. In the latter system, for example, each of the one or more physical computers may execute predetermined software so that the one or more physical computers are constructed as SDx (Software-Defined anything). As SDx, for example, SDS (Software-Defined Storage) or SDDC (Software-defined Data Center) can be adopted.

[0022] Also, in the following description, #(identification number) and ID are adopted as examples of identification information of elements, but the identification information may be information that can identify an element, such as a name.

[0023] Also, in the following description, when describing without distinguishing between elements of the same type, common reference signs among the reference signs are used, and when describing while distinguishing between elements of the same type, reference signs may be used. Also, in the following description, an element X with #n (identification number n) may be denoted as "X#n".

[0024] Also, in the following description, snapshots are created in the RoW method, but the RoW method may be an example of a copy-free method of data. [Embodiment 1]

[0025] Figure 1 shows an overview of the storage control of the storage system according to Embodiment 1. In Figure 1, data represented by uppercase alphabets (Data A, B, C, …) is block data, and data represented by lowercase alphabets (data a, b, c, …) is sub-block data. The block data may be data in block units. The block may be a fixed-length logical storage area (logical address range). The sub-block data is compressed data of the block data, and a sub-block group (one or more sub-blocks) is the data to be stored. The sub-block may be a logical storage area smaller than the block. For example, an integer multiple of the sub-block may be the block.

[0026] A storage system having a storage device and a processor includes an SS-Family (Snapshot Family) 9, an SS-VDEV (Snapshot Virtual Device) 11S, a Dedup-VDEV (Deduplication Virtual Device) 11V, a CR-VDEV (Compression Append Virtual Device) 11C, and a pool 13.

[0027] The SS-Family 9 is a VOL group including a PVOL 10P and an SVOL 10S which is a snapshot of the PVOL 10P.

[0028] The SS-VDEV 11S is a virtual device as a logical address space, and is the storage destination of data whose storage destination is any VOL 10 in the SS-Family 9.

[0029] The Dedup-VDEV 11D is a virtual device as a logical address space different from the SS-VDEV 11S, and is the storage destination of duplicate data of two or more SS-VDEVs 11S.

[0030] The CR-VDEV 11C is a virtual device as a logical address space different from the SS-VDEV 11S and the Dedup-VDEV 11D, and is the storage destination of compressed data.

[0031] Each of the plurality of CR-VDEV11Cs is associated with either the SS-VDEV11S or the Dedup-VDEV11D, and is not associated with both the VDEV11S and 11D. That is, each CR-VDEV11C serves as the storage destination for the data whose corresponding VDEV (virtual device) is the storage destination, and does not serve as the storage destination for the data whose non-corresponding VDEV is the storage destination. The compressed data for which the CR-VDEV11C is the storage destination is stored in the pool 13.

[0032] The pool 13 is a logical address space based on at least a part of a storage device (e.g., a persistent storage device) of the storage system. The pool 13 may be based on at least a part of an external storage device (e.g., a persistent storage device) of the storage system instead of or in addition to at least a part of the storage device of the storage system. The pool 13 has a plurality of pages 14 which are a plurality of logical regions. The compressed data for which the CR-VDEV11C is the storage destination is stored in the page 14 in the pool 13. The mapping between the address in the CR-VDEV11C and the address in the pool 13 is 1:1. The pool 13 is composed of one or more pool VOLs.

[0033] According to the example shown in FIG. 1, the following storage control is performed.

[0034] The processor creates the SVOL10S0 as a snapshot of the PVOL10P0, thereby creating the SS-Family9-0 with the PVOL10P0 as the root VOL. Also, the processor creates the SVOL10S1 as a snapshot of the PVOL10P1, thereby creating the SS-Family9-1 with the PVOL10P1 as the root VOL. According to FIG. 1, as examples of a plurality of SS-Families, there are the SS-Family9-0 and 9-1.

[0035] The storage system has one or more SS-VDEV11S for each of the plurality of SS-Family9. For each SS-Family9, for the data whose any VOL10 in the SS-Family9 is the storage destination, among the plurality of SS-VDEV11S, the SS-VDEV11S corresponding to the SS-Family9 is set as the storage destination. Taking SS-Family9-0 as an example, specifically, for example, it is as follows. · For the data A whose storage destination is SVOL10S0 of SS-Family9-0, the processor uses SS-VDEV11S0 as the storage destination. The processor maps the address corresponding to data A among SVOL10S0 to the address corresponding to data A among the SS-VDEV11S corresponding to SS-Family9-0. · When there is the same data B in the plurality of VOLs (PVOL10P0 and SVOL10S0) of SS-Family9-0, the processor maps the plurality of addresses of the same data B (the address in PVOL10P0 and the address in SVOL10S0) among the plurality of VOL10s to the address of SS-VDEV11S0 of SS-Family9-0 (the address corresponding to data B).

[0036] For each of SS-Family9-0 and 9-1 (an example of two or more SS-Family9), the storage destination of non-duplicate data is the CR-VDEV11C corresponding to the SS-Family9, and the storage destination of duplicate data is the Dedup-VDEV11D.

[0037] That is, since data C overlaps in SS-VDEV11S0 and 11S1 of SS-Family9-0 and 9-1 (an example of two or more SS-VDEVs), the processor maps the two addresses of the overlapping data C among SS-VDEV11S0 and 11S1 to the address corresponding to the overlapping data C in Dedup-VDEV11D. Then, the processor compresses the overlapping data C and sets the storage destination of the compressed data c to CR-VDEV11CC corresponding to Dedup-VDEV11D. That is, the processor maps the address (block address) of the overlapping data C in Dedup-VDEV11D to the address (sub-block address) of the compressed data c in CR-VDEV11CC. Also, the processor allocates page 14B to CR-VDEV11CC and stores the compressed data c in page 14B. The address of the compressed data c in CR-VDEV11CC is mapped to the address in page 14B of pool 13.

[0038] On the other hand, since data A in SS-VDEV11S0 does not overlap with data in other SS-VDEV11S1, the processor compresses the non-overlapping data A and sets the compressed data a to CR-VDEV11C0 corresponding to SS-VDEV11S0. That is, the processor maps the address (block address) of the non-overlapping data A in SS-VDEV11S0 to the address (sub-block address) of the compressed data a in CR-VDEV11CC. Also, the processor allocates page 14A to CR-VDEV11C0 and stores the compressed data a in page 14a. The address of the compressed data a in CR-VDEV11C0 is mapped to the address in page 14A of pool 13.

[0039] CR-VDEV11C is an append-type VDEV. Therefore, the processor updates the address mapping when the CR-VDEV11C corresponding to SS-VDEV11S is set as the storage destination of the update data and when the CR-VDEV11C corresponding to Dedup-VDEV11D is set as the storage destination of the update data. Specifically, the processor performs, for example, the following storage control. · When the data to be stored in CR-VDEV11C is the updated data a' of the compressed data a, the processor sets the storage destination of the updated data a' as the free address in CR-VDEV11C and invalidates the address of the compressed data a before the update. The processor remaps the address that was mapped to the address of the compressed data a and is the address in SS-VDEV11S0 to the storage destination address of the updated data a' in CR-VDEV11C instead of the address of the compressed data a in CR-VDEV11C. Also, the processor remaps the address that was mapped to the address of the compressed data a and is the address in page 14A to the storage destination address of the updated data a' in CR-VDEV11C instead of the address of the compressed data a in CR-VDEV11C. · When the data to be stored in CR-VDEV11CC is the updated data c' of the compressed data c, the processor sets the storage destination of the updated data c' as the free address in CR-VDEV11CC and invalidates the address of the compressed data c before the update. The processor remaps the address that was mapped to the address of the compressed data c and is the address in Dedup-VDEV11D to the storage destination address of the updated data c' in CR-VDEV11CC instead of the address of the compressed data c in CR-VDEV11CC. Also, the processor remaps the address that was mapped to the address of the compressed data c and is the address in page 14B to the storage destination address of the updated data c' in CR-VDEV11CC instead of the address of the compressed data a in CR-VDEV11CC.

[0040] CR-VDEV11C is the append-type VDEV as described above, and garbage collection is performed. That is, the processor can make the valid addresses (the addresses of the latest data) continuous and make the addresses of the free areas continuous by performing garbage collection on CR-VDEV11C.

[0041] According to the example shown in FIG. 1, apart from SS-VDEV11S which is the storage destination of data in SS-Family9, Dedup-VDEV11D which is the storage destination of duplicate data in two or more SS-Family9s is prepared. Therefore, even if the address of duplicate data C in Dedup-VDEV11D changes, the address mapping to be changed only requires two mappings (mappings for each of the two addresses in SS-VDEV11S0 and 11S1). On the other hand, in a comparative example, the storage destination of data in SS-Family9 and the storage destination of duplicate data in two or more SS-Family9s are the same VDEV. In this case, the address mapping to be changed for duplicate data C requires four mappings (mappings for each of the four addresses in VOL10P0, 10S0, 10P1, and 10S1). In this embodiment, it is expected that the address mapping can be changed in a shorter time than in the comparative example.

[0042] Separate from Dedup-VDEV11D, CR-VDEV11CC is prepared for Dedup-VDEV11D. Therefore, even if the address of compressed data c in CR-VDEV11CC changes, the address mapping to be changed only requires one mapping (mapping for one address in Dedup-VDEV11D). On the other hand, in a comparative example, the storage destination of compressed data of data in SS-Family9 and the storage destination of compressed data of duplicate data in two or more SS-Family9s are the same VDEV. In this case, the address mapping to be changed for compressed data c of duplicate data C requires four mappings (mappings for each of the four addresses in VOL10P0, 10S0, 10P1, and 10S1). In this embodiment, it is expected that the address mapping can be changed in a shorter time than in the comparative example.

[0043] The address of the compressed data c is changed in the garbage collection of CR-VDEV11CC corresponding to Dedup-VDEV11D. For example, in the garbage collection of CR-VDEV11CC, the processor changes the address of the updated data c' (the updated data of the compressed data c) in CR-VDEV11CC, and maps the address that was mapped to the address before the change and is the address in Dedup-VDEV11D to the address after the change in CR-VDEV11CC. Since it is expected to change the address mapping for the compressed data of duplicate data in a short time, it is expected to perform garbage collection in a short time. Note that the garbage collection of CR-VDEV11C corresponding to SS-VDEV11S includes, for example, the following processing. That is, the processor changes the address of the updated data a' (the updated data of the compressed data a) in CR-VDEV11C0, and maps the address that was mapped to the address before the change and is the address in SS-VDEV11S0 to the address after the change in CR-VDEV11C0.

[0044] For at least one CR-VDEV11C, an append-type VDEV in which uncompressed data is stored instead of CR-VDEV11C may be adopted. However, in this embodiment, CR-VDEV11C is adopted as the append-type VDEV. Therefore, the data finally stored in the storage device is compressed data, and thus the consumed storage capacity can be reduced.

[0045] Figure 2 shows an example of the range of the owner CPU.

[0046] The processor is a plurality of CPUs as an example of a plurality of processor devices. By limiting the exclusive ownership (access right (I / O right)) to the data and control information regarding VOL10 to a specific CPU, the processing time required for the exclusive processing in the I / O processing and the communication time between CPUs are reduced, and thus performance improvement is expected.

[0047] For each of the plurality of CPUs, if the owner CPU (identification number of the CPU with the authority) is associated with VOL10 or VDEV11, then if the CPU corresponds to the owner CPU, the CPU performs I / O on the VOL10 or the VDEV11, or updates the address mapping information about the VOL10 or the VDEV11. In other words, if the CPU does not correspond to the owner CPU, communication between CPUs is required for updating I / O and address mapping information.

[0048] In this embodiment, the owner CPU of SS-Family9 (specifically, each of all VOL10s within SS-Family9), the owner CPU of SS-VDEV11S for the SS-Family9, and the owner CPU of CR-VDEV11C corresponding to the SS-VDEV11S are the same CPU. As a result, for non-duplicate data, CPU communication for transferring processing to the CPU with the ownership right for I / O and address mapping changes can be made unnecessary, and thus, performance improvement of processing including short-time change of address mapping is expected. For example, for non-duplicate data A (see FIG. 1) in SS-Family9-0, both I / O and address mapping change are performed only by CPU#0 among the plurality of CPUs. Similarly, for non-duplicate data F (see FIG. 1) in SS-Family9-1, both I / O and address mapping change are performed only by CPU#1 among the plurality of CPUs.

[0049] Also, in this embodiment, the owner CPU of Dedup-VDEV11D and the owner CPU of CR-VDEV11C corresponding to the Dedup-VDEV11D are the same CPU. As a result, for duplicate data, with respect to Dedup-VDEV11D and its corresponding CR-VDEV11C, CPU-to-CPU communication for transferring processing to the CPU with ownership for I / O and address mapping changes can be made unnecessary. For example, assume that the owner CPU of Dedup-VDEV#10x (x is an integer from 0 to 7) and its CR-VDEV#110x is CPU#x. When the storage destinations of duplicate data C and D (see FIG. 1) are Dedup-VDEV#100, for duplicate data C, CPU-to-CPU communication for I / O and address mapping changes is unnecessary, but for duplicate data D, with respect to Dedup-VDEV#100 and its CR-VDEV#1100, communication from CPU#0 to CPU#1 is required.

[0050] Hereinafter, this embodiment will be described in detail.

[0051] FIG. 3 shows an overview of the management of the mapping between the addresses in SS-Family9 and the addresses in SS-VDEV11S. In the drawings, "GX" (X is an integer greater than or equal to 0) means generation X. Also, FIG. 3 takes SS-Family9-0 and SS-VDEV11S0 as examples.

[0052] The processor can manage the mapping between the addresses in VOL10 in SS-Family9-0 and the addresses in SS-VDEV11S0 using meta information. The meta information includes Dir-Info (directory information) and SS-Mapping-Info (snapshot mapping information). The processor manages the data of PVOL10P0 and SVOL10S0 by associating Dir-Info with SS-Mapping-Info. For the data stored with VOL10 as the storage destination, Dir-Info has information representing the source address (the address in VOL10), and the corresponding SS-Mapping-Info for the data has information representing the destination address (the address in SS-VDEV11S0).

[0053] Furthermore, the processor manages the time series of PVOL10P0 and SVOL10S0 with generation information associated with Dir-Info, and for each data with SS-VDEV11S0 as the storage destination, manages by associating with SS-Mapping-Info the generation information indicating the generation in which the data was created. In addition, the processor manages the latest generation information at that time as the latest generation.

[0054] Before the snapshot is acquired, assume there is data A0, B0, and C0 with PVOL10P0 as the storage destination. Also assume that the latest generation is "0".

[0055] The Dir-Info associated with PVOL10P0 has "0" associated as generation # (the number representing the generation), and includes reference information indicating the destinations of all the data A0, B0, and C0 in PVOL10P0. Hereafter, when the generation # associated with Dir-Info is "X", it can be expressed that Dir-Info is generation X.

[0056] SS-VDEV11S0 is the storage destination for data A0, B0, and C0, and SS-Mapping-Info is associated with each of data A0, B0, and C0. Also, each of the SS-Mapping-Info has "0" associated as the generation #. When the generation # associated with the SS-Mapping-Info represents "X", it can be stated that the data corresponding to the SS-Mapping-Info is the data of generation X.

[0057] In the state before snapshot acquisition, for each of data A0, B0, and C0, the information in Dir-Info refers to the SS-Mapping-Info corresponding to the data. By associating Dir-Info and SS-Mapping-Info in this way, PVOL10P0 and SS-VDEV11S0 can be associated, and data processing for PVOL10P0 can be realized.

[0058] To acquire a snapshot, the processor makes a copy of the Dir-Info as the read-only Dir-Info of SVOL10S0. Then, the processor increments the generation of the Dir-Info of PVOL10P0 and also increments the latest generation. As a result, for each of data A0, B0, and C0, the SS-Mapping-Info is referred to from both the Dir-Info of generation 0 and the Dir-Info of generation 1.

[0059] In this way, a snapshot can be created by replicating the Dir-Info, and a snapshot can be created without increasing the data and SS-Mapping-Info on SS-VDEV11S0.

[0060] Here, when a snapshot is acquired, a snapshot (SVOL10S0) in which writing is prohibited and data is fixed at the acquisition time becomes generation 0, and PVOL10P0 in which data can be written even after the acquisition becomes generation 1. Generation 0 is "a generation one generation older in the direct line" with respect to generation 1, and is referred to as "parent" for convenience. Similarly, generation 1 is "a generation one generation newer in the direct line" with respect to generation 0, and is referred to as "child" for convenience. The storage system manages the parent-child relationship of generations as the Dir-Info generation management tree 70. Also, the generation # of Dir-Info is the same as the generation # of VOL10 corresponding to the Dir-Info. Also, the generation # of SS-Mapping-Info is the oldest generation # among the generation #s of one or more Dir-Info that reference the SS-Mapping-Info.

[0061] Figure 4 shows an overview of the management of the mapping between the address in SS-VDEV11S and the address in Dedup-VDEV11D, the management of the mapping between the address in SS-VDEV11S and the address in CR-VDEV11C, and the management of the mapping between the address in Dedup-VDEV11D and the address in CR-VDEV11C. Figure 4 takes SS-VDEV11S0, CR-VDEV11C0, and 11CC as examples.

[0062] The processor can manage the mapping between the address in SS-VDEV11S0 and the address in Dedup-VDEV11D, the mapping between the address in SS-VDEV11S0 and the address in CR-VDEV11C0, and the mapping between the address in Dedup-VDEV11D and the address in CR-VDEV11CC by using the meta information. As described above, the meta information includes Dir-Info and CR-Mapping-Info. The processor manages the data of SS-VDEV11S0 and Dedup-VDEV11D by associating Dir-Info with CR-Mapping-Info. For the data with SS-VDEV11S0 as the storage destination, Dir-Info has information representing the source address (the address in SS-VDEV11S0), and the corresponding CR-Mapping-Info for the data has information representing the destination address (the address in CR-VDEV11C0 or the address in Dedup-VDEV11D). For the data with Dedup-VDEV11D as the storage destination, Dir-Info has information representing the source address (the address in Dedup-VDEV11D), and the corresponding CR-Mapping-Info for the data has information representing the destination address (the address in CR-VDEV11CC). By referring to the compression allocation information, the processor can identify the address in SS-VDEV11S or Dedup-VDEV11D from the address in CR-VDEV11C.

[0063] Figure 5 shows an overview of the reverse mapping from the address in CR-VDEV11C to the address in SS-VDEV11S or Dedup-VDEV11D.

[0064] The storage system stores compression allocation information for each piece of data stored in SS-VDEV11S or Dedup-VDEV11D as the storage destination. The compression allocation information represents the mapping between the source address (the address in CR-VDEV11C) and the destination address (the address in SS-VDEV11S or Dedup-VDEV11D). By referring to the compression allocation information, the processor can identify the address in SS-VDEV11S or Dedup-VDEV11D from the address in CR-VDEV11C.

[0065] Figure 6 shows an overview of the reverse mapping from the address in Dedup-VDEV11D to the address in SS-VDEV11S.

[0066] The storage system has Dedup-Dir-Info, which is Dir-Info for reverse mapping, for each Dedup-VDEV11D. For each piece of data stored in Dedup-VDEV11D, the information in Dedup-Dir-Info refers to the Dedup allocation information corresponding to that data. The Dedup allocation information represents the address of the duplicate data for each SS-VDEV11S where the duplicate data is stored as the storage destination. By referring to Dedup-Dir-Info and the Dedup allocation information, the processor can identify the addresses in each of SS-VDEV11S0 and 11S1 from the address in Dedup-VDEV11D.

[0067] Figure 7 shows an overview of the management of the mapping between the address in CR-VDEV11C and the address in page 14 of pool 13.

[0068] Page 14 as a continuous area is allocated to CR-VDEV11C. The storage system stores Pool-Mapping-Info for each page 14. Pool-Mapping-Info represents the source address (the address in CR-VDEV11C) and the destination address (the address in page 14). By referring to Pool-Mapping-Info, the processor can identify the address in page 14 from the address in CR-VDEV11C.

[0069] Figure 8 shows the hardware configuration of a computer system.

[0070] The computer system 100 includes a storage system 201, a server system 202, and a management system 203. The storage system 201 and the server system 202 are connected via a storage network 204 using FC (Fiber Channel) or the like. The storage system 201 and the management system 203 are connected via a management network 205 using IP (Internet Protocol) or the like. Note that the storage network 204 and the management network 205 may be the same communication network.

[0071] The storage system 201 includes a plurality of storage controllers 210 and a plurality of SSDs 220. A plurality of SSDs 220 are connected to the storage controller 210. The plurality of SSDs 220 are an example of a persistent storage device. A pool 13 is configured based on the plurality of SSDs 220. The data stored in page 14 of pool 13 is stored in one or more SSDs 220.

[0072] The storage controller 210 includes a CPU 211, a memory 212, a backend interface 213, a frontend interface 214, and a management interface 215.

[0073] The CPU 211 executes a program stored in the memory 212.

[0074] The memory 212 stores the programs executed by the CPU 211 and the data used by the CPU 211, etc. The memory may be duplicated by the combination of the memory 212 and the CPU 211.

[0075] The back-end interface 213, the front-end interface 214, and the management interface 215 are examples of interface devices.

[0076] The back-end interface 213 is a communication interface device that mediates the data exchange between the SSD 220 and the storage controller 210. A plurality of SSDs 220 are connected to the back-end interface 213.

[0077] The front-end interface 214 is a communication interface device that mediates the data exchange between the server system 202 and the storage controller 210. The server system 202 is connected to the front-end interface 214 via the storage network 204.

[0078] The management interface 215 is a communication interface device that mediates the data exchange between the management system 203 and the storage controller 210. The management system 203 is connected to the management interface 215 via the management network 205.

[0079] The server system 202 is configured to include one or more host devices. The server system 202 transmits an I / O request (write request or read request) specifying an I / O destination to the storage controller 210. The I / O destination is, for example, a logical volume number such as an LUN (Logical Unit Number), a logical address such as an LBA (Logical Block Address), or the like.

[0080] The management system 203 is configured to include one or more management devices. The management system 203 manages the storage system 201.

[0081] FIG. 9 shows the configuration of the memory 212.

[0082] The memory 212 has a control information section 901 in which control information (which may also be referred to as management information) is stored, a program section 902 in which programs are stored, and a cache section 903 in which data is temporarily stored.

[0083] FIG. 10 shows the information stored in the control information section 901.

[0084] The control information section 901 stores an ownership management table 1001, a CR-VDEV management table 1002, a snapshot management table 1003, a VOL-Dir management table 1004, a latest generation table 1005, a recovery management table 1006, a generation management tree table 1007, a snapshot allocation management table 1008, a Dir management table 1009, an SS-Mapping management table 1010, a compression allocation management table 1011, a CR-Mapping management table 1012, a Dedup-Dir management table 1013, a Dedup allocation management table 1014, a Pool-Mapping management table 1015, and a Pool allocation management table 1016.

[0085] FIG. 11 shows the programs stored in the program section 902.

[0086] The program section 902 stores a snapshot acquisition program 1101, a snapshot restore program 1102, a snapshot deletion program 1103, an asynchronous recovery program 1104, a read / write program 1105, a snapshot append program 1106, a Dedup append program 1107, a compression append program 1108, a destage program 1109, a GC (garbage collection) program 1110, a CPU determination program 1111, and an ownership transfer program 1112.

[0087] FIG. 12 shows the configuration of the owner right management table 1001.

[0088] The owner right management table 1001 manages the owner rights of VOL10 or VDEV11. For example, the owner right management table 1001 has entries for each of VOL10 and VDEV11. The entry has information such as VOL# / VDEV# 1201 and owner CPU# 1202.

[0089] VOL# / VDEV# 1201 represents the identification number of VOL10 or VDEV11. Owner CPU# 1202 represents the identification number of the CPU as the owner CPU of VOL10 or VDEV11 (the CPU having the owner right of VOL10 or VDEV11).

[0090] Note that the owner CPU may be assigned in units of CPU groups instead of in units of CPU211, or may be assigned in units of storage controller 210.

[0091] FIG. 13 shows the configuration of the CR-VDEV management table 1002.

[0092] The CR-VDEV management table 1002 represents the CR-VDEV 11C associated with SS-VDEV11S or Dedup-VDEV11D. For example, the CR-VDEV management table 1002 has entries for each of SS-VDEV11S and Dedup-VDEV11D. The entry has information such as VDEV# 1301 and CR-VDEV# 1302.

[0093] VDEV# 1301 represents the identification number of SS-VDEV11S or Dedup-VDEV11D. CR-VDEV# 1302 represents the identification number of CR-VDEV11C.

[0094] FIG. 14 shows the configuration of the snapshot management table 1003.

[0095] The snapshot management table 1003 exists for each PVOL10P (for each SS-Family9). The snapshot management table 1003 represents the acquisition time of each snapshot (SVOL10S). For example, the snapshot management table 1003 has an entry for each SVOL10S. The entry has information such as PVOL#1401, SVOL#1402, and acquisition time 1403.

[0096] PVOL#1401 represents the identification number of PVOL10P. SVOL#1402 represents the identification number of SVOL10S. Acquisition time 1403 represents the acquisition time of SVOL10S.

[0097] Figure 15 shows the configuration of the VOL-Dir management table 1004.

[0098] The VOL-Dir management table 1004 represents the correspondence between VOL and Dir-Info. For example, the VOL-Dir management table 1004 has an entry for each VOL10. The entry has information such as VOL#1501, Root-VOL#1502, and Dir-Info#1503.

[0099] VOL#1501 represents the identification number of PVOL10P or SVOL10S. Root-VOL#1502 represents the identification number of the Root-VOL. If VOL10 is PVOL10P, then the Root-VOL is the PVOL10P. If VOL10 is SVOL10S, then the Root-VOL is the PVOL10P corresponding to the SVOL10S. Dir-Info#1503 represents the identification number of the Dir-Info corresponding to VOL10.

[0100] Figure 16 shows the configuration of the latest generation table 1005.

[0101] The latest generation table 1005 exists for each PVOL10P (for each SS-Family9) and represents the generation (generation #) of the PVOL10P.

[0102] FIG. 17 shows the configuration of the recovery management table 1006.

[0103] The recovery management table 1006 may be, for example, a bitmap, and exists for each PVOL10P (for each SS-Family9), in other words, exists for each Dir-Info generation management tree 70. The recovery management table 1006 has an entry for each Dir-Info. The entry has information such as Dir-Info#1701 and recovery request 1702.

[0104] Dir-Info#1701 represents the identification number of the Dir-Info. The recovery request 1702 indicates whether to request the recovery of the Dir-Info. "1" means to request recovery, and "0" means not to request recovery.

[0105] FIG. 18 shows the configuration of the generation management tree table 1007.

[0106] The generation management tree table 1007 exists for each PVOL10P (for each SS-Family9), in other words, exists for each Dir-Info generation management tree 70. The generation management tree table 1007 has an entry for each Dir-Info. The entry has information such as Dir-Info#1801, generation#1802, Prev1803, and Next1804.

[0107] Dir-Info# represents the identification number of the Dir-Info. Generation#1802 represents the generation of VOL10 corresponding to the Dir-Info. Prev1803 represents the Dir-Info of the parent (one level above) of the Dir-Info. Next1804 represents the Dir-Info of the child (one level below) of the Dir-Info. The number of Next1804 may be the same as the number of child Dir-Info. In FIG. 18, since there are two child Dir-Info, there are two Next1804 (Next-A1804A and Next-B1804B).

[0108] FIG. 19 shows the configuration of the snapshot allocation management table 1008.

[0109] The snapshot allocation management table 1008 exists for each SS-VDEV11S and represents the mapping from the address in SS-VDEV11S to the address in VOL10. The snapshot allocation management table 1008 has an entry for each address in SS-VDEV11S. The entry has information such as the block address 1901, the status 1902, the destination VOL#1903, and the destination address 1904.

[0110] The block address 1901 represents the address of the block in SS-VDEV11S. The status 1902 indicates whether the block is assigned to the address of any VOL ("1" means assigned, "0" means free). The destination VOL#1903 represents the identification number of the VOL10 (PVOL10P or SVOL10S) that has the destination address of the block ("n / a" means unassigned). The destination address 1904 represents the destination address (block address) of the block ("n / a" means unassigned).

[0111] Figure 20 shows the configuration of the Dir management table 1009.

[0112] The Dir management table 1009 exists for each Dir-Info and represents the reference destination Mapping-Info for each data (for each block data). For example, the Dir management table 1009 has an entry for each address (block address). The entry has information such as the VOL / VDEV address 2001 and the reference destination Mapping-Info#2002.

[0113] The VOL / VDEV address 2001 represents the address (block address) in VOL10 (PVOL10P or SVOL10S), or the address in VDEV11 (SS-VDEV11S or Dedup-VDEV11D). The reference destination Mapping-Info#2002 represents the identification number of the reference destination Mapping-Info.

[0114] Figure 21 shows the configuration of the SS-Mapping management table 1010.

[0115] The SS-Mapping management table 1010 exists for each Dir-Info of VOL10. The SS-Mapping management table 1010 has an entry for each SS-Mapping-Info corresponding to the Dir-Info of VOL10. The entry has information such as Mapping-Info#2101, reference address 2102, reference destination SS-VDEV#2103, and generation #2104.

[0116] Mapping-Info#2101 represents the identification number of the SS-Mapping-Info. The reference address 2102 represents the address (the address in SS-VDEV11S) referred to by the SS-Mapping-Info. The reference destination SS-VDEV#2103 represents the identification number of the SS-VDEV11S that has the address referred to by the SS-Mapping-Info. The generation #2104 represents the generation of the data corresponding to the SS-Mapping-Info.

[0117] Figure 22 shows the configuration of the compression allocation management table 1011.

[0118] The compression allocation management table 1011 exists for each CR-VDEV11C and has compression allocation information for each sub-block in CR-VDEV11C. The compression allocation management table 1011 has an entry corresponding to the compression allocation information for each sub-block in CR-VDEV11C. The entry has information such as sub-block address 2201, data length 2202, status 2203, start sub-block address 2204, allocation destination VDEV#2205, and allocation destination address 2206.

[0119] The sub-block address 2201 represents the address of the sub-block. The data length 2202 represents the number of sub-blocks that make up a group of sub-blocks (one or more sub-blocks) in which the compressed data is stored (for example, "2" means that the compressed data exists in two sub-blocks). The status 2203 represents the status of the sub-block ("0" means free, "1" means allocated, and "2" means a target for GC (garbage collection)). The start sub-block address 2204 represents the address of the start sub-block of one or more sub-blocks (one or more sub-blocks containing compressed data). The assigned VDEV# 2205 represents the identification number of the VDEV11 (SS-VDEV11S or Dedup-VDEV11D) that has the block assigned to the sub-block. The assigned address 2206 represents the address of the block assigned to the sub-block (the block address in SS-VDEV11S or Dedup-VDEV11D).

[0120] Figure 23 shows the configuration of the CR-Mapping management table 1012.

[0121] The CR-Mapping management table 1012 exists for each Dir-Info of Dedup-VDEV11D. The CR-Mapping management table 1012 has an entry for each CR-Mapping-Info corresponding to the Dir-Info of Dedup-VDEV11D. The entry has information such as Mapping-Info# 2301, the reference address 2302, the reference CR-VDEV# 2303, and the data length 2304.

[0122] Mapping-Info#2301 represents the identification number of the CR-Mapping-Info. The reference address 2302 represents the address (the address of the first sub-block among the sub-block groups) referred to by the CR-Mapping-Info. The reference destination CR-VDEV#2303 represents the identification number of the CR-VDEV11C having the sub-block address referred to by the CR-Mapping-Info. The data length 2304 represents the number of blocks (blocks in the Dedup-VDEV11D) referred to by the CR-Mapping-Info, or the number of sub-blocks constituting the sub-block group referred to by the CR-Mapping-Info.

[0123] Figure 24 shows the configuration of the Dedup-Dir management table 1013.

[0124] The Dedup-Dir management table 1013 exists for each Dedup-VDEV11D and corresponds to the Dedup-Dir-Info. The Dedup-Dir management table 1013 has an entry for each address in the Dedup-VDEV11D. The entry has information such as the Dedup-VDEV address 2401 and the reference destination allocation information#2402.

[0125] The Dedup-VDEV address 2401 represents the address (block address) in the Dedup-VDEV11D. The reference destination allocation information#2402 represents the identification number of the reference destination Dedup allocation information.

[0126] Figure 25 shows the configuration of the Dedup allocation management table 1014.

[0127] The Dedup allocation management table 1014 exists for each Dedup-VDEV11D (for each Dedup-Dir-Info), and represents the mapping from the address in the Dedup-VDEV11D to the address in the SS-VDEV11S based on the Dedup allocation information corresponding to the address. The Dedup allocation management table 1014 has an entry for each Dedup allocation information. The entry has information such as allocation information #2501, target SS-VDEV #2502, target address 2503, and linked allocation information #2504.

[0128] The allocation information #2501 represents the identification number of the Dedup allocation information. The target SS-VDEV #2502 represents the identification number of the SS-VDEV11S having the address referred to by the Dedup allocation information. The target address 2503 represents the address (block address in the SS-VDEV11S) referred to by the Dedup allocation information. The linked allocation information #2504 represents the identification number of the Dedup allocation information linked to the Dedup allocation information.

[0129] According to FIG. 25, Dedup allocation information #3 is linked to Dedup allocation information #1, and there is no Dedup allocation information linked to Dedup allocation information #3. Therefore, it can be seen that the duplicate data at the Dedup-VDEV address corresponding to Dedup allocation information #1 is the duplicate data in the SS-VDEV11S referred to by Dedup allocation information #1 and the SS-VDEV11S referred to by Dedup allocation information #3. Since the number of duplicate data is indefinite, the Dedup allocation information is linked according to the number of duplicate data. When there is duplicate data N in the SS-VDEV11S, sequential N Dedup allocation information is prepared.

[0130] FIG. 26 shows the configuration of the Pool-Mapping management table 1015.

[0131] The Pool-Mapping management table 1015 exists for each CR-VDEV11C. The Pool-Mapping management table 1015 has entries for each area in units of page size in the CR-VDEV11C. The entry has information such as the VDEV address 2601 and page #2602.

[0132] The VDEV address 2601 represents the start address of an area in units of page size (for example, a plurality of blocks). The page #2602 represents the identification number of the allocated page 14 (for example, the address in the pool 13 of page 14). When there are a plurality of pools 13, the page #2602 may include the identification number of the pool 13 having the page 14.

[0133] Figure 27 shows the configuration of the Pool allocation management table 1016.

[0134] The Pool allocation management table 1016 exists for each pool 13, for example, when there are a plurality of pools 13. The Pool allocation management table 1016 represents the correspondence between the page 14 and the area in the CR-VDEV11C. The Pool allocation management table 1016 has entries for each page 14. The entry has information such as the page #2701, RG#2702, start address 2703, status 2704, allocated VDEV#2705, and allocated address 2706.

[0135] The page #2701 represents the identification number of the page 14. The RG#2702 represents the identification number of the RAID group (in this embodiment, a RAID group composed of two or more SSDs 220) on which the page 14 is based. The start address 2703 represents the start address of the page 14. The status 2704 represents the status of the page 14 ("1" means allocated, and "0" means free). The allocated VDEV#2705 represents the identification number of the CR-VDEV11C to which the page 14 is allocated ("n / a" means unallocated). The allocated address 2706 represents the allocated address of the page 14 (the address in the CR-VDEV11C) ("n / a" means unallocated).

[0136] Hereinafter, an example of the processing performed in this embodiment will be described.

[0137] FIG. 28 shows the flow of the snapshot acquisition process. The snapshot acquisition process is executed by the snapshot acquisition program 1101 in response to a snapshot acquisition instruction from the management system 203 (or another system such as the server system 202). In the snapshot acquisition instruction, for example, the target PVOL 10P is specified.

[0138] First, the snapshot acquisition program 1101 allocates the Dir management table 1009 as the copy destination and updates the VOL-Dir management table 1004 (S2801).

[0139] The snapshot acquisition program 1101 increments the latest generation # (S2802) and updates the generation management tree table 1007 (Dir-Info generation management tree 70) (S2803). At this time, the snapshot acquisition program 1101 sets the latest generation # as the copy source and the generation # before increment as the copy destination.

[0140] The snapshot acquisition program 1101 determines whether there is cache dirty data for the target PVOL 10P (S2804). The "cache dirty data" may be data in the data stored in the cache unit 903 that has not yet been written to the pool 13.

[0141] If the determination result in S2804 is true (S2804: Yes), the snapshot acquisition program 1101 causes the snapshot append program 1106 to execute the snapshot append process (S2805).

[0142] If the determination result of S2804 is false (S2804: No), or after S2805, the snapshot acquisition program 1101 copies the Dir management table 10009 of the target PVOL10P to the Dir management table 1009 at the copy destination (S2806).

[0143] Thereafter, the snapshot acquisition program 1101 updates the snapshot management table 1003 (S2807) and ends the process. In S2807, an entry having PVOL#1401 representing the identification number of the target PVOL10P, SVOL#1402 representing the identification number of the acquired snapshot (SVOL10S), and acquisition time 1403 representing the acquisition time is added.

[0144] FIG. 29 shows the flow of the snapshot restore process. The snapshot restore process is executed by the snapshot restore program 1102 in response to a restore instruction from the management system 203 (or another system such as the server system 202). In the restore instruction, for example, the restore source SVOL and the restore destination PVOL are specified.

[0145] First, the snapshot restore program 1102 allocates the Dir management table 1009 as the restore destination and updates the VOL-Dir management table 1004 (S2901).

[0146] The snapshot restore program 1102 increments the latest generation # (S2902) and updates the generation management tree table 1007 (Dir-Info generation management tree 70) (S2903). At this time, the snapshot restore program 1102 sets the generation # before increment as the copy source and sets the latest generation # as the copy destination.

[0147] The snapshot restore program 1102 purges the cache area of the restore destination PVOL (the area in the cache unit 903) (S2904).

[0148] The snapshot restore program 1102 copies the directory management table 1009 of the source SVOL for restoration to the directory management table 1009 of the destination PVOL (S2905).

[0149] After that, the snapshot restore program 1102 registers the directory information number of the old directory information at the destination in the recovery management table 1006 (S2906) and ends the process. In S2906, the recovery request 1702 corresponding to the directory information number is set to "1".

[0150] Figure 30 shows the flow of the snapshot deletion process. The snapshot deletion process is executed by the snapshot deletion program 1103 in response to a snapshot deletion instruction from the management system 203 (or another system such as the server system 202). In the restore instruction, for example, the target SVOL is specified.

[0151] First, the snapshot deletion program 1103 refers to the VOL-Dir management table 1004 and invalidates the directory information (directory information number 1503) of the target SVOL (S3001).

[0152] Then, the snapshot deletion program 1103 updates the snapshot management table 1003 (S3002), registers the old directory information number of the target SVOL in the recovery management table 1006 (S3003), and ends the process. In S3003, the recovery request 1702 corresponding to the directory information number is set to "1".

[0153] Figure 31 shows the flow of the asynchronous recovery process. The asynchronous recovery process is executed periodically, for example, by the asynchronous recovery program 414.

[0154] First, the asynchronous recovery program 1104 identifies the Dir-Info# to be recovered from the recovery management table 1006 (S3101). The "Dir-Info# to be recovered" is the Dir-Info# for which the recovery request 1702 is "1". The asynchronous recovery program 1104 refers to the generation management tree table 1007, checks the entry of the Dir-Info# to be recovered, and does not select a Dir-Info having two or more children.

[0155] Thereafter, the asynchronous recovery program 1104 determines whether there is an unprocessed entry (S3102). The "unprocessed entry" here refers to one Mapping-Info referred to by the Dir-Info identified in S3101.

[0156] If the determination result in S3102 is true (S3102: Yes), the asynchronous recovery program 1104 determines a processing target entry (an entry including the recovery request 1702 "1") from one or more unprocessed entries (S3103), and identifies the reference destination Mapping-Info#2002 from the Dir management table 1009 corresponding to the target Dir-Info (the Dir-Info identified from Dir-Info#1701 in the processing target entry) (S3104).

[0157] The asynchronous recovery program 1104 refers to the generation management tree table 1007 and determines whether there is a Dir-Info in the child generation of the target Dir-Info (S3105).

[0158] If the determination result of S3105 is true (S3105: Yes), the asynchronous recovery program 1104 identifies the reference Mapping-Info#2002 from the Dir management table 1009 corresponding to the Dir-Info of the child generation, and determines whether the reference Mapping-Info#2002 of the target Dir-Info matches the reference Mapping-Info#2002 of the Dir-Info of the child generation (S3106). If the determination result of S3106 is true (S3106: Yes), the process returns to S3102. The entry obtained in S3103 is one entry within the Dir-Info. In contrast, the entry determined for matching in S3106 is the entry corresponding to the address within the same SVOL in the child's Dir-Info.

[0159] If the determination result of S3106 is false (S3106: No), or if the determination result of S3105 is false (S3105: No), the asynchronous recovery program 1104 determines whether the generation # of the Dir-Info of the parent generation of the target Dir-Info is older than the generation #2104 of the reference Mapping-Info of the target Dir-Info (see Figure 21) (S3107). If the determination result of S3107 is false (S3107: No), the process returns to S3102. The determination in S3107 is made for the Mapping-Info corresponding to the same address as the address within the VOL managed by the entry specified in S3103 among the entries within the target Dir-Info.

[0160] If the determination result of S3107 is true (S3107: Yes), the asynchronous recovery program 1104 initializes the target entry in the SS-Mapping management table 1010 and releases the target entry in the snapshot allocation management table 1008 (S3108). Thereafter, the process returns to S3102. The release in S3108 corresponds to the release of blocks in the SS-VDEV. S3108 corresponds to the invalidation in units of Mapping-Info. Also, for S3108, the "target entry" is the entry corresponding to the block address referred from the SS-Mapping-Info for which the determination result of S3107 is Yes.

[0161] When the determination result of S3102 is false (S3102: No), the asynchronous recovery program 1104 updates the recovery management table 1006 (S3109), and also updates the generation management tree table 1007 (Dir-Info generation management tree 70) (S3110), and ends the process. S3109 is the recovery of Dir-Info, and the recovery request 1702 is updated from "1" to "0". When the Mapping-Info referred to by the Dir-Info to be recovered is also referred to by other Dir-Info, the Mapping-Info remains and the target Dir-Info is recovered. In S3110, the invalidated (recovered) Dir-Info is pulled out of the tree, and the connection relationship of the tree is updated.

[0162] FIG. 32 shows the flow of the Write process (front end). The Write process (front end) is executed by the read / write program 1105 when a write request from the server system 202 is received.

[0163] First, the read / write program 1105 determines whether the target data of the write request causes a cache hit (S3201). "Cache hit" means that a cache area corresponding to the write destination VOL address (the VOL address specified in the write request) of the target data has been secured. When the determination result of S3201 is false (S3201: No), the read / write program 1105 secures a cache area corresponding to the write destination VOL address of the target data from the cache unit 903 (S3202). Then, the process proceeds to S3206.

[0164] When the determination result of S3201 is true (S3201: Yes), the read / write program 1105 determines whether the cached data (data in the secured cache area) is dirty data (data not reflected (not written) in the pool 13) (S3203). When the determination result of S3203 is false (S3203: No), the process proceeds to S3206.

[0165] When the determination result of S3203 is true (S3203: Yes), the read / write program 1105 determines whether the WR (Write) generation # of the dirty data matches the generation # of the data targeted by the write request (S3204). The "WR generation #" is the latest generation # of the snapshot when the data was written to the cache, and is held, for example, in the management information (not shown) of the cache data. Also, the generation # of the data targeted by the write request is obtained from the latest generation #403. S3204 is a process for preventing the dirty data from being updated with the data targeted by the write request and overwriting the snapshot data before the append process for the target data (dirty data) of the immediately preceding snapshot is completed. If data written before the host write (write following the write request from the server system 202) already exists in the cache and a snapshot is taken before the host write, the data in the cache becomes the snapshot data. In this state, when the host write request is received, the WR generation # and the latest generation # will not match.

[0166] When the determination result of S3204 is false (S3204: No), the read / write program 1105 causes the snapshot append program 1106 to execute the snapshot append process (S3205).

[0167] After S3202, or when the determination result of S3204 is true (S3204: Yes), the read / write program 1105 writes the data targeted by the write request to the cache area secured in S3202 or to the cache area obtained through S3205 (S3206). After that, the read / write program 1105 sets the WR generation # of the data written in S3206 to the latest generation # compared in S3204 (S3207), and returns a normal response (Good response) to the server system 202 (S3208).

[0168] Figure 33 shows the flow of the Write process (backend). The Write process (backend) is a process of writing unreflected data (dirty data) to the pool 13 when the unreflected data is on the cache unit 903. The Write process (backend) is performed synchronously or asynchronously with the Write process (frontend). The Write process (backend) is executed by the read / write program 1105.

[0169] The read / write program 1105 determines whether there is dirty data on the cache unit 903 (S3301). If the determination result of S3301 is true (S3301: Yes), the read / write program 1105 causes the snapshot append program 1106 to execute the snapshot append process (S3302).

[0170] Figure 34 shows the flow of the snapshot append process. The snapshot append process is executed by the snapshot append program 1106 called from the snapshot acquisition program 1101 or the read / write program 1105.

[0171] The snapshot append program 1106 secures a new area (block address with status 1902 being "0") in the target SS-VDEV11S (SS-VDEV11S corresponding to SS-Family 9 including the target VOL (e.g., the SVOL to be acquired or the VOL where data is written)) by updating the snapshot allocation management table 1008 (S3401). Then, the snapshot append program 1106 causes the Dedup append program 1107 to execute the Dedup append process (S3402).

[0172] Thereafter, the snapshot append program 1106 updates the SS-Mapping management table 1010 (S3403). In S3403, for example, the snapshot append program 1106 sets the latest generation # (the generation # represented by the latest generation table 1005) to the generation #2104 corresponding to the Mapping-Info# of the target SS-Mapping-Info. The "target SS-Mapping-Info" mentioned here is the SS-Mapping-Info corresponding to the data in the target VOL.

[0173] The snapshot append program 1106 updates the Dir management table 1009 corresponding to the Dir-Info of the target SS-VDEV11S (S3404). In S3404, the SS-Mapping-Info (information representing the reference address in the SS-VDEV) for the data to be written is associated with the address in the VOL10 of the data.

[0174] The snapshot append program 1106 refers to the generation management tree table 1007 (Dir-Info generation management tree 70) (S3405) and determines whether the generation # of the Dir-Info of the target VOL (the write destination VOL of the data) matches the generation # of the SS-Mapping-Info before the append (S3406). The "SS-Mapping-Info before the append" refers to the Mapping-info that manages the data before the update (that is, among the Mapping-info referred to by the Dir-Info, the SS-Mapping-Info corresponding to the data before the update is the determination target in S3406). "Before the append" means before the snapshot append process is performed.

[0175] When the determination result of S3406 is true (S3406: Yes), the snapshot append program 1106 initializes the target entry in the SS-Mapping management table 1010 before appending (for example, sets an invalid value in the target entry), releases the target entry in the snapshot allocation management table 1008 (specifically, the entry corresponding to the block address represented by the reference address 2102 in the SS-Mapping management table 1010 before appending) (S3407), and ends the process. By performing S3407 when S3406: Yes, the unreferenced area can be garbage-collected and made reusable.

[0176] Figure 35 shows the flow of the Dedup append process. The Dedup append process is executed by the Dedup append program 1107 called from the snapshot append program 1106.

[0177] The Dedup append program 1107 determines whether duplicate data exists (S3501). When the determination result of S3501 is false (S3501: No), the Dedup append program 1107 causes the compression append program 1108 to execute the compression append process (S3508). As a result, the compressed data of the data whose storage destination is SS-VDEV11S is stored in CR-VDEV11C without passing through Dedup-VDEV11D and without the CPU 211 that is the processing subject being changed. Note that when S3501: No, as indicated by the dashed arrow, after S3508, the Dedup append process ends. Also, the Dedup append program 1107 manages the directory of the hash values of the data stored in each block in Dedup-VDEV11D, and in S3501, it may determine whether a hash value that matches the hash value of the written data exists in the directory.

[0178] When the determination result of S3501 is true (S3501: Yes), the Dedup append program 1107 causes the CPU determination program 1111 to execute CPU determination processing (S3502). S3502 is implemented because the CPU 211 executing the Dedup append program 1107 may not match the owner CPU of the target Dedup-VDEV 11D (the Dedup-VDEV 11D that stores duplicate data). If the CPU 211 executing the Dedup append program 1107 does not match the owner CPU of the target Dedup-VDEV 11D, the process is taken over by the CPU 211 as the owner CPU from the CPU 211. That is, S3503 to S3507 are performed by the Dedup append program 1107, but the CPU 211 executing the Dedup append program 1107 may be the same CPU 211 as the CPU 211 that performed S3501 or a different CPU 211. In other words, the CPU 211 that performs S3503 to S3507 is the owner CPU of the Dedup-VDEV 11D where the data is stored.

[0179] After S3502, the Dedup append program 1107 updates the Dedup allocation management table 1014 (S3503). In S3503, an entry for the Dedup allocation information corresponding to the data for which the target Dedup-VDEV 11D is the storage destination is added to the Dedup allocation management table 1014.

[0180] After S3503, S3508 is performed. Then, the Dedup append program 1107 updates the CR-Mapping management table 1012 (S3504). In S3504, an entry for the CR-Mapping-Info corresponding to the data for which the Dedup-VDEV 11D is the storage destination is added to the CR-Mapping management table 1012.

[0181] The Dedup append program 1107 updates the Dir management table 1009 corresponding to the target Dedup-VDEV 11D (S3505). In S3505, the CR-Mapping-Info (information representing the reference address in CR-VDEV 11CC) for duplicate data is associated with the address in the target Dedup-VDEV 11D of the data.

[0182] The Dedup append program 1107 initializes the CR-Mapping management table 1012 before appending (S3506). Note that the tables to be updated in S3504 and S3506 are different. Specifically, in S3504, the entry indicating the address where data will be written is the update target, and in S3506, the entry indicating the address where the data before the update is stored is the update target. Until S3506 is executed, there are two pieces of CR-Mapping-Info. By S3505, the connection destination of the Dir-Info switches from the CR-Mapping-Info pointing to the old address to the CR-Mapping-Info pointing to the new address. After that, in S3506, the CR-Mapping-Info pointing to the old address is released, and the released CR-Mapping-Info becomes reusable as the CR-Mapping-Info in another area.

[0183] The Dedup append program 1107 invalidates the pre-update allocation information (S3507). The invalidation in S3507 is a process for releasing the VDEV area (the area in the VDEV) to which the pre-update data was allocated. In S3507, the Dedup append program 1107 updates the Dedup allocation management table 1014 and the compression allocation management table 1011. If the destination of the pre-update data is Dedup-VDEV11D, the entry in the Dedup allocation management table 1014 (the entry for the destination of the pre-update data) is released. If the destination of the pre-update data is CR-VDEV11C, the entry in the pre-update data of the compression allocation management table 1011 (the entry for the destination of the pre-update data) is released and garbage-collected. Also, in S3507, for the area where the number of destinations in the Dedup allocation management table 1014 becomes zero (the state where the value of the reference destination allocation information #2402 corresponding to the Dedup-VDEV address 2401 is not registered (invalid value)), the target entry of the compression allocation information is also garbage-collected (by updating the status 2203 in the compression allocation management table 1011 to "2").

[0184] Figure 36 shows the flow of the compression append process. The compression append process is executed by the compression append program 1108 called from the Dedup append program 1107.

[0185] The compression append program 1108 compresses the data to be written (S3601). The compression append program 1108 updates the compression allocation management table 1011 (S3602). In S3602, for each of the one or more sub-blocks that are the storage destinations of the compressed data in S3601, the entry corresponding to the sub-block is updated.

[0186] The compression append program 1108 causes the destage program 1109 to execute the destage process (S3603).

[0187] The compression append program 1108 updates the CR-Mapping management table 1012 (S3604). The CR-Mapping management table 1012 has an entry indicating the reference destination of the data before the update and an entry indicating the reference destination after the update, and the entry is switched by the Dir management table 1009. The reference destination address where the updated data is stored is registered in the empty entry of the CR-Mapping management table 1012.

[0188] The compression append program 1108 updates the Dir management table 1009 (S3605). In the case of S3501: No, the Dir management table 1009 corresponding to the SS-VDEV is updated. In the case of S3501: Yes, the Dir management table 1009 corresponding to the Dedup-VDEV is updated.

[0189] The compression append program 1108 invalidates the pre-update allocation information (S3607). In S3607, the Dedup append program 1107 updates the Dedup allocation management table 1014 and the compression allocation management table 1011. The invalidation in S3607 is a process for releasing the area where the data before the update is stored and making the released area allocable when other data is stored. If the data before the update is mapped to the Dedup-VDEV 11D, the release destination is the Dedup allocation management table 1014. If the data before the update is mapped to the CR-VDEV 11C, the release destination is the compression allocation management table 1011. Also, in S3607, for the area where the number of allocation destinations in the Dedup allocation management table 1014 becomes 0 (the state where the value of the reference destination allocation information #2402 corresponding to the Dedup-VDEV address 2401 is not registered (invalid value)), the target entry of the compression allocation information is also garbage-collected (by updating the status 2203 of the compression allocation management table 1011 to "2").

[0190] Figure 37 shows the flow of the destaging process. The destaging process is executed by the destaging program 1109 called from the compression append program 1108.

[0191] The destage program 1109 determines whether the append data (one or more compressed data) for the RAID stripe is in the cache unit 903 (S3701). A "RAID stripe" is a stripe in a RAID group (a storage area spanning multiple SSDs 220 that make up the RAID group). When the RAID level of the RAID group requires parity, the size of the "append data for the RAID stripe" may be the size obtained by subtracting the size of the parity from the size of the stripe. If the determination result of S3701 is false (S3701: No), the process ends.

[0192] If the determination result of S3701 is true (S3701: Yes), the destage program 1109 refers to the Pool-Mapping management table 1015 and determines whether page 14 has been allocated to the storage destination (the address in CR-VDEV11C) of the append data for the RAID stripe (S3702). If the determination result of S3702 is false (S3702: No), the process proceeds to S3705.

[0193] If the determination result of S3702 is true (S3702: Yes), the destage program 1109 updates the Pool allocation management table 1016 (S3703). Specifically, the destage program 1109 allocates page 14. In S3703, among the Pool allocation management table 1016, the entry corresponding to the allocated page 14 (for example, status 2704, allocation destination VDEV #2705, and allocation destination address 2706) is updated.

[0194] The destage program 1109 registers the page #2602 of the allocated page in the entry corresponding to the storage destination of the append data for the RAID stripe in the Pool allocation management table 1016 (S3704).

[0195] The destage program 1109 writes the appended data for the RAID stripe to the stripe that is the basis of the page (S3705). When the RAID level is a RAID level that requires parity, the destage program 1109 generates parity based on the appended data for the RAID stripe and also writes the parity to the stripe.

[0196] FIG. 38 shows the flow of the GC (garbage collection) process. The GC process is executed by the GC program 1110, for example, periodically (or in response to an instruction from the management system 203).

[0197] The GC program 1110 refers to the Pool-Mapping management table 1015 and the compression allocation management table 1011 to identify the pages having sub-blocks in the garbage state (status 2203 “2”) (S3801). If there are no pages having sub-blocks in the garbage state, the GC process may end. Also, in S3801, the GC program 1110 may preferentially select the CR-VDEV11C having the least free space among the plurality of CR-VDEVs 11C. Also, in S3801, the GC program 1110 may preferentially identify the page having the most sub-blocks in the garbage state among the CR-VDEVs 11C. Also, the GC may be performed in a unit of area different from that of page 14.

[0198] The GC program 1110 determines whether there are unprocessed (not yet determined in S3803) sub-blocks in the page identified in S3801 (S3802).

[0199] When the determination result of S3802 is true (S3802: Yes), the GC program 1110 refers to the compression allocation management table 1011 to determine the sub-block to be processed (S3803). The GC program 1110 determines whether the status 2203 corresponding to the sub-block to be processed is "1" (allocated) (S3804). When the determination result of S3804 is false (S3804: No), the process returns to S3802. When the determination result of S3804 is true (S3804: Yes), the GC program 1110 appends the sub-block to be processed to another area (S3805). This "another area" may be a free sub-block (a sub-block with status 2203 being "0") in a CR-VDEV11C different from the CR-VDEV11C (the CR-VDEV11C having the sub-block before the allocation destination of the page specified in S3801) which is the target of the GC process. Also, this "different CR-VDEV11C" may be a CR-VDEV11C where all sub-blocks are flow sub-blocks. Also, page 14 may be allocated to this "another area", and the compressed data in the sub-block to be processed may be written to the page 14 (in other words, the compressed data may be moved from the page allocated to the sub-block to be processed to the page allocated to another area).

[0200] When the determination result of S3802 is false (S3802: No), the GC program 1110 updates all entries of the compression allocation management table 1011 corresponding to the CR-VDEV11C which is the target of the GC process (S3806). In S3806, for example, the status 2203 of all entries becomes "0".

[0201] Also, the GC program 1110 updates the Pool-Mapping management table 1015 and the Pool allocation management table 1016 (S3807). In S3806, for example, the page #2602 of all entries of the Pool-Mapping management table 1015 corresponding to the CR-VDEV11C which is the target of the GC process is initialized, and the status 2704 corresponding to all pages allocated to the CR-VDEV11C which is the target of the GC process may be set to "0" (free).

[0202] Thus, the GC process according to this embodiment may transfer valid compressed data (compressed data in the allocated sub-blocks) between CR-VDEV11Cs to make a plurality of allocated sub-blocks in a discontinuous state continuous. Note that making a plurality of allocated sub-blocks in a discontinuous state continuous may be performed without data transfer between CR-VDEV11Cs.

[0203] FIG. 39 shows the flow of the CPU determination process. The CPU determination process is executed by a CPU determination program 1111 called from a Dedup append program 1107.

[0204] The CPU determination program 1111 refers to the owner right management table 1001 and determines whether the owner CPU of the Dedup-VDEV11D to be processed is the own CPU211 (the CPU211 that made the determination in S3501) (S3901).

[0205] If the determination result in S3901 is false (S3901: No), the CPU determination program 1111 immediately passes the process of the own CPU to the owner CPU (S3902). As a result, the owner CPU takes over the process, and as a result, the CPU that performs the processes after S3503 in FIG. 35 becomes the owner CPU (the CPU to which the process is taken over) of the Dedup-VDEV11D.

[0206] The CPU determination process may be performed in processes other than the Dedup append process described in FIG. 35. However, in this embodiment, in the process from VOL10 to CR-VDEV11C, the CPU determination process may be performed only when the owner CPUs may be different, specifically, only when Dedup-VDEV11D is the write destination. In other words, when the write destinations are VOL10, SS-VDEV11S, and CR-VDEV11C, since their owner CPUs are the same, the CPU determination process may not be performed. As a result, an improvement in the performance of the write process is expected.

[0207] Figure 40 shows the flow of the owner right transfer process. The owner right transfer process is performed by the owner right transfer program 1112, for example, in response to an instruction from the management system 203 or when a failure occurs in the CPU 211. The CPU 211 that executes the owner right transfer program 1112 may be any normal CPU 211 or the CPU 211 with the lowest load.

[0208] The owner right transfer program 1112 acquires all the SVOL #1402 within the target SS-Family 9 from the snapshot management table 1003 (S4001). The "target SS-Family 9" is the SS-Family 9 that has the target PVOL 10P. The "target PVOL 10P" may be the PVOL 10P specified in the instruction from the management system 203 or the PVOL 10P for which the CPU 211 in which a failure has occurred is the owner CPU.

[0209] Also, the owner right transfer program 1112 refers to the Dir management table 1009, the CR-VDEV management table 1002, and the CR-Mapping management table 1012, and acquires the CR-VDEV #1302 of all the CR-VDEV 11C related to the target PVOL 10P (S4002). The "CR-VDEV 11C related to the target PVOL 10P" is the CR-VDEV 11C whose allocation destination is the SS-VDEV 11S of the target SS-Family 9.

[0210] The owner right transfer program 1112 updates the owner CPU #1202 of the SVOL 10S of the SVOL # acquired in S4001, the CR-VDEV 11C of the CR-VDEV # acquired in S4002, the target PVOL 10P, and the SS-VDEV 11S of the target SS-Family 9 (S4003). The updated owner CPU #1202 has the same identification number. That is, the owner rights of the SVOL 10S of the SVOL # acquired in S4001, the CR-VDEV 11C of the CR-VDEV # acquired in S4002, the target PVOL 10P, and the SS-VDEV 11S of the target SS-Family 9 are all held by the same CPU. [Embodiment 2]

[0211] Embodiment 2 will be described. In doing so, the differences from Embodiment 1 will be mainly described, and the description of the common points with Embodiment 1 will be omitted or simplified.

[0212] In Embodiment 2, SVOL10S may be a writable snapshot. Hereinafter, a writable SVOL (snapshot) will be abbreviated as a "WR-SVOL", and a Read Only SVOL (snapshot) will be referred to as a "RO-SVOL".

[0213] Figure 41 shows an overview of the acquisition of a WR-SVOL.

[0214] When creating a WR-SVOL, the processor prepares RO-Dir-Info and R / W-Dir-Info for the WR-SVOL. RO-Dir-Info is Dir-Info that prohibits writing (Read Only). R / W-Dir-Info is Dir-Info that permits writing (for Read / Write).

[0215] The processor sets the latest generation # before snapshot creation (the generation # represented by the latest generation table 1005) as the generation # of RO-Dir-Info, and sets the generation # obtained by incrementing the generation # of RO-Dir-Info as the generation # of R / W-Dir-Info. The processor sets the generation # obtained by incrementing the generation # of R / W-Dir-Info as the latest generation # and the generation # of the Dir-Info of PVOL.

[0216] According to the example shown in Figure 41, the latest generation # before snapshot creation is "0". Therefore, the generation # of RO-Dir-Info is "0", the generation # of R / W-Dir-Info is "1", and the latest generation # after snapshot creation and the generation # of the Dir-Info of PVOL are both "2".

[0217] In addition, as the Dir-Info generation management tree 70, the RO-SVOL (Read Only SVOL) corresponding to generation 0 is the parent, and the RW-SVOL corresponding to generation 1 and the PVOL of generation 2 are the children.

[0218] Figure 42 shows an overview of writing to the WR-SVOL.

[0219] When rewriting data A0 to data A1 in the WR-SVOL, the processor secures a new area in the SS-VDEV and uses this area as the storage destination for data A1. For the new data A1 in the SS-VDEV, the processor generates new SS-Mapping-Info and associates generation information representing the generation # of the WR-SVOL with the SS-Mapping-Info of data A1. Therefore, the generation # of the SS-Mapping-Info of data A1 becomes "1".

[0220] The processor associates the write destination address in the PVOL and the data A1 to be written by switching the reference relationship (corresponding relationship) between the Dir-Info and SS-Mapping-Info of generation 1.

[0221] By this switching of the reference destination, the SS-Mapping-Info of data A0 is no longer referenced from generation 1. However, the SS-Mapping-Info of data A0 remains referenced from the RO-Dir-Info (generation 0) of the WR-SVOL. Therefore, the SS-Mapping-Info of data A0 should not be invalidated.

[0222] The processor makes a determination on whether to invalidate this. This determination includes a comparison between the generation # of the Mapping-Info to be determined for invalidation and the generation # of the R / W-Dir-Info of the write destination VOL (here, the WR-SVOL). If those generation #s match, the processor determines that invalidation is possible. On the other hand, if the generation # of the Mapping-Info is older, the processor determines that invalidation is not possible.

[0223] Figure 43 shows an overview of the restore from the WR-SVOL.

[0224] When performing a restore from the WR-SVOL to the PVOL (e.g., the PVOL of generation 2), the processor newly creates RO-Dir-Info and R / W-Dir-Info for the destination PVOL. Both the RO-Dir-Info and the R / W-Dir-Info are copies of the R / W-Dir-Info of the WR-SVOL.

[0225] The generation # of the new RO-Dir-Info is the same generation # as the generation # of the R / W-Dir-Info of the source WR-SVOL. On the other hand, the generation # of the new R / W-Dir-Info is the generation # obtained by incrementing the generation # of the destination PVOL by two. The generation # of the R / W-Dir-Info of the source WR-SVOL is the generation # obtained by incrementing the original generation # by two.

[0226] As a result, as illustrated in Figure 43, the latest generation # becomes "4" as a result of incrementing the generation # of the destination PVOL by two.

[0227] The generation # of the source RO-Dir-Info is "0". The generation # of the destination RO-Dir-Info is "1", and the generation # of the old Dir-Info of the destination is "2", both of which are children of generation # "0".

[0228] The generation # of the source R / W-Dir-Info becomes "3", and the generation # of the destination R / W-Dir-Info becomes "4", both of which are children of generation # "1".

[0229] By the restore, the old Dir-Info of generation 2 is detached from the correspondence with the PVOL and becomes an object for asynchronous recovery with no reference from the PVOL or the SVOL (snapshot). That is, the processor invalidates this generation 2 Dir-Info. Also, the processor designates generation # "2" as the generation to be invalidated.

[0230] Although several embodiments have been described above, these are examples for explaining the present invention and are not intended to limit the scope of the present invention only to these embodiments. The present invention can be implemented in various other forms.

[0231] In addition, the above description can be summarized as follows, for example. The following summary may include supplementary explanations or explanations of modified examples of the above description.

[0232] In a storage system 201 including a storage device and a processor, the storage device stores first mapping information and second mapping information.

[0233] The first mapping information includes information representing the mapping between the address of VOL10 in each SS-Family9 and the address in SS-VDEV11S. For example, the first mapping information includes Dir-Info (an example of the first control information) for each VOL10 and SS-Mapping-Info (an example of the second control information) for each data in SS-VDEV11S. The Dir-Info for each VOL is associated with the generation # of the VOL and represents which SS-Mapping-Info to refer to for the address in the VOL. The SS-Mapping-Info for each data is associated with the oldest generation # among the generation #s of the Dir-Info that refers to the SS-Mapping-Info and represents the address where the data is located in SS-VDEV11S.

[0234] The second mapping information includes information representing the mapping between the address in SS-VDEV11S and the address in Dedup-VDEV10D. For example, the second mapping information includes Dir-Info (an example of the third control information) prepared for each of SS-VDEV11S0 and Dedup-VDEV11D, and CR-Mapping-Info (an example of the fourth control information) that is associated on a one-to-one basis with the data stored in SS-VDEV11S or Dedup-VDEV10D.

[0235] For each of the plurality of SS-Family9, there is one or more SS-VDEV11S. Each SS-VDEV11S is a logical address space (virtual device) that serves as the storage destination for the data of VOL10 in the SS-Family9 corresponding to the SS-VDEV11S. Dedup-VDEV10D is a logical address space (virtual device) different from the SS-VDEV11S.

[0236] When there is the same data in a plurality of VOL10 of SS-Family9, the processor updates the first mapping information so as to map the plurality of addresses of the same data among the plurality of VOL10 to the address of the SS-VDEV11S of the SS-Family9. When there is duplicate data in two or more SS-VDEV11S of two or more SS-Family9, the processor updates the second mapping information so as to map the two or more addresses of the duplicate data among the two or more SS-VDEV11S to the address corresponding to the duplicate data in Dedup-VDEV10D.

[0237] Thereby, even if the address of the duplicate data is changed, the address mapping can be changed in a short time.

[0238] The second mapping information may further include the following information. · A pair of the Dir-Info of the SS-VDEV11S and the CR-Mapping-Info that refers to the address in the CR-VDEV11C (an example of an append virtual device) referred to from the Dir-Info. This pair is an example of information representing the mapping between the address in the SS-VDEV11S and the address in the CR-VDEV11C. · A pair of the Dir-Info of the Dedup-VDEV10D and the CR-Mapping-Info that refers to the address in the CR-VDEV11C referred to from the Dir-Info. This pair is an example of information representing the mapping between the address in the Dedup-VDEV10D and the address in the CR-VDEV11C.

[0239] Each CR-VDEV11C may have a logical address space (virtual device) corresponding to either the SS-VDEV11S or the Dedup-VDEV10D. Each CR-VDEV11C may be the storage destination of the data whose virtual device corresponds to the CR-VDEV11C, and may not be the storage destination of the data whose virtual device does not correspond to the CR-VDEV11C.

[0240] The processor may store the data for which the CR-VDEV11C is the storage destination in the pool 13. The pool 13 may be a logical address space based on at least one of at least a part of the storage device of the storage system 201 and at least a part of the external storage device of the storage system 201. When the data to be stored for the CR-VDEV11C is updated data, the processor may use the free address in the CR-VDEV11C as the storage destination of the updated data, invalidate the address of the data before the update of the updated data, and update the second mapping information so as to map the address mapped to the address of the data before the update and being an address in the SS-VDEV11S or the Dedup-VDEV10D to the storage destination address of the updated data in the CR-VDEV11C.

[0241] The mapping between the address in the CR-VDEV11C and the address in the pool 13 may be 1:1.

[0242] Thereby, even if the address of the compressed data of the duplicate data in the CR-VDEV11C is changed (for example, due to a GC process), the address mapping can be changed in a short time.

[0243] The processor may be a plurality of CPUs 211 (an example of a plurality of processor devices). When the CPU 211 corresponds to the owner CPU of the VOL or VDEV, the CPU 211 may perform I / O for the VOL or the VDEV and update the address mapping information of the VOL or the VDEV among the first mapping information and the second mapping information. The owner CPU of SS-Family9, the owner CPU of SS-VDEV11S for the SS-Family9, and the owner CPU of CR-VDEV11C corresponding to the SS-VDEV11S may be the same CPU 211. Thereby, in the write process of non-duplicate data, since the owner CPU is consistently the same CPU, the transfer of ownership between the CPU 211s (communication for process handover between the CPU 211s) is unnecessary.

[0244] The owner CPU of Dedup-VDEV10D and the owner CPU of CR-VDEV11CC corresponding to the Dedup-VDEV10D may be the same CPU 211. Thereby, in the writing to the Dedup-VDEV10D and the corresponding CR-VDEV11CC, the transfer of ownership between the CPU 211s is unnecessary.

[0245] When Dedup-VDEV10D is the write destination, each CPU 211 may perform a CPU determination as to whether the CPU 211 corresponds to the owner CPU of the Dedup-VDEV10D. In other words, when Dedup-VDEV10D is not the write destination, the CPU 211 may not perform the CPU determination. Thereby, an improvement in processing performance can be expected.

[0246] In the process of a write request designating any PVOL10P, when the following conditions (a) and (b) are satisfied, the processor may perform the process in the snapshot append process (for example, writing to the CR-VDEV11C associated with the SS-VDEV11S of the SS-Family9 including the PVOL10P, or writing to the Dedup-VDEV10D). (a) The cache area corresponding to the address specified in the write request is in the cache unit 903, and the data in the cache area is dirty data that is not reflected in the pool 13. (b) The generation of the dirty data is different from the latest generation of the PVOL10P.

[0247] The processor may determine, for each SS-Family9, whether to invalidate the Dir-Info and / or SS-Mapping-Info based on the generation of the Dir-Info and the generation of the SS-Mapping-Info, asynchronously with the processing of the I / O request of the VOL in the SS-Family9, and invalidate the Dir-Info and / or SS-Mapping-Info for which invalidation is determined to be possible. Even without the meta information of the reverse reference system (for example, the reference information from the pool 13 to the VOL10), it is possible to efficiently determine whether invalidation is possible.

[0248] The processor may invalidate the SS-Mapping-Info referred to by the target Dir-Info when the following conditions (x) and (y) are satisfied. (x) The generation of the SS-Mapping-Info referred to by the Dir-Info of the generation one newer than the target Dir-Info does not match the generation of the SS-Mapping-Info referred to by the target Dir-Info. (y) The generation of the Dir-Info of the generation one older than the target Dir-Info is older than the generation of the SS-Mapping-Info referred to by the target Dir-Info.

[0249] When creating a WR-SVOL, the processor may create RO-Dir-Info and R / W-Dir-Info for the SVOL, set the latest generation before the creation of the SVOL as the generation of the RO-Dir-Info, and set the generation obtained by incrementing the generation of the RO-Dir-Info as the generation of the R / W-Dir-Info.

[0250] The processor may compare the generation of the SS-Mapping-Info with the generation of the R / W-Dir-Info of the WR-SVOL. When the generation of the SS-Mapping-Info matches the generation of the R / W-Dir-Info, the processor may invalidate the SS-Mapping-Info. When the generation of the SS-Mapping-Info is older than the generation of the R / W-Dir-Info, the processor may not invalidate the SS-Mapping-Info.

[0251] When the processor performs a restore from the WR-SVOL to the PVOL, it may create the RO-Dir-Info and R / W-Dir-Info for the destination PVOL as a copy of the R / W-Dir-Info of the WR-SVOL. At that time, the processor may set the generation of the RO-Dir-Info for the destination PVOL to the generation of the R / W-Dir-Info of the WR-SVOL, set the generation of the R / W-Dir-Info for the destination PVOL to two generations newer than the generation of the destination PVOL, and set the generation of the R / W-Dir-Info for the WR-SVOL to two generations newer than the original generation of the R / W-Dir-Info. If the original Dir-Info of the destination PVOL is not referenced from any PVOL and SVOL, the processor may invalidate the Dir-Info.

[0252] When the generation of the Dir-Info to be invalidated is specified as the target generation, the processor may determine whether to invalidate it based on the reference status in the generation one older and the reference status in the generation one newer in the direct line of the target generation. This allows the determination of whether to invalidate without having to look at all generations.

[0253] When the processor creates the SVOL10S, it may set the latest generation before creation as the generation of the SVOL10S and increment the latest generation by one. When performing a restore from the SVOL10S to any PVOL10P, the processor may also increment the latest generation.

[0254] When the processor writes to PVOL10P, it uses a new area of SS-VDEV11S as the storage destination for the data to be written, and by switching the correspondence relationship between Dir-Info and SS-Mapping-Info, it associates the write destination address in PVOL10P with the data to be written, associates the generation of PVOL10P with SS-Mapping-Info, and may target for invalidation the SS-Mapping-Info whose correspondence relationship with Dir-Info has been eliminated by the switching of the correspondence relationship. If the generation of the SS-Mapping-Info targeted for invalidation matches the generation associated with the Dir-Info of PVOL10P, the processor may determine that invalidation is possible. Thereby, it is possible to determine the invalidation of the already stored data at the timing of the write process, and it is possible to efficiently determine whether invalidation is possible. When the processor invalidates SS-Mapping-Info, it may invalidate the data referenced from the SS-Mapping-Info.

[0255] When the processor performs a restore from SVOL10S to PVOL10P, it associates a copy of the Dir-Info of the source SVOL10S of the restore with PVOL10P, increments the latest generation, and may specify the generation of the Dir-Info associated with PVOL10P before the restore as the target generation for invalidation. For the SS-Mapping-Info associated with the Dir-Info of the target generation, if the generation associated with the SS-Mapping-Info is newer than the generation that is one generation older in the direct line of the target generation and is not referenced from the generation that is one generation newer in the direct line of the target generation, the processor may determine that invalidation is possible.

[0256] Further, when the processor deletes SVOL10S, it may identify the generation associated with the Dir-Info of the deleted SVOL10S as the target generation for invalidation. Regarding the SS-Mapping-Info associated with the Dir-Info of the target generation, when the generation associated with the SS-Mapping-Info is newer compared to the generation that is one older in the direct line of the target generation and is not also referenced from the generation that is one newer in the direct line of the target generation, the processor may determine that invalidation is possible.

Explanation of Signs

[0257] 201: Storage System

Claims

1. A storage system comprising a memory device and a processor, wherein the memory device stores first mapping information and second mapping information, the first mapping information includes information representing a mapping between an address of a PVOL (Primary Volume) or an SVOL (Secondary Volume), which is a VOL (Volume) in a snapshot family composed of a PVOL and a snapshot of the PVOL, and an address in a snapshot virtual device for each of a plurality of snapshot families, the second mapping information includes information representing a mapping between an address in a snapshot virtual device and an address in a deduplication virtual device, for each of the plurality of snapshot families, there is one or more snapshot virtual devices, each snapshot virtual device is a virtual device as a logical address space serving as a storage destination of data of a VOL in a snapshot family corresponding to the snapshot virtual device, the deduplication virtual device is a virtual device as a logical address space different from the snapshot virtual device, when there is the same data in a plurality of VOLs of a snapshot family, the processor updates the first mapping information so as to map a plurality of addresses of the same data among the plurality of VOLs to an address of the snapshot virtual device of the snapshot family, when there is duplicate data in two or more snapshot virtual devices of two or more snapshot families, the processor updates the second mapping information so as to map two or more addresses of the duplicate data among the two or more snapshot virtual devices to an address corresponding to the duplicate data in the deduplication virtual device, A storage system.

2. The second mapping information further includes the following information: information representing a mapping between an address in a snapshot virtual device and an address in an append virtual device, and information representing a mapping between an address in a deduplication virtual device and an address in an append virtual device. Each of the plurality of append virtual devices is a virtual device as a logical address space corresponding to either a snapshot virtual device or a deduplication virtual device, each append virtual device serves as a storage destination for data whose storage destination is the virtual device corresponding to the append virtual device, and does not serve as a storage destination for data whose storage destination is a virtual device not corresponding to the append virtual device, the processor is configured to store data whose storage destination is an append virtual device in a pool, the pool is a logical address space based on at least one of at least a part of the storage device and at least a part of an external storage device of the storage system, when the storage target for an append virtual device is updated data, the processor sets the storage destination of the updated data to an empty address in the append virtual device, invalidates the address of the data before the update of the updated data, and updates the second mapping information so that an address mapped to the address of the data before the update and being an address in a snapshot virtual device or a deduplication virtual device is mapped to the storage destination address of the updated data in the append virtual device, the mapping between the address in the append virtual device and the address in the pool is 1:1, The storage system according to claim 1.

3. In the garbage collection of the append virtual device corresponding to the deduplication virtual device, the processor changes the address of the data whose storage destination is the append virtual device, and updates the second mapping information so that an address mapped to the address before the change and being an address in the deduplication virtual device is mapped to the address after the change in the append virtual device, The storage system according to claim 2.

4. The storage target for each append virtual device is compressed data. The storage system according to claim 2.

5. The processor is a plurality of processor devices, Each of the plurality of processor devices is configured to perform I / O operations on the VOL or the virtual device and update the address mapping information for the VOL or the virtual device among the first mapping information and the second mapping information when the processor device corresponds to the owner processor device of the VOL or the virtual device. The owner processor device of the snapshot family, the owner processor device of the snapshot virtual device for the snapshot family, and the owner processor device of the append virtual device corresponding to the snapshot virtual device are the same processor device. The storage system according to claim 2.

6. The owner processor device of the deduplication virtual device and the owner processor device of the append virtual device corresponding to the deduplication virtual device are the same processor device. The storage system according to claim 5.

7. Each of the plurality of processor devices when the deduplication virtual device is the write destination, performs a processor device determination as to whether the processor device corresponds to the owner processor device of the deduplication virtual device, when the deduplication virtual device is not the write destination, does not perform the processor device determination. The storage system according to claim 5.

8. The processor is a plurality of processor devices, each of the plurality of processor devices is configured to perform I / O operations on the VOL or the virtual device and update the address mapping information for the VOL or the virtual device among the first mapping information and the second mapping information when the processor device corresponds to the owner processor device associated with the VOL or the virtual device, the owner processor device of the snapshot family and the owner processor device of the snapshot virtual device for the snapshot family are the same processor device. The storage system according to claim 1.

9. The storage device includes a cache portion in which data is temporarily stored. When processing a write request that designates any PVOL, if the following conditions (a) and (b) are satisfied, the processor writes to the append virtual device associated with the snapshot virtual device of the snapshot family including the PVOL, or writes to the deduplication virtual device: (a) The cache area corresponding to the address specified in the write request is in the cache unit, and the data in the cache area is dirty data that is data not yet reflected in the pool. (b) The generation of the dirty data is different from the latest generation of the PVOL. The storage system according to claim 2.

10. The first mapping information includes first control information for each VOL in the plurality of snapshot families and second control information associated with the data in the snapshot virtual device for each of the plurality of snapshot families. The first control information for each VOL is associated with the generation of the VOL and indicates which second control information to refer to for the addresses in the VOL. The second control information for each data is associated with the oldest generation among the generations of the first control information that refers to the second control information, and represents the address where the data is located in the snapshot virtual device. For each snapshot family, the processor determines whether the first control information and / or the second control information can be invalidated based on the generation of the first control information and the generation of the second control information, asynchronously with the processing of the I / O request for the VOL in the snapshot family, and invalidates the control information determined to be invalid. The storage system according to claim 1.

11. When the following conditions (x) and (y) are satisfied, the processor invalidates the second control information referred to by the target first control information. (x) The generation of the second control information referred to by the first control information of the generation one newer than the target first control information does not match the generation of the second control information referred to by the target first control information. (y) The generation of the first control information of the generation one older than the target first control information is older than the generation of the second control information referred to by the target first control information. The storage system according to claim 10.

12. When creating a writable SVOL, the processor creates, for the SVOL, first control information prohibiting writing and first control information permitting writing, sets the latest generation before creating the SVOL as the generation of the first control information prohibiting writing, and sets the generation obtained by incrementing the generation of the first control information prohibiting writing as the generation of the first control information permitting writing. The storage system according to claim 10.

13. The processor compares the generation of second control information with the generation of the first control information permitting writing of the writable SVOL, invalidates the second control information when the generation of the second control information matches the generation of the first control information permitting writing, and does not invalidate the second control information when the generation of the second control information is older than the generation of the first control information permitting writing. The storage system according to claim 12.

14. When performing a restore from a writable SVOL to a PVOL, the processor creates, as a copy of the first control information permitting writing of the writable SVOL, first control information prohibiting writing and first control information permitting writing for the PVOL to be restored, sets the generation of the first control information prohibiting writing for the PVOL to be restored as the generation of the first control information permitting writing of the writable SVOL, sets the generation of the first control information permitting writing for the PVOL to be restored as two generations newer than the generation of the PVOL to be restored, sets the generation of the first control information permitting writing for the writable SVOL as two generations newer than the original generation of the first control information permitting writing, and invalidates the original first control information of the PVOL to be restored when it is not referenced from any PVOL and SVOL. The storage system according to claim 12.

15. A storage control method for a storage system, wherein the processor of the storage system stores first mapping information and second mapping information in the storage device of the storage system. The first mapping information includes information representing the mapping between the address of a PVOL (Primary Volume) or an SVOL (Secondary Volume), which is a snapshot of the PVOL, in each of a plurality of snapshot families each composed of a PVOL and an SVOL, and the address in a snapshot virtual device. The second mapping information includes information representing the mapping between the address in a snapshot virtual device and the address in a deduplication virtual device. The processor of the storage system prepares one or more snapshot virtual devices for each snapshot family. Each snapshot virtual device is a virtual device serving as a logical address space for storing data of a VOL in the snapshot family corresponding to the snapshot virtual device. The deduplication virtual device is a virtual device serving as a logical address space different from the snapshot virtual device. When there is identical data in a plurality of VOLs of a snapshot family, the processor of the storage system updates the first mapping information so as to map the plurality of addresses of the identical data among the plurality of VOLs to the address of the snapshot virtual device of the snapshot family. When there is duplicate data in two or more snapshot virtual devices of two or more snapshot families, the processor of the storage system updates the second mapping information so as to map the two or more addresses of the duplicate data among the two or more snapshot virtual devices to the address corresponding to the duplicate data in the deduplication virtual device. A storage control method.

Citation Information

Patent Citations

  • Storage controller and its control method

    JP2008299434A

  • Storage controller and storage control method

    JP2020047036A

  • Storage system and data duplication method in storage system

    JP2022026812A

  • Threshold based incremental flashcopy backup of a raid protected array

    US20170161153A1

  • Dedupe as an infrastructure to avoid data movement for snapshot copy-on-writes

    US20190108100A1