STORAGE SYSTEM AND DATA PROCESSING METHOD
The storage system addresses the throughput reduction in RoW methods by managing snapshot volumes with a snapshot virtual device, optimizing resource allocation and compression, enhancing efficiency and throughput.
Patent Information
- Application Number
- JP2023101392
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-06-21
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2043-06-21
AI Technical Summary
The Redirect on Write (RoW) method for implementing snapshots in storage systems leads to increased load on the storage controller due to increased mapping information and metadata, reducing throughput and resource allocation efficiency, particularly in All Flash Array devices.
A storage system that manages primary and snapshot volumes using a snapshot virtual device, switching between overwrite and new allocation processes based on reference status and distribution of address ranges, with a compressed append virtual device for efficient resource utilization.
This approach optimizes mapping information and data transfer, achieving high throughput and resource efficiency by minimizing metadata overhead and reducing storage media requirements.
Smart Images

Figure 0007680497000001 
Figure 0007680497000002 
Figure 0007680497000003
Abstract
Description
[Technical field]
[0001] The present invention relates to a storage system and a data processing method. [Background technology]
[0002] In recent years, the need for data utilization has increased, and the opportunities for data duplication are on the rise. Accordingly, the snapshot function has become increasingly important in storage systems. Traditionally, the typical method for implementing snapshots has been the Redirect on Write (RoW) method. The RoW method has the advantage of having little impact on I / O performance when creating a snapshot, since there is no copying of data or meta-information. The RoW method is widely adopted in AFA (All Flash Array) devices. The RoW method is a method of overwriting data. Overwriting is a data storage method in which, when data is written to a storage system, the data to be written is stored in a new area, and the meta-information is rewritten to refer to the data stored in the new area, without overwriting the data stored before the write.
[0003] When using a data storage method that uses additional writing, repeated writes from the host computer tend to fragment free space. As fragmentation progresses, data of the requested size cannot be allocated to a continuous area, causing problems such as a decrease in read and write performance over time. Fragmentation can be reduced by implementing garbage collection processing, which moves the location of data, but this processing has a particularly large impact on performance.
[0004] Patent document 1 presents a method of reducing the number of references to duplicate data and improving the efficiency of garbage collection processing by allocating data shared by multiple VOLs through deduplication or snapshots to a virtual space separate from the storage location of individual data that is referenced only by one VOL. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] U.S. Pat. No. 1,081,7209 Summary of the Invention [Problem to be solved by the invention]
[0006] However, when adding a virtual device space as in Patent Document 1, the amount of mapping information to be referenced and updated when performing read / write IO processing increases, which increases the load on the storage controller and reduces throughput.Allocation of virtual device space also increases the amount of metadata relative to the amount of data to be stored, reducing the reduction effect, and there are problems such as a shortage of virtual device resources themselves, limiting the number of VOLs that can be created, but no solutions to these problems have been presented. Therefore, an object of the present invention is to realize high throughput by efficiently utilizing the resources of a virtual device. [Means for solving the problem]
[0007] In order to achieve the above object, one representative storage system of the present invention comprises a storage device and a processor that accesses the storage device, the processor manages a primary volume that is the subject of read / write by a host and a snapshot volume generated from the primary volume as a snapshot family, the processor uses a snapshot virtual device, which is a logical address space associated with the snapshot family, as a storage destination for data of the primary volume and the snapshot volume, and when the processor receives a write request from the host, the processor switches between an overwrite process that overwrites an allocated area on the snapshot virtual device and a new allocation process that allocates a new area on the snapshot virtual device to the address range of the write destination, based on the reference status from the snapshot volume to the address range of the write destination and the degree of distribution of the address range of the write destination in the snapshot virtual device. Furthermore, one representative data processing method of the present invention is a data processing method for a storage system comprising a storage device and a processor that accesses the storage device, wherein the processor manages a primary volume that is the subject of read / write by a host and a snapshot volume generated from the primary volume as a snapshot family, and the processor uses a snapshot virtual device, which is a logical address space associated with the snapshot family, as a storage destination for data of the primary volume and the snapshot volume, and when the processor receives a write request from the host, switches between an overwrite process that overwrites an area on the snapshot virtual device that has already been allocated and a new allocation process that allocates a new area on the snapshot virtual device to the address range of the write destination, based on the reference status from the snapshot volume to the address range of the write destination and the degree of distribution of the address range of the write destination in the snapshot virtual device. Effect of the Invention
[0008] According to the present invention, it is possible to efficiently utilize the resources of a virtual device and achieve high throughput. Problems, configurations and effects other than those described above will become apparent from the following description of the embodiments. [Brief description of the drawings]
[0009] [Figure 1] 1 is a block diagram showing a hardware configuration of a storage system according to the present invention. [Diagram 2] FIG. 2 is a schematic diagram of storage control of a storage system according to the present invention. [Diagram 3] 1 is a schematic diagram of management of mapping between addresses in the SS-Family and addresses in the SS-VDEV. [Figure 4] This is an overview diagram of mapping management between addresses in an SS-VDEV and addresses in a Dedup-VDEV, mapping management between addresses in an SS-VDEV and addresses in a CR-VDEV, and mapping management between addresses in a Dedup-VDEV and addresses in a CR-VDEV. [Diagram 5] FIG. 2 is an explanatory diagram showing the configuration of a memory included in a storage controller according to the present invention; [Figure 6] 2 is an explanatory diagram showing the configuration of a control information area in a memory of a storage controller in the present invention. FIG. [Figure 7] 2 is an explanatory diagram showing the configuration of a program area in a memory of a storage controller in the present invention. FIG. [Figure 8] FIG. 13 is a diagram showing the configuration of an ownership management table. [Figure 9] FIG. 13 is a diagram showing the configuration of a CR-VDEV management table. [Figure 10] FIG. 13 is a diagram showing the configuration of a snapshot management table. [Figure 11] FIG. 13 is a diagram showing the configuration of a VOL-Dir management table. [Figure 12] FIG. 13 is a diagram showing the configuration of a latest generation table. [Figure 13] FIG. 13 is a diagram showing the configuration of a collection management table. [Figure 14] FIG. 13 is a diagram showing the configuration of a generation management tree table. [Figure 15] FIG. 13 is a diagram showing a configuration of a snapshot allocation management table. [Figure 16] FIG. 13 is a diagram illustrating a configuration of a Dir management table. [Figure 17] FIG. 13 is a diagram showing the configuration of an SS-Mapping management table. [Figure 18] FIG. 13 is a diagram showing the configuration of a compression allocation management table. [Figure 19] FIG. 13 is a diagram showing the configuration of a CR-Mapping management table. [Figure 20] FIG. 13 is a diagram illustrating the configuration of a Dedup-Dir management table. [Figure 21] FIG. 13 is a diagram showing the configuration of a Dedup allocation management table. [Figure 22] FIG. 13 is a diagram illustrating a configuration of a Pool-Mapping management table. [Figure 23] 13 is a diagram showing the configuration of a Pool allocation management table. [Figure 24] FIG. 11 is a diagram showing the flow of a snapshot acquisition process. [Diagram 25] FIG. 13 is a diagram showing the flow of a snapshot restore process. [Figure 26] FIG. 13 is a diagram showing the flow of a snapshot deletion process. [Figure 27] FIG. 13 is a diagram showing the flow of an asynchronous collection process. [Figure 28] FIG. 13 is a diagram showing the flow of Write processing (front end). [Figure 29] FIG. 13 is a diagram showing the flow of Write processing (backend). [Diagram 30] FIG. 13 is a diagram showing the flow of a snapshot allocation determination process. [Diagram 31] FIG. 11 is a diagram showing the flow of a snapshot append process. [Diagram 32] FIG. 13 is a diagram showing the flow of Dedup additional writing processing. [Diagram 33] FIG. 11 is a diagram showing the flow of a compression and appending process. [Diagram 34] FIG. 13 is a diagram showing the flow of a destage process. [Diagram 35] FIG. 13 is a diagram showing the flow of a read process. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0010] Hereinafter, the embodiments will be described with reference to the drawings. EXAMPLES
[0011] FIG. 1 shows the hardware configuration of a computer system. The computer system 100 includes a storage system 201, a server system 202, and a management system 203. The storage system 201 and the server system 202 are connected via a storage network 204 using FC (Fiber Channel) or the like. The storage system 201 and the management system 203 are connected via a management network 205 using IP (Internet Protocol) or the like. The storage network 204 and the management network 205 may be the same communication network.
[0012] The storage system 201 includes a plurality of storage controllers 210 and a plurality of SSDs 220. The plurality of SSDs 220 are connected to the storage controller 210. The plurality of SSDs 220 are an example of a persistent storage device. A pool 13 is configured based on the plurality of SSDs 220. Data stored in a page 14 of the pool 13 is stored in one or more SSDs 220.
[0013] The storage controller 210 includes a CPU 211 , a memory 212 , a back-end interface 213 , a front-end interface 214 , and a management interface 215 .
[0014] The CPU 211 executes a program stored in the memory 212 . The memory 212 stores programs executed by the CPU 211 and data used by the CPU 211. The memory 212 and the CPU 211 may be duplicated.
[0015] The back-end interface 213, the front-end interface 214, and the management interface 215 are examples of interface devices. The back-end interface 213 is a communication interface device that mediates data exchange between the SSD 220 and the storage controller 210. To the back-end interface 213, a plurality of SSDs 220 are connected. The front-end interface 214 is a communication interface device that mediates data exchange between the server system 202 and the storage controller 210. The server system 202 is connected to the front-end interface 214 via the storage network 204. The management interface 215 is a communication interface device that mediates data exchange between the management system 203 and the storage controller 210. The management system 203 is connected to the management interface 215 via a management network 205.
[0016] The server system 202 is configured to include one or more host devices. The server system 202 transmits an I / O request (write request or read request) specifying an I / O destination to the storage controller 210. The I / O destination is, for example, a logical volume number such as a LUN (Logical Unit Number), a logical address such as a LBA (Logical Block Address), etc. The management system 203 includes one or more management devices and manages the storage system 201.
[0017] FIG. 2 shows an overview of storage control of a storage system. In FIG. 2, data written in uppercase alphabetic characters (data A, B, C, ...) is block data, and data written in lowercase alphabetic characters (data a, b, c, ...) is sub-block data. Block data may be data in units of blocks. A block may be a fixed-length logical storage area (logical address range). Sub-block data is compressed data of block data, and a group of sub-blocks (one or more sub-blocks) is the data storage destination. A sub-block may be a logical storage area of a size smaller than a block. For example, a block may be an integer multiple of a sub-block.
[0018] A storage system having a storage device and a processor includes an SS-Family (snapshot family) 9, an SS-VDEV (snapshot virtual device) 11S, a Dedup-VDEV (deduplicated virtual device) 11D, a CR-VDEV (compressed append virtual device) 11C, and a pool 13.
[0019] SS-Family9 is a VOL group that includes PVOL10P and SVOL10S, which is a snapshot of PVOL10P. SS-VDEV11S is a virtual device serving as a logical address space, and serves as a storage destination for data whose storage destination is any one of VOL10 in SS-Family9. The Dedup-VDEV11D is a virtual device that serves as a logical address space separate from the SS-VDEV11S, and serves as a storage destination for duplicate data of two or more SS-VDEV11S.
[0020] The CR-VDEV 11C is a virtual device that serves as a logical address space separate from the SS-VDEV 11S and Dedup-VDEV 11D, and is used as a storage destination for compressed data. Each of the multiple CR-VDEV11C is associated with either the SS-VDEV11S or the Dedup-VDEV11D, and is not associated with both the VDEV11S and 11D. That is, each CR-VDEV11C becomes the storage destination of data whose storage destination is a VDEV (virtual device) corresponding to the CR-VDEV11C, and does not become the storage destination of data whose storage destination is a VDEV not corresponding to the CR-VDEV11C. The compressed data whose storage destination is the CR-VDEV11C becomes the storage destination of the pool 13.
[0021] The pool 13 is a logical address space based on at least a part of a storage device (e.g., a persistent storage device) that the storage system has. The pool 13 may be based on at least a part of an external storage device (e.g., a persistent storage device) of the storage system instead of or in addition to at least a part of the storage device that the storage system has. The pool 13 has a plurality of pages 14 that are a plurality of logical areas. Compressed data whose storage destination is the CR-VDEV 11C is stored in the pages 14 in the pool 13. There is a 1:1 mapping between an address in the CR-VDEV 11C and an address in the pool 13. The pool 13 is composed of one or more pool VOLs.
[0022] According to the example shown in FIG. 2, the following storage control is performed. The processor creates SVOL10S0 as a snapshot of PVOL10P0, which creates SS-Family 9-0 with PVOL10P0 as the root VOL. The processor also creates SVOL10S1 as a snapshot of PVOL10P1, which creates SS-Family 9-1 with PVOL10P1 as the root VOL. According to Figure 2, examples of multiple SS-Families include SS-Family 9-0 and 9-1.
[0023] The storage system has one or more SS-VDEV11S for each of a plurality of SS-Family 9. For each SS-Family 9, the SS-VDEV11S corresponding to the SS-Family 9 is set as the storage destination for data for which any VOL10 in the SS-Family 9 is the storage destination among the plurality of SS-VDEV11S. Taking SS-Family 9-0 as an example, the specific example is as follows. The processor specifies SS-VDEV11S0 as the storage destination for data A, whose storage destination is SVOL10S0 of SS-Family9-0. The processor maps the address in SVOL10S0 corresponding to data A to the address in SS-VDEV11S0 corresponding to SS-Family9-0 corresponding to data A. ·If the same data B exists in multiple VOLs (PVOL10P0 and SVOL10S0) of SS-Family9-0, the processor maps multiple addresses of the same data B among the multiple VOL10s (addresses in PVOL10P0 and addresses in SVOL10S0) to an address of SS-VDEV11S0 of SS-Family9-0 (address corresponding to data B).
[0024] For each of SS-Family 9-0 and 9-1 (an example of two or more SS-Family 9), the storage destination of non-duplicate data is the CR-VDEV 11C corresponding to that SS-Family 9, and the storage destination of duplicate data is the Dedup-VDEV 11D. That is, since data C is duplicated in SS-VDEV11S0 and 11S1 (an example of two or more SS-VDEV11S) of SS-Family 9-0 and 9-1, the processor maps the two addresses of the duplicate data C in SS-VDEV11S0 and 11S1 to an address in Dedup-VDEV11D corresponding to the duplicate data C. Then, the processor compresses the duplicate data C and sets the storage destination of the compressed data c to CR-VDEV11CC corresponding to Dedup-VDEV11D. That is, the processor maps the address (block address) of the duplicate data C in Dedup-VDEV11D to the address (sub-block address) of the compressed data c in CR-VDEV11CC. Also, the processor assigns page 14B to CR-VDEV11CC and stores the compressed data c in page 14B. The address of the compressed data c in the CR-VDEV 11 CC is mapped to an address in page 14 B of the pool 13 .
[0025] On the other hand, since data A in SS-VDEV11S0 is non-duplicate with data in other SS-VDEV11S1, the processor compresses the non-duplicate data A and sets the compressed data a to CR-VDEV11C0 corresponding to SS-VDEV11S0. That is, the processor maps the address (block address) of the non-duplicate data A in SS-VDEV11S0 to the address (sub-block address) of the compressed data a in CR-VDEV11CC. In addition, the processor assigns page 14A to CR-VDEV11C0 and stores the compressed data a in page 14A. The address of the compressed data a in CR-VDEV11C0 is mapped to an address in page 14A of pool 13.
[0026] The CR-VDEV11C is a write-once VDEV. Therefore, the processor updates the address mapping when the CR-VDEV11C corresponding to the SS-VDEV11S is set as the storage destination for update data, and when the CR-VDEV11C corresponding to the Dedup-VDEV11D is set as the storage destination for update data. Specifically, the processor performs the following storage control, for example. When the storage target in CR-VDEV11C0 is update data a' of compressed data a, the processor sets the storage destination of the update data a' to an empty address in CR-VDEV11C0 and invalidates the address of the compressed data a before the update. The processor maps the address in SS-VDEV11S0 that was mapped to the address of compressed data a to the storage destination address of the update data a' in CR-VDEV11C0 instead of the address of compressed data a in CR-VDEV11C0. The processor also maps the address in page 14A that was mapped to the address of compressed data a to the storage destination address of the update data a' in CR-VDEV11C0 instead of the address of compressed data a in CR-VDEV11C0. When the storage target in CR-VDEV11CC is update data c' of compressed data c, the processor sets the storage destination of the update data c' to an empty address in CR-VDEV11CC and invalidates the address of the compressed data c before the update. The processor maps the address in Dedup-VDEV11D that was mapped to the address of compressed data c to the storage destination address of the update data c' in CR-VDEV11CC instead of the address of compressed data c in CR-VDEV11CC. The processor also maps the address in page 14B that was mapped to the address of compressed data c to the storage destination address of the update data c' in CR-VDEV11CC instead of the address of compressed data a in CR-VDEV11CC.
[0027] The CR-VDEV11C is an append type VDEV as described above, and garbage collection is performed. That is, the processor performs garbage collection on the CR-VDEV11C, so that valid addresses (addresses of the latest data) are continuous, and also, addresses of free areas are continuous.
[0028] According to the example shown in FIG. 2, in addition to SS-VDEV11S which is the storage destination of data in SS-Family9, Dedup-VDEV11D which is the storage destination of duplicated data in two or more SS-Family9 is prepared. Therefore, even if the address of duplicated data C in Dedup-VDEV11D is changed, the address mapping to be changed is only two mappings (mapping for each of the two addresses in SS-VDEV11S0 and 11S1). On the other hand, in one comparative example, the storage destination of data in SS-Family9 and the storage destination of duplicated data in two or more SS-Family9 are the same VDEV, and in this case, the address mapping to be changed for duplicated data C is four mappings (mapping for each of the four addresses in VOL10P0, 10S0, 10P1 and 10S1). In this embodiment, it is expected that the address mapping can be changed in a shorter time than in the one comparative example.
[0029] A CR-VDEV11CC is prepared for the Dedup-VDEV11D separately from the Dedup-VDEV11D. Therefore, even if the address of the compressed data c in the CR-VDEV11CC is changed, the address mapping to be changed is only one mapping (mapping for one address in the Dedup-VDEV11D). On the other hand, in one comparative example, the storage destination of the compressed data of the data in the SS-Family9 and the storage destination of the compressed data of the duplicated data in two or more SS-Family9 are the same VDEV, and in this case, the address mapping to be changed for the compressed data c of the duplicated data C is four mappings (mapping for each of the four addresses in VOL10P0, 10S0, 10P1, and 10S1). In this embodiment, it is expected that the address mapping can be changed in a shorter time than in the comparative example.
[0030] The address of the compressed data c is changed during garbage collection of the CR-VDEV11CC corresponding to the Dedup-VDEV11D. For example, during garbage collection of the CR-VDEV11CC, the processor changes the address of the update data c' (update data of the compressed data c) in the CR-VDEV11CC, and maps the address in the Dedup-VDEV11D that was mapped to the address before the change to the address after the change in the CR-VDEV11CC. Since it is expected that the address mapping for the compressed data of duplicate data can be changed in a short time, it is expected that the garbage collection can be performed in a short time. Note that the garbage collection of the CR-VDEV11C corresponding to the SS-VDEV11S includes, for example, the following processes. That is, the processor changes the address of the update data a' (update data of the compressed data a) in CR-VDEV11C0, and maps the address in SS-VDEV11S0 that was mapped to the address before the change to the address after the change in CR-VDEV11C0.
[0031] For at least one CR-VDEV11C, a write-once VDEV in which uncompressed data is stored may be used instead of the CR-VDEV11C, but in this embodiment, the CR-VDEV11C is used as the write-once VDEV. Therefore, the data that is ultimately stored in the storage device is compressed data, and therefore the consumed storage capacity can be reduced.
[0032] Fig. 3 shows an overview of mapping management between addresses in SS-Family9 and addresses in SS-VDEV11S. In the drawing, "GX" (X is an integer equal to or greater than 0) means generation X. Fig. 3 also takes SS-Family9-0 and SS-VDEV11S0 as examples.
[0033] The processor can use the meta-information to manage mapping between addresses in VOL10 in SS-Family9-0 and addresses in SS-VDEV11S0. The meta-information includes Dir-Info (directory information) and SS-Mapping-Info (snapshot mapping information). The processor manages data in PVOL10P0 and SVOL10S0 by associating Dir-Info with SS-Mapping-Info. For data stored in VOL10, Dir-Info has information indicating the address of the reference source (address in VOL10), and SS-Mapping-Info corresponding to the data has information indicating the address of the reference destination (address in SS-VDEV11S0).
[0034] Furthermore, the processor manages the time series of PVOL10P0 and SVOL10S0 by generation information associated with Dir-Info, and for each piece of data stored in SS-VDEV11S0, manages generation information indicating the generation in which the data was created by associating it with SS-Mapping-Info. In addition, the processor manages the latest generation information at that time as the latest generation.
[0035] It is assumed that, before a snapshot is taken, there are data A0, B0, and C0 whose storage destination is PVOL10P0, and the latest generation is "0". Dir-Info associated with PVOL10P0 is associated with "0" as the generation # (a number indicating a generation), and includes reference information indicating the reference destination of all data A0, B0, and C0 of PVOL10P0. Hereinafter, when the generation # associated with Dir-Info is "X", it can be expressed as Dir-Info being generation X.
[0036] SS-VDEV11S0 is the storage destination for data A0, B0, and C0, and SS-Mapping-Info is associated with each of the data A0, B0, and C0. Furthermore, each SS-Mapping-Info is associated with a generation # of "0." If the generation # associated with SS-Mapping-Info represents "X," it can be expressed that the data corresponding to SS-Mapping-Info is generation X data.
[0037] Before the snapshot is taken, the information in Dir-Info for each of data A0, B0, and C0 references the SS-Mapping-Info corresponding to that data. By associating Dir-Info with SS-Mapping-Info in this way, PVOL10P0 and SS-VDEV11S0 are associated with each other, and data processing for PVOL10P0 can be realized.
[0038] To acquire a snapshot, the processor copies the Dir-Info to the read-only Dir-Info of SVOL10S0. The processor then increments the generation of the Dir-Info of PVOL10P0 and also increments the latest generation. As a result, for each of data A0, B0, and C0, the SS-Mapping-Info is referenced from both the Dir-Info of generation 0 and the Dir-Info of generation 1.
[0039] In this way, a snapshot can be created by replicating Dir-Info, and a snapshot can be created without increasing the amount of data or SS-Mapping-Info on SS-VDEV11S0.
[0040] Here, when a snapshot is taken, the snapshot (SVOL10S0) in which writing is prohibited and data is fixed at the time of taking the snapshot becomes generation 0, and PVOL10P0 in which data can be written even after taking the snapshot becomes generation 1. Generation 0 is the "generation one generation older in the direct line" of generation 1, and is called the "parent" for convenience. Similarly, generation 1 is the "generation one generation newer in the direct line" of generation 0, and is called the "child" for convenience. The storage system manages the parent-child relationship of the generations as a Dir-Info generation management tree 70. Furthermore, the generation # of Dir-Info is the same as the generation # of the VOL10 corresponding to the Dir-Info. Furthermore, the generation # of SS-Mapping-Info is the oldest generation # of the generation # of one or more Dir-Info that reference the SS-Mapping-Info.
[0041] Fig. 4 shows an overview of mapping management between addresses in SS-VDEV11S and addresses in Dedup-VDEV11D, mapping management between addresses in SS-VDEV11S and addresses in CR-VDEV11C, and mapping management between addresses in Dedup-VDEV11D and addresses in CR-VDEV11C. Fig. 4 takes SS-VDEV11S0, CR-VDEV11C0, and 11CC as examples.
[0042] The processor can use the meta-information to manage mapping between addresses in SS-VDEV11S0 and addresses in Dedup-VDEV11D, mapping between addresses in SS-VDEV11S0 and addresses in CR-VDEV11C0, and mapping between addresses in Dedup-VDEV11D and addresses in CR-VDEV11CC. As described above, the meta-information includes Dir-Info and CR-Mapping-Info. The processor manages data in SS-VDEV11S0 and Dedup-VDEV11D by associating Dir-Info with CR-Mapping-Info. For data stored in SS-VDEV11S0, Dir-Info has information indicating the address of the reference source (address in SS-VDEV11S0), and CR-Mapping-Info corresponding to the data has information indicating the address of the reference destination (address in CR-VDEV11C0 or address in Dedup-VDEV11D). For data stored in Dedup-VDEV11D, Dir-Info has information indicating the address of the reference source (address in Dedup-VDEV11D), and CR-Mapping-Info corresponding to the data has information indicating the address of the reference destination (address in CR-VDEV11CC). The processor can specify the address in SS-VDEV11S or Dedup-VDEV11D from the address in CR-VDEV11C by referring to the compression allocation information. Although omitted in the drawing, the storage system 201 stores a compression allocation management table 1011 in memory 212 as information on reverse mapping from addresses in the CR-VDEV to addresses in the SS-VDEV or Dedup-VDEV in order to move valid data on the CR-VDEV and secure continuous free space.
[0043] FIG. 5 shows the configuration of the memory 212. The memory 212 has a control information section 901 in which control information (which may be called management information) is stored, a program section 902 in which programs are stored, and a cache section 903 in which data is temporarily stored.
[0044] FIG. 6 shows information stored in the control information section 901. The control information unit 901 stores an ownership management table 1001, a CR-VDEV management table 1002, a snapshot management table 1003, a VOL-Dir management table 1004, a latest generation table 1005, a recovery management table 1006, a generation management tree table 1007, a snapshot allocation management table 1008, a Dir management table 1009, an SS-Mapping management table 1010, a compression allocation management table 1011, a CR-Mapping management table 1012, a Dedup-Dir management table 1013, a Dedup allocation management table 1014, a Pool-Mapping management table 1015 and a Pool allocation management table 1016.
[0045] FIG. 7 shows the programs stored in the program section 902. The program section 902 stores a snapshot acquisition program 1101, a snapshot restore program 1102, a snapshot deletion program 1103, an asynchronous recovery program 1104, a read / write program 1105, a snapshot append program 1106, a Dedup append program 1107, a compression append program 1108, a destaging program 1109, a GC (garbage collection) program 1110, a CPU determination program 1111, an ownership transfer program 1112, and a Snapshot allocation determination program 1113.
[0046] FIG. 8 shows the configuration of the ownership management table 1001 . The ownership management table 1001 manages the ownership of a VOL 10 or a VDEV 11. For example, the ownership management table 1001 has an entry for each VOL 10 and each VDEV 11. The entry has information such as a VOL# / VDEV# 1201 and an owner CPU# 1202.
[0047] VOL# / VDEV# 1201 indicates the identification number of VOL10 or VDEV11. Owner CPU# 1202 indicates the identification number of a CPU serving as the owner CPU of VOL10 or VDEV11 (a CPU that has ownership of VOL10 or VDEV11). Incidentally, instead of allocation in units of CPU 211, the owner CPU may be allocated in units of a CPU group, or may be allocated in units of a storage controller 210.
[0048] FIG. 9 shows the configuration of the CR-VDEV management table 1002 . The CR-VDEV management table 1002 indicates a CR-VDEV 11C associated with an SS-VDEV 11S or a Dedup-VDEV 11D. For example, the CR-VDEV management table 1002 has an entry for each SS-VDEV 11S and each Dedup-VDEV 11D. The entry has information such as VDEV#1301 and CR-VDEV#1302. VDEV#1301 indicates the identification number of the SS-VDEV11S or the Dedup-VDEV11D. CR-VDEV#1302 indicates the identification number of the CR-VDEV11C.
[0049] FIG. 10 shows the configuration of the snapshot management table 1003 . A snapshot management table 1003 exists for each PVOL10P (for each SS-Family9). The snapshot management table 1003 indicates the acquisition time of each snapshot (SVOL10S). For example, the snapshot management table 1003 has an entry for each SVOL10S. The entry has information such as PVOL#1401, SVOL#1402, and acquisition time 1403. PVOL# 1401 indicates the identification number of PVOL10P. SVOL# 1402 indicates the identification number of SVOL10S. Acquisition time 1403 indicates the acquisition time of SVOL10S.
[0050] FIG. 11 shows the configuration of the VOL-Dir management table 1004 . The VOL-Dir management table 1004 indicates the correspondence between VOLs and Dir-Info. For example, the VOL-Dir management table 1004 has an entry for each VOL 10. The entry has information such as VOL#1501, Root-VOL#1502, and Dir-Info#1503.
[0051] VOL#1501 represents the identification number of PVOL10P or SVOL10S. Root-VOL#1502 represents the identification number of Root-VOL. If VOL10 is PVOL10P, the Root-VOL is the PVOL10P, and if VOL10 is SVOL10S, the Root-VOL is the PVOL10P corresponding to the SVOL10S. Dir-Info#1503 represents the identification number of Dir-Info corresponding to VOL10.
[0052] FIG. 12 shows the configuration of the latest generation table 1005 . The latest generation table 1005 exists for each PVOL 10P (for each SS-Family 9) and indicates the generation (generation #) of the PVOL 10P.
[0053] FIG. 13 shows the configuration of the collection management table 1006 . The collection management table 1006 may be, for example, a bitmap, and exists for each PVOL 10P (for each SS-Family 9), in other words, for each Dir-Info generation management tree 70. The collection management table 1006 has an entry for each Dir-Info. The entry has information such as Dir-Info# 1701 and a collection request 1702. Dir-Info# 1701 indicates the identification number of Dir-Info. Collection request 1702 indicates whether or not collection of Dir-Info is requested. "1" means that collection is requested, and "0" means that collection is not requested.
[0054] FIG. 14 shows the configuration of the generation management tree table 1007 . A generation management tree table 1007 exists for each PVOL 10P (for each SS-Family 9), in other words, for each Dir-Info generation management tree 70. The generation management tree table 1007 has an entry for each Dir-Info. The entry has information such as Dir-Info #1801, Generation #1802, Prev 1803, and Next 1804. Dir-Info #1801 represents the identification number of Dir-Info. Generation #1802 represents the generation of VOL10 corresponding to Dir-Info. Prev 1803 represents the parent Dir-Info of Dir-Info (one level above). Next 1804 represents the child Dir-Info of Dir-Info (one level below). The number of Next 1804s can be the same as the number of child Dir-Infos. In Figure 14, there are two child Dir-Infos, so there are two Next 1804s (Next-A 1804A and Next-B 1804B).
[0055] FIG. 15 shows the configuration of the snapshot allocation management table 1008 . A snapshot allocation management table 1008 exists for each SS-VDEV11S, and indicates mapping from an address in the SS-VDEV11S to an address in a VOL 10. The snapshot allocation management table 1008 has an entry for each address in the SS-VDEV11S. The entry has information such as a block address 1901, a status 1902, an allocation destination VOL# 1903, and an allocation destination address 1904.
[0056] Block address 1901 indicates the address of the block in SS-VDEV11S. Status 1902 indicates whether the block is assigned to the address of any VOL ("1" means assigned, and "0" means free). Assigned VOL# 1903 indicates the identification number of the VOL10 (PVOL10P or SVOL10S) that has the address to which the block is assigned ("n / a" means unassigned). Assigned address 1904 indicates the address (block address) to which the block is assigned ("n / a" means unassigned).
[0057] FIG. 16 shows the configuration of the Dir management table 1009 . A Dir management table 1009 exists for each Dir-Info, and indicates the Mapping-Info of the reference destination for each data (for each block data). For example, the Dir management table 1009 has an entry for each address (block address). The entry has information such as a VOL / VDEV address 2001 and a reference destination Mapping-Info# 2002. The VOL / VDEV address 2001 indicates an address (block address) in VOL10 (PVOL10P or SVOL10S) or an address in VDEV11 (SS-VDEV11S or Dedup-VDEV11D). The referenced Mapping-Info# 2002 indicates the identification number of the referenced Mapping-Info.
[0058] FIG. 17 shows the configuration of the SS-Mapping management table 1010. An SS-Mapping management table 1010 exists for each Dir-Info of the VOL 10. The SS-Mapping management table 1010 has an entry for each SS-Mapping-Info corresponding to the Dir-Info of the VOL 10. The entry has information such as Mapping-Info#2101, reference address 2102, reference SS-VDEV#2103, and generation#2104.
[0059] Mapping-Info #2101 indicates the identification number of SS-Mapping-Info. Reference address 2102 indicates the address referenced by SS-Mapping-Info (address in SS-VDEV11S). Reference SS-VDEV #2103 indicates the identification number of SS-VDEV11S having the address referenced by SS-Mapping-Info. Generation #2104 indicates the generation of data corresponding to SS-Mapping-Info.
[0060] FIG. 18 shows the configuration of the compression allocation management table 1011. The compression allocation management table 1011 exists for each CR-VDEV 11C, and has compression allocation information for each sub-block in the CR-VDEV 11C. The compression allocation management table 1011 has an entry corresponding to the compression allocation information for each sub-block in the CR-VDEV 11C. The entry has information such as a sub-block address 2201, a data length 2202, a status 2203, a first sub-block address 2204, an allocation destination VDEV# 2205, and an allocation destination address 2206.
[0061] The subblock address 2201 indicates the address of the subblock. The data length 2202 indicates the number of subblocks constituting the subblock group (one or more subblocks) in which compressed data is stored (for example, "2" means that the compressed data exists in two subblocks). The status 2203 indicates the status of the subblock ("0" means free, "1" means allocated, and "2" means that it is a target for GC (garbage collection)). The head subblock address 2204 indicates the address of the head subblock of one or more subblocks (one or more subblocks in which compressed data is stored) that include the subblock. The assigned VDEV# 2205 indicates the identification number of the VDEV11 (SS-VDEV11S or Dedup-VDEV11D) that has the block to which the subblock is assigned. The assigned address 2206 indicates the address of the block to which the subblock is assigned (block address in the SS-VDEV11S or Dedup-VDEV11D).
[0062] FIG. 19 shows the configuration of the CR-Mapping management table 1012. A CR-Mapping management table 1012 exists for each Dir-Info of Dedup-VDEV11D and for each Dir-Info of SS-VDEV11S. The CR-Mapping management table 1012 has an entry for each CR-Mapping-Info corresponding to the Dir-Info of Dedup-VDEV11D and for each CR-Mapping-Info corresponding to the Dir-Info of SS-VDEV11S. The entry has information such as Mapping-Info#2301, reference address 2302, reference CR-VDEV#2303, and data length 2304.
[0063] Mapping-Info#2301 indicates the identification number of CR-Mapping-Info. Reference address 2302 indicates the address referenced by CR-Mapping-Info (the address of the first sub-block in the sub-block group). Reference CR-VDEV#2303 indicates the identification number of the CR-VDEV11C having the sub-block address referenced by CR-Mapping-Info. Data length 2304 indicates the number of blocks referenced by CR-Mapping-Info (blocks in Dedup-VDEV11D) or the number of sub-blocks constituting the sub-block group referenced by CR-Mapping-Info.
[0064] FIG. 20 shows the configuration of the Dedup-Dir management table 1013. The Dedup-Dir management table 1013 exists for each Dedup-VDEV 11D and corresponds to Dedup-Dir-Info. The Dedup-Dir management table 1013 has an entry for each address in the Dedup-VDEV 11D. The entry has information such as a Dedup-VDEV address 2401 and reference destination allocation information #2402. The Dedup-VDEV address 2401 indicates an address (block address) in the Dedup-VDEV 11D. The reference destination allocation information #2402 indicates the identification number of the reference destination Dedup allocation information.
[0065] FIG. 21 shows the configuration of the Dedup allocation management table 1014. The Dedup allocation management table 1014 exists for each Dedup-VDEV 11D (for each Dedup-Dir-Info) and indicates a reverse reference mapping from the Dedup allocation information corresponding to an address in the Dedup-VDEV 11D to an address in the SS-VDEV 11S. The Dedup allocation management table 1014 has an entry for each Dedup allocation information. The entry has information such as allocation information #2501, allocation destination SS-VDEV #2502, allocation destination address 2503, and linked allocation information #2504. Allocation information #2501 indicates the identification number of the Dedup allocation information. Allocation destination SS-VDEV #2502 indicates the identification number of the SS-VDEV11S having the address referenced by the Dedup allocation information. Allocation destination address 2503 indicates the address referenced by the Dedup allocation information (block address in the SS-VDEV11S). Linked allocation information #2504 indicates the identification number of the Dedup allocation information linked to the Dedup allocation information.
[0066] According to Fig. 21, Dedup allocation information # "3" is linked to Dedup allocation information # "1", and there is no Dedup allocation information linked to Dedup allocation information # "3". Therefore, it can be seen that the duplicated data in the Dedup-VDEV address corresponding to Dedup allocation information # "1" is duplicated data in the SS-VDEV11S referenced by Dedup allocation information # "1" and the SS-VDEV11S referenced by Dedup allocation information # "3". Since the number of duplicated data is indefinite, Dedup allocation information is linked according to the number of duplicated data. When duplicated data exists in N SS-VDEV11S, N sequential Dedup allocation information is prepared.
[0067] FIG. 22 shows the configuration of the Pool-Mapping management table 1015. A Pool-Mapping management table 1015 exists for each CR-VDEV 11C. The Pool-Mapping management table 1015 has an entry for each area in page size units in the CR-VDEV 11C. The entry has information such as a VDEV address 2601 and a page #2602. The VDEV address 2601 indicates the start address of an area (e.g., multiple blocks) in page size units. The page #2602 indicates the identification number of the allocated page 14 (e.g., the address of the page 14 in the pool 13). Note that if there are multiple pools 13, the page #2602 may include the identification number of the pool 13 that has the page 14.
[0068] FIG. 23 shows the configuration of the Pool allocation management table 1016 . For example, if there are multiple pools 13, a pool allocation management table 1016 exists for each pool 13. The pool allocation management table 1016 indicates the correspondence between pages 14 and areas in the CR-VDEV 11C. The pool allocation management table 1016 has an entry for each page 14. The entry has information such as a page #2701, RG #2702, a top address 2703, a status 2704, an assigned VDEV #2705, and an assigned address 2706.
[0069] Page #2701 indicates the identification number of page 14. RG#2702 indicates the identification number of the RAID group that is the basis of page 14 (in this embodiment, a RAID group composed of two or more SSDs 220). First address 2703 indicates the first address of page 14. Status 2704 indicates the status of page 14 ("1" means allocated, and "0" means free). Allocated VDEV#2705 indicates the identification number of the CR-VDEV11C to which page 14 is allocated ("n / a" means unallocated). Allocated address 2706 indicates the allocated address of page 14 (address in CR-VDEV11C) ("n / a" means unallocated).
[0070] 24 shows the flow of snapshot acquisition processing. The snapshot acquisition processing is executed by the snapshot acquisition program 1101 in response to a snapshot acquisition instruction from the management system 203 (or another system such as the server system 202). In the snapshot acquisition instruction, for example, the target PVOL10P is specified.
[0071] First, the snapshot acquisition program 1101 assigns the Dir management table 1009 that is to be the copy destination, and updates the VOL-Dir management table 1004 (S2401). The snapshot acquisition program 1101 increments the latest generation # (S2402) and updates the generation management tree table 1007 (Dir-Info generation management tree 70) (S2403). At this time, the snapshot acquisition program 1101 sets the latest generation # in the copy source, and sets the generation # before the increment in the copy destination.
[0072] The snapshot acquisition program 1101 judges whether or not there is cache dirty data for the target PVOL 10P (S2404). "Cache dirty data" may be data stored in the cache unit 903 that has not yet been written to the pool 13.
[0073] If the determination result of S2404 is true (S2404: Yes), the snapshot acquisition program 1101 causes the snapshot addition program 1106 to execute a snapshot addition process (S2405). If the determination result in S2404 is false (S2404: No), or after S2405, the snapshot acquisition program 1101 copies the Dir management table 1009 of the target PVOL 10P to the Dir management table 1009 of the copy destination (S2406).
[0074] Thereafter, the snapshot acquisition program 1101 updates the snapshot management table 1003 (S2407) and terminates the process. In S2407, an entry is added having PVOL#1401 indicating the identification number of the target PVOL10P, SVOL#1402 indicating the identification number of the acquired snapshot (SVOL10S), and acquisition time 1403 indicating the acquisition time.
[0075] 25 shows the flow of snapshot restore processing. Snapshot restore processing is executed by the snapshot restore program 1102 in response to a restore instruction from the management system 203 (or another system such as the server system 202). In the restore instruction, for example, a restore source SVOL and a restore destination PVOL are specified.
[0076] First, the snapshot restore program 1102 assigns the Dir management table 1009 that is to be the restore destination, and updates the VOL-Dir management table 1004 (S2501). The snapshot restore program 1102 increments the latest generation # (S2502) and updates the generation management tree table 1007 (Dir-Info generation management tree 70) (S2503). At this time, the snapshot restore program 1102 sets the generation # before the increment in the copy source, and sets the latest generation # in the copy destination.
[0077] The snapshot restore program 1102 purges the cache area (area in the cache unit 903) of the restore destination PVOL (S2504). The snapshot restore program 1102 copies the Dir management table 1009 of the restore source SVOL to the Dir management table 1009 of the restore destination PVOL (S2505).
[0078] Thereafter, the snapshot restore program 1102 registers the Dir-Info# of the old Dir-Info of the restore destination in the collection management table 1006 (S2506), and ends the process. In S2506, the collection request 1702 corresponding to the Dir-Info# is set to "1".
[0079] 26 shows the flow of snapshot deletion processing. The snapshot deletion processing is executed by the snapshot deletion program 1103 in response to a snapshot deletion instruction from the management system 203 (or another system such as the server system 202). In the restore instruction, for example, the target SVOL is specified.
[0080] First, the snapshot deletion program 1103 references the VOL-Dir management table 1004 and invalidates the Dir-Info (Dir-Info #1503) of the target SVOL (S2601). Then, the snapshot deletion program 1103 updates the snapshot management table 1003 (S2602), registers the old Dir-Info# of the target SVOL in the recovery management table 1006 (S2603), and terminates the processing. In S2603, the recovery request 1702 corresponding to that Dir-Info# is set to "1".
[0081] 27 shows the flow of asynchronous collection processing. The asynchronous collection processing is executed by the asynchronous collection program 1104, for example, periodically. First, the asynchronous collection program 1104 identifies the Dir-Info# to be collected from the collection management table 1006 (S2701). The "Dir-Info# to be collected" is the Dir-Info# for which the collection request 1702 is "1". The asynchronous collection program 1104 refers to the generation management tree table 1007, checks the entries for the Dir-Info# to be collected, and does not select a Dir-Info that has two or more children.
[0082] Thereafter, the asynchronous collection program 1104 judges whether or not there is an unprocessed entry (S2702). The "unprocessed entry" here refers to an entry in the collection management table 1006 for which the collection request 1702 is "1" and for which the asynchronous collection process has not been processed.
[0083] If the judgment result of S2702 is true (S2702: Yes), the asynchronous collection program 1104 determines the entry to be processed (an entry containing collection request 1702 “1”) from one or more unprocessed entries (S2703), and identifies the referenced Mapping-Info#2002 from the Dir management table 1009 corresponding to the target Dir-Info (Dir-Info identified from Dir-Info#1701 in the entry to be processed) (S2704).
[0084] The asynchronous collection program 1104 refers to the generation management tree table 1007 and judges whether or not Dir-Info of a child generation of the target Dir-Info exists (S2705). If the determination result of S2705 is true (S2705: Yes), the asynchronous collection program 1104 identifies the referenced Mapping-Info#2002 from the Dir management table 1009 corresponding to the child generation Dir-Info, and determines whether the referenced Mapping-Info#2002 of the target Dir-Info matches the referenced Mapping-Info#2002 of the child generation Dir-Info (S2706). If the determination result of S2706 is true (S2706: Yes), the process returns to S2702.
[0085] If the result of the determination in S2706 is false (S2706: No) or if the result of the determination in S2705 is false (S2705: No), the asynchronous collection program 1104 determines whether the generation # of the Dir-Info of the parent generation of the target Dir-Info is older than the generation #2104 of the Mapping-Info referenced by the target Dir-Info (see FIG. 21) (S2707). If the result of the determination in S2707 is false (S2707: No), the process returns to S2702.
[0086] If the determination result of S2707 is true (S2707: Yes), the asynchronous collection program 1104 initializes the target entry in the SS-Mapping management table 1010, and releases the target entry in the snapshot allocation management table 1008 (S2708). Thereafter, the process returns to S2702. The release in S2708 corresponds to the release of blocks in the SS-VDEV.
[0087] If the determination result of S2702 is false (S2702: No), the asynchronous collection program 1104 updates the collection management table 1006 (S2709) and updates the generation management tree table 1007 (Dir-Info generation management tree 70) (S2710), and terminates the processing.
[0088] 28 shows the flow of the write process (front end). The write process (front end) is executed by the read / write program 1105 when a write request is received from the server system 202.
[0089] First, the read / write program 1105 judges whether the target data of the write request is a cache hit (S2801). "Cache hit" means that a cache area corresponding to the write destination VOL address of the target data (the VOL address specified in the write request) has been secured. If the judgment result of S2801 is false (S2801: No), the read / write program 1105 secures a cache area corresponding to the write destination VOL address of the target data from the cache unit 903 (S2802). After that, the process proceeds to S2806.
[0090] If the result of the determination in S2801 is true (S2801: Yes), the read / write program 1105 determines whether the cache hit data (data in the allocated cache area) is dirty data (data not reflected (not written) in the pool 13) (S2803). If the result of the determination in S2803 is false (S2803: No), the process proceeds to S2806.
[0091] If the determination result of S2803 is true (S2803: Yes), the read / write program 1105 determines whether the WR (Write) generation # of the dirty data matches the generation # of the target data of the write request (S2804). The "WR generation #" is held, for example, in cache data management information (not shown). The generation # of the target data of the write request is obtained from the latest generation #403. S2804 is a process for preventing the target data of the most recently acquired snapshot (dirty data) from being updated with the target data of the write request before the append process has been performed, thereby overwriting the data of the snapshot.
[0092] If the determination result of S2804 is false (S2804: No), the read / write program 1105 causes the snapshot addition program 1106 to execute a snapshot addition process (S2805).
[0093] After S2802, or if the determination result of S2804 is true (S2804: Yes), the read / write program 1105 writes the target data of the write request to the cache area secured in S2802, or to the cache area obtained through S2805 (S2806). Thereafter, the read / write program 1105 sets the WR generation # of the data written in S2806 to the latest generation # compared in S2804 (S2807), and returns a normal response (Good response) to the server system 202 (S2808).
[0094] 29 shows the flow of write processing (backend). The write processing (backend) is a process in which, when unreflected data (dirty data) exists in the cache unit 903, the unreflected data is written to the pool 13. The write processing (backend) is performed synchronously or asynchronously with the write processing (frontend). The write processing (backend) is executed by the read / write program 1105.
[0095] The read / write program 1105 judges (S2901) whether or not there is dirty data in the cache unit 903. If the judgment result of S2901 is true (S2901: Yes), the read / write program 1105 causes the snapshot append program 1106 to execute a snapshot allocation judgment process (S2902).
[0096] 30 shows the flow of the Snapshot allocation judgment process. In the Snapshot allocation judgment process, it is determined whether to newly allocate an area on SS-VDEV11S to the area of PVOL10P that is the target of a write request from the host, or to overwrite the area already allocated on SS-VDEV11S. The Snapshot allocation judgment process is executed by the Snapshot allocation judgment program 1113 called from the read / write program 1105.
[0097] The Snapshot allocation judgment program 1113 first judges whether the generation # of the Dir-Info of the target VOL (the VOL to which data is written) matches the generation # of the SS-Mapping-Info before the append (S3001). If the judgment result of S3001 is false (S3001: No), the Snapshot allocation judgment program 1113 executes a Snapshot append process (S3009).
[0098] If the result of the determination in S3001 is true (S3001: Yes), the process proceeds to S3002. If the result of the determination in S3001 is true, the data in the range for which a write request is received from the host is not referenced by other snapshots (hereinafter referred to as an independent state), so by overwriting the data at the same address on the SS-VDEV, it is possible to eliminate the need to update the Snapshot allocation management table and the SS-Mapping management table.
[0099] Next, the Snapshot allocation decision program 1113 refers to SS-Mapping-Info and acquires how much data in the write target range is distributed and stored on the SS-VDEV (S3002). This number is hereafter called the distribution number. For example, when 256 KB of data is written from the host, if the pre-update SS-Mapping is stored in three chunks of 32 KB, 128 KB, and 96 KB, each of which is stored in a different address on the SS-VDEV, the distribution number is counted as 3. This situation can occur when only a part of the data is updated after the snapshot is created. If the distribution number is large, it is necessary to separately perform the process of updating the mapping from the SS-VDEV to the CR-VDEV or Dedup-VDEV, so the write processing overhead increases and the throughput decreases. For this reason, when a write of a size larger than the management unit of the mapping information is requested and distributed on the SS-VDEV, it is expected that the effect of reducing the amount of processing can be expected. In this case, the data is in a single state, and the snapshot append process is performed and the continuous area on the SS-VDEV is collectively allocated.
[0100] Next, the Snapshot allocation decision program 1113 decides whether the number of distributions found in S3002 is equal to or greater than a threshold (S3003). If the number of distributions is equal to or greater than the threshold (S3003: Yes), the program proceeds to S3004. If the number of distributions is equal to or less than the threshold (S3003: No), the program proceeds to S3007. The threshold value here is set in order to switch between updating the mapping information multiple times in a distributed state and allocating new areas and updating the mapping information collectively, which is more advantageous in terms of performance. The specific threshold value is set according to the implementation and performance characteristics of the actual program. For example, if the number of distributions is up to 2 and overwriting is more effective than allocating a new area in the Snapshot space, the threshold value can be set to 3.
[0101] Next, the Snapshot allocation judgment program 1113 judges whether or not a new area can be allocated on the SS-VDEV for the area to be written that is equal to or less than the current distribution number (S3004). Specifically, it refers to the Snapshot allocation management table 1008 and judges whether or not an unallocated continuous area can be secured. If it is judged that allocation is possible (S3004: Yes), the Snapshot allocation judgment program 1113 executes a Snapshot append process (S3008). If it is judged that allocation is not possible (S3004: No), the program proceeds to S3005.
[0102] Next, the Snapshot allocation judgment program 1113 judges whether a new VDEV can be allocated to the SS-VDEV (S3005). If it is judged that allocation is possible (S3005: Yes), the process proceeds to S3006, and the new VDEV is allocated to the SS-VDEV. Next, the Snapshot allocation judgment program 1113 executes a Snapshot append process (S3008), and allocates a new area on the SS-VDEV. If it is judged that allocation of a new VDEV is not possible (S3005: No), the process proceeds to S3007, and overwriting is performed on the SS-VDEV.
[0103] When proceeding to S3007, the Snapshot allocation judgment program 1113 executes the Dedup append process. In this case, unlike when proceeding to the snapshot append process of S3008, the snapshot append process is skipped and overwriting is performed at the same address on the SS-VDEV. This makes it possible to eliminate the need to update the Snapshot allocation management table and the SS-Mapping management table.
[0104] 31 shows the flow of snapshot append processing. Snapshot append processing is processing for newly allocating an area on SS-VDEV11S to the area of PVOL10P that is the target of a write request from the host. Snapshot append processing is executed by a snapshot append program 1106 called from a snapshot acquisition program 1101 or a read / write program 1105.
[0105] The snapshot append program 1106 updates the snapshot allocation management table 1008 to secure a new area (a block address where the status 1902 is "0") in the target SS-VDEV11S (SS-VDEV11S corresponding to SS-Family 9 including the target VOL (e.g., the SVOL to be acquired or the VOL to which data is written)) (S3101). Then, the snapshot append program 1106 causes the Dedup append program 1107 to execute the Dedup append process (S3102).
[0106] Thereafter, the snapshot append program 1106 updates the SS-Mapping management table 1010 (S3103). In S3103, for example, the snapshot append program 1106 sets the latest generation # (the generation # indicated by the latest generation table 1005) to the generation # 2104 corresponding to the Mapping-Info # of the target SS-Mapping-Info. The "target SS-Mapping-Info" referred to here is the SS-Mapping-Info corresponding to the data in the target VOL.
[0107] The snapshot appending program 1106 updates the Dir management table 1009 corresponding to the Dir-Info of the target SS-VDEV 11S (S3104). In S3104, the SS-Mapping-Info (information indicating the reference address in the SS-VDEV) for the data to be written is associated with the address of the data in VOL 10.
[0108] The snapshot append program 1106 refers to the generation management tree table 1007 (Dir-Info generation management tree 70) (S3105) and determines whether the generation # of the Dir-Info of the target VOL (VOL to which data is written) matches the generation # of the SS-Mapping-Info before appending (S3106).
[0109] If the judgment result of S3106 is true (S3106: Yes), it means that after data was stored in the area on the SS-VDEV pointed to by the SS-Mapping-Info before the addition, no snapshot was created that shares the area. In other words, it can be determined that the area on the SS-VDEV pointed to by the SS-Mapping-Info before the addition has become garbage. Therefore, the snapshot addition program 1106 initializes the target entry in the SS-Mapping management table 1010 before the addition, releases the target entry in the snapshot allocation management table 1008 (S3107), and ends the process.
[0110] If the result of the determination in S3106 is false (S3106: No), it means that after data was stored in the area on the SS-VDEV pointed to by the SS-Mapping-Info before the addition, a snapshot that shares that area was created and the generation # of the DIR-Info was incremented. In this case, the area on the SS-VDEV before the addition does not become garbage, so the SS-Mapping management table 1010 before the addition is left as it is and the process ends.
[0111] 32 shows the flow of the Dedup append process. The Dedup append process is executed by the Dedup append program 1107 called from the snapshot append program 1106.
[0112] The Dedup append program 1107 judges whether or not duplicate data exists among the data stored in the Pool (S3201). Although the details are not described in the figure, if all data is checked to see if there is data with the same contents, the amount of calculation becomes enormous, so a method is adopted in which a representative value of the data, such as a hash value, is calculated by performing an operation using a hash function for each data, and a comparison process is performed only between data whose representative values match. If the judgment result of S3201 is false (S3201: No), the Dedup append program 1107 causes the compressed append program 1108 to execute a compressed append process (S3207). As a result, the compressed data of the data whose storage destination is the SS-VDEV11S is stored in the CR-VDEV11C without passing through the Dedup-VDEV11D and without changing the CPU 211 which is the processing subject.
[0113] If the determination result of S3201 is true (S3201: Yes), the Dedup appending program 1107 updates the Dedup allocation management table 1014 (S3202). In S3202, an entry for Dedup allocation information corresponding to data whose storage destination is the target Dedup-VDEV 11D is added to the Dedup allocation management table 1014.
[0114] The Dedup appending program 1107 updates the CR-Mapping management table 1012 (S3203). In S3203, an entry for CR-Mapping-Info corresponding to the data whose storage destination is the Dedup-VDEV 11D is added to the CR-Mapping management table 1012.
[0115] The Dedup append program 1107 updates the Dir management table 1009 corresponding to the Dir-Info of the target Dedup-VDEV 11D (S3204). In S3204, the CR-Mapping-Info (information indicating the reference address in the CR-VDEV 11CC) for the duplicate data is associated with the address of the data in the target Dedup-VDEV 11D. In this way, the capacity of the Pool used is reduced by associating the duplicate data with data that has already been stored.
[0116] The Dedup appending program 1107 initializes the CR-Mapping management table 1012 before appending (S3205). The Dedup append program 1107 invalidates the pre-update allocation information (S3206). In S3206, the Dedup append program 1107 updates the Dedup allocation management table 1014 and the compression allocation management table 1011. Also, in S3206, for an area in the Dedup allocation management table 1014 where the number of allocation destinations has become zero, the target entry of the compression allocation information is garbage-generated.
[0117] 33 shows the flow of the compression and appending process. The compression and appending process is executed by a compression and appending program 1108 called from a Dedup appending program 1107. The compression and append program 1108 compresses the data to be written (S3301). The compression and append program 1108 updates the compression allocation management table 1011 (S3302). In S3302, for each of one or more sub-blocks that are storage destinations for the compressed data in S3301, an entry corresponding to that sub-block is updated.
[0118] The compression and appending program 1108 causes the destaging program 1109 to execute the destaging process (S3303). The compression and appending program 1108 updates the CR-Mapping management table 1012 after the appending (S3304). The compression and appending program 1108 updates the Dir management table 1009 (S3305).
[0119] The compression and appending program 1108 initializes the target entry in the CR-Mapping management table 1012 before appending (S3306). The compression and appending program 1108 invalidates the pre-update allocation information (S3307). In S3307, the Dedup appending program 1107 updates the Dedup allocation management table 1014 and the compression allocation management table 1011. Also, in S3307, for an area where the number of allocation destinations in the Dedup allocation management table 1014 has become zero, the target entry of the compression allocation information is garbage-generated.
[0120] 34 shows the flow of the destage process. The destage process is executed by the destage program 1109 called from the compression and append program 1108. The destaging program 1109 determines whether or not there is a RAID stripe's worth of append data (one or more compressed data) in the cache unit 903 (S3401). A "RAID stripe" is a stripe in a RAID group (a storage area spanning multiple SSDs 220 that make up the RAID group). If the RAID level of the RAID group requires parity, the size of the "RAID stripe's worth of append data" may be the size of the stripe minus the size of the parity. If the determination result of S3401 is false (S3401: No), the processing ends.
[0121] If the result of the determination in S3401 is true (S3401: Yes), the destaging program 1109 references the Pool-Mapping management table 1015 and determines whether or not page 14 has been allocated to the storage destination (address in the CR-VDEV 11C) of the additional data for the RAID stripe (S3402). If the result of the determination in S3402 is false (S3402: No), the process proceeds to S3405.
[0122] If the determination result of S3402 is true (S3402: Yes), the destaging program 1109 updates the Pool allocation management table 1016 (S3403). Specifically, the destaging program 1109 allocates page 14. In S3403, entries (for example, status 2704, allocated VDEV# 2705, and allocated address 2706) corresponding to the allocated page 14 in the Pool allocation management table 1016 are updated.
[0123] The destaging program 1109 registers the page #2602 of the allocated page in the entry corresponding to the storage destination of the additional data for the RAID stripe in the Pool allocation management table 1016 (S3404).
[0124] The destaging program 1109 writes the additional data for the RAID stripe to the stripe that is the basis of the page (S3405). If the RAID level requires parity, the destaging program 1109 generates parity based on the additional data for the RAID stripe, and writes the parity to the stripe as well.
[0125] 35 is a flowchart showing the processing steps of the read process. The read process is executed by the Read / Write program 1105 in response to a read request from the host device. First, in S3500, the Read / Write program 1105 acquires the address in the PVOL or Snapshot of the data targeted by the read request from the server system 202. Next, in S3701, the Read / Write program 1105 determines whether the targeted data of the read request is a cache hit. If the targeted data of the read request is a cache hit (S3501: Yes), the Read / Write program 1105 moves the process to S3508, and if there is no cache hit (S3501: No), the process moves to S3502.
[0126] In S3502, the Read / Write program 1105 references the Dir management table 1009 and the SS-Mapping management table 1010 to obtain an address on the SS-VDEV as the reference destination based on the address in the PVOL / Snapshot obtained in S3500. At this time, if the size of the target data of the read request is larger than the management unit of the SS-Mapping management table 1010, all entries are referenced to obtain an address on the SS-VDEV.
[0127] Next, in S3503, the Read / Write program 1105 refers to the Dir management table 1009 and the CR-Mapping management table 1012, and acquires an address within the CR-VDEV or the Dedup-VDEV based on the address within the SS-VDEV acquired in S3502.
[0128] Next, in S3504, it is determined whether the reference address acquired in S3503 is on the Dedup-VDEV. Specifically, the reference CR-VDEV#2303 in the CR-Mapping management table 1012 is acquired, and the CR-VDEV management table 1002 is further referenced to identify the VDEV#1301 whose CR-VDEV#2303 and CR-VDEV#1302 match. If the result of the determination in S3504 is false (S3504: No), the Read / Write program 1105 proceeds to S3506. On the other hand, if the result of the determination in S3504 is true (S3504: Yes), the Read / Write program 1105 refers to the Dir management table 1009 and the CR-Mapping management table 1012 based on the address in the Dedup-VDEV identified in S3503, and acquires an address in the CR-VDEV (S3505).
[0129] Next, the Read / Write program 1105 decompresses the data stored at the address in the CR-VDEV identified in S3503 or S3505 and stages it in the cache memory.
[0130] Next, the Read / Write program 1105 judges whether all data in the range requested by the host device has been read onto the cache (S3507). If the result of the judgment is true, the Read / Write program 1105 proceeds to S3508, transfers the data that was a cache hit in S3501 or the data that was staged in S3506 to the host device, and ends the process.
[0131] On the other hand, if the result of S3507 is false (S3507: No), the Read / Write program 1105 returns to S3502 and re-stages the missing data on the cache. In this way, if the range of data requested by the host is stored in a continuous area on the CR-VDEV, the data can be arranged on the cache by staging it once from the drive. However, if the data is stored in a distributed manner on the SS-VDEV, Dedup-VDEV, or CR-VDEV, multiple metadata references or staging operations from the drive occur, which reduces throughput performance.
[0132] As described above, the disclosed storage system 201 comprises an SSD 220 as a storage device, and a processor 211 that accesses the storage device, and the processor 211 manages a primary volume 10P that is the subject of read / write by the host, and a snapshot volume 10S generated from the primary volume 10P, as a snapshot family 9, and the processor 211 uses a snapshot virtual device 11S, which is a logical address space corresponding to the snapshot family 9, as a storage destination for data of the primary volume 10P and the snapshot volume 10S, and when a write request is received from the host, the processor 211 switches between an overwrite process that overwrites an allocated area on the snapshot virtual device 11S and a new allocation process that allocates a new area on the snapshot virtual device 11S to the address range of the write destination, based on the reference status from the snapshot volume 10S to the address range of the write destination and the degree of distribution of the address range of the write destination in the snapshot virtual device 11S. For this reason, in a storage system that provides snapshots, the amount of mapping information updated and the amount of data transferred can be optimized according to the data length written from the host and the mapping state, and the allocation of virtual device areas can be switched to achieve high throughput performance.
[0133] In addition, since the RoW method is used for the snapshot function, snapshot data can be compressed in the same way as normal data, resulting in high capacity efficiency and reducing the amount of storage media such as flash memory and the power used, thereby conserving resources and power.
[0134] Furthermore, the processor 211 performs the overwrite process on the condition that there is no reference from the snapshot volume 10S to the address range of the write destination, and the address range of the write destination has a distribution number less than a threshold value on the snapshot virtual device 11S. For this reason, when the write destination is independent data and sufficiently contiguous, overwrite processing is performed, and when the write destination is not independent data or when the data is independent but has been written in a dispersed manner over discontinuous areas, new allocation processing can be performed.
[0135] In addition, in the disclosed storage system, when multiple volumes belonging to the same snapshot family 9 have the same data, a specified area on the snapshot virtual device 11S is allocated to the same data, and the multiple volumes reference the specified area. This allows efficient management of data referenced by snapshots.
[0136] Furthermore, in the case where multiple volumes belonging to different snapshot families 9 contain the same data, the disclosed storage system allocates a predetermined area on the deduplication virtual device 11D referenced by the snapshot virtual device 11S to the same data. This allows for efficient management of duplicate data.
[0137] In addition, the processor 211 includes information regarding the generation of the snapshot in the mapping information that manages the correspondence between addresses in volumes belonging to the snapshot family 9 and areas on the snapshot virtual device 11S, and performs the new allocation process if the generation of the address range of the write destination does not match the latest generation. Therefore, it is possible to efficiently manage whether the write destination is independent data or not.
[0138] Furthermore, when there is no reference from the snapshot volume 10S to the address range of the write destination and the address range of the write destination has a distribution number on the snapshot virtual device 11S that is equal to or greater than a threshold value, the processor 211 determines whether the area required for the new allocation process is present on the snapshot virtual device 11S, and if the necessary area is present, performs the allocation, or if the necessary area is not present, expands the snapshot virtual device 11S and then performs the new allocation process. Therefore, efficient memory operation can be achieved while expanding the snapshot virtual device 11S as required.
[0139] Furthermore, the processor 211 provides a compressed append virtual device 11C between the snapshot virtual device 11S and the storage device, and compresses data for which a write request has been made and writes the data to the storage device. Therefore, in a storage system that compresses and manages data, it is possible to efficiently utilize the resources of the virtual device and achieve high throughput.
[0140] The present invention is not limited to the above-mentioned embodiment, and various modifications are included. For example, the above-mentioned embodiment is described in detail to easily explain the present invention, and is not necessarily limited to the embodiment having all the described configurations. Moreover, the present invention is not limited to the deletion of the configurations, and it is also possible to replace or add the configurations. [Explanation of symbols]
[0141] 9: Snapshot family, 10P: Primary volume, 10S: Snapshot volume, 11C: Compressed append virtual device, 11D: Deduplication virtual device, 11S: Snapshot virtual device, 13: Pool, 70: Dir-Info generation management tree, 100: Computer system, 201: Storage system, 202: Server system, 203: Management system, 204: Storage network, 205: Management network, 210: Storage controller, 211: CPU, 212: Memory, 213: Back-end interface, 214: Front-end interface, 215: Management interface, 901: Control information unit, 902: Program unit, 903: Cache unit
Claims
1. A storage device; a processor for accessing the storage device; The processor manages a primary volume that is a target of read / write by a host and a snapshot volume generated from the primary volume as a snapshot family; the processor uses a snapshot virtual device, which is a logical address space associated with the snapshot family, as a storage destination for data of the primary volume and the snapshot volume; a storage device that stores data in the snapshot virtual device and a write request from the host to the write destination address range, and a new allocation process that allocates a new area on the snapshot virtual device to the write destination address range, based on the reference status from the snapshot volume to the write destination address range and the degree of distribution of the write destination address range in the snapshot virtual device.
2. 2. The storage system according to claim 1, A storage system characterized in that the processor performs the overwrite processing under the condition that there is no reference from the snapshot volume to the address range of the write destination and the distribution number of the address range of the write destination on the snapshot virtual device is less than a threshold value.
3. 2. The storage system according to claim 1, A storage system characterized in that, when multiple volumes belonging to the same snapshot family have the same data, a specified area on the snapshot virtual device is allocated to the same data, and the multiple volumes reference the specified area.
4. 2. The storage system according to claim 1, A storage system characterized in that, when multiple volumes belonging to different snapshot families contain identical data, a specified area on a deduplication virtual device referenced by the snapshot virtual device is allocated to the identical data.
5. 2. The storage system according to claim 1, The processor, Mapping information for managing the correspondence between the addresses in the volumes belonging to the snapshot family and the areas on the snapshot virtual device includes information about the generation of the snapshot; If the generation of the address range of the write destination does not match the latest generation, the storage system performs the new allocation process.
6. 3. The storage system according to claim 2, a processor that, when there is no reference from the snapshot volume to the address range of the write destination and the address range of the write destination has a distribution number on the snapshot virtual device that is equal to or greater than a threshold, determines whether the area required for the new allocation process exists on the snapshot virtual device, and if the required area exists, performs the allocation, or if the required area does not exist, expands the snapshot virtual device and then performs the new allocation process.
7. 2. The storage system according to claim 1, The processor, A storage system comprising: a compressed append virtual device provided between the snapshot virtual device and the storage device; and compressing data requested for writing and writing the data to the storage device.
8. A data processing method for a storage system including a storage device and a processor that accesses the storage device, comprising: The processor manages a primary volume that is a target of read / write by a host and a snapshot volume generated from the primary volume as a snapshot family; the processor uses a snapshot virtual device, which is a logical address space associated with the snapshot family, as a storage destination for data of the primary volume and the snapshot volume; A data processing method characterized in that, when the processor receives a write request from the host, it switches between an overwrite process that overwrites an allocated area on the snapshot virtual device and a new allocation process that allocates a new area on the snapshot virtual device to the address range of the write destination, based on the reference status from the snapshot volume to the address range of the write destination and the degree of distribution of the address range of the write destination in the snapshot virtual device.
Citation Information
Patent Citations
Storage controller and storage control method
JP2020047036A
Deduplication as infrastructure to avoid data movement for copy-on-write snapshots
JP2020536342A
Storage system and data copying method for storage system
JP2023056222A
Storage controller and storage control method
US10817209B2
Method and apparatus to manage groups for deduplication
US20110258404A1