Operation management system and operation management method
The operation management system accelerates data recovery in cloud storage systems by using high-speed virtual drives for interim storage and parallel migration, reducing the time to make cloud storage systems accessible from a host.
Patent Information
- Application Number
- JP2024001239
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-09
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-01-09
AI Technical Summary
In Active/Passive type disaster recovery systems, the recovery time for cloud storage systems to become available from a host is prolonged due to the need to start computer resources, which are typically not operated until needed, leading to delays in data restoration.
An operation management system and method that stores restored data in a high-speed virtual drive and migrates it to a logical volume while maintaining access, utilizing an external volume for faster access and parallel processing to reduce recovery time.
This approach significantly shortens the recovery time by enabling immediate host access during data migration, optimizing the use of high-speed virtual drives for faster data availability.
Smart Images

Figure 2025107796000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an operation management system and an operation management method, and is suitable for application to an operation management system related to a technology for executing input / output processing of data with a host, for example.
Background Art
[0002] In recent years, an operation mode called a hybrid cloud has emerged, in which on-premises IT (Information Technology) assets and a public cloud are combined and used according to cost and usage. This public cloud is characterized in that, compared with on-premises IT assets, necessary computer resources can be flexibly used on a pay-per-use basis. For example, a virtual machine, which is one of the computer resources provided by a public cloud (hereinafter simply referred to as "cloud"), incurs a charge only while it is running, and no charge is incurred while it is stopped.
[0003] Under such a charging system, for example, a system called Active / Passive type disaster recovery (hereinafter abbreviated as "DR") has emerged. The Active / Passive type DR only creates a backup of data, and at the time of a disaster, allocates necessary computer resources to restore the data and operates a recovery site (hereinafter referred to as a "secondary site").
[0004] Non-Patent Document 1 discloses a technique related to the above-described Active / Passive type DR. The technique disclosed in Non-Patent Document 1 can also be applied to a storage system. The storage system stores data to be protected. The data is backed up to the cloud in advance. At the time of a disaster, the storage system is constructed using the computer resources of the cloud, and the data backed up as described above is restored. By doing so, it is possible to recover the storage system at the time of a disaster while suppressing the cost of the secondary site in normal times.
Prior Art Documents
Non-Patent Literature
[0005]
Non-Patent Literature 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] As described above, in the Active / Passive type DR, in order to reduce the operation cost, it is common not to operate computer resources until they are put into use. Therefore, in order to make the storage system on the cloud (hereinafter referred to as the cloud storage system) available from the host, it is necessary to start the computer resources, so it takes a certain recovery time.
[0007] The present invention has been made in consideration of the above points, and intends to propose an operation management system and an operation management method capable of shortening the recovery time until it becomes available from the host.
Means for Solving the Problems
[0008] In order to solve such problems, in the present invention, there is provided an operation management system for a computer, wherein the computer includes a backup data store for storing backup data, and a restore computer resource for interpreting the backup data and restoring the data. When the operation management system restores the backup data to a logical volume in a virtual drive, the operation management system stores the restored data of the backup data in an external volume of a high-speed virtual drive that can be accessed at a higher speed than the virtual drive, controls the restore computer resource so as to enable access to the restored data, and migrates the restored data from the external volume of the high-speed virtual drive to the logical volume of the virtual drive while maintaining an accessible state to the restored data in the external volume of the high-speed virtual drive.
[0009] Further, in the present invention, there is provided an operation management method for a computer, wherein the computer includes a backup data store for storing backup data, and a restore computer resource for interpreting the backup data and restoring the data. When the operation management system restores the backup data to a logical volume in a virtual drive, the operation management system stores the restored data of the backup data in an external volume of a high-speed virtual drive that can be accessed at a higher speed than the virtual drive, and includes an access control step of controlling the restore computer resource so as to enable access to the restored data, and a migration processing step of migrating the restored data from the external volume of the high-speed virtual drive to the logical volume of the virtual drive while maintaining an accessible state to the restored data in the external volume of the high-speed virtual drive.
Advantages of the Invention
[0010] According to the present invention, it is possible to shorten the recovery time until it becomes available from the host.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5A
Figure 5B
Figure 5C
Figure 6A
Figure 6B
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14A
Figure 14B
Figure 14C
Figure 15
Figure 16A
Figure 16B
Figure 17A
Figure 17B
Figure 18
Embodiments for Carrying Out the Invention
[0012] Hereinafter, based on the drawings, one embodiment of the present invention will be described in detail.
[0013] (1) First Embodiment In the first embodiment, a configuration for shortening the recovery time until the host becomes accessible by restoring the backed-up data (hereinafter referred to as "backup data") to a logical volume on a cloud storage system in a short time will be described. In the following embodiments, regarding such a recovery time, the target recovery time is also referred to as the "target recovery time".
[0014] FIG. 1 is a block diagram showing a configuration example of the entire system according to the first embodiment. In the illustrated example, a data center 1 that performs main operations during normal times, a data center 2 that serves as a backup destination and a recovery destination in the event of a disaster, a terminal 80, and a network 70 are provided.
[0015] The data center 1 is, for example, an on-premises system owned by a user as an IT (Information Technology) asset, and includes a storage system 10 and at least one host. In the data center 1, the storage system 10 performs input / output processing of data with the host 11.
[0016] The data center 2 is, for example, a virtual data center provided by a provider of public cloud services.
[0017] The data center 2 is an example of a computer. In the present embodiment, the data center 2 has at least a restore processing instance 30, a backup data store 40, and an operation management system 50, and preferably has a cloud storage system 20, a virtual computer resource providing service 60, and a host 21. Among these, the cloud storage system 20 and the restore processing instance 30 shown by dotted lines in FIG. 1 are stopped or deleted during normal times and are started or constructed when necessary.
[0018] The backup data store 40 has a storage area in which backup data of data being operated in the storage system 10 in the data center 1 is stored. The backup data store 40 is realized, for example, by object storage in a public cloud service. The storage area of the backup data store 40 is configured by, for example, an inexpensive object storage device in order to suppress costs. For this reason, in the present embodiment, the backup data stored in the storage area is stored in the backup data store 40 in a mode that is difficult to access at high speed from the host, for example. The storage mode of the backup data will be described later.
[0019] The restore processing instance 30 is an example of computer resources and is activated by the operation management system 50 as needed. The restore processing instance 30 executes a restore process for restoring the backup data of such a backup data store 40. The restore processing instance 30 is an example of restore computer resources, appropriately accesses the backup data store 40 in which the backup data is stored, interprets the backup data, restores the backup data, and finally stores the restored data in a logical volume in a virtual drive (restore process). Note that a plurality of restore computer resources may be executed as needed.
[0020] The cloud storage system 20 is a virtual storage system constructed as software using a virtual machine group and a virtual drive group of a public cloud service. The cloud storage system 20 is an example of a virtual storage system and is activated by the operation management system 50 as needed. The cloud storage system 20 configures an external volume 280, which will be described later, as at least one restore source volume in which backup data can be stored as a part of a plurality of logical volumes in a virtual drive, and executes a migration process for moving the backup data from the external volume 280.
[0021] The operation management system 50 is a computer in which at least one program operates. The operation management system 50 is a system that manages the backup process of the storage system 10 and the data being operated in the data center 1 during normal times, and controls the restore process when necessary or during a disaster. In this embodiment, the operation management system 50 is provided in the data center 2, but it is not limited to this, and it may be provided in the data center 1. Also, in this embodiment, the operation management system 50 is described as an independent element, but for example, it may be configured as a part of the cloud storage system 20, or may be configured as a part of the storage system 10 or the host.
[0022] In this embodiment, when the operation management system 50 restores backup data to a logical volume (corresponding to the logical volume 250 described later) in a virtual drive by the restore processing instance 30, the operation management system 50 stores the data restored from the backup data in an external volume of a high-speed virtual drive (corresponding to the external volume 280 described later) that can be accessed faster than the virtual drive, controls the restore processing instance 30 so that the restored data can be accessed, and migrates (migration processing) the restored data from the external volume of the high-speed virtual drive to the logical volume of the virtual drive while maintaining a state in which the external volume of the high-speed virtual drive can access the restored data.
[0023] The virtual computer resource providing service 60 is a front end for providing virtual machines and virtual drives. The virtual computer resource providing service 60 provides virtual computer resources required in the data center 2 including the cloud storage system 20 according to requests and manages charging.
[0024] The data center 2 provides a plurality of types (lineups) of virtual machines and virtual drives depending on performance and cost. The virtual computer resource providing service 60 also has a function of changing the type of virtual drive to be used according to requests.
[0025] The data center 1 and the data center 2 are interconnected via the network 70. The network 70 is, for example, the Internet or a dedicated Ethernet (registered trademark) line. The terminal 80 is, for example, a computer or a mobile terminal. The user can access the systems and services of both data centers 1 and 2 using the terminal 80. Note that the terminal 80 may be arranged in either of the data centers 1 and 2.
[0026] Although illustration is omitted, in data centers 1 and 2, each device and system are connected by a network and can communicate with each other within the range permitted by security.
[0027] FIG. 2 is a block diagram showing a configuration example of the cloud storage system 20 shown in FIG. 1. The cloud storage system 20 is a computer environment for operating storage control software having functions common to or similar to those of the above-described storage system 10. The cloud storage system 20 includes a virtual machine group 210, a virtual drive group 220, a virtual machine startup image 230, configuration information 240, and redundant protection logical volumes 250a, 250b, 250c (hereinafter collectively referred to as logical volumes 250) provided through the control thereof, a host connection I / F (InterFace) 260, an external virtual drive 270, and an external volume 280 that presents the data of the external virtual drive 270 as a logical volume.
[0028] The virtual machine group 210 is composed of virtual computers provided by data center 2, and is a so-called storage controller in which storage control software operates using the CPU (Central Processing Unit) and memory of the virtual computer. The storage control software will be described with reference to FIG. 3 below.
[0029] The virtual drive group 220 is virtual drives (hereinafter collectively referred to as “virtual drives”) provided by data center 2, and is used to provide logical drives through the storage control software operating in the virtual machine group 210.
[0030] Note that a plurality of types (lineups) of virtual drives having different costs and performances are prepared. In the cloud storage system 20 of the present embodiment, a standard SSD (Solid State Drive; hereinafter collectively referred to as “Standard SSD”) is used.
[0031] The virtual machine startup image 230 is a machine image that includes an OS (Operating System) and storage control software for starting up and operating the virtual machine group 210.
[0032] The configuration information 240 is an area where various setting information, reference information, operation logs, etc. for the operation of this cloud storage system 20 are stored, and is managed by, for example, a database.
[0033] The logical volume 250 is a logical capacity resource provided through the control of the virtual machine group 210, and is a logical volume configured in the virtual drive group 220. The logical volume 250 is a logical capacity unit recognized by the host. The data stored in the logical volume 250 is redundant for data protection using technologies such as RAID (Redundant Array of Independent Disks) and Erasure Coding.
[0034] The host I / F 260 is an interface for the host to access the logical volume 250 to the external volume 280. The host I / F 260 is provided so that a logical volume can be identified, for example, by an IP (Interna Protocol) address and an iSCSI (Internet Small Computer System Interface) name. Multiple host I / F 260s may be provided, and it may be equipped with a security function so that only an arbitrary host can recognize or access the logical volume.
[0035] The external virtual drive 270 is a virtual drive provided through the virtual computer resource providing service 60. The external virtual drive 270 is a single drive different from the virtual drive group 220 that constitutes the logical volume 250. In this embodiment, in order to distinguish it from the virtual drive group 220 within the cloud storage system 20, it is collectively referred to as the "external virtual drive 270". Note that multiple external virtual drives 270 may exist.
[0036] The external volume 280 is a volume virtualized so that data can be accessed from a host in the same manner as the logical volume 250 by the function of storage control software described later. The external volume 280 is an example of a high-speed virtual drive that can be accessed at a higher speed than the virtual drive that constitutes the logical volume 250. Note that a plurality of external volumes 280 can be configured in the same manner as the external virtual drive 270.
[0037] FIG. 3 is a block diagram showing a configuration example of the storage control software P120 of the cloud storage system 20 shown in FIG. 1. The storage control software P120 is included in the virtual machine startup image 230 and is executed using the CPU (Central Processing Unit) and memory of the virtual machine group 210.
[0038] The storage control software P120 is composed of a host I / F control unit 2110, a logical volume control and capacity pool control unit 2120, a data redundancy and distributed storage control unit 2130, a configuration management and monitoring control unit 2140, a migration control unit 2150, an external volume control unit 2160, and a control API group 2170.
[0039] The host I / F control unit 2110 is a control program that processes I / O requests from a host via the host I / F 260. The host I / F control unit 2110 controls so that the first storage area of the external volume 280, which is an example of the restore source volume to be migrated, and the second storage area to which data is written to the external volume 280 do not overlap during the execution of the migration process.
[0040] The logical volume control and capacity pool control unit 2120 is a program that processes data read and write access from the corresponding logical volume in response to a request from a host and manages the storage free capacity held by the cloud storage system 20 as a capacity pool.
[0041] The data redundancy and distributed storage control unit 2130 is a program that performs address conversion, data compression, redundancy processing for data protection, etc. on accesses to logical volumes, and controls the data stored in the virtual drive group 220.
[0042] The configuration management and monitoring control unit 2140 reflects the configuration and setting information of the host I / F, logical volumes, and virtual drive group 220 in the configuration information 240, and monitors the usage status of the CPU, memory, network, etc., as well as the health status of the virtual machine group 210 and virtual drive group 220, etc. It is a program that performs control.
[0043] The configuration management and monitoring control unit 2140 is an example of a configuration control unit, and configures a logical volume 250 as at least one restore destination volume as part of a plurality of logical volumes in a virtual drive.
[0044] The migration control unit 2150 is a control program that controls the process of copying the restored data from the external volume 280 to the logical volume 250 (that is, migrating the volume to which the host accesses) while continuing the input / output processing with the host (hereinafter also collectively referred to as the "migration process"). In this control, without making the host aware that the access destination has changed, the access destination volume is internally switched from the external volume 280 to the logical volume 250 after the completion of the migration process. Note that the migration process can be executed for a plurality of sets simultaneously, and the target can be an external volume instead of a logical volume.
[0045] The external volume control unit 2160 is a control program that connects the external virtual drive 270 via the virtual machine group 210 and configures an external volume 280 that can be treated in the same way as the logical volume 250. Note that a plurality of external volumes 280 can be configured.
[0046] The operation control API group 2170 is an interface program for controlling instructions and responses from the operation management system 50 and the terminal 80.
[0047] The outline of the system configuration example of the data center 2 according to this embodiment is as above. Next, the operation management method and the like according to this embodiment will be described. First, as described above, the data center 2 includes a backup data store 40 for storing backup data, and a restore processing instance 30 as an example of a restore computer resource for interpreting backup data and restoring data. In this operation management method, when the operation management system 50 restores backup data to the logical volume 250 in the virtual drive, the data restored from the backup data is stored in the external volume 280 of the high-speed virtual drive that can be accessed faster than the virtual drive, and an access control step of controlling the restore processing instance 30 so that the restored data can be accessed, and a migration processing step of migrating the restored data from the external volume 280 of the high-speed virtual drive to the logical volume 250 of the virtual drive while maintaining a state where the restored data in the external volume of the high-speed virtual drive can be accessed are executed. Hereinafter, the management mode of the backup data in the backup data store 40 will be further described. Note that in the backup data store 40, since it has the following management mode, as described above, it is difficult to access it from the host at high speed as it is.
[0048] FIG. 4 is a block diagram showing a configuration example of the backup data store 40 shown in FIG. 1. The backup data store 40 includes a plurality of management areas called "buckets". In the illustrated example, a Volume100-Backup01 bucket 4110 for storing the first-generation backup of Volume100, a Volume100-Backup02 bucket 4120 for storing the second-generation backup thereof, and a Volume100-Backup03 bucket 4130 for storing the third generation thereof. Similarly, a Volume200-Backup01 bucket 4210, a Volume200-Backup02 bucket 4220, and a Volume200-Backup03 bucket 4230 for storing the first, second, and third generations of Volume200 are configured. In the following description, unless otherwise specified, the various buckets described above are also collectively referred to as "buckets".
[0049] The buckets 4110 to 4230 store a set of backup data (hereinafter referred to as "backup data set") for restoring the volume of the corresponding generation. The backup data set includes a differential bitmap 410 indicating the presence or absence of backups for each management block size on the volume, a data block group 420 in which the backed-up blocks are pre-packed, and a backup catalog 430 recording configuration information such as the identification number, device number, capacity, backup date and time of the backup source volume, and the parent-child relationship of the incremental backup generations.
[0050] Note that there is no differential bitmap 410 in the Volume100-Backup01 bucket 4110 and the Volume200-Backup01 bucket 4210, which means that the backup is not an incremental backup but a full backup targeting the entire volume capacity.
[0051] In the following description, the backup data of each volume stored in the backup data store 40 may be described as "restored volume", which is synonymous.
[0052] Figures 5A to 5C are diagrams each showing an example schematically depicting how backup data of a volume is created. The backup is performed, for example, by the cloud backup function incorporated in the storage system 10.
[0053] Figure 5A shows the state of the first-generation backup (full backup) of the volume. The volume is in state 100A where blocks indicated by "A", "B", and "C" exist. By the full backup, a data block group 420A storing "A", "B", and "C" is stored in the backup data store 40. Since it is a full backup, the differential bitmap 410A is empty (that is, not created).
[0054] Figure 5B is a diagram showing an example of the state of the second-generation backup (incremental backup) of the volume. The volume is in state 100B where block "A" has been rewritten to "A'" from the previous state 100A. By the incremental backup, a data block group 420B storing only the changed block "A'" is created in the backup data store 40. Also, a differential bitmap 410B indicating the storage position of the updated block is created.
[0055] Figure 5C is a diagram showing an example of the state of the third-generation backup (incremental backup) of the volume. The volume is in state 100C where blocks "B" and "C" have been rewritten to "B'" and "C'", respectively, from the previous state 100B. By the incremental backup, a data block group 420C storing the changed blocks "B'" and "C'" is created in the backup data store 40. Also, a differential bitmap 410C indicating the storage position of the updated block is created.
[0056] Figures 6A and 6B are diagrams each showing an example of a backup catalog stored in the backup data store 40. Figure 6A is the backup catalog 430A at the time of the first-generation backup (full backup) of the volume.
[0057] The backup catalog 430A is stored with an object name (file name) of "Volume100 - Backup01.catalog", and information regarding the serial number R610 of the backup source, volume number R611, original volume capacity R612, volume name R613, backup generation number R614, backup date and time R515, parent catalog information R616 indicating the parent - child relationship during incremental backup, and backup capacity R617 is recorded.
[0058] In the example shown in FIG. 6A, the serial number of the backup source is "VSP56432", the volume number is "100", the original volume capacity is "16.0TB", the volume name is "Volume B", the backup generation is "the first generation", the backup date and time is "2023 - 08 - 01 00:00". Since it is the first full backup, there is no parent catalog indicating the parent - child relationship of the backup, and it is shown that the backed - up capacity is "16.0TB".
[0059] FIG. 6B is the backup catalog 430B during the second - generation backup (incremental backup) of the volume. The backup catalog 430B is stored with an object name (file name) of "Volume100 - Backup02.catalog".
[0060] Regarding the serial number R610 of the backup source, volume number R611, original volume capacity R612, volume name R613, backup generation number R614, backup date and time R515, parent catalog information R616 indicating the parent - child relationship during incremental backup, and backup capacity R617, since they are the same as those shown in FIG. 6A respectively, the description is omitted.
[0061] In the example shown in FIG. 6B, the serial number of the backup source is "VSP56432", the volume number is "100", the original volume capacity is "16.0TB", the volume name is "Volume B", the backup generation is the 02nd generation, and the backup date and time is 2023-08-01 03:00. Further, in the example shown in FIG. 6B, for incremental backup, it is shown that there exists Volume100-Backup01.catalog, i.e., backup catalog 430A, as the parent catalog indicating the parent-child relationship of the backup.
[0062] Therefore, it can be understood that when restoring the backup, it is necessary to first restore the full backup of the parent backup 430A. Also, the backup size is 3.2TB.
[0063] FIG. 7 is a block diagram showing a configuration example of the restore processing program P30. The restore processing program P30 is a group of programs that operate using the CPU and memory on the restore processing instance 30. The restore processing program P30 includes a virtual drive I / O control unit 3010, a logical volume I / O control unit 3020, a backup data store I / O control unit 3030, a backup data analysis control unit 3040, and a control API group 3050.
[0064] The virtual drive I / O control unit 3010 is a program that mounts a single virtual drive and reads and writes block data. Note that the virtual drive is also used as the external virtual drive 270 shown in FIG. 2.
[0065] The logical volume I / O control unit 3020 is a program that connects to the cloud storage system 20, mounts a logical volume, and reads and writes block data.
[0066] The backup data store I / O control unit 3030 is a program that reads and writes the backup data set of the backup data store 40.
[0067] The backup data analysis control unit 3040 is a program that identifies the parent-child relationships of full backups and incremental backups based on the backup catalog 430, constructs the processing steps necessary for restoring the target generation, and controls the writing of block data to the appropriate addresses of the virtual drive or logical volume at the restore destination while referring to the differential bitmap 410.
[0068] The control API group 3050 is an interface program for controlling instructions and responses from the operation management system 50 and the terminal 80.
[0069] Figure 8 is a block diagram showing a configuration example of the operation management system 50 shown in Figure 1. The operation management system 50 includes a GUI (Graphic User Interface) providing control unit 5010, a virtual computer resource control API 5020, a cloud storage system control API 5030, a backup / restore management unit 5040, an information management DB group 5050, and a monitoring control unit 5060.
[0070] The GUI providing control unit 5010 is a program for providing a Graphical User Interface (GUI) when the user operates from the terminal 80.
[0071] The virtual computer resource control API (Application Programming Interface) 5020 is a program that operates the API of the computer resources provided by the virtual computer resource providing service 60. The virtual computer resource control API 5020, for example, requests the virtual computer resource providing service 60 to create a virtual drive of a predetermined capacity or starts the virtual machine group 210 that constitutes the cloud storage system 20.
[0072] The cloud storage control API (Application Programming Interface) 5030 is a program for instructing the storage control software P120 operating in the cloud storage system 20. The cloud storage system control API 5030 performs, for example, creation of a logical volume 250 and instruction for execution of migration processing.
[0073] The backup / restore management unit 5040 is a program for managing the access destination URL (Uniform Resource Locator) of the backup data store 40, access rights, the backup schedule of the storage system 10, device information of the restore destination, etc. The backup / restore management unit 5040 also starts a restore processing instance 30 and gives an instruction to restore to a logical volume or a virtual drive. In the present embodiment, when the restore processing instance 30 restores backup data to the logical volume 250 in the virtual drive, the restore processing instance 30 stores the restored data of the backup data in the external volume 280 of the high-speed virtual drive that can be accessed faster than the virtual drive, and controls so that the host can access the data that has been restored. The above-described migration control unit 2150 migrates the restored data from the external volume of the high-speed virtual drive to the logical volume of the virtual drive while continuing to make the restored data in the external volume of the high-speed virtual drive accessible.
[0074] The information management DB group 5050 is a group of databases that manage various data necessary for the operation management system 50 to control.
[0075] The monitoring control unit 5060 is a program for acquiring various information of the storage system 10 and the restore processing instance 30 to be managed and periodically monitoring them.
[0076] FIG. 9 is a diagram showing a configuration example of a performance management table T501 included in the information management DB group 5050 shown in FIG. 8. The performance management table T501 is an example of a performance management table, and manages information on the performance of a plurality of virtual drives including a normal virtual drive available in the data center 2 and at least one high-speed virtual drive (hereinafter, the virtual drive with the fastest access speed among the high-speed virtual drives is also referred to as the "fastest virtual drive") that can be accessed faster than the normal virtual drive. Further, the performance management table T501 is also a table that manages information on the performance of the cloud storage system 20. The virtual drive is, for example, an SSD (Solid State Drive).
[0077] The performance management table T501 manages the type C511, the maximum IOPS (Input / Output Per Second) performance C512, and the maximum throughput performance C513 of the cloud storage system 20 and each virtual drive. The unit of the information on the maximum throughput performance C513 is, for example, MB / s. By referring to the performance management table T501, information on the performance of each virtual drive available in the data center 2 can be obtained.
[0078] The above-described configuration management and monitoring control unit 2140 is an example of a configuration control unit, refers to the performance management table T501, selects the fastest virtual drive as an example of a high-speed virtual drive from among a plurality of virtual drives, and adopts the fastest virtual drive instead of the virtual drive to be used. In this embodiment, when there are a plurality of high-speed virtual drives, among the plurality of high-speed virtual drives, a virtual drive with lower performance than the fastest virtual drive but higher speed may be selected.
[0079] In the example of FIG. 7 described above, it can be seen that the cloud storage system 20 has a maximum IOPS of 50,000 and a throughput of 1,000 MB / s, the cloud storage system "No. 21" has a maximum IOPS performance of "80,000" and a maximum throughput of "2,000 MB / s", and the cloud storage system "No. 22" has a maximum IOPS performance of "256,000" and a maximum throughput of "10,000 MB / s".
[0080] Also, it can be seen that the maximum IOPS performance of "Standard SSD" is "3,000" and the maximum throughput is "125 MB / s", and similarly, "Performance SSD" has a maximum IOPS performance of "16,000" and a maximum throughput of "500 MB / s", and Ultra SSD has a maximum IOPS performance of "160,000" and a maximum throughput of "4,000 MB / s".
[0081] FIG. 10 is a flowchart showing an example of the procedure of storage recovery control processing. This storage recovery control processing is executed by the operation management system 50 based on an instruction from the terminal 80 or a determination by the operation management system 50 or the like.
[0082] First, the backup / restore management unit 5040 of the operation management system 50 refers to the performance management table T501 (see FIG. 9) and acquires the performance information of the cloud storage system 20 that is the restoration destination (step S1000). In FIG. 10, the cloud storage system 20 that is the restoration destination is abbreviated as the "restoration destination storage system" for notation.
[0083] Next, the backup / restore management unit 5040 of the operation management system 50 refers to the performance management table T501 and acquires information regarding the performance of the fastest virtual drive that can be accessed faster than the logical volume of the cloud storage system 20 (step S1010).
[0084] The backup / restore management unit 5040 of the operation management system 50 refers to the backup catalog 430 stored in the backup data store 40. The backup / restore management unit 5040 of the operation management system 50 calculates the total values of the amount of fully backed-up data (i.e., the capacity for full restore) and the amount of incrementally backed-up data (i.e., the capacity for incremental restore) required for the restore of the target generation of the volume to be restored (step S1020).
[0085] For example, in the example shown in FIG. 4, when restoring the third generation of the volume, the backup capacity of the first generation for full restore becomes the full restore capacity, and the sum of the backup capacities of the second and third generations for incremental restore becomes the incremental restore capacity.
[0086] Next, the backup / restore management unit 5040 of the operation management system 50 calculates an estimated restore time A in the destination cloud storage system 20 using the previously calculated values of the restore capacity (full restore capacity and incremental restore capacity) (step S1030).
[0087] The estimated restore time A is obtained, for example, as follows. First, for the restore of the full restore, which is generally sequential write, the required time (in minutes) for the full restore can be obtained by calculating "Full restore capacity (MB) ÷ Throughput (MB / s) ÷ 60s".
[0088] Next, for the incremental restore, which is generally random write, the number of restore blocks (i.e., the number of restore I / Os) can be obtained by calculating "Incremental restore capacity (MB) ÷ Management block size (MB)". Further, the required time (in minutes) for the incremental restore can be obtained by calculating "Number of restore I / Os ÷ IOPS performance (IO / s) ÷ 60s". In this embodiment, the estimated restore time A can be calculated from the sum of the required time for the full restore and the required time for the incremental restore. Note that the above calculation method is an example, and the estimated restore time A may be obtained by another method such as estimation by machine learning.
[0089] Next, the backup / restore management unit 5040 of the operation management system 50 calculates an estimated restoration time B when restoring to the fastest virtual drive (step S1040). The estimated restoration time B can be obtained, for example, by performing the same calculation as the estimated restoration time A. Note that for the estimated restoration time B, it may be calculated using a method different from that of the estimated restoration time A.
[0090] The backup / restore management unit 5040 of the operation management system 50 checks whether it has calculated for all restoration targets (step S1050). If there is still another restoration target volume (step S1050: No), it returns to step S1020 to repeat the process. This is, for example, a case where in addition to Volume100 in FIG. 4, Volume200 is also to be restored.
[0091] On the other hand, when there is no other restoration target volume (step S1050: Yes), the backup / restore management unit 5040 of the operation management system 50 acquires the estimated preparation time C (step S1060) of the destination cloud storage system 20 and the restore process instance startup time D (step S1065).
[0092] The estimated preparation time C is the time required until the logical volume for restoration can be prepared through the creation of the virtual drive group 220 and the creation of the capacity pool after the virtual machine group 210 is started. The estimated preparation time C can be calculated, for example, by previously holding the construction time per unit of each phase in the information management DB group 5050 of the operation management system 50 and calculating from them. Alternatively, another method such as defining the estimated preparation time C as a fixed value may be used.
[0093] The restore process instance startup time D is the time from when the restore process instance 30 is started until the restore process program P30 that starts operating according to that instruction becomes available. Since the restore process instance startup time D can generally be expected to be constant, the operation management system 50 may hold it as a fixed value, or another method such as referring to the startup times of other virtual machines in the data center 2 may be used.
[0094] Next, the backup / restore management unit 5040 of the operation management system 50 compares which is faster between the case of performing the restore process using the virtual drive that the cloud storage system 20 plans to use and the case of performing the restore process using the fastest virtual drive (step S1070). That is, it compares "the estimated restoration time of all volumes in the cloud storage system 20 (total value of A) + the estimated preparation time C of the cloud storage system 20" with "the estimated restoration time of the volume that takes the longest restoration time when restoring all volumes in parallel using the virtual drive (maximum value of B) + the restore process instance startup time D".
[0095] Because, as shown in FIG. 7, the maximum performance of the cloud storage system 20 depends on the system performance, the total required time does not change whether the volume restoration is processed in parallel or sequentially. On the other hand, since the virtual drive is the performance of a single drive, when prepared in parallel by the number of volumes, the required time is limited by the volume that takes the longest restoration time.
[0096] If the backup / restore management unit 5040 of the operation management system 50 determines, based on the result of the above comparison (step S1070), that the fastest virtual drive can execute the restore process faster in a shorter time than the virtual drive planned to be used (step S1070: Yes), it executes the restore process using the fastest virtual drive (hereinafter also collectively referred to as "high-speed recovery process") (step S1080) and ends the process. The high-speed recovery process using the fastest virtual drive will be described later.
[0097] On the other hand, when the backup / restore management unit 5040 of the operation management system 50 determines that the virtual drive scheduled to be used is as fast as or faster than the fastest virtual drive in a short time (step S1070: No), that is, when it determines that the high-speed recovery process cannot be performed, it executes a restore process (normal recovery process) using the virtual drive scheduled to be used (step S1090) and ends the process. The normal recovery process will be described later.
[0098] FIG. 11 is a flowchart showing an example of the procedure of the high-speed recovery process using a virtual drive. This high-speed recovery process is executed by the operation management system 50 when it is determined that the restoration using the fastest virtual drive can shorten the recovery time required based on the branch of step S1080 in FIG. 8.
[0099] The backup / restore management unit 5040 of the operation management system 50 first checks the number and each capacity of the volumes to be restored from the backup data (step S2000). Next, the backup / restore management unit 5040 of the operation management system 50 creates a new virtual drive based on the previous number and capacity of the restore target volumes and starts the restore process (step S2010). The details of the restore process to the virtual drive will be described in FIG. 12, which will be described later.
[0100] In parallel with the restore process to the virtual drive (step S2010), the backup / restore management unit 5040 of the operation management system 50 executes the following steps S2020 to S2100 to make preparations so that the cloud storage system 20 can be used from the host.
[0101] The backup / restore management unit 5040 of the operation management system 50 checks whether the virtual machine group 210 has been constructed as a storage controller (step S2020). If it has been constructed (step S2020: Yes), the virtual machine group 210 is started (step S2030). On the other hand, if it has not been constructed (step S2020: No), the backup / restore management unit 5040 of the operation management system 50 requests the virtual computer resource providing service 60 to allocate a predetermined number of virtual machines, and then starts the virtual machine group 210 using the virtual machine startup image 230 (step S2040), and performs the initial settings of the storage control software P120 (step S2050). The initial settings include, for example, setting the system name, setting the management subnet, setting the administrator, license authentication, etc.
[0102] Next, the backup / restore management unit 5040 of the operation management system 50 creates a virtual drive group 220 composed of virtual drives with a predetermined capacity and number (step S2060), and attaches it to the virtual machine group 210 (step S2070).
[0103] After that, the backup / restore management unit 5040 of the operation management system 50 instructs the storage control software through the cloud storage system control API 5030 to perform a format (for example, RAID format) for distributed storage and redundancy configuration on the virtual drive group 220 (step S2080), constructs a capacity pool (step S2090), and creates a logical volume 250 based on the number and capacity of the volumes confirmed in step S2000 (step S2100).
[0104] The backup / restore management unit 5040 of the operation management system 50 checks whether the restore process to the virtual drive performed in step S2010 has been completed (step S2110). If it has not been completed (step S2110: No), it waits until it is completed.
[0105] When the restore process to the virtual drive is completed in the backup / restore management unit 5040 of the operation management system 50 (step S2110: Yes), the storage control software P120 is instructed to attach the virtual drive to the cloud storage system 20 as an external virtual drive 270 and start the migration process (step S2120). Details of the migration process will be described with reference to FIG. 13 below.
[0106] When the backup / restore management unit 5040 of the operation management system 50 starts the migration process, it sets the host connection path so that the host can access the logical volume 250 through the host I / F 260 (step S2130). The migration process can be executed while receiving the I / O requests of the host. In the operation management system 50, the backup / restore management unit 5040 instructs the storage control software P120 to start the I / O access reception process from the host (step S2140), and this process ends.
[0107] FIG. 12 is a flowchart showing an example of the procedure for the restore process to the virtual drive in step S2010 of FIG. 11. First, the backup / restore management unit 5040 of the operation management system 50 creates a plurality of fastest virtual drives according to the number and capacity of the volumes to be restored based on the information (the number of volumes to be restored from the backup data and each capacity) obtained in step S2000 of FIG. 11 (step S3000). For example, in this embodiment, assume that an Ultra SSD is used from the performance management table T501 shown in FIG. 9.
[0108] The backup / restore management unit 5040 of the operation management system 50 starts a plurality of restore process instances 30 according to the number of volumes to be restored (step S3010), and attaches each fastest virtual drive to each restore process instance 30 (step S3020).
[0109] Next, the backup / restore management unit 5040 of the operation management system 50 instructs the restore processing program P30 of each restore processing instance 30 to restore the specified backup data to the fastest virtual drive (step S3030). As a result, the restore processing instances 30 operate in parallel according to the number of restore volumes, and the restore processing is performed. If the restore processing instance 30 has sufficient processing capabilities (computing performance and transfer bandwidth), the restore processing to a plurality of the fastest virtual drives may be assigned to one restore processing instance 30.
[0110] The backup / restore management unit 5040 of the operation management system 50 checks whether the processing of each restore processing instance 30 has been completed (step S3040). If there is a restore processing instance 30 for which the restore processing has been completed (step S3040: Yes), the fastest virtual drive is detached from the restore processing instance 30 (step S3050), and the restore processing instance 30 is terminated.
[0111] In order to reduce the subsequent pay-per-use cost, the fastest virtual drive for which the restore processing has been completed is requested from the virtual computer resource providing service 60, and the type of the virtual drive is changed by the operation management system 50 from the fastest type to an inexpensive type. For example, it is changed to the same "Standard SSD" as that used in the cloud storage system 20 (step S3070).
[0112] The backup / restore management unit 5040 of the operation management system 50 checks whether all the restore processing instances 30 have completed the restore processing (step S3080). If there is a restore processing instance 30 for which the restore processing has not been completed (step S3080: No), the monitoring is continued (step S3040). On the other hand, if all the restore processing instances 30 have completed the restore processing (step S3080: Yes), the process is terminated, and the process returns to the process of FIG. 9 described above and is executed.
[0113] FIG. 13 is a flowchart showing an example of the procedure of the migration process as step S2120 shown in FIG. 11. The backup / restore management unit 5040 of the operation management system 50 attaches the virtual drive in which data has been restored by the restore process shown in FIG. 12 to the cloud storage system 20 as an external virtual drive 270 (step S4000).
[0114] The backup / restore management unit 5040 of the operation management system 50 instructs the storage control software P120 to set it to use the external virtual drive 270 instead of the external volume 280 (step S4010).
[0115] Thereafter, the backup / restore management unit 5040 of the operation management system 50 starts the migration process between the external volume 280 and the logical volume 250 having the same capacity as the external volume 280 (step S4020). Note that the migration process is executed in the background.
[0116] The backup / restore management unit 5040 of the operation management system 50 checks whether the same processing has been performed on all the virtual drives in which data has been restored created in step S2010 (step S4030). If there is still an unprocessed virtual drive (step S4030: No), steps S4000 to S4020 are similarly repeated. On the other hand, if the backup / restore management unit 5040 of the operation management system 50 can start the migration process for all the virtual drives in which data has been restored (step S4030: Yes), the migration process is terminated and the process returns to the process of FIG. 9 for execution.
[0117] FIGS. 14A to 14C are diagrams showing an example of a state in which the storage control software P120 performs the migration process in the background without stopping the input / output process with the host. That is, FIGS. 14A to 14C show the state of control when the input / output process with the host is started for the logical volume 250 in step S2140 shown in FIG. 11.
[0118] FIG. 14A is an example of a state when there is a read I / O request from a host to the logical volume 250 during the execution of the migration process. In the external volume 280, data restored from backup data is stored. The data is sequentially read by the storage control software P120 and copied to the logical volume 250.
[0119] During that time, when there is a read request from the host to the logical volume 250, the storage control software P120 reads data from the external volume 280, which is the copy source, and responds to the host, thereby providing access to the restored data while continuing the migration process in the background.
[0120] FIG. 14B is an example of a state when there is a write I / O request from a host to the logical volume 250 during the execution of the migration process. As described above, the data in the external volume 280 is sequentially read by the storage control software P120 and copied to the logical volume 250.
[0121] During that time, when there is a write request from the host to the logical volume 250, the storage control software P120 writes data corresponding to the write request to both the external volume 280 and the logical volume 250, thereby continuing the migration process in the background without stopping the input / output processing with the host.
[0122] FIG. 14C is an example of a state when a read request or a write request is made from a host to the logical volume 250 after the migration process is completed. Since all the data in the external volume 280 has been copied to the logical volume 250, the external volume 280 is no longer needed hereafter, and all I / O requests are made to the logical volume 250.
[0123] FIG. 15 is a flowchart showing an example of post-processing performed by the operation management system 50 when the migration process that has been executed in the background is completed.
[0124] When the migration is completed, the backup / restore management unit 5040 of the operation management system 50 disconnects the connection to the external volume 280 that is no longer accessed (step S5000). Thereafter, the backup / restore management unit 5040 of the operation management system 50 detaches (step S5010) the external virtual drive 270 that has been controlled as an external volume and then deletes it (step S5020), and ends the process.
[0125] FIGS. 16A and 16B are flowcharts each showing an example of normal recovery processing. This normal recovery processing is executed by the operation management system 50 when the high-speed recovery processing using a virtual drive cannot be applied based on the branch of step S1070 shown in FIG. 10.
[0126] First, the backup / restore management unit 5040 of the operation management system 50 checks the number of volumes to be restored from the backup data and their respective capacities (step S6000). Next, the backup / restore management unit 5040 of the operation management system 50 makes preparations (steps S6020 to S6100) so that the cloud storage system 20 can be used by the host.
[0127] The backup / restore management unit 5040 of the operation management system 50 checks whether the virtual machine group 210 has been constructed (step S6020). If it has been constructed (step S6020: Yes), the virtual machine group 210 is started (step S6030).
[0128] On the other hand, if it has not been constructed (step S6020: No), the backup / restore management unit 5040 of the operation management system 50 requests the virtual computer resource providing service 60 to allocate a predetermined number of virtual machines, and then starts the virtual machine group 210 using the virtual machine startup image 230 (step S6040), and performs the initial setting of the storage control software P120 (step S6050). The initial setting includes, for example, setting the system name, setting the management subnet, setting the administrator, license authentication, etc.
[0129] Next, the backup / restore management unit 5040 of the operation management system 50 creates a virtual drive group 220 with a predetermined capacity and number (step S6060) and attaches it to the virtual machine group 210 (step S6070).
[0130] After that, the backup / restore management unit 5040 of the operation management system 50 instructs the storage control software P120 through the cloud storage system control API 5030 to perform a format (e.g., RAID format) for distributed storage and redundancy configuration on the virtual drive group 220 (step S6070), constructs a capacity pool (step S6090), and creates a logical volume 250 based on the number and capacity of the volumes confirmed in step S2000 (step S6100).
[0131] After that, the backup / restore management unit 5040 of the operation management system 50 instructs the storage control software P120 to set the connection host path for the restore processing instance 30 to access the logical volume 250 (step S6110).
[0132] Next, the backup / restore management unit 5040 of the operation management system 50 activates one restore processing instance 30 (step S6120). Then, the backup / restore management unit 5040 of the operation management system 50 instructs the restore processing program P30 to connect a logical volume corresponding to the capacity of the restore volume (backup data) to be restored to the restore processing instance (step S6130). Then, based on the instruction of the backup / restore management unit 5040 of the operation management system 50, the restore processing program P30 restores the backup of the specified volume to the logical volume.
[0133] The backup / restore management unit 5040 of the operation management system 50 waits until the restore processing program P30 completes the restore processing (steps S6150 and S6150: No). When it confirms completion (step S6150: Yes), the backup / restore management unit 5040 of the operation management system 50 disconnects the connection of the logical volume to which the data has been restored (step S6170).
[0134] If there is a restore volume that has not yet been restored (step S6170: Yes), the backup / restore management unit 5040 of the operation management system 50 repeats steps S6130 to S6160 until the restoration of all volumes is completed. On the other hand, when the restoration of all restore volumes is completed (step S6160: No), the backup / restore management unit 5040 of the operation management system 50 terminates the restore processing instance 30 (step S6180).
[0135] Next, the backup / restore management unit 5040 of the operation management system 50 sets the host connection path so that the host can access the logical volume 250 through the host I / F 260 (step S6190), instructs the storage control software to start the I / O access reception process from the host (step S6200), and ends this process.
[0136] 17A is a diagram showing an example of the relationship between the time required for each process in the high-speed recovery process (see FIG. 9) using a virtual drive. In the example shown, the horizontal axis represents elapsed time.
[0137] In this embodiment, restore processing by a restore processing instance 30 as an example of a restore computer resource (corresponding to processing spanning time t1 shown in the figure) and configuration of a logical volume 250 by a cloud storage system 20 (corresponding to processing spanning time t0 shown in the figure) are executed in parallel.
[0138] In this embodiment, the host I / F control unit 2110 resumes input / output processing with the host when migration processing begins.
[0139] A more detailed explanation is given below. In this high-speed recovery process, while restoration (step S2010) to the virtual drive (external volume 280) is being performed at time t1, preparation of the cloud storage system 20 (steps S2020 to S2100) is performed in parallel at time t0. Then, when restoration to the virtual drive (external volume 280) is completed, migration processing (step S2120) starts at time t1 and continues for time t2. During the migration processing, the host can access the restored data from time t1, so it can be seen that the time required for the host to start accessing is t1 (<time t2). Note that, from the branch judgment in step S1070 in FIG. 10, it is known that the time required for t1 is t1 (<time t2). <t2であることは自明である。
[0140] FIG. 17B is a diagram showing an example of the relationship between the required times of respective processes in the normal recovery process shown in FIG. 14. The horizontal axis represents the elapsed time. In this normal recovery process, in order to directly restore to the logical volume 250 without going through the external volume 280 of the virtual drive, first, the preparation of the cloud storage (steps S6020 to S6100) is carried out at time t0. Thereafter, the restoration to the logical volume 250 (steps S6110 to S6160) is carried out at time t2, and host access is started after the restoration is completed. Therefore, it can be understood that the time required until the start of host access is t0 + t2.
[0141] From the comparison between FIG. 17A and FIG. 17B, it can be seen that the high-speed recovery process (see FIG. 17A) using the virtual drive (the external volume 280 thereof) has a shorter required time until the start of host access than the normal recovery process (see FIG. 17B) by the amount that can conceal time t0 and time t2. Thus, according to the present embodiment, it is possible to shorten the recovery time until the host becomes accessible, and the target of the backup applicable by the present method can be expanded to a logical volume with a short target recovery time.
[0142] As described above, the operation management system 50 of the data center 2 according to the present embodiment includes the backup data store 40 that stores backup data, and the restore processing instance 30 as an example of the restore computer resources that interprets the backup data and restores the data. When restoring the backup data to the logical volume in the virtual drive, the operation management system 50 stores the data restored from the backup data in the external volume of the high-speed virtual drive that can be accessed faster than the virtual drive, controls the restore processing instance 30 so that the host can access the restored data, and migrates the restored data from the external volume of the high-speed virtual drive to the logical volume of the virtual drive while maintaining the state where the restored data in the external volume of the high-speed virtual drive can be accessed.
[0143] By doing so, even though the migration process is started immediately after the restoration process by the restoration process instance 30 (corresponding to "restore virtual drive" during the time t1 shown in FIG. 17A) and the configuration of the logical volume 250 by the cloud storage system 20 (corresponding to "cloud storage preparation" during the time t0 shown in FIG. 17A) is completed, the input / output process (host access) with the host can be started. Therefore, the recovery time until it becomes available from the host can be shortened.
[0144] In the data center 2 according to this embodiment, the restoration process by the restoration process instance 30 and the configuration of the logical volume 250 by the cloud storage system 20 are executed in parallel (see FIG. 17A for example). By doing so, since it is not necessary to wait until the end of the migration process (corresponding to "migration" shown in FIG. 17A) to start the input / output process with the host, the input / output process with the host can be restarted earlier, and the recovery time can be further shortened.
[0145] In this embodiment, the host I / F control unit 2110 is an example of an interface control unit, and upon the start of the migration process, resumes the input / output process with the host. By doing so, the input / output process with the host can be started without waiting for the end of the migration process.
[0146] (2) Second Embodiment Since the second embodiment has the same configuration and operation as the first embodiment, in the second embodiment, the description of the same configuration and operation as the first embodiment will be omitted, and the following will focus on the differences.
[0147] In the first embodiment, when the processing time can be shortened, the above high-speed recovery process was unconditionally performed as a restoration process using the fastest virtual drive. However, it is also conceivable that there is a margin in the overall system recovery target time.
[0148] Therefore, in the second embodiment, in consideration of this, the recovery processing method and the virtual drive are selected according to the recovery target time. Below, the description will mainly focus on the differences from the first embodiment. For other configurations without description, they are the same as those in the first embodiment.
[0149] FIG. 18 is a flowchart showing an example of the procedure of storage recovery control processing in the second embodiment. The storage recovery control processing is executed by the operation management system 50 based on an instruction from the terminal 80 or a determination by the operation management system 50, etc., and is basically the same as FIG. 10 in the first embodiment, but some processes (step S1010A, step S1040A, step S1069, step S1070A, FIG. 1075A, step S7000, step S7100, FIG. 7200) are different.
[0150] The backup / restore management unit 5040 of the operation management system 50 refers to the performance management table T501 (see FIG. 9) in the same way as step S1000 in FIG. 10, and acquires the performance information of the cloud storage system 20 that is the restoration destination (step S1000). Next, referring to the performance management table T501 (see FIG. 9), the performance information of the virtual drive i available in this data center is acquired (step S1010A). In this embodiment, there are three types in the performance management table T501 regarding the performance of the virtual drive, from "Standard SSD" to "Ultra SSD". Therefore, the information with i = 1 to 3 is acquired.
[0151] Next, the backup / restore management unit 5040 of the operation management system 50 refers to the backup catalog 430 stored in the backup data store 40. Furthermore, the backup / restore management unit 5040 of the operation management system 50 calculates the total values of the fully backed-up backup data (i.e., the capacity for full restore) and the incrementally backed-up backup data (i.e., the capacity for incremental restore) required for restoring the target generation of the volume to be recovered (step S1020).
[0152] The backup / restore management unit 5040 of the operation management system 50 calculates an estimated restoration time A in the restoration destination cloud storage system 20 using the previous calculated value (step S1030). Specific examples of steps S1020 and S1030 have already been described in FIG. 10, so they will be omitted.
[0153] Next, the operation management system 50 calculates an estimated restoration time Bi when each virtual drive i is the restoration destination (step S1040A). Since the calculation method of the estimated restoration time Bi is the same as the calculation method of the restoration time B by the fastest virtual drive in step S1040 of FIG. 10, the description will be omitted.
[0154] The operation management system 50 checks whether it has calculated for all restoration targets (step S1050). If there is still another restoration target volume (step S1050: No), it returns to step S1020 and repeats the process.
[0155] On the other hand, when there is no other restoration target volume (step S1050: Yes), the operation management system 50 acquires the estimated preparation time C (step S1060) and the restore process instance startup time D (step S1065) of the restoration destination cloud storage system 20. The acquisition methods of the estimated preparation time C and the restore process instance startup time D are the same as the processes in FIG. 10.
[0156] After that, the operation management system 50 acquires the system recovery target time T (step S1069). Then, the operation management system 50 compares whether the time required to directly restore to the cloud storage system 20, that is, "the total estimated restoration time of all volumes in the cloud storage system 20 (total value of A) + the estimated preparation time C of the cloud storage system 20", is within the recovery target time T (step S1070A).
[0157] As a result of the comparison, if the recovery target time T can be achieved even by direct restoration (step S1070A: Yes), the operation management system 50 performs normal recovery processing (step S1090). Since the details of the normal recovery processing are the same as those in the first embodiment described above (see FIG. 16), the description thereof is omitted.
[0158] On the other hand, if the recovery target time T cannot be achieved by direct restoration (step S1070A: No), when the operation management system 50 restores using each virtual drive i, that is, when "when restoring all volumes in parallel using the virtual drive i, the estimated restoration time of the volume that takes the longest restoration time (the maximum value of Bi) + the restoration process instance startup time D" is compared with the recovery target time T (step S1075A).
[0159] If there is no virtual drive i that satisfies the specific conditions regarding the recovery target time T (step S1075A: No), since the operation management system 50 cannot recover within the recovery target time T, it notifies a target setting error (step S7200) and ends the process.
[0160] On the other hand, if there are a plurality of virtual drives i that satisfy the specific conditions (step S1075A: Yes), the operation management system 50 selects a plurality of high-speed virtual drives k (k ∈ i) for restoring data based on the recovery target time T of the restoration and the performance and usage charges of the virtual drives, that is, selects the virtual drive k (k ∈ i) with the lowest usage charge that satisfies the specific conditions among these plurality of virtual drives (step S7000), and causes a restoration process (high-speed recovery process) using the virtual drive k to be performed (step S7100). By doing so, it may be possible to complete the restoration process (high-speed recovery process) within the recovery target time T while suppressing costs.
[0161] Note that the high-speed recovery process using the virtual drive k is equivalent to the process in which "the fastest virtual drive" is replaced with "virtual drive k" in the operations of FIGS. 11 to 15 in the first embodiment, so the illustration and description are omitted.
[0162] According to the present embodiment, it is possible to shorten the recovery time until the host becomes accessible while suppressing the cost for switching the recovery method based on the recovery target time T, and it is possible to expand the target to which the backup by this method is applicable to a logical volume with a short target recovery time. Further, according to the present embodiment, since it is determined whether the recovery target time T can be satisfied before the data recovery process, it is possible to perform a pretest without incurring the cost and time required for the actual data recovery process.
[0163] Note that the present invention is not limited to the above-described embodiments, and includes various modifications and equivalent configurations within the scope of the appended claims. For example, the above-described embodiments have been described in detail for easy understanding of the present invention, and the present invention is not necessarily limited to those having all the configurations described. Also, each element described in parallel in the present embodiment may be in a mode in which at least one of the elements is connected in series to another element.
[0164] In the above-described embodiments, the processing may be described with "system" or "program" as the subject. However, the system is a computer resource including a processor (e.g., CPU (Central Processing Unit)), a storage resource (e.g., memory), and a communication interface device (e.g., NIC (Network Interface Card)), and the program is executed by the processor and appropriately uses the storage resource and / or the communication interface device. Therefore, the subject of the processing may be the processor or the processing performed by a computer having the processor.
Industrial Applicability
[0165] The present invention can be applied to an operation management system related to a technique for executing input / output processing of data with a host.
Explanation of Signs
[0166] 1... Data center, 2... Data center, 20... Cloud storage system, 30... Restore processing instance, 40... Backup data store, 50... Operation management system, 60... Virtual computer resource providing service
Claims
1. A computer operation management system, The computer includes: A backup data store for storing backup data; a restore computing resource for interpreting the backup data and restoring the data; The operation management system includes: When restoring the backup data to a logical volume in a virtual drive, storing data restored from the backup data in an external volume of a high-speed virtual drive that can be accessed at high speed from the virtual drive, and controlling the restore computer resource so that the restored data can be accessed; Migrating the restored data from the external volume of the high-speed virtual drive to a logical volume of the virtual drive while maintaining an accessible state of the restored data in the external volume of the high-speed virtual drive. An operation management system comprising:
2. The restoration of the data by the restoration computer resource and the configuration of the logical volume for storing the data are executed in parallel.
2. The operation management system according to claim 1 .
3. The operation management system includes: Using the start of the migration as a trigger, input / output processing with the host is resumed.
2. The operation management system according to claim 1 .
4. The operation management system includes: A high-speed virtual drive for restoring the data is selected from among a plurality of virtual drives based on a target time for restoration, the performance of the virtual drives, and a usage charge.
2. The operation management system according to claim 1 .
5. A computer operation management method, comprising: The computer includes: A backup data store for storing backup data; a restore computing resource for interpreting the backup data and restoring the data; When the operation management system restores the backup data to a logical volume in a virtual drive, an access control step of storing data restored from the backup data in an external volume of a high-speed virtual drive that can be accessed at high speed from the virtual drive, and controlling the restore computer resource so that the restored data can be accessed; A migration processing step of migrating the restored data from the external volume of the high-speed virtual drive to the logical volume of the virtual drive while continuing to enable access to the restored data in the external volume of the high-speed virtual drive; An operation management method characterized by executing the above.
Citation Information
Patent Citations
Disk array device and data control method
JP2011159150A
Asynchronous cross-region block volume replication
JP2023548373A