Operation management system and operation management method
By storing and accessing backup data in the external volume of the virtual drive and performing data migration, the long recovery time in active/passive disaster recovery systems is solved, enabling faster data recovery and access.
Patent Information
- Application Number
- CN202411250655.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-09
- Filing Date
- 2024-09-06
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In active/passive disaster recovery systems, the prior art has a problem of long recovery time, which makes it impossible to quickly recover data and enable computer resources to run when a disaster occurs.
Provides a management system and method to use the external volume of a high-speed virtual drive to access and migrate data by storing and accessing backup data in an external volume of a virtual drive, and then migrating it after recovering data in a short time.
It effectively shortens the recovery time of data that can be accessed from the host and improves the efficiency and speed of the disaster recovery process.
Smart Images

Figure CN120295832A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an operation management system and an operation management method, and is suitable for use in, for example, an operation management system related to a technology for performing input / output processing of data between a host and a device. Background Art
[0002] In recent years, there has emerged an operation mode called a hybrid cloud that combines the use of on-premises information technology (IT) assets and a public cloud according to cost and usage. The public cloud is characterized in that it can flexibly utilize the required computing resources on a pay-per-use basis compared to on-premises IT assets. For example, a virtual machine, which is one of the computing resources provided by the public cloud (hereinafter simply referred to as "the cloud"), incurs a cost only during startup and does not incur a cost during shutdown.
[0003] Under such a cost system, for example, there has emerged a system called active / passive disaster recovery (hereinafter abbreviated as "DR"). The active / passive DR only generates a backup of data, and in the event of a disaster, allocates the necessary computing resources to restore the data and makes the recovery site (hereinafter referred to as the "secondary site") operational.
[0004] The technology related to the above-described active / passive DR is disclosed in Non-Patent Document 1. The technology disclosed in Non-Patent Document 1 can also be applied to a storage system. The storage system stores data to be protected. The data is pre-backed up in the cloud. In the event of a disaster, the computing resources of the cloud are used to construct a storage system, and the backed-up data is restored as described above. Thereby, the cost of the secondary site during normal times can be suppressed, and the storage system can be restored in the event of a disaster.
[0005] As described above, in the active / passive DR, in order to reduce the operation cost, the computing resources are usually not run before use. Therefore, in order to be able to use the cloud storage system (hereinafter referred to as the "cloud storage system") from the host, it is necessary to start the computing resources, which thus takes a certain recovery time.
[0006] Non-Patent Document 1: Amazon Web Services, Inc: Disaster Recovery (DR) Architecture on AWS, Part I: Strategies for Recovery in the Cloud. 2021 / 4 / 5. https: / / aws.amazon.com / jp / blogs / architecture / disaster-recovery-dr-architecture-on-aws-part-i-strategies-for-recovery-in-the-cloud / Summary of the Invention
[0007] The present invention has been completed in consideration of the above points, and provides an operation management system and an operation management method capable of shortening the recovery time until utilization from a host is possible.
[0008] To solve this problem, in the present invention, there is provided an operation management system for a computer, the computer including: a backup data repository that stores backup data; and a recovery computer resource that interprets the backup data and restores data. When restoring the backup data to a logical volume within a virtual drive, the operation management system controls the recovery computer resource so that data obtained by restoring the backup data is stored in an external volume of a high-speed virtual drive that can be accessed at a higher speed than the virtual drive, and the restored data can be accessed, and the state of being able to access the data restored to the external volume of the high-speed virtual drive is continued, and the restored data is migrated from the external volume of the high-speed virtual drive to the logical volume of the virtual drive.
[0009] In addition, in the present invention, there is provided an operation management method for a computer, the computer including: a backup data repository that stores backup data; and a recovery computer resource that interprets the backup data and restores data. When restoring the backup data to a logical volume within a virtual drive, the operation management method performs the following steps: an access control step of controlling the recovery computer resource so that data obtained by restoring the backup data is stored in an external volume of a high-speed virtual drive that can be accessed at a higher speed than the virtual drive, and the restored data can be accessed; and a migration processing step of continuing the state of being able to access the data restored to the external volume of the high-speed virtual drive, and migrating the restored data from the external volume of the high-speed virtual drive to the logical volume of the virtual drive.
[0010] According to the present invention, the recovery time until utilization from a host is possible can be shortened. Description of the Drawings
[0011] Figure 1 is a block diagram showing a structural example of the entire system according to the first embodiment.
[0012] Figure 2 is showing Figure 1 a block diagram showing a structural example of the cloud storage system shown.
[0013] Figure 3 is showing Figure 1 a block diagram showing a structural example of the storage control software of the cloud storage system shown.
[0014] Figure 4 is showing Figure 1 a block diagram showing a structural example of the backup data repository shown.
[0015] Figure 5A is a diagram schematically showing an example of a method for generating backup data of a volume.
[0016] Figure 5B is a diagram schematically showing an example of a method for generating backup data of a volume.
[0017] Figure 5C is a diagram schematically showing an example of a method for generating backup data of a volume.
[0018] Figure 6A is a diagram showing an example of a backup directory stored in the backup data repository.
[0019] Figure 6B is a diagram showing an example of a backup directory stored in the backup data repository.
[0020] Figure 7 is a block diagram showing a structural example of the recovery processing program.
[0021] Figure 8 is showing Figure 1 a block diagram showing a structural example of the operation management system shown.
[0022] Figure 9 is showing Figure 8 a diagram showing a structural example of the performance management table included in the information management DB group shown.
[0023] Figure 10 is a flowchart showing an example of the steps of the storage recovery control process.
[0024] Figure 11 is a flowchart showing an example of the steps of the high-speed recovery process using a virtual drive.
[0025] Figure 12It is a flowchart showing an example of the steps of the recovery process to the virtual drive.
[0026] Figure 13 It is a flowchart showing an example of the steps of the migration process.
[0027] Figure 14A It is a diagram showing an example of the manner of performing the migration process in the background.
[0028] Figure 14B It is a diagram showing an example of the manner of performing the migration process in the background.
[0029] Figure 14C It is a diagram showing an example of the manner of performing the migration process in the background.
[0030] Figure 15 It is a flowchart showing an example of the post-processing performed by the management system when the migration process performed in the background is completed.
[0031] Figure 16A It is a flowchart showing an example of the normal recovery process.
[0032] Figure 16B It is shown Figure 16A A flowchart showing an example of the follow-up of the normal recovery process shown.
[0033] Figure 17A It is a diagram showing an example of the relationship of the required time for each process in the high-speed recovery process using the virtual drive.
[0034] Figure 17B It is a diagram showing an example of the relationship of the required time for each process in the normal recovery process shown in FIG. 14.
[0035] Figure 18 It is a flowchart showing an example of the overall picture of the storage recovery control process in the second embodiment. Detailed Embodiment
[0036] Hereinafter, an embodiment of the present invention will be described in detail based on the accompanying drawings.
[0037] (1) First Embodiment
[0038] In the first embodiment, a structure for shortening the recovery time until the host can access by recovering backup data (hereinafter referred to as "backup data") in a short time in a logical volume on a cloud storage system will be described. In the following embodiments, regarding such a recovery time, the target recovery time is also referred to as the "target recovery time".
[0039] Figure 1It is a block diagram showing a structural example of the entire system of the first embodiment. In the illustrated example, there are provided a data center 1 that conducts main operations during normal times, a data center 2 that serves as a backup destination and a recovery destination in the event of a disaster, terminals 80, and a network 70.
[0040] The data center 1 is, for example, an on-premises system owned by a user as an IT (Information Technology) asset, and includes a storage system 10 and at least one host. In the data center 1, input / output processing of data is performed between the storage system 10 and the host 11.
[0041] The data center 2 is, for example, a virtual data center provided by a provider of public cloud services.
[0042] The data center 2 is an example of a computer. In the present embodiment, the data center 2 at least has a recovery processing instance 30, a backup data repository 40, and an operation management system 50, and preferably has a cloud storage system 20, a virtual computer resource providing service 60, and a host 21. Among them, Figure 1 the cloud storage system 20 and the recovery processing instance 30 shown by the dashed line are stopped or deleted during normal times and started or constructed when needed.
[0043] The backup data repository 40 has a storage area for storing backup data of the data being used in the storage system 10 in the data center 1. The backup data repository 40 is implemented, for example, by an object storage in a public cloud service. The storage area of the backup data repository 40 is composed of inexpensive object storage devices, for example, in order to reduce costs. Therefore, in the present embodiment, the backup data stored in this storage area is stored in the backup data repository 40 in a manner that makes it difficult to access at high speed from the host. The storage method of the backup data will be described later.
[0044] The recovery processing instance 30 is an example of computer resources and is started by the operation management system 50 as needed. The recovery processing instance 30 performs a recovery process of recovering the backup data in the backup data repository 40 like this. The recovery processing instance 30 is an example of recovery computer resources, appropriately accesses the backup data repository 40 storing the backup data, interprets the backup data, recovers the backup data, and finally stores the recovered data in a logical volume in a virtual drive (recovery process). In addition, multiple recovery computer resources are executed as needed.
[0045] The cloud storage system 20 is a virtual storage system constructed using a virtual machine cluster and a virtual drive cluster of a public cloud service as software. The cloud storage system 20 is an example of a virtual storage system and is started by the operation management system 50 as needed. The cloud storage system 20 forms at least one external volume 280 described later, which can store backup data, as a recovery source volume in the virtual drive, as part of a plurality of logical volumes, and performs a migration process for moving the backup data from the external volume 280.
[0046] The operation management system 50 is a computer in which at least one program operates. The operation management system 50 is a system that, in addition to managing the storage system 10 in the data center 1 and the backup process of the data being used therein during normal times, also controls the recovery process when necessary or during a disaster. In addition, in the present embodiment, the operation management system 50 is provided in the data center 2, but is not limited thereto and may also be provided in the data center 1. Further, in the present embodiment, the operation management system 50 is described as an independent element, but for example, it may also be configured as part of the cloud storage system 20, or may be configured as part of the storage system 10 or the host.
[0047] In the present embodiment, when the recovery processing instance 30 recovers backup data in a logical volume (corresponding to the logical volume 250 described later) in the virtual drive, the operation management system 50 controls the recovery processing instance 30 so that the data obtained by storing the recovered backup data is stored in an external volume of a high-speed virtual drive (corresponding to the external volume 280 described later) that can be accessed faster than the virtual drive, and the recovered data can be accessed, and the state where the data recovered to the external volume of the high-speed virtual drive can be continuously accessed is maintained, and the recovered data is migrated from the external volume of the high-speed virtual drive to the logical volume of the virtual drive (migration process).
[0048] The virtual computer resource providing service 60 is a front end for providing virtual machines and virtual drives. The virtual computer resource providing service 60 provides virtual computer resources required in the data center 2, including the cloud storage system 20, according to a request, and manages the fees.
[0049] The data center 2 provides multiple types (specifications) of virtual machines and virtual drives according to performance and cost. The virtual computer resource providing service 60 also has a function of changing the type of virtual drive to be used according to a request.
[0050] The data center 1 and the data center 2 are interconnected via the network 70. The network 70 is, for example, the Internet or an Ethernet (registered trademark) dedicated line. The terminal 80 is, for example, a computer, a portable terminal, or the like. A user can use the terminal 80 to access the systems and services of the two data centers 1 and 2. In addition, the terminal 80 may also be configured in either of the data centers 1 and 2.
[0051] In addition, although not shown in the figure, in Data Centers 1 and 2, each device and system are connected via a network and can communicate with each other within a safe allowable range.
[0052] Figure 2 It represents Figure 1 A block diagram showing a structural example of the cloud storage system 20. The cloud storage system 20 is a computer environment for operating storage control software having functions common or similar to those of the above-described storage system 10. The cloud storage system 20 includes: a virtual machine group 210, a virtual drive group 220, a virtual machine startup image 230, configuration information 240, and redundant protected logical volumes 250a, 250b, 250c (hereinafter collectively referred to as logical volumes 250) provided through their control, a host connection I / F (Interface) 260, an external virtual drive 270, and an external volume 280 that regards the data of the external virtual drive 270 as a logical volume.
[0053] The virtual machine group 210 is composed of virtual computers provided by Data Center 2 and is a so-called storage controller in which the storage control software operates using the CPU (Central Processing Unit) and memory of the virtual computer. In addition, the storage control software will be described in Figure 3 described later.
[0054] The virtual drive group 220 is virtual drives provided by Data Center 2 (hereinafter collectively referred to as "virtual drives") and is used to provide logical drives through the storage control software operating in the virtual machine group 210.
[0055] In addition, multiple types (specifications) of virtual drives with different preparation costs and performances are prepared. In the cloud storage system 20 of the present embodiment, standard SSDs (Solid State Drives, hereinafter collectively referred to as "standard SSDs") are used.
[0056] The virtual machine startup image 230 is a machine image including an OS (Operating System) and storage control software for starting up and operating the virtual machine group 210.
[0057] The configuration information 240 is an area for storing various setting information, reference information, operation logs, etc. for the operation of this cloud storage system 20 and is managed by a database, for example.
[0058] The logical volume 250 is a logical capacity resource provided under the control of the virtual machine group 210 and is a logical volume formed in the virtual drive group 220. The logical volume 250 is a logical capacity unit recognized by the host. Data stored in the logical volume 250 is, for example, made redundant for data protection using technologies such as RAID (Redundant Array of Independent Disks) or Erasure Coding.
[0059] The host I / F 260 is an interface for the host to access the logical volume 250 or the external volume 280. The host I / F 260 is provided, for example, in such a way that the logical volume can be identified by an IP (Internet Protocol) address and an iSCSI (Internet Small Computer System Interface) name, etc. Multiple host I / F 260s can be provided, and furthermore, a security function that allows only arbitrary hosts to identify or access the logical volume can also be provided.
[0060] The external virtual drive 270 is a virtual drive provided by the virtual computer resource providing service 60. The external virtual drive 270 is a single drive different from the virtual drive group 220 that forms the logical volume 250. In addition, in the present embodiment, in order to distinguish it from the virtual drive group 220 within the cloud storage system 20, it is collectively referred to as the "external virtual drive 270". In addition, multiple external virtual drives 270 can exist.
[0061] The external volume 280 is virtualized in the same way as the logical volume 250 through the function of the storage control software described later so that data can be accessed from the host. The external volume 280 is an example of a high-speed virtual drive that can be accessed at a higher speed than the virtual drives that form the logical volume 250. In addition, multiple external volumes 280 can be configured in the same way as the external virtual drive 270.
[0062] Figure 3 is a block diagram showing Figure 1 an example of the structure of the storage control software P120 of the cloud storage system 20 shown. The storage control software P120 is included in the virtual machine startup image 230 and is executed using the CPU (Central Processing Unit) and memory of the virtual machine group 210.
[0063] The storage control software P120 is composed of a host I / F control unit 2110, a logical volume control and capacity pool control unit 2120, a data redundancy and distributed storage control unit 2130, a structure management and monitoring control unit 2140, a migration control unit 2150, an external volume control unit 2160, and a control API group 2170.
[0064] The host I / F control unit 2110 is a control program that processes I / O requests from the host via the host I / F 260. The host I / F control unit 2110 controls so that during the execution of the migration process, the first storage area of the external volume 280, which is an example of a restoration source volume that is the object of the migration process, and the second storage area that is the object of writing data to the external volume 280 do not overlap.
[0065] The logical volume control and capacity pool control unit 2120 is a program that processes read and write accesses to data from the corresponding logical volume according to requests from the host, and manages the memory free capacity held by the cloud storage system 20 as a capacity pool.
[0066] The data redundancy and distributed storage control unit 2130 is a program that performs redundancy processing such as address translation, data compression, and data protection on accesses to the logical volume, and controls the data stored in the virtual drive group 220.
[0067] The structure management and monitoring control unit 2140 reflects the structures and setting information of the host I / F, logical volumes, and virtual drive group 220 into the structure information 240, and is a control program that monitors the usage status of the CPU, memory, network, etc., and the health status of the virtual machine group 210, virtual drive group 220, etc.
[0068] The structure management and monitoring control unit 2140 is an example of a structure control unit, and as a part of multiple logical volumes, at least one logical volume 250 as a restoration destination volume is constituted in the virtual drive.
[0069] The migration control unit 2150 is a control program that controls the process of copying the restored data (i.e., transferring the volume that becomes the access destination of the host) from the external volume 280 to the logical volume 250 (hereinafter also collectively referred to as "migration process") while continuing the input / output process with the host. In this control, the host is not made aware of the change in the access destination, and after the migration process is completed, the access destination volume is switched from inside the external volume 280 to the logical volume 250. In addition, multiple groups of migration processes can be executed simultaneously, and the object can also be an external volume instead of a logical volume.
[0070] The external volume control unit 2160 is a control program that connects to an external virtual drive 270 via a virtual machine group 210 to form an external volume 280 that can be processed in the same way as a logical volume 250. In addition, multiple external volumes 280 can be formed.
[0071] The control API group 2170 is an interface program for controlling instructions or responses from the operation management system 50 or the terminal 80.
[0072] The outline of the system configuration example of the data center 2 in this embodiment is as described above. Next, the operation management method and the like of this embodiment will be described. First, as described above, the data center 2 includes a backup data repository 40 that stores backup data, and a recovery processing instance 30 that interprets the backup data and restores the data, which is an example of a recovery computer resource. In this operation management method, the operation management system 50 performs the following steps: an access control step of controlling the recovery processing instance 30 so that when restoring backup data to a logical volume 250 in a virtual drive, the data obtained by restoring the backup data is stored in an external volume 280 of a high-speed virtual drive that can be accessed faster than the virtual drive, and the restored data can be accessed; and a migration processing step of maintaining a state in which the data restored to the external volume of the high-speed virtual drive can be accessed, and migrating the restored data from the external volume 280 of the high-speed virtual drive to the logical volume 250 of the virtual drive. Hereinafter, the management method of the backup data in the backup data repository 40 will be further described. In addition, in the backup data repository 40, since the management method is as follows, it is difficult to directly perform high-speed access from the host as described above.
[0073] Figure 4 It represents Figure 1 A block diagram showing a structural example of the backup data repository 40 shown. The backup data repository 40 includes a plurality of management areas called "Buckets". In the illustrated example, a Volume100 - Backup01 bucket 4110 that stores the first-generation backup of the storage volume (Volume) 100, a Volume100 - Backup02 bucket 4120 that stores its second-generation backup, and a Volume100 - Backup03 bucket 4130 that stores its third-generation backup are formed. Similarly, a Volume200 - Backup01 bucket 4210, a Volume200 - Backup02 bucket 4220, and a Volume200 - Backup03 bucket 4230 that store the first-generation, second-generation, and third-generation of the storage volume (Volume) 200 are formed respectively. In addition, in the following description, unless a specific type of bucket is mentioned, the above various buckets are also simply referred to as "buckets".
[0074] In the storage buckets 4110 to 4230, a set of backup data for restoring the volumes of this generation (hereinafter referred to as "backup data set") is stored. The backup data set includes: a differential bitmap 410 indicating the presence or absence of backups for each management block size on the volume; a data block group 420 that stores the backed-up blocks in the front; and a backup directory 430 that records structural information such as the identification number, device number, capacity, backup date and time, and the parent-child relationship of the incremental backup generations of the backup source volume.
[0075] In addition, the differential bitmap 410 does not exist in the Volume100 - Backup01 storage bucket 4110 and the Volume200 - Backup01 storage bucket 4210. This means that this backup is not an incremental backup but a full backup targeting the total volume capacity.
[0076] In addition, in the following description, the backup data of each volume stored in the backup data repository 40 is sometimes referred to as a "restored volume", but they have the same meaning.
[0077] Figures 5A - 5C These are diagrams schematically showing an example of the method of generating backup data for a volume. The backup is implemented, for example, through the cloud backup function built into the storage system 10.
[0078] Figure 5A This shows the method of the first-generation backup (full backup) of the volume. The volume is in the state 100A with blocks shown as "A", "B", and "C". Through the full backup, a data block group 420A storing "A", "B", and "C" is stored in the backup data repository 40. In addition, since it is a full backup, the differential bitmap 410A is empty (that is, not generated).
[0079] Figure 5B This is a diagram showing an example of the method of the second-generation backup (incremental backup) of the volume. The volume is in the state 100B where the block "A" has been rewritten as "A'" from the previous state 100A. Through the incremental backup, a data block group 420B storing only the changed block "A'" is generated in the backup data repository 40. In addition, a differential bitmap 410B indicating the storage location of the updated block is generated.
[0080] Figure 5C This is a diagram showing an example of the method of the third-generation backup (incremental backup) of the volume. The volume is in the state 100C where the blocks "B" and "C" have been rewritten as "B'" and "C'" respectively from the previous state 100B. Through the incremental backup, a data block group 420C storing the changed blocks "B'" and "C'" is generated in the backup data repository 40. In addition, a differential bitmap 410C indicating the storage location of the updated block is generated.
[0081] Figure 6A and Figure 6BThis is a diagram showing examples of backup directories stored in the backup data repository 40. Figure 6A This is the backup directory 430A during the first-generation backup (full backup) of the volume.
[0082] The backup directory 430A is stored with an object name (file name) such as "Volume100-Backup01.catalog", and records information related to the serial number R610 of the backup source, the volume number R611, the original volume capacity R612, the volume name R613, the backup generation number R614, the backup date and time R515, the parent directory information R616 indicating the parent-child relationship during incremental backup, and the backup capacity R617.
[0083] In Figure 6A In the example shown, the serial number of the backup source is "VSP56432", the volume number is "100", the original volume capacity is "16.0TB", the volume name is "Volume B", the backup generation is "the 01st generation", the backup date and time is "2023-08-01 00:00". Since it is the first full backup, there is no parent directory indicating the parent-child relationship of the backup, and the backup capacity is "16.0TB".
[0084] Figure 6B This is the backup directory 430B during the second-generation backup (incremental backup) of the volume. The backup directory 430B is stored with the object name (file name) "Volume100-Backup02.catalog".
[0085] In addition, regarding the serial number R610 of the backup source, the volume number R611, the original volume capacity R612, the volume name R613, the backup generation number R614, the backup date and time R515, the parent directory information R616 indicating the parent-child relationship during incremental backup, and the backup capacity R617, they are the same as those shown in Figure 6A and thus the description is omitted.
[0086] In Figure 6B In the example shown, the serial number of the backup source is "VSP56432", the volume number is "100", the original volume capacity is "16.0TB", the volume name is "Volume B", the backup generation is the 02nd generation, and the backup date and time is 2023-08-01 03:00. Moreover, in Figure 6B In the example shown, it is shown that due to the incremental backup, Volume100-Backup01.catalog, that is, the backup directory 430A, exists as the parent directory indicating the parent-child relationship of the backup.
[0087] Therefore, it can be known that when restoring the backup, the full backup of 430A as the parent backup needs to be returned to the previous state. In addition, the backup capacity is 3.2TB.
[0088] Figure 7 FIG. is a block diagram showing a structural example of the recovery processing program P30. The recovery processing program P30 is a group of programs that operate using the CPU and memory on the recovery processing instance 30. The recovery processing program P30 includes a virtual drive I / O control unit 3010, a logical volume I / O control unit 3020, a backup data repository I / O control unit 3030, a backup data parsing control unit 3040, and a control API group 3050.
[0089] The virtual drive I / O control unit 3010 is a program that mounts a separate virtual drive to read and write block data. In addition, this virtual drive is also used as Figure 2 the external virtual drive 270 shown in FIG.
[0090] The logical volume I / O control unit 3020 is a program that connects to the cloud storage system 20 and mounts a logical volume to read and write block data.
[0091] The backup data repository I / O control unit 3030 is a program that reads and writes the backup data set of the backup data repository 40.
[0092] The backup data parsing control unit 3040 is a program that performs the following control: determines the parent-child relationship between the full backup and the incremental backup according to the backup directory 430, constructs the processing steps required for restoring the generation as the object, or writes block data to the appropriate address of the virtual drive or logical volume at the restoration destination while referring to the differential bitmap 410.
[0093] The control API group 3050 is an interface program for controlling instructions or responses from the operation management system 50 or the terminal 80.
[0094] Figure 8 is a diagram showing Figure 1 a structural example of the operation management system 50 shown in FIG. The operation management system 50 includes a GUI (Graphic User Interface) providing control unit 5010, a virtual computer resource control API 5020, a cloud storage system control API 5030, a backup / recovery management unit 5040, an information management DB group 5050, and a monitoring control unit 5060.
[0095] The GUI providing control unit 5010 is a program for providing a graphical user interface (GUI) when the user operates from the terminal 80.
[0096] The virtual computer resource control API (Application Programming Interface) 5020 is a program for operating on the computer resources provided by the virtual computer resource providing service 60. The virtual computer resource control API 5020, for example, requests the virtual computer resource providing service 60 to generate a virtual drive of a predetermined capacity, or starts the virtual machine group 210 that constitutes the cloud storage system 20.
[0097] The cloud storage control API (Application Programming Interface) 5030 is a program for instructing the storage control software P120 operating in the cloud storage system 20. The cloud storage system control API 5030, for example, gives instructions for generating the logical volume 250 and executing migration processing.
[0098] The backup / restore management unit 5040 is a program for managing the access destination URL (Uniform Resource Locator) or access rights of the backup data repository 40, scheduling the backup of the storage system 10, managing device information of the restore destination, etc. The backup / restore management unit 5040 starts the restore processing instance 30 and also gives restore instructions to the logical volume or virtual drive. In the present embodiment, when the restore processing instance 30 restores the backup data to the logical volume 250 in the virtual drive, it controls so that the data obtained by storing the restored backup data in the external volume 280 of the high-speed virtual drive that can be accessed faster than the virtual drive can be accessed by the host. The above-mentioned migration control unit 2150 continues the state where the restored data in the external volume of the high-speed virtual drive can be accessed, and migrates the restored data from the external volume of the high-speed virtual drive to the logical volume of the virtual drive.
[0099] The information management DB group 5050 is a group of databases for various data required for the management and control of the operation management system 50.
[0100] The monitoring control unit 5060 is a program for acquiring various information on the storage system 10 and the restore processing instance 30 to be managed and periodically monitoring them.
[0101] Figure 9 It represents Figure 8FIG. 5 is a diagram showing a structural example of a performance management table T501 included in the information management DB group 5050 shown. The performance management table T501 is an example of a performance management table, and is a table that manages information related to the performance of multiple virtual drives, including a normal virtual drive that can be used in the data center 2 and at least one high-speed virtual drive that can be accessed at a higher speed than the normal virtual drive (hereinafter, the virtual drive that can be accessed fastest among the high-speed virtual drives is also referred to as the "fastest virtual drive"). In addition, the performance management table T501 is a table that also manages information related to the performance of the cloud storage system 20. The virtual drive is, for example, an SSD (Solid State Drive).
[0102] The performance management table T501 manages the cloud storage system 20 and the type C511 of each virtual drive, the maximum IOPS (Input / Output Per Second) performance C512, and the maximum throughput performance C513. The unit of the information of the maximum throughput performance C513 is, for example, MB / s. When referring to the performance management table T501, information related to the performance of each virtual drive that can be used in the data center 2 can be obtained.
[0103] The structure management and monitoring control unit 2140 is an example of a structure control unit, and refers to the performance management table T501 to select the fastest virtual drive from multiple virtual drives as an example of a high-speed virtual drive, and use the fastest virtual drive instead of the virtual drive to be used. In addition, in this embodiment, when there are multiple high-speed virtual drives, a virtual drive with lower performance than the fastest virtual drive but higher speed can also be selected from the multiple high-speed virtual drives.
[0104] In the above Figure 7 In the example, the maximum IOPS of cloud storage system 20 is 50,000 and the throughput is 1,000MB / s, the maximum IOPS performance of cloud storage system "No. 21" is "80,000" and the maximum throughput is "2,000MB / s", and the maximum IOPS performance of cloud storage system "No. 22" is "256,000" and the maximum throughput is "10,000MB / s".
[0105] In addition, it can be seen that the maximum IOPS performance of the "Standard SSD" is "3,000" and the maximum throughput is "125MB / s". Similarly, the maximum IOPS performance of the "Performance SSD" is "16,000" and the maximum throughput is "500MB / s". The maximum IOPS performance of the ultra SSD is "160,000" and the maximum throughput is "4,000MB / s".
[0106] Figure 10 This is a flowchart showing an example of the steps of a storage recovery control process. This storage recovery control process is executed by the operation management system 50 based on an instruction from the terminal 80, or a judgment of the operation management system 50, etc.
[0107] First, the backup / recovery management unit 5040 of the operation management system 50 refers to the performance management table T501 (refer to Figure 9 ), and obtains the performance information of the cloud storage system 20 that is the restoration destination (step S1000). Additionally, Figure 10 in this, the cloud storage system 20 that is the restoration destination is abbreviated as the "restoration destination storage system".
[0108] Next, the backup / recovery management unit 5040 of the operation management system 50 refers to the performance management table T501, and obtains information related to the performance of the fastest virtual drive that can access faster than the logical volume of the cloud storage system 20 (step S1010).
[0109] The backup / recovery management unit 5040 of the operation management system 50 refers to the backup directory 430 stored in the backup data repository 40. The backup / recovery management unit 5040 of the operation management system 50 calculates the total values of the data amounts of all backups required for the recovery of the object generation of the volume to be recovered (i.e., the capacity for full recovery) and the data amount of the incremental backup (i.e., the capacity for incremental recovery) (step S1020).
[0110] For example, in the example shown in Figure 4 , in the case of restoring the third generation of the volume, the backup capacity of the first generation for full recovery becomes the full recovery capacity, and the sum of the backup capacities of the second and third generations for incremental recovery becomes the incremental recovery capacity.
[0111] Next, the backup / recovery management unit 5040 of the operation management system 50 uses the previously calculated values of the recovery capacity (full recovery capacity and incremental recovery capacity) to calculate the estimated recovery time A in the cloud storage system 20 that is the restoration destination (step S1030).
[0112] The estimated recovery time A is obtained as follows, for example. First, for the restoration of full recovery that is roughly sequential writing, the time required for full recovery (in minutes) is obtained by calculating "full recovery capacity (MB) ÷ throughput (MB / s) ÷ 60s".
[0113] Next, for incremental recovery that is roughly a random write, the number of recovery blocks (i.e., the number of recovery I / Os) is obtained by calculating "incremental recovery capacity (MB) ÷ management block size (MB)". Moreover, the time required for incremental recovery (in minutes) is obtained by calculating "number of recovery I / Os ÷ IOPS performance (I / Os per second) ÷ 60 s". In the present embodiment, the estimated recovery time A can be calculated based on the sum of the time required for full recovery and the time required for incremental recovery. In addition, the above calculation method is an example, and the estimated recovery time A can also be obtained by other methods such as speculation based on machine learning.
[0114] Next, the backup / recovery management unit 5040 of the management system 50 calculates the estimated recovery time B when the fastest virtual drive is used as the recovery destination (step S1040). The estimated recovery time B can be obtained, for example, by performing the same calculation as the estimated recovery time A. In addition, a method different from the estimated recovery time A can also be used to calculate the estimated recovery time B.
[0115] The backup / recovery management unit 5040 of the management system 50 confirms whether all recovery targets have been calculated (step S1050). If there are other recovery target volumes (step S1050: No), the process returns to step S1020 and the processing is repeated. This is, for example, a case where in addition to volume 100, volume 200 is also recovered. Figure 4 in which volume 200 is also recovered in addition to volume (Volume) 100.
[0116] On the other hand, when there are no other recovery target volumes (step S1050: Yes), the backup / recovery management unit 5040 of the management system 50 obtains the estimated preparation time C (step S1060) and the recovery process instance start time D (step S1065) of the recovery destination cloud storage system 20.
[0117] The estimated preparation time C is the time required from the start of the virtual machine group 210 until a logical volume for recovery can be prepared through the generation of the virtual drive group 220, the generation of the capacity pool, etc. The estimated preparation time C can be calculated, for example, by previously storing the construction time per unit of each stage in the information management DB group 5050 of the management system 50 and calculating based on them. Or, other methods such as defining the estimated preparation time C as a fixed value can also be used.
[0118] The recovery process instance start time D is the time from the start of the recovery process instance 30 until the recovery process program P30 that starts working according to this instruction can be utilized. Since the recovery process instance start time D can be expected to be roughly fixed, it can also be possessed by the management system 50 as a fixed value, or other methods such as obtaining it by referring to the start time of other virtual machines in the data center 2 can also be used.
[0119] Next, when the backup / recovery management unit 5040 of the management system 50 performs a recovery process using the virtual drive to be used as scheduled in the cloud storage system 20 and when performing a recovery process using the fastest virtual drive, it compares which one is faster (step S1070). That is, it compares "the estimated recovery time of all volumes in the cloud storage system 20 (the total value of A) + the estimated preparation time C of the cloud storage system 20" with "the estimated recovery time of the volume that takes the longest recovery time when recovering all volumes in parallel in the virtual drive (the maximum value of B) + the recovery process instance startup time D".
[0120] This is because, as Figure 7 shown, the maximum performance of the cloud storage system 20 depends on the system performance. Therefore, even if the recovery of volumes is processed in parallel or sequentially, the total required time remains unchanged. On the other hand, the virtual drive is the performance of a single drive. Therefore, when preparing in parallel according to the number of volumes, the required time is limited by the volume that takes the longest recovery time.
[0121] When the backup / recovery management unit 5040 of the management system 50 determines that the fastest virtual drive can perform the recovery process at high speed in a short time compared to the virtual drive to be used as scheduled based on the result of the above comparison (step S1070: Yes), it performs a recovery process using the fastest virtual drive (hereinafter, also collectively referred to as "high-speed recovery process") (step S1080), and ends this process. The high-speed recovery process using the fastest virtual drive will be described later.
[0122] On the other hand, when the backup / recovery management unit 5040 of the management system 50 determines that the virtual drive to be used as scheduled is the same as the fastest virtual drive at high speed in a short time (step S1070: No), that is, when it determines that the high-speed recovery process cannot be implemented, it performs a recovery process using the virtual drive to be used as scheduled (normal recovery process) (step S1090), and ends the process. The normal recovery process will be described later.
[0123] Figure 11 is a flowchart showing an example of the steps of the high-speed recovery process using a virtual drive. This high-speed recovery process is executed by the management system 50 when the branch in step S1080 based on Figure 8 determines that using the fastest virtual drive for recovery can shorten the required recovery time.
[0124] First, the backup / recovery management unit 5040 of the management system 50 confirms the number of volumes restored from the backup data and their respective capacities (step S2000). Next, the backup / recovery management unit 5040 of the management system 50 generates new virtual drives based on the previous number and capacities of the volumes to be restored, and starts the recovery process (step S2010). Details of the recovery process to the virtual drives are described in the subsequent Figure 12 below.
[0125] In parallel with the recovery process to the virtual drives (step S2010), the backup / recovery management unit 5040 of the management system 50 executes the subsequent steps S2020 to S2100 to make preparations so that the cloud storage system 20 can be utilized from the host.
[0126] The backup / recovery management unit 5040 of the management system 50 confirms whether the virtual machine cluster 210 has been constructed as a storage controller (step S2020). If it has been constructed (step S2020: Yes), the virtual machine cluster 210 is started (step S2030). On the other hand, if it has not been constructed (step S2020: No), after the backup / recovery management unit 5040 of the management system 50 requests and allocates a predetermined number of virtual machines from the virtual computer resource providing service 60, the virtual machine cluster 210 is started using the virtual machine startup image 230 (step S2040), and the initial setting of the storage control software P120 is implemented (step S2050). The initial setting includes, for example, the setting of the system name, the setting of the management subnet, the setting of the administrator, the authentication of the license, etc.
[0127] Next, the backup / recovery management unit 5040 of the management system 50 generates a virtual drive group 220 composed of virtual drives with a predetermined capacity and number, and attaches it to the virtual machine cluster 210 (step S2070).
[0128] After that, the backup / recovery management unit 5040 of the management system 50 instructs the storage control software through the cloud storage system control API 5030 to perform formatting (e.g., RAID formatting) for distributed storage and redundant structure on the virtual drive group 220 (step S2080), constructs a capacity pool (step S2090), and generates a logical volume 250 based on the number and capacity of the volumes confirmed in step S2000 (step S2100).
[0129] The backup / recovery management unit 5040 of the management system 50 confirms whether the recovery process to the virtual drives implemented in step S2010 is completed (step S2110). If it is not completed (step S2110: "No"), it waits until it is completed.
[0130] When the backup / restore management unit 5040 of the operation management system 50 finishes the restore process to the virtual drive (step S2110: Yes), it instructs the storage control software P120 to attach the virtual drive as an external virtual drive 270 to the cloud storage system 20 and starts the migration process (step S2120). Details of the migration process will be described in Figure 13 described later.
[0131] When starting the migration process, the backup / restore management unit 5040 of the operation management system 50 sets the host connection path so that the host can access the logical volume 250 through the host I / F 260 (step S2130). The migration process can be executed while accepting I / O requests from the host. In the operation management system 50, the backup / restore management unit 5040 instructs the storage control software P120 to start the I / O access acceptance process from the host (step S2140), and this process ends.
[0132] Figure 12 is a flowchart showing an example of the steps of the restore process to the virtual drive in Figure 11 step S2010. First, the backup / restore management unit 5040 of the operation management system 50 generates a plurality of fastest virtual drives corresponding to the number and capacity of the volumes to be the object of the restore process based on the information (the number and each capacity of the volumes restored from the backup data) obtained in Figure 11 step S2000 (step S3000). For example, in this embodiment, Ultra SSDs are used from the Figure 9 shown performance management table T501.
[0133] The backup / restore management unit 5040 of the operation management system 50 starts a plurality of restore process instances 30 corresponding to the number of volumes to be restored (step S3010), and attaches each fastest virtual drive to each restore process instance 30 (step S3020).
[0134] Next, the backup / restore management unit 5040 of the operation management system 50 instructs the restore process program P30 of each restore process instance 30 to restore the respectively specified backup data to the fastest virtual drive (step S3030). Thus, the restore process instances 30 work in parallel according to the number of restored volumes and perform the restore process. In addition, when the restore process instances 30 have sufficient processing capabilities (computing performance, transmission bandwidth), multiple restore processes to the fastest virtual drive can also be assigned to one restore process instance 30.
[0135] The backup / recovery management unit 5040 of the operation management system 50 confirms whether the processing of each recovery processing instance 30 is completed (step S3040). If there is a recovery processing instance 30 for which the recovery processing has been completed (step S3040: Yes), the fastest virtual drive is detached from the recovery processing instance 30 (step S3050), and the recovery processing instance 30 is terminated.
[0136] In order to suppress subsequent metering and charging costs, the fastest virtual drive for which the recovery processing has been completed requests the virtual computer resource providing service 60, and the type of the virtual drive is changed from the fastest to an inexpensive type through the operation management system 50. For example, it is changed to the same "standard SSD" as the "standard SSD" used in the cloud storage system 20 (step S3070).
[0137] The backup / recovery management unit 5040 of the operation management system 50 confirms whether all the recovery processing instances 30 have completed the recovery processing (step S3080). If there is a recovery processing instance 30 for which the recovery processing has not been completed (step S3080: No), the monitoring is continued (step S3040). On the other hand, if all the recovery processing instances 30 have completed the recovery processing (step S3080: Yes), the processing is terminated and the above Figure 9 processing is executed.
[0138] Figure 13 is a flowchart showing an example of the order of the migration processing as Figure 11 shown in step S2120. The backup / recovery management unit 5040 of the operation management system 50 attaches the virtual drive that has had its data restored through the Figure 12 recovery processing shown as an external virtual drive 270 to the cloud storage system 20 (step S4000).
[0139] The backup / recovery management unit 5040 of the operation management system 50 instructs the storage control software P120 to be set to use the external virtual drive 270 instead of the external volume 280 (step S4010).
[0140] After that, the backup / recovery management unit 5040 of the operation management system 50 starts the migration processing between the external volume 280 and the logical volume 250 having the same capacity as the external volume 280 (step S4020). In addition, the migration processing is executed in the background.
[0141] The backup / restore management unit 5040 of the operation management system 50 confirms whether the same processing has been performed on the virtual drives of all the restored data generated in step S2010 (step S4030). If there are still virtual drives that have not been processed (step S4030: No), steps S4000 to S4020 are similarly repeated. On the other hand, if the backup / restore management unit 5040 of the operation management system 50 can start the migration process for the virtual drives of all the restored data (step S4030: Yes), the migration process is ended and the process returns Figure 9 is executed by the processing of
[0142] Figures 14A - 14C are diagrams showing an example of a method in which the storage control software P120 does not stop the input / output processing with the host and performs the migration process in the background. That is, Figures 14A - 14C shows the control method when starting to accept the input / output processing with the host for the logical volume 250 in step S2140 shown in Figure 11 .
[0143] Figure 14A is an example of the method when there is a read I / O request from the host to the logical volume 250 during the execution of the migration process. The data restored from the backup data is stored in the external volume 280. This data is sequentially read out by the storage control software P120 and copied to the logical volume 250.
[0144] During this period, when there is a read request from the host to the logical volume 250, the storage control software P120 reads the data from the external volume 280 as the copy source and responds to the host, while providing access to the restored data and continuing the migration process in the background.
[0145] Figure 14B is an example of the method when there is a write I / O request from the host to the logical volume 250 during the execution of the migration process. As described above, the data in the external volume 280 is sequentially read out by the storage control software P120 and copied to the logical volume 250.
[0146] During this period, when there is a write request from the host to the logical volume 250, the storage control software P120 can continue the migration process in the background without stopping the input / output processing with the host by writing the data corresponding to the above write request to both the external volume 280 and the logical volume 250.
[0147] Figure 14C is an example of the method when there is a read request or a write request from the host to the logical volume 250 after the migration process is completed. Since all the data in the external volume 280 has been copied to the logical volume 250, the external volume 280 is no longer needed later, and all I / O requests are made to the logical volume 250.
[0148] Figure 15 It is a flowchart showing an example of post-processing implemented by the management system 50 when the migration process implemented in the background is completed.
[0149] When the migration is completed, the backup / restore management unit 5040 of the management system 50 disconnects the external volume 280 that is no longer accessed (step S5000). After that, the backup / restore management unit 5040 of the management system 50 separates (step S5010) and then deletes (step S5020) the external virtual drive 270 that is controlled as an external volume, and the process ends.
[0150] Figure 16A and Figure 16B are flowcharts showing an example of normal recovery processing respectively. This normal recovery processing is executed by the management system 50 based on the branch of step S1070 shown in Figure 10 when the high-speed recovery processing using the virtual drive cannot be applied.
[0151] First, the backup / restore management unit 5040 of the management system 50 confirms the number and respective capacities of the volumes restored from the backup data (step S6000). Next, the backup / restore management unit 5040 of the management system 50 makes preparations so that the cloud storage system 20 can be utilized from the host (steps S6020 to S6100).
[0152] The backup / restore management unit 5040 of the management system 50 confirms whether the virtual machine group 210 has been constructed (step S6020). If it has been constructed (step S6020: Yes), the virtual machine group 210 is started (step S6030).
[0153] On the other hand, if it has not been constructed (step S6020: No), the backup / restore management unit 5040 of the management system 50 requests and allocates a predetermined number of virtual machines from the virtual computer resource providing service 60, and then starts the virtual machine group 210 using the virtual machine startup image 230 (step S6040), and performs the initial setting of the storage control software P120 (step S6050). The initial setting includes, for example, the setting of the system name, the setting of the management subnet, the setting of the administrator, the authentication of the license, etc.
[0154] Next, the backup / restore management unit 5040 of the management system 50 generates a virtual drive group 220 with a predetermined capacity and number, and attaches it to the virtual machine group 210 (step S6060).
[0155] After that, the backup / recovery management unit 5040 of the management system 50 instructs the storage control software P120 through the cloud storage system control API 5030 to perform formatting for distributed storage and redundant structure (e.g., RAID formatting) on the virtual drive group 220 (step S6070), constructs a capacity pool (step S6090), and generates a logical volume 250 based on the number of volumes and capacity confirmed in step S2000 (step S6100).
[0156] After that, the backup / recovery management unit 5040 of the management system 50 instructs the storage control software P120 to set a connection host path for the recovery processing instance 30 to access the logical volume 250 (step S6110).
[0157] Next, the backup / recovery management unit 5040 of the management system 50 starts a recovery processing instance 30 (step S6120). Then, the backup / recovery management unit 5040 of the management system 50 instructs the recovery processing program P30 to connect a logical volume corresponding to the capacity of the restoration volume (backup data) to be restored to the recovery processing instance (step S6130). Then, according to the instruction of the backup / recovery management unit 5040 of the management system 50, the recovery processing program P30 restores the backup of the specified volume to the logical volume.
[0158] The backup / recovery management unit 5040 of the management system 50 waits until the recovery processing program P30 completes the recovery processing (steps S6150 and S6150: No). If it is confirmed that the recovery is completed (step S6150: Yes), the backup / recovery management unit 5040 of the management system 50 disconnects the logical volume with the restored data (step S6170).
[0159] If there is a restoration volume that has not been restored yet (step S6170: Yes), the backup / recovery management unit 5040 of the management system 50 repeats steps S6130 to S6160 until the recovery of all volumes is completed. On the other hand, when the recovery of all restoration volumes is completed (step S6160: No), the backup / recovery management unit 5040 of the management system 50 ends the recovery processing instance 30 (step S6180).
[0160] Next, the backup / recovery management unit 5040 of the management system 50 sets a host connection path so that the host can access the logical volume 250 through the host I / F 260 (step S6190), instructs the storage control software to start the I / O access acceptance processing from the host (step S6200), and ends this processing.
[0161] Figure 17A It represents a high-speed recovery process using virtual drives (refer to Figure 9) A diagram showing an example of the relationship between the required times of the respective processes. In the illustrated example, the horizontal axis represents the elapsed time.
[0162] In the present embodiment, the recovery process of the recovery process instance 30, which is an example of recovering computer resources, is executed in parallel (corresponding to the process within the illustrated time t1), and the configuration of the logical volume 250 of the cloud storage system 20 is performed (corresponding to the process within the illustrated time t0).
[0163] In the present embodiment, the host I / F control unit 2110 restarts the input / output process with the host on the occasion of the start of the migration process.
[0164] Hereinafter, a specific description will be given. In this high-speed recovery process, during the execution of the recovery to the virtual drive (external volume 280) at time t1 (step S2010), the preparation of the cloud storage system 20 is performed in parallel at time t0 (steps S2020 to S2100). Then, if the recovery to the virtual drive (external volume 280) is completed, the migration process is started from time t1 (step S2120), and the migration process continues for a duration of t2. During the execution of the migration process, the host can access the restored data from time t1, so it can be known that the time required until the host can start accessing is t1 (< time t2). In addition, according to Figure 10 the branch determination in step S1070, it can be known that t1 < t2.
[0165] Figure 17B A diagram showing an example of the relationship between the required times of the respective processes in the normal recovery process shown in FIG. 14. The horizontal axis represents the elapsed time. In this normal recovery process, in order to directly recover to the logical volume 250 without passing through the external volume 280 of the virtual drive, first, the preparation of the cloud storage is performed at time t0 (steps S6020 to S6100). After that, the recovery to the logical volume 250 is performed at time t2 (steps S6110 to S6160), and host access starts after the recovery is completed. Therefore, it can be known that the time required until the start of host access is t0 + t2.
[0166] According to Figure 17A and Figure 17B the comparison, it can be seen that the high-speed recovery process using the virtual drive (external volume 280) (refer to Figure 17A ) compared with the normal recovery process (refer to Figure 17B ) can hide time t0 and time t2, and correspondingly, the time required until the start of host access is shorter. Thus, according to the present embodiment, the recovery time until the host can access can be shortened, and the objects of backup to which this method can be applied can be expanded to logical volumes with a short target recovery time.
[0167] In the operation management system 50 of the data center 2 in this embodiment, as described above, the data center 2 includes: a backup data repository 40 that stores backup data; and a recovery processing instance 30 that serves as an example of a recovery computer resource that interprets the backup data and recovers the data. When the operation management system 50 restores the backup data to a logical volume in a virtual drive, it controls the recovery processing instance 30 so that the data obtained by restoring the backup data is stored in an external volume of a high-speed virtual drive that can be accessed faster than the virtual drive, and the restored data can be accessed from the host, and the state where the data restored to the external volume of the high-speed virtual drive can continue to be accessed is maintained, and the restored data is migrated from the external volume of the high-speed virtual drive to the logical volume of the virtual drive.
[0168] In this way, in the recovery processing of the recovery processing instance 30 (corresponding to the "recovery of the virtual drive" within the time t1 shown in Figure 17A ), and the configuration of the logical volume 250 of the cloud storage system 20 (corresponding to the "cloud storage preparation" within the time t0 shown in Figure 17A ), at the moment when they are completed (for example, the time t1 shown in Figure 17A ), although the migration process starts immediately, the input / output process (host access) with the host can be started, so the recovery time until the data can be utilized from the host can be shortened.
[0169] In this embodiment, the data center 2 of this embodiment executes the recovery processing of the recovery processing instance 30 and the configuration of the logical volume 250 of the cloud storage system 20 in parallel (for example, refer to Figure 17A ). In this way, in order to start the input / output process with the host, it is not necessary to wait for the end of the migration process (corresponding to the "migration" shown in Figure 17A ), so the input / output process with the host can be restarted earlier, and the recovery time can be further shortened.
[0170] In this embodiment, the host I / F control unit 2110 is an example of an interface control unit, and takes the start of the migration process as an opportunity to restart the input / output process with the host. In this way, the input / output process with the host can be started without waiting for the end of the migration process.
[0171] (2) Second Embodiment
[0172] The second embodiment has the same structure and operation as the first embodiment. Therefore, in the second embodiment, the description of the same structure and operation as the first embodiment is omitted, and the following will focus on the differences.
[0173] In the first embodiment, when the processing time can be shortened, the above-mentioned high-speed recovery process is unconditionally performed as a recovery process using the fastest virtual drive, but it is also considered that there is room for the overall system recovery target time.
[0174] Therefore, in the second embodiment, taking this into account, the recovery processing method or virtual drive is selected according to the recovery target time. Hereinafter, the description will mainly focus on the differences from the first embodiment. For other structures not described, they are the same as those in the first embodiment.
[0175] Figure 18 It is a flowchart showing an example of the steps of the storage recovery control process in the second embodiment. This storage recovery control process is a process executed by the operation management system 50 according to an instruction from the terminal 80 or a judgment of the operation management system 50, etc., and is basically the same as that in the first embodiment Figure 10 but some processes (step S1010A, step S1040A, step S1069, step S1070A, Figure 1075A, step S7000, step S7100, Figure 7200) are different.
[0176] The backup / recovery management unit 5040 of the operation management system 50 Figure 10 similarly refers to the performance management table T501 in the step S1000 of Figure 9 ), and obtains the performance information of the cloud storage system 20 that is the restoration destination (step S1000). Next, referring to the performance management table T501 (refer to Figure 9 ), the performance information of the virtual drive i that can be used in this data center is obtained (step S1010A). In this embodiment, in the performance management table T501, there are three types of information related to the performance of the virtual drive, namely "standard SSD" to "super SSD". Therefore, the information of i = 1 to 3 is obtained.
[0177] Next, the backup / recovery management unit 5040 of the operation management system 50 refers to the backup directory 430 stored in the backup data repository 40. Moreover, the backup / recovery management unit 5040 of the operation management system 50 calculates the respective total values of the backup data of the full backup (that is, the capacity for complete recovery) and the backup data of the incremental backup (that is, the capacity for incremental recovery) required for the recovery of the object generation of the volume to be recovered (step S1020).
[0178] The backup / recovery management unit 5040 of the operation management system 50 uses the previous calculated values to calculate the estimated recovery time A in the cloud storage system 20 that is the restoration destination (step S1030). For specific examples of steps S1020 and S1030 respectively, they have been described in Figure 10 and are therefore omitted.
[0179] Next, the management system 50 calculates the estimated recovery time Bi when each virtual drive i is used as the recovery destination (step S1040A). Regarding the calculation method of the estimated recovery time Bi, it is the same as the calculation method of the recovery time B of the fastest virtual drive in step S1040 of Figure 10 , so the description is omitted.
[0180] The management system 50 confirms whether all recovery targets have been calculated (step S1050). If there are other recovery target volumes (step S1050: No), the process returns to step S1020 and repeats.
[0181] On the other hand, when there are no other recovery target volumes (step S1050: Yes), the management system 50 obtains the estimated preparation time C (step S1060) and the recovery process instance start time D (step S1065) of the recovery destination cloud storage system 20. The methods for obtaining the estimated preparation time C and the recovery process instance start time D are the same as those in Figure 10 .
[0182] After that, the management system 50 obtains the system recovery target time T (step S1069). Then, the management system 50 compares whether the time required for direct recovery in the cloud storage system 20, that is, "the sum of the estimated recovery times of all volumes in the cloud storage system 20 (the total value of A) + the estimated preparation time C of the cloud storage system 20" converges within the recovery target time T (step S1070A).
[0183] When the comparison result is that the recovery target time T can be achieved even by direct recovery (step S1070A: Yes), the management system 50 performs the normal recovery process (step S1090). The detailed content of the normal recovery process is the same as that of the above first embodiment (refer to FIG. 16), so the description is omitted.
[0184] On the other hand, when the recovery target time T cannot be achieved by direct recovery (step S1070A: No), the management system 50 compares "the estimated recovery time of the volume that takes the most recovery time (the maximum value of Bi) + the recovery process instance start time D when all volumes are recovered in parallel in the virtual drive i" and the recovery target time T when recovering using each virtual drive i (step S1075A).
[0185] When there is no virtual drive i that satisfies the specific conditions related to the recovery target time T (step S1075A: No), the management system 50 cannot recover within the recovery target time T, so it notifies a target setting error (step S7200) and ends this process.
[0186] On the other hand, when there are multiple virtual drives i that satisfy the specific condition (step S1075A: Yes), the management system 50 selects, based on the recovery target time T of the recovery, the performance of the virtual drives, and utilization billing, a high-speed virtual drive k (k ∈ i) for recovering data from the multiple virtual drives, that is, selects the virtual drive k (k ∈ i) with the cheapest utilization billing that satisfies the specific condition among these multiple virtual drives (step S7000), and performs a recovery process (high-speed recovery process) using the virtual drive k (step S7100). In this way, it is sometimes possible to complete the recovery process (high-speed recovery process) within the recovery target time T while suppressing costs.
[0187] In addition, the high-speed recovery process using the virtual drive k is the same as the process of replacing "the fastest virtual drive" with "virtual drive k" in the operation of Figures 11 - 15 in the first embodiment, and thus the illustration and description are omitted.
[0188] According to the present embodiment, since the recovery method is switched according to the recovery target time T, it is possible to suppress costs and shorten the recovery time until the host can access, and the objects of backup to which this method can be applied can be expanded to logical volumes with a short target recovery time. In addition, according to the present embodiment, since it is determined whether the recovery target time T is satisfied before the data recovery process, it is possible to perform a preliminary test without incurring the costs and time required for the actual data recovery process.
[0189] In addition, the present invention is not limited to the above-described embodiments, and includes various modifications and equivalent structures within the scope of the claimed invention. For example, the foregoing embodiments are examples described in detail for easy understanding of the present invention, and the present invention is not limited to having all the structures described. In addition, each of the elements described in parallel in the present embodiment may also be in a manner in which at least one of the elements is connected in series with other elements.
[0190] In addition, in the above-described embodiment, the "system" and "program" are sometimes used as the subject to describe the process, but the system is a computer resource having a processor (e.g., CPU (Central Processing Unit)), a storage resource (e.g., memory), and a communication interface device (e.g., NIC (Network Interface Card)), and the program is executed by the processor, and thus appropriately uses the storage resource and / or the communication interface device to perform, and therefore the subject of the process may also be set to a process performed by the processor or a computer having a processor.
[0191] Industrial Applicability
[0192] The present invention can be applied to an operation management system related to a technology for performing input / output processing of data with a host computer.
[0193] Symbol Explanation
[0194] 1 Data Center
[0195] 2 Data Center
[0196] 20 Cloud Storage System
[0197] 30 Recovery Processing Instance
[0198] 40 Backup Data Repository
[0199] 50 Operation Management System
[0200] 60 Virtual Computer Resource Provisioning Service
Claims
1. An operation management system for a computer, characterized in that: The computer includes: A backup data repository that stores backup data; And A restore computer resource that interprets the backup data and restores the data, When the operation management system restores the backup data to a logical volume in a virtual drive, it controls the restore computer resource so that the data obtained by restoring the backup data is stored in an external volume of a high-speed virtual drive that can be accessed faster than the virtual drive, and the restored data can be accessed, Continuing the state where the data restored to the external volume of the high-speed virtual drive can be accessed, and migrating the restored data from the external volume of the high-speed virtual drive to the logical volume of the virtual drive.
2. The operation management system according to claim 1, characterized in that: The restoration of the data based on the restore computer resource and the configuration of the logical volume for storing the data are performed in parallel.
3. The operation management system according to claim 1, characterized in that: The operation management system restarts the input / output processing with the host when the migration starts.
4. The operation management system according to claim 1, characterized in that: The operation management system selects a high-speed virtual drive for restoring the data from multiple virtual drives based on the restoration target time, the performance of the virtual drive, and the usage cost.
5. An operation management method for a computer, characterized in that: The computer includes: A backup data repository that stores backup data; And A restore computer resource that interprets the backup data and restores the data, When the operation management method restores the backup data to a logical volume in a virtual drive, the following steps are performed: An access control step of controlling the restore computer resource so that the data obtained by restoring the backup data is stored in an external volume of a high-speed virtual drive that can be accessed faster than the virtual drive, and the restored data can be accessed; And A migration processing step of continuing the state where the data restored to the external volume of the high-speed virtual drive can be accessed, and migrating the restored data from the external volume of the high-speed virtual drive to the logical volume of the virtual drive.