Integrating mirring storage to remote

By saving operating system-specific information in the storage-level descriptor area and managing active-active configuration information with remote site controllers, the integration problem of different types of data storage systems of multiple replication sites is solved, and high data availability and disaster recovery capabilities are achieved.

CN119923636APending Publication Date: 2025-05-02INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380061652.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-08-25
Filing Date
2023-07-26
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

It is difficult to effectively integrate different types of data storage systems for multiple replication sites in the prior art, especially in disaster recovery scenarios, where high availability and diversified recovery solutions of data are difficult to implement.

Method used

By saving operating system-specific information in the storage-level descriptor area, retrieving and managing proactive-active configuration information with the remote site controller, another remote replica creation and update is implemented to ensure high availability and rapid recovery of data in disaster situations.

Benefits of technology

The integration of different data replication systems is achieved, data availability and disaster recovery capabilities are improved, ensuring that remote replicas can remain up to date when active-active mirror replicas change, supporting high availability data storage for critical workloads.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119923636A_ABST
    Figure CN119923636A_ABST
Patent Text Reader

Abstract

A method, a computer system, and a computer program product are provided. The computer sends a query command to a storage descriptor region of the first disk. The first disk belongs to a dual-site data replication system. A dual site data replication system provides proactive-proactive access to a data volume stored in a proactive disk and replicated in a backup disk. The computer receives a response to the query command. The response indicates an active disk and a backup disk of the dual site data replication system. The computer controls an additional copy of the data volume at another remote site based on the active disk.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] The present invention generally relates to data storage for disaster recovery and to integrating different types of data storage for multiple replication sites. Summary of the invention

[0002] According to an exemplary embodiment, a computer-implemented method is provided. A computer sends a query command to a storage descriptor area of ​​a first disk. The first disk belongs to a dual-site data replication system. The dual-site data replication system provides active-active access to a data volume stored in an active disk and replicated in a backup disk. The computer receives a response to the query command. The response indicates an active disk and a backup disk for the dual-site data replication system. The computer controls an additional copy of the data volume at another remote site based on the active disk.

[0003] A computer system and a computer program product corresponding to the above method are also provided. BRIEF DESCRIPTION OF THE DRAWINGS

[0004] These and other objects, features and advantages of the present invention will become apparent from the following detailed description of illustrative embodiments of the present invention, which is to be read in conjunction with the accompanying drawings. The various features of the accompanying drawings are not drawn to scale, as the illustrations are intended to clearly facilitate understanding of the present invention by those skilled in the art in conjunction with the detailed description. In the drawings:

[0005] Figure 1 is a block diagram illustrating a replication system architecture according to at least one embodiment;

[0006] Figure 2 is an operational flow diagram illustrating a process for integrating mirrored active-active storage to a remote site in accordance with at least one embodiment;

[0007] Figure 3 is a diagram illustrating a storage description area for replication integration according to at least one embodiment;

[0008] Figure 4 According to at least one embodiment Figure 1 a block diagram of the internal and external components of a computer and server depicted in;

[0009] Figure 5 According to an embodiment of the present disclosure, Figure 1 and Figure 4 A block diagram of an illustrative cloud computing environment of computers depicted in; and

[0010] Figure 6 According to the embodiment of the present disclosure Figure 5 A block diagram of the functional layers of an illustrative cloud computing environment. DETAILED DESCRIPTION

[0011] Detailed embodiments of the claimed structures and methods are disclosed herein; however, it is understood that the disclosed embodiments are merely illustrative of the claimed structures and methods that may be implemented in various forms. The present invention may be implemented in many different forms and should not be construed as limited to the exemplary embodiments set forth herein. Instead, these exemplary embodiments are provided so that this disclosure will be thorough and complete and will fully convey the scope of the invention to those skilled in the art. In the specification, details of well-known features and techniques may be omitted to avoid unnecessarily obscuring the presented embodiments.

[0012] In a data mirroring system, data may be stored in a volume pair including a primary volume associated with a primary memory storage and a corresponding secondary volume associated with a secondary memory storage. The primary memory storage and the secondary memory storage may be set at different sites to provide backup protection when a problem occurs at the first site. For example, a hurricane, earthquake, or other natural disaster may cause a power outage, which may cause the first site to be unavailable for an extended period of time and result in the need to access the secondary volume. For an active-active memory storage architecture, both sites can serve applications / workloads at any time, so each site and volume is used as an active application site that allows access. The secondary volume may be a copy of the data maintained in the primary volume. The primary and secondary volumes are identified by a replication relationship in which the data of the primary volume (also referred to as the source volume) is copied to the secondary volume (also referred to as the target volume). The primary storage controller and the secondary storage controller may be used to control access to the primary memory storage and the secondary memory storage. The secondary site may be set at any distance from the primary site. If the secondary sites are located less than 100km apart, the advantage of faster access to the backup can be presented to the user, since the data still has to flow over the physical link, and the shorter physical distance will result in a faster response. These storage systems allow input / output operations to be swapped from a first set of disks to a second set of disks to store operational data. The first set of disks can be used for primary volumes, while the other set of disks is used for secondary volumes.

[0013] In some examples, referred to herein as a controller, a computer may be configured to allow for management of planned and unplanned interruptions of storage. The controller is configured to detect failures at a primary storage subsystem that may be at a local site. Such failures may include problems writing to or accessing primary storage volumes at the local site, or other problems as discussed herein. When a controller (e.g., an operating system) detects such a failure, the controller may invoke or cause invocation of a storage unit swap function, an example of which is (IBM, HyperSwap and all IBM-based trademarks and logos are trademarks or registered trademarks of International Business Machines Corporation and / or its affiliates). The swap function can be used to automatically swap the handling of all data volumes in a mirrored configuration from the primary site to the secondary site. The swap can include peer-to-peer remote replication failover. Application / workload input / output can be transparently redirected to the secondary storage subsystem, allowing the application / workload to continue running as part of high availability.

[0014] As a result of the swap, the storage volume at the secondary site that was originally configured as the secondary volume in the original replication relationship is reconfigured as the primary volume in the new replication relationship. Similarly, once the volume at the local site is operational again, the storage volume at the first site that was originally configured as the primary volume in the original replication relationship can be reconfigured as the secondary volume in the new replication relationship. In anticipation of an unplanned swap, the controller can pass information so that one or more swap controllers can automatically detect its own failure and call a swap when a failure is detected. In various cases, the swap can include switching input / output (I / O) operations for the workload and directing one or more volumes of the data storage to corresponding copies of the volumes of the secondary source storage without affecting I / O production work. The swap can be a peer-to-peer remote replication failover. One or more nodes (e.g., virtual machines) that are executing the workload did not previously know which of the two data volumes was operating as the primary copy; instead, the node only knew that its workload was supported. Application I / O is redirected to the secondary storage subsystem, allowing the application to continue running.

[0015] Such an active-active site can support more than one input / output group. Data written to a volume can be automatically sent to replicas at both sites. If one site is no longer available, the other site can provide access to the volume. An active-active relationship can be established between replicas at each site. Data flows automatically run and switch directions based on which replica or replicas are online, up-to-date, and available. These relationships help clarify which replica should be provided to the workload / application. The latest replica can be selected by a single volume. A single volume can have a unique ID. Relationships can be grouped into consistency groups similar to those used for synchronous and asynchronous mirroring relationships. Based on the status of all replicas in the group, the consistency group fails over consistently as a group. Mirrors that can be used for disaster recovery are maintained at each site.

[0016] When the system topology is set to such an active-active configuration, each node, controller, and host in the system configuration can have a site attribute set to 1 or 2. Two node containers of an input / output group can be in the same site. The site can be the same site as the controller that provides the managed disks to the input / output group. When managed disks are added to the storage pool, their site attributes can be matched. This matching ensures that each replica in the active-active relationship is completely independent and in different sites.

[0017] The Small Computer System Interface (SCSI) protocol allows storage devices to indicate a preferred port for a host to use when submitting input / output requests. Using the Asymmetric Logical Unit Access (ALUA) state of a volume, the storage controller can inform the host which paths are active and which paths are preferred. The system can suggest that the host use a "local" node instead of a remote node. A "local" node is a node that is configured at the same site as the host.

[0018] Active-active relationships can be used to manage synchronous replication of volume data between two sites. The primary volume can be accessed through the input / output group. The synchronization process begins after the change volume is added to the active-active relationship.

[0019] In other storage systems sometimes referred to as disaster recovery systems, it is desirable that the physical distance between the primary data volume storage and the secondary storage storage for the backup of the primary data volume be greater. The primary storage and the secondary storage may include disks referred to as production disks and backup disks, respectively. In instances where a large event destroys the storage function and accessibility, for example, in the case of a fire, earthquake, destruction, power outage, or other catastrophic event, a greater physical distance may be useful. Utilizing a greater physical distance, for example, a distance greater than 100 km, reduces the chance that the first destructive large event will damage both the primary storage and the secondary storage. This storage for geographically dispersed disaster recovery will typically be implemented using a one-to-one mapping for volume data. Such a disaster recovery technical solution may be considered an active-passive data architecture, and is generally not robust enough to provide high availability such as provided by an active-active system. This reduced robustness level is acceptable because disaster recovery backups will be less often needed.

[0020] The exemplary embodiment described below provides a method for integrating such different data replication systems for improving data availability and disaster recovery. The present embodiment integrates a one-to-one mapped active-passive memory storage system with a swappable active-active mirrored memory storage system. Such integration has previously been a challenge because (1) the active disks used by the virtual machines in the active-active mirrored memory system change without user intervention during operation, and (2) the further remote memory storage used for disaster recovery includes a number of disks that matches the number of disks at one of the two sites of the active-active mirrored storage, and there are not enough disks to match the total number of disks from the two sites of the active-active mirrored storage. The present embodiment includes using and setting storage class descriptor bits and / or storage class descriptor areas to save operating system specific information that can be retrieved and used by the controller to handle disaster recovery and organize the creation and / or update of the further remote copy in a manner that allows the further remote copy to remain up-to-date regardless of the active-active mirrored copy pair changing its active-active relationship during operation. Therefore, the present embodiment achieves improvements in memory storage for supporting high-availability data storage used by critical workloads whose downtime is highly disruptive to the organization.

[0021] Figure 1 An enhanced replication system architecture 100 is shown in accordance with at least one embodiment. Figure 1 An active-active data replication system including a first site 102a and a second site 102b is shown in the left portion of the screen. The two sites 102a, 102b can allow local data replication and the above-mentioned active-active volume replication relationship. Since, for example, the distance between them is 100km or less, the two sites can be considered local relative to each other. The first site 102a includes a plurality of virtual machines configured to operate / run workloads / applications. The first virtual machine 104 is marked at the first site 102a. Active-active data storage for supporting the workload of the first virtual machine 104 is provided via a disk at the first site 102a and at the second site 102b. A first disk controller 106a and a second disk controller 106b are arranged at the first site 102a to control a first site disk 108a also arranged at the first site 102a. Data can be stored on the first site disk 108a. A third disk controller 106c is arranged at the second site 102b to control a second site disk 108b also arranged at the second site 102b. The data may be stored on the second site disk 108b. The various disk controllers may send control transmissions, such as the first control transmission 109, to one or more of the disks in order to save the data and update the saved data.

[0022] In some embodiments, a first data volume for supporting the workload of the first virtual machine 104 is stored in the first site disk 108a. In some embodiments, a copy of the data volume is stored in the second site disk 108b. Data generated via the workload executed by the first virtual machine 104 can be automatically sent to be stored as a first volume copy at the first site disk 108a and as a second volume copy at the second site disk 108b. If the first site disk 108a or the second site disk 108b becomes unavailable, the other of the two sites can make the data / volume available for the workload / application of the first virtual machine 104. An active-active relationship can be established between the copies so that the data flow automatically runs and switches direction according to which copy or copies are online and up-to-date. If the first site storage fails, the system can continue to support the operation of the first virtual machine 104 by transferring the operation to the storage of the second site 102b (e.g., to the third disk controller 106c and the second site disk 108b). The virtual machine is unaware of which storage is supporting its operations and is unaware that the designation of primary storage is transferred from the first site 102a to the second site 102b, but from the perspective of the first virtual machine 104, the system continues to operate.

[0023] The first site 102a and the second site 102b and their internal components can communicate with each other and exchange data via a communication network. The communication network that allows communication between HyperSwap sites can include various types of communication networks, such as the Internet, a wide area network (WAN), a local area network (LAN), a telecommunications network, a wireless network, a public switched telephone network (PTSN), and / or a satellite network. The communication network can include connections such as wired, wireless communication links, and / or fiber optic cables.

[0024] Figure 1, further remote replication site 112 is shown on the right side of the diagram, which is arranged at a great distance from the first site 102a and the second site 102b and is geographically dispersed. The further remote replication site 112 can host a further copy of the data volume that was first stored at the first site 102a and / or the second site 102b. The further remote replication site 112 includes a disaster recovery disk 114, which can store a further copy of the data volume. The remote site controller 110 controls the operation of the further remote replication site 112 and facilitates integration with the active-active data replication system including the first site 102a and the second site 102b. The further remote replication site 112 can be arranged at a second distance from the first site 102a, which is greater than the first distance between the first site 102a and the second site 102b. For example, the first site 102a and the second site 102b can be in the same region or metropolitan area, while the further remote replication site 112 can be arranged in a completely different region in the country. For example, the further remote replication site 112 may be in the Chicago metropolitan area, while the first site 102a and the second site 102b may both be in the Austin, Texas metropolitan area.

[0025] Building a data replication system including both a data replication backup at the second site 102b and a data replication backup at the further remote replication site 112 allows for diversification of data recovery and enhanced data recovery when a primary system failure occurs. Such a primary system failure may occur for a variety of reasons. In this embodiment, the data replication backup at the further remote replication site 112 is integrated with the active-active replication between the first site 102a and the second site 102b, and the following challenges are overcome: the virtual machine at the further remote replication site 112 may have a disk that is only a copy of the disk used for primary storage at the first site 102a or the disk used for secondary storage at the second site 102b, depending on which of the two sites has currently designated its disk as the leading disk for active-active backup.

[0026] The disaster recovery disk 114 can store another copy of the data volume and is set to actively mirror the first copy in the first site disk 108a or the second copy in the second site disk 108b, but not both. The storage at the further remote replication site 112 here can be referred to as active-passive storage, because the data will not be relied upon unless the primary system fails and the virtual machine at the further remote replication site 112 is activated using the volume copy at the disaster recovery disk 114 supporting the newly started remote virtual machine. This embodiment helps the further remote replication site 112 know which of the first two copies to actively mirror. The first site disk 108a and the second site disk 108b that store data for the volume copy can be twice the number of disaster recovery disks 114 that store data for the volume copy.

[0027] Another remote replication site 112 may include one or more on-demand virtual machines 116 that are activated as needed. If a failure occurs at the first site 102a and the second site 102b, such activation may occur so that support for the workload of operating the first virtual machine 104 may be transferred to the volume copy in the further remote replication site 112 and the disaster recovery disk 114. In this type of failover, the on-demand virtual machine 116 may be activated to support the workload / operation and better integrate with the data copy being used from the disaster recovery disk 114. Due to this inactivation during normal operation of the first site 102a and / or the second site 102b, the data replication at the further remote replication site 112 here may be considered as active-passive replication. If a failure occurs at the first site 102a and the second site 102b, such on-demand virtual machines 116 may need to be activated to continue the operation of the workload / application. Another remote replication site 112 may also include a virtual input / output server 118 that facilitates sharing of physical input / output resources between client logical partitions within the server.

[0028] The virtual input / output server 118 may receive information from the remote site controller 110 for starting and updating other copies of the data volumes in the disaster recovery disk 114. The remote site controller 110 may be a server, such as a computer, that communicates between (1) an active-active zone, i.e., a zone including the first site 102a and the second site 102b, and (2) an active-passive zone of another remote replication site 112. The remote site controller 110 may provide a single point of control for the entire environment managed by the disaster recovery technology solution including active-passive replication. To be successful, the remote site controller 110 must not be affected by errors that may cause an outage in the production system. Therefore, the remote site controller 110 must be self-contained and share a minimum number of resources with the production system. For example, the remote site controller 110 may be deployed in an alternate site to the first site 102a and the second site 102b, so that the remote site controller 110 is isolated from any problems or failures in the active site (the first site 102a and / or the second site 102b). In some embodiments, the remote site controller 110 may have an out-of-band deployment in its own logical partition running on an operating system.

[0029] If a disaster or potential disaster occurs that disables the first site 102a and the second site 102b, the remote site controller 110 is responsible for recovery actions. Therefore, the availability of the remote site controller 110 is a basic requirement of the technical solution. The remote site controller 110 is deployed in the alternative site and must remain operational even if the active site fails or if the disk located in the active site fails. The remote site controller 110 can constantly monitor the production environment to discover any unplanned interruptions that affect the production site or the disk system. If an unexpected power outage occurs, the remote site controller 110 can analyze the situation to determine the status of the production environment. When a site fails, in some embodiments, the remote site controller 110 can notify the administrator of the failure. If the failure is serious, the administrator can initiate a site takeover. Optionally, if the remote site controller 110 senses that the primary site has failed, the remote site controller 110 can initiate a failover to another remote replication site 112 by itself. The remote site controller 110 can suspend the processing of data replication to ensure secondary data consistency and handle site takeover.

[0030] The remote site controller 110 may handle discovery, authentication, monitoring, notification, and recovery operations to support disaster recovery of a technical solution that invokes the use of yet another remote replication site 112. The remote site controller 110 may interact with a hardware management console to collect configuration information of the managed system. The remote site controller 110 may interact with the first disk controller 106a, the second disk controller 106b, the third disk controller 106c, and / or the virtual input / output server 118, and may do so through the hardware management console to obtain storage configuration information of the virtual machine. The remote site controller 110 provides storage replication management, and may also provide management of computing / storage capacity required on demand.

[0031] The remote site controller 110 may run in an operating system logical partition. The operating system logical partition may include customized security according to the operating system requirements of the corresponding organization. In some embodiments, management of the remote site controller 110 may be enabled only for the root user in the operating system logical partition. In some embodiments, the remote site controller 110 may be restricted to not communicate with any external system except the hardware management console. The remote site controller 110 may use one or more application programming interfaces ("APIs") to communicate with the hardware management console. These application programming interfaces may include those application programming interfaces that conform to the design principles of the representational state transfer architectural style, for example, those application programming interfaces that require HTTPS to be enabled in the enhanced replication system architecture 100.

[0032] In this embodiment, the data replication backup at another remote replication site 112 is integrated with the active-active replication between the first site 102a and the second site 102b, and the following challenges are overcome: the virtual machine at the another remote replication site 112 may have a disk that is only a copy of the disk used for primary storage at the first site 102a or the disk used for secondary storage at the second site 102b, depending on which of the two sites is currently designated as the leading disk for active-active backup. In this embodiment, the storage class descriptor bit and / or storage class descriptor area in the disk of the active-active system is used to store active-active configuration information. This information can be retrieved and then used to select the correct disk for another remote replication site 112 to replicate as an additional copy of the volume data for supporting workloads / applications. The correct disk can be part of a consistency group and is replicated to be stored at another remote replication site 112 for use in the event of a disaster. The active-active configuration information in the storage class descriptor bit and / or area may include an indication that the active-active feature is enabled and which disk in the active-active pair is currently configured as the leading disk. The remote site controller 110 can retrieve the information and manage control of the additional replica, such as creation and update, at the further remote replication site 112 based on the retrieved information. The additional replica is created and / or updated to match the replica in the dominant disk in the active-active pairing, i.e., the first site disk 108a or the second site disk 108b. The remote site controller 110 controls the further remote replica by creating the further remote replica and by updating the further remote replica to correspond to changes, additions, deletions, and / or other updates in the designated dominant disk. The remote site controller 110 can access the data of the determined active disk to send to the further remote replication site 112 for controlling the additional replica of the data stored at the further remote replication site 112.

[0033] The remote site controller 110 may retrieve active-active information in out-of-band communications with the first site 102a and / or the second site 102b, in particular with the first disk controller 106a, the second disk controller 106b, and / or the third disk controller 106c. Information is then appropriately retrieved from the corresponding storage class descriptor area in the first site disk 108a and / or the second site disk 108b via the corresponding disk controller. When the active / leading disk changes during operation (e.g., from the first site disk 108a to the second site disk 108b), the kernel extension of the active-active system may update the information in the storage class descriptor bits for the disk. The remote site controller 110 may recognize the update and modify the replication / update of the additional replica at the yet another remote replication site 112 in the disaster recovery disk 114 accordingly. Thus, in the above example where the first site disk 108a is the active / dominant disk for an active-active configuration and a change subsequently occurs, where the second site disk 108b becomes the active / dominant disk for the active-active configuration, the additional replica in the disaster recovery disk will begin receiving updates based on the changes in the second site disk 108b rather than based on the changes in the first site disk 108a. The remote site controller 110 that retrieves the active-active information is an example of active-active details that are dynamically queried and utilized outside of the virtual machine (e.g., outside the first virtual machine 104). However, in existing active-active configurations, the operating system of the supported virtual machine does not know which of the two disks is the dominant disk and which is the backup, and the relationship information and configuration information are now exposed to the remote site controller 110. Out-of-band communication as referred to herein may refer to communication that is not accompanied by and / or is not part of conventional data transfer. dscli command communication is an example of such out-of-band communication.

[0034] Figure 1 The connections between the first disk controller 106a, the second disk controller 106b, and the third disk controller 106c and the disaster recovery disk 114 are shown respectively, and the disaster recovery disk 114 is used to send volume data for creating and / or updating additional volume copies at another remote replication site 112. These connections can occur via a communication network such as the Internet, a wide area network (WAN), a local area network (LAN), a telecommunications network, a wireless network, a public switched telephone network (PTSN), and / or a satellite network. The communication network can include connections such as wired, wireless communication links, and / or fiber optic cables. In these connections, Figure 1 A first remote storage transfer 120 between the first disk controller 106a and the disaster recovery disk 114 is labeled.

[0035] The remote site controller 110 may retrieve active-active information in out-of-band communication with the first site 102a and / or the second site 102b, and in particular with the first disk controller 106a, the second disk controller 106b, and / or the third disk controller 106c. Figure 1 Arrows between the remote site controller 110 and the third disk controller 106c are included to illustrate an example of such out-of-band communication for retrieving active-active enablement and configuration information from a disk (in this case, from the storage descriptor area of ​​the second site disk 108b). Although arrows between the remote site controller 110 and the first disk controller 106a and between the remote site controller 110 and the second disk controller 106b are not shown in the figures for the purpose of simplicity, in at least some embodiments, the remote site controller 110 performs such out-of-band communication with the first disk controller 106a and / or with the second disk controller 106b.

[0036] Through the tight integration of the operating system and storage technology, software with permission to perform out-of-band communication with a specific memory storage can recognize that active-active exchangeable data replication is configured in the system. In the event that the active-active setting is enabled and a production site virtual machine fails, the application can continue after the virtual machine at the remote site is restarted. This restart support overcomes the current situation where a failure at the production site is fatal even if the remote site has sufficient capacity and support. This embodiment only requires half of the production site disks to make all applications / systems work again. The consistency group membership will be automatically adjusted on the fly based on the dynamic behavior of the active disks used by the corresponding virtual machines.

[0037] It should be understood that Figure 1 An illustration of one implementation is provided and does not imply any limitation on other embodiments in which the replication integration system and method may be implemented.Many modifications may be made to the depicted environments, structures, and components based on design and implementation requirements.

[0038] Reference now Figure 2 , an operational flow chart depicting an operation that may be used according to at least one embodiment Figure 1 The integrated process 200 is performed by the enhanced replication system architecture 100 shown in FIG. Various modules, user interfaces, and services, and data stores can be used to perform the integrated process 200.

[0039] In step 202 of the integrated process 200, in response to the active-active relationship being activated, a storage descriptor area is modified. The modification may include setting one or more bits to indicate the enablement of the active-active replication relationship and also indicating which disk is the lead disk in the active-active replication relationship. Details of the mirrored active-active storage relationship may be stored in a storage class descriptor area on the disks of the mirrored volume.

[0040] For example, when Figure 1 When the active-active replication relationship is activated in the enhanced replication system architecture 100 shown to support the workload running on the first virtual machine 104, in some embodiments, the first site disk 108a is designated as the leading disk, and the second site disk 108b is designated as the backup disk. Therefore, when the workload / application runs on the first virtual machine 104, the volume copy is started and updated in the first site disk 108a. The backup copy of the volume copy is started and updated in the second site disk 108b. The transmission from the first site 102a to the second site 102b on the communication network is used to send instructions for starting and / or updating the backup volume copy in the second site disk 108b.

[0041] A kernel extension for active-active programs may be provided on the first disk controller 106a to set these information bits in the storage class descriptor area on the first site disk 108a. Another kernel extension for active-active programs may be provided on the third disk controller 106c to set these information bits in the storage class descriptor area on the second site disk 108b. Figure 3 Describe the bit settings in more detail.

[0042] If the second disk controller 106b is used to support the workload on another virtual machine at the first site 102a, a similar bit setting may occur, wherein the first site disk 108a is also used as the leading disk for storing the data volume copy for the workload run by the other virtual machine. As in the previous example, the second site disk 108b may also be used as a backup disk in an active-active replication relationship with the first volume copy at the first site 102a controlled via the second disk controller 106b.

[0043] If the third disk controller 106c is used to support the workload on another virtual machine at the second site 102b, another bit setting may occur, wherein the second site disk 108b is used as the leading disk for storing the data volume copy for the workload run by the other virtual machine. In this embodiment, the first site disk 108a may be used as a backup disk in an active-active replication relationship with the first volume copy at the second site 102b controlled via the third disk controller 106c.

[0044] Step 202 may initiate a SCSI query command to the associated storage disk via an operating system path control module in the corresponding disk controller and occur via a listening unit. When selecting a data transfer path from two data stores, the path control module may set a bit in the storage class descriptor area. The bit indicates whether the disk is part of an active-active replication relationship and whether the disk is an active (dominant) disk. The query command from the path control module may be an in-band communication, such as a communication accompanying or part of a conventional data transfer.

[0045] In step 204 of the integration process 200, the remote site controller retrieves the modified information via communication with the disk. The remote site controller may use an application programming interface to retrieve the disk relationship and boot status stored in step 202. The retrieval may occur via out-of-band communication to one or more disks of the mirrored volume. In some embodiments, out-of-band communication to one or more disks occurs via communication through the corresponding disk controllers of those disks. Out-of-band communication may refer to control messages.

[0046] For example, in Figure 1 In the enhanced replication system architecture 100 shown, in some embodiments, the remote site controller 110 sends an application programming interface to one or more of the disk controllers, such as Figure 1 As shown by the arrow to the third disk controller 106c. Additionally and / or alternatively, a similar application programming interface query may be sent from the remote site controller 110 to the first disk controller 106a and / or the second disk controller 106b.

[0047] The retrieved modified information may include indicators as to whether active-active replication status is enabled and which disk is hosting the leading volume copy for supporting the workload.

[0048] In embodiments where the storage class descriptor area is configured to indicate active-active information at both the primary volume replica hosting disk and the backup volume replica hosting disk, the query here as part of step 204 may be sent to one or both of the two disks / two disk controllers. In some embodiments, the remote site controller 110 may send a first query to the first controller / disk and then send a second query to the second controller / disk to confirm the information retrieved in the first query.

[0049] In some embodiments, the remote station controller will send a first query to request information if an active-active relationship is currently activated, and then if the first query is positive (active-active is activated), a second query will be sent to learn which disk is the dominant disk. The information set in step 202 can include these two different information bits in different bytes, so that a two-step (first step to ask for enable, second step to ask which is the dominant disk) query can be applied in some instances.

[0050] In step 206 of the integration process 200, the remote site controller creates and / or updates the additional volume replica at the remote site based on the master disk determined in the retrieved information. The additional volume replica can be established as an asynchronous replica and is in an active-passive replication relationship with respect to the primary volume replica at the master disk.

[0051] For example, in Figure 1 In the enhanced replication system architecture 100 shown, in some embodiments, the remote site controller 110 sends instructions to the further remote replication site 112 and / or to the first site 102a or the second site 102b (whichever site hosts the disk currently designated as the lead disk in the active-active replication relationship) so that an additional volume copy will be started, updated and / or populated in the disaster recovery disk 114 at the further remote replication site 112. The data of the volume copy and / or its updates can be transmitted via a communication network between the site hosting the lead disk and the further remote replication site 112.

[0052] In step 208 of the integration process 200, it is determined whether any updates to the master disk designation have occurred. In response to the determination of step 208 being affirmative, i.e., an update to the master disk designation has occurred, the integration process 200 proceeds to step 210. In response to the determination of step 208 being negative, i.e., no update to the master disk designation has occurred, the integration process 200 proceeds to step 212.

[0053] The determination of step 208 may occur via the path control module of the corresponding disk controller snooping the unit attention to the disks involved in the active-active relationship. The snooping may occur continuously and / or intermittently, for example, on a scheduled basis with uniform periods between snooping sessions.

[0054] In step 210 of integration process 200, the storage descriptor area is modified to reflect any changed designations. In some embodiments, the active-active relationship and the kernel extension in the corresponding disk controller can be used to update the descriptor area when the active-active state and / or relationship changes. After the path control module receives the response with the changed information, the path control module can notify the kernel extension, which causes the kernel extension to update the bits in the storage class descriptor area. In some embodiments, the corresponding disk controller can send a notification signal to the remote site controller 110 in response to any changes in the active-active information (e.g., bits).

[0055] Regarding the above Figure 1 In the first example described in the enhanced replication system architecture 100 shown in FIG, an active-active replication relationship is used to support a workload running on a first virtual machine 104, wherein a first site disk 108a is designated as a leading disk, and wherein a second site disk 108b is designated as a backup disk, if some problem causes the backup disk to become the leading disk, the information in the storage class descriptor area is updated. Specifically, the information may be updated to indicate that the second site disk 108b is now the leading disk, and the first site disk 108a is a backup disk in an active-active relationship. The following will describe the Figure 3 Describe the bit changes in more detail.

[0056] After step 210, the integration process 200 returns to step 204 to repeat the above steps 204, 206 and 208 in the integration process. Since the goal of the embodiment is to provide high-availability data storage for workloads operating on virtual machines, the integration process 200 is provided with a repeating loop so that the data will continue to be available, whether provided by the original primary disk, by the active-active backup disk or by another remote (active-passive backup) disk.

[0057] In step 212 of the integration process 200 that occurs after no lead disk designation is identified in step 208, a determination is made as to whether any failover-causing site failures have occurred. In response to the determination in step 212 being affirmative, since a failover-causing site failure has occurred, the integration process 200 proceeds to step 214. In response to the determination in step 212 being negative, since a failover-causing site failure has not occurred, the integration process 200 returns to step 208 to repeat step 208. If one of the supporting disks becomes unresponsive, unavailable, or otherwise corrupted, the disk controller and / or virtual machine may initiate a failover process.

[0058] In step 214 of the integration process 200, workload support is adjusted using the replicated backup as needed. If an active-active relationship is activated, a failover to the backup copy in the active-active relationship will occur as a first attempt. For example, in the case of Figure 1 In the first example described, if the first site disk 108a becomes unavailable, support for the operation of the workload on the first virtual machine can be transferred to the second site disk 108b that already has a complete or substantial copy of the volume replica. In another example, the second site disk 108b is the leading disk but becomes unavailable, and the first site disk 108a can then be used as the primary support option for supporting the workload at one of the virtual machines. If in another example, both the first site disk 108a and the second site disk 108b become unavailable, support for the operation of the workload can be transferred to the disaster recovery disk 114 at another remote replication site 112.

[0059] In step 216 of the integrated process 200, it is determined whether the failed site has been restored. In response to the determination of step 216 being affirmative, since the failed site has been restored, the integrated process 200 proceeds to step 218. In response to the determination of step 212 being negative, i.e., the failed site has not been restored, the integrated process 200 returns to step 208 to repeat step 208.

[0060] In step 218 of the integration process 200, workload support is adjusted via the remote site controller. This step 218 may include falling back to the original production site after the original production site is restored. The remote site controller 110 and / or one of the disk controllers may track the storage relationship and status of the most recent query. The remote site controller 110 may determine the production site storage status and the ability to fall back. Fallback from another remote replication site 112 may occur even if not all mirrored disks are available, for example, if the first site disk 108a is restored but the second site disk 108b is not restored, or vice versa, if the second site disk 108b is restored but the first site disk 108a is not restored. Therefore, in these instances, although the active-active replication relationship has not yet been rebuilt, support for application operations at one of the two primary sites can be rebuilt.

[0061] After step 218, the integration process 200 returns to step 210 to repeat the above steps 210, 204, 206 and 208 in the integration process. Since the goal of the embodiment is to enhance high-availability data storage for workloads operating on virtual machines, the integration process 200 is provided with a repeating loop so that data will continue to be available, whether provided by the original primary disk, by an active-active backup disk, or by another remote (active-passive backup) disk.

[0062] Understandably, Figure 2An illustration of some embodiments is provided, and no limitation on how different embodiments may be implemented is implied. Many modifications may be made to the depicted embodiments (e.g., to the depicted sequence of steps) based on design and implementation requirements. Steps and features from various processes may be combined into other processes described in other figures or embodiments.

[0063] Figure 3 is a diagram illustrating a storage descriptor area for replication integration according to at least one embodiment. Figure 3 A first storage descriptor area 300 including byte 14 is shown, which indicates that the disk is currently in an active-active replication relationship using one of its subcommands (e.g., using its subcommand 0x67). The first storage descriptor area 300 also includes an indication in its byte 11 whether the disk is a leading disk or a backup disk in an active-active replication relationship. The path control module of the disk controller can initiate a SCSI command with an operation code 0xED and a new subcommand such as 0x67 to set these bits in response to the active-active replication relationship being started and / or changed. The first storage descriptor area 300 can be stored in the first site disk 108a or the second site disk 108b. A second similar storage descriptor area can be stored in the other of the two storage disks. However, if the first storage descriptor area 300 indicates that these corresponding disks are leading disks using byte 11, the second similar storage descriptor area will indicate that its corresponding disk is currently a backup disk in an active-active configuration using its byte 11.

[0064] The retrieval of the bits may occur via a command such as a dscli command that includes a request for disk storage. The command may use out-of-band extended communications to retrieve details of active storage from disk storage, such as information for bytes 11 and 14. The remote site controller 110 has access to the storage disk via a network and has the ability to issue commands such as dscli commands. The dscli command may include a command line interface that receives commands in text form. The remote site controller 110 may issue query commands at periodic intervals and may store information locally.

[0065] Using the embodiments described herein, if a production site virtual machine or host in an active-active group fails, the applications on the virtual machine or host can continue after the virtual machine is restarted at the disaster recovery site. This embodiment brings more transparency to the active-active relationship by indicating which disks are currently the leading disks and / or which disks are currently the backup disks. At the virtual input / output server level, different disks returned as part of the active-active relationship can trigger a request for an output query command. The first storage disk and the second storage disk currently used by the virtual machine will both have corresponding replica disks in the remote active-passive storage.

[0066] Figure 4 is a block diagram 400 of internal and external components of a computer that may be used in accordance with an illustrative embodiment of the present invention. Figure 1 or otherwise used in the above-described integrated process 200. It should be understood that Figure 4 This only provides an illustration of one embodiment and does not imply any limitation with regard to the environments in which different embodiments may be implemented. Many modifications may be made to the depicted environment based on design and implementation requirements.

[0067] Data processing systems 402a, 402b, 404a, 404b represent any electronic device capable of executing machine-readable program instructions. Data processing systems 402a, 402b, 404a, 404b may represent smart phones, computer systems, PDAs, or other electronic devices. Examples of computing systems, environments, and / or configurations that may be represented by data processing systems 402a, 402b, 404a, 404b include, but are not limited to, personal computer systems, server computer systems, thin clients, fat clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, network PCs, minicomputer systems, and distributed cloud computing environments that include any of the above systems or devices.

[0068] A computer engaged with a virtual machine (e.g., with the first virtual machine 104) of a workload and / or other computers involved in remote replication operating at the first site 102a, at the second site 102b, and at the further remote replication site 112a, and a computer operating the remote site controller 110 may include Figure 4 404a, 404b. Each of the sets of internal components 402a, 402b includes one or more processors 406, one or more computer-readable RAMs 408 and one or more computer-readable ROMs 410 on one or more buses 412, one or more operating systems 414, and one or more computer-readable tangible storage devices 416. Programs for controlling one or more of the remote site controllers 110 and the disk controllers may be stored on the one or more computer-readable tangible storage devices 416 for execution by the one or more processors 406 via the one or more RAMs 408 (which typically include cache memory). Figure 4 In the illustrated embodiment, each of the computer-readable tangible storage devices 416 is a disk storage device of an internal hard drive. Alternatively, each of the computer-readable tangible storage devices 416 is a semiconductor memory device such as ROM 410, EPROM, flash memory, or any other computer-readable tangible storage device that can store computer programs and digital information.

[0069] Each set of internal components 402a, 402b also includes a R / W drive or interface 418 to read from and write to one or more portable computer readable tangible storage devices 420, such as CD-ROMs, memory sticks, tapes, magnetic disks, optical disks, or semiconductor storage devices. Software programs, such as implemented via the remote site controller 110 and / or via one or more of the disk controllers, can be stored on one or more of the corresponding portable computer readable tangible storage devices 420, read via the corresponding R / W drive or interface 418, and loaded into the corresponding hard drive (e.g., tangible storage device 416).

[0070] Each set of internal components 402a, 402b may also include a network adapter (or switch port card) or interface TCP 422, such as an A / IP adapter card, a wireless wi-fi interface card, or a 3G, 4G, or 5G wireless interface card, or other wired or wireless communication link. Programs for the remote site controller 110 and / or for one or more other disk controllers may be downloaded from an external computer (e.g., a server) via a network (e.g., the Internet, a local area network, or other wide area network) and a corresponding network adapter or interface 422. From the network adapter (or switch port adapter) or interface 422, the program for the remote site controller 110 may be loaded into a corresponding hard drive, such as a tangible storage device 416. The network may include copper wire, optical fiber, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers.

[0071] Each of the set of external components 404a, 404b may include a computer display monitor 424, a keyboard 426, and a computer mouse 428. External components 404a, 404b may also include touch screens, virtual keyboards, touch pads, pointing devices, and other human-machine interface devices. Each of the set of internal components 402a, 402b also includes a device driver 430 to interface with the computer display monitor 424, keyboard 426, and computer mouse 428. Device driver 430, R / W driver or interface 418, and network adapter or interface 422 include hardware and software (stored in storage device 416 and / or ROM 410).

[0072] The present invention may be a system, method and / or computer program product at any possible level of technical detail integration. The computer program product may include a computer-readable storage medium (or medium) having computer-readable program instructions thereon for causing a processor to perform various aspects of the present invention.

[0073] A computer-readable storage medium may be a tangible device that can retain and store instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device (such as a punched card or a raised structure in a groove on which instructions are recorded), and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be interpreted as a temporary signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through a wire.

[0074] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network). The network can include copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network, and forwards the computer-readable program instructions to be stored in a computer-readable storage medium in the corresponding computing / processing device.

[0075] The computer-readable program instructions for performing the operation of the present invention can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, and procedural programming languages ​​such as "C" programming language or similar programming languages. The computer-readable program instructions can be executed completely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on a remote computer, or completely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using the Internet of an Internet service provider). In some embodiments, the electronic circuit including, for example, a programmable logic circuit, a field programmable gate array (FPGA) or a programmable logic array (PLA) can execute the computer-readable program instructions to personalize the electronic circuit by utilizing the state information of the computer-readable program instructions, so as to perform various aspects of the present invention.

[0076] Various aspects of the present invention are described herein with reference to the flow charts and / or block diagrams of the methods, devices (systems) and computer program products according to embodiments of the present invention. It should be understood that each frame of the flow chart diagram and / or block diagram and the combination of frames in the flow chart diagram and / or block diagram can be implemented by computer-readable program instructions.

[0077] These computer-readable program instructions can be provided to a processor of a computer or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device create components for implementing the functions / actions specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, which can instruct a computer, a programmable data processing device, and / or other devices to function in a specific manner, so that the computer-readable storage medium having the instructions stored therein includes an article of manufacture, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes of the flowchart and / or block diagram.

[0078] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing device, or other apparatus, so that a series of operational steps are performed on the computer, other programmable device, or other apparatus to produce a computer-implemented process, so that the instructions executed on the computer, other programmable device, or other apparatus implement the functions / actions specified in one or more boxes of the flowchart and / or block diagram.

[0079] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present invention.In this regard, each frame in the flow chart or block diagram can represent a module, segment or part of an instruction, which includes one or more executable instructions for realizing the specified logical function.In some optional embodiments, the function marked in the frame may not occur in the order marked in the figure.For example, the two frames shown in succession can in fact be completed as a step, executed simultaneously, substantially simultaneously, executed in a partially or completely time-overlapping manner, or these frames can sometimes be executed in reverse order, depending on the functions involved.It will also be noted that each frame of the block diagram and / or flow chart illustration and the combination of frames in the block diagram and / or flow chart illustration can be implemented by a system based on dedicated hardware that performs a specified function or action or performs a combination of dedicated hardware and computer instructions.

[0080] It should be understood that although the present disclosure includes detailed descriptions about cloud computing, the implementation of the teachings recorded herein is not limited to cloud computing environments. Instead, embodiments of the present invention can be implemented in conjunction with any other type of computing environment now known or later developed.

[0081] Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be quickly provisioned and released with minimal management effort or interaction with the service provider. The cloud model may include at least five characteristics, at least three service models, and at least four deployment models.

[0082] Features are as follows:

[0083] On-demand self-service: Cloud consumers can automatically and unilaterally provision computing capabilities, such as server time and network storage, on demand without requiring manual interaction with the service provider.

[0084] Broad network access: Capabilities are available over the network and accessed through standard mechanisms that facilitate use by heterogeneous thin-client platforms or thick-client platforms (e.g., mobile phones, laptops, and PDAs).

[0085] Resource pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, where different physical and virtual resources are dynamically allocated and reallocated based on demand. There is a sense of location independence, as consumers generally have no control or knowledge of the exact location of the resources provided, but may be able to specify a location at a higher level of abstraction (e.g., country, state, or data center).

[0086] Rapid Elasticity: Capacity can be provisioned quickly and elastically, in some cases automatically, to scale up quickly and released quickly to scale down quickly. To the consumer, the capacity available for provisioning often appears to be unlimited and can be purchased at any time in any quantity.

[0087] Metered Services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at some level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported to provide transparency to both providers and consumers of the utilized services.

[0088] The service model is as follows:

[0089] Software as a Service (SaaS): The capability provided to consumers is to use the provider's applications running on a cloud infrastructure. The applications can be accessed by a variety of client devices through a thin client interface such as a web browser (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even individual application capabilities, with the possible exception of limited user-specific application configuration settings.

[0090] Platform as a Service (PaaS): The capability provided to consumers is to deploy consumer-created or acquired applications on cloud infrastructure using programming languages ​​and tools supported by the provider. Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, or storage, but have control over deployed applications and possibly the configuration of the application hosting environment.

[0091] Infrastructure as a Service (IaaS): The capability provided to consumers is the provisioning of processing, storage, networking, and other basic computing resources, where consumers can deploy and run arbitrary software, which may include operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but have control over the operating system, storage, deployed applications, and may have limited control over select networking components (e.g., host firewalls).

[0092] The deployment model is as follows:

[0093] Private cloud: A cloud infrastructure is operated for only one organization. It can be managed by that organization or a third party and can exist on-premises or off-premises.

[0094] Community cloud: The cloud infrastructure is shared by several organizations and supports a specific community with common concerns (e.g., mission, security requirements, policies, and compliance considerations). It can be managed by these organizations or a third party and can exist on-premises or off-premises.

[0095] Public cloud: Cloud infrastructure is available to the general public or large industry groups and is owned by an organization that sells cloud services.

[0096] Hybrid cloud: A cloud infrastructure is a combination of two or more clouds (private, community, or public) that remain independent entities but are bound together by standardized or proprietary technologies that enable data and application portability (e.g., cloud bursting for load balancing between clouds).

[0097] The cloud computing environment is service-oriented, focusing on statelessness, low coupling, modularity, and semantic interoperability. The core of cloud computing is the infrastructure consisting of a network of interconnected nodes.

[0098] Reference now Figure 5 , depicts an illustrative cloud computing environment 500. As shown, the cloud computing environment 500 includes one or more cloud computing nodes 50 with which a local computing device used by a cloud computing consumer (e.g., a personal digital assistant (PDA) or cellular phone 50A, a desktop computer 50B, a laptop computer 50C, and / or an automobile computer system 50N) can communicate. The nodes 50 can communicate with each other and can include personal computers for accessing data in the cloud. They can be physically or virtually grouped in one or more networks (not shown) such as a private cloud, community cloud, public cloud, or hybrid cloud, or a combination thereof, as described above. This allows the cloud computing environment 500 to provide infrastructure as a service, platform as a service, and / or software as a service, for which the cloud consumer does not need to maintain resources on a local computing device. It should be understood that Figure 5 The types of computing devices 50A-N shown in are intended to be illustrative only, and computing node 50 and cloud computing environment 500 may communicate with any type of computerized device over any type of network and / or network addressable connection (eg, using a web browser).

[0099] Reference now Figure 6 , shows a set of functional abstraction layers 600 provided by the cloud computing environment 500. First, it should be understood that Figure 6 The components, layers, and functions shown in are intended to be illustrative only, and embodiments of the present invention are not limited thereto. As shown, the following layers and corresponding functions are provided:

[0100] The hardware and software layer 602 includes hardware and software components. Examples of hardware components include: mainframes 604; servers based on RISC (Reduced Instruction Set Computer) architecture 606; servers 608; blade servers 610; storage devices 612; and network and networking components 614. In some embodiments, software components include network application server software 616 and database software 618.

[0101] Virtualization layer 620 provides an abstraction layer through which the following examples of virtual entities may be provided: virtual servers 622 ; virtual storage 624 ; virtual networks 626 , including virtual private networks; virtual applications and operating systems 628 ; and virtual clients 630 .

[0102] In one example, the management layer 632 may provide the functionality described below. Resource provisioning 634 provides dynamic procurement of computing resources and other resources for performing tasks within a cloud computing environment. Metering and pricing 636 provides cost tracking when utilizing resources within a cloud computing environment, as well as billing or invoicing for the consumption of these resources. In one example, these resources may include application software licenses. Security provides authentication for cloud consumers and tasks, as well as protection for data and other resources. User portal 638 provides access to the cloud computing environment for consumers and system administrators. Service level management 640 provides cloud computing resource allocation and management so that the required service levels are met. Service level agreement (SLA) planning and fulfillment 642 provides pre-arrangement and procurement of cloud computing resources in anticipation of future demand according to the SLA.

[0103] Workload layer 644 provides examples of functionality that can take advantage of a cloud computing environment. Examples of workloads and functionality that can be provided through this layer include: mapping and navigation 646; software development and lifecycle management 648; virtual classroom instruction delivery 650; data analytics processing 652; transaction processing 654; and replication system integration 656. Performing active-active system integration with active-passive systems using remote site controller 110 and / or one or more disk controllers by updating storage class descriptor region information in storage disks and retrieving that information provides a way to integrate these different replication systems.

[0104] The terms used herein are only used for the purpose of describing specific embodiments and are not intended to limit the present invention. As used herein, the singular forms "a", "an" and "the" are intended to also include plural forms, unless the context clearly indicates otherwise. It will be further understood that when used in this specification, the terms "comprises", "comprising", "includes", "including", "has", "have", "having", "with" and the like specify the presence of stated features, integers, steps, operations, elements and / or parts, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, parts and / or groups thereof.

[0105] The description of various embodiments of the present invention has been given for the purpose of illustration, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, practical applications, or technical improvements over technologies found in the marketplace, or to enable those of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A computer-implemented method comprising: sending, via a computer, a query command to a storage descriptor area of ​​a first disk, the first disk belonging to a dual-site data replication system that provides active-active access to a data volume stored in an active disk and replicated in a backup disk; receiving, via the computer, a response to the query command, the response indicating the active disk and the backup disk used for the dual-site data replication system; as well as An additional copy of the data volume at another remote site is controlled via the computer based on the active disk.

2. The method according to claim 1, wherein: The sending occurs as an out-of-band communication with the first disk.

3. The method according to claim 1, further comprising: In response to determining that the relationship between the active disk and the backup disk has changed, at least one indicator of the relationship is updated in the storage descriptor area.

4. The method according to claim 1, further comprising: In response to a failure of the dual-site data replication system, a workload is transferred to the additional copy of the volume at the further remote site.

5. The method according to claim 4, further comprising: When the dual-site data replication system is restored, the workload is returned to the dual-site data replication system.

6. The method according to claim 5, wherein: Recovery of the dual-site data replication system occurs only for the first disk hosting the data volume and not for the second disk of the dual-site data replication system.

7. The method according to claim 1, wherein: The storage descriptor area of ​​the first disk includes a custom status bit, and the response to the query command is based on the custom status bit.

8. A computer system comprising: one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage media, and program instructions stored on at least one of the one or more computer-readable tangible storage media for execution by at least one of the one or more processors via at least one of the one or more computer-readable memories to cause the computer system to: sending a query command to a storage descriptor area of ​​a first disk, the first disk belonging to a dual-site data replication system, the dual-site data replication system providing active-active access to a data volume; receiving a response to the query command, the response indicating an active disk and a backup disk for the dual-site data replication system; as well as Based on the active disk, an additional copy of the data volume at another remote site is controlled.

9. The computer system according to claim 8, wherein: The sending occurs as an out-of-band communication with the first disk.

10. The computer system of claim 8, wherein the program instructions are further configured to execute to cause the computer system to: In response to determining that the relationship between the active disk and the backup disk has changed, at least one indicator of the relationship is updated in the storage descriptor area.

11. The computer system of claim 8, wherein the program instructions are further configured to execute to cause the computer system to: In response to a failure of the dual-site data replication system, a workload is transferred to the additional copy of the volume at the further remote site.

12. The computer system of claim 11, wherein the program instructions are further configured to execute to cause the computer system to: When the dual-site data replication system is restored, the workload is returned to the dual-site data replication system.

13. The computer system according to claim 12, wherein: Recovery of the dual-site data replication system occurs only for the first disk hosting the data volume and not for the second disk of the dual-site data replication system.

14. The computer system according to claim 8, wherein: The storage descriptor area of ​​the first disk includes a custom status bit, and the response to the query command is based on the custom status bit.

15. A computer program product comprising a computer readable storage medium having program instructions embodied therewith, wherein the program instructions are executable by a computer system to cause the computer system to: sending a query command to a storage descriptor area of ​​a first disk, the first disk belonging to a dual-site data replication system, the dual-site data replication system providing active-active access to a data volume; receiving a response to the query command, the response indicating an active disk and a backup disk for the dual-site data replication system; as well as Based on the active disk, an additional copy of the data volume at another remote site is controlled.

16. The computer program product of claim 15, wherein: The sending occurs as an out-of-band communication with the first disk.

17. The computer program product of claim 15, wherein: The program instructions are also used to execute so that the computer system: In response to determining that the relationship between the active disk and the backup disk has changed, at least one indicator of the relationship is updated in the storage descriptor area.

18. The computer program product of claim 15, wherein: The program instructions are also used to execute so that the computer system: In response to a failure of the dual-site data replication system, a workload is transferred to the additional copy of the volume at the further remote site.

19. The computer program product of claim 18, wherein: The program instructions are also used to execute so that the computer system: When the dual-site data replication system is restored, the workload is returned to the dual-site data replication system.

20. The computer program product of claim 19, wherein: Recovery of the dual-site data replication system occurs only for the first disk hosting the data volume and not for the second disk of the dual-site data replication system.