Asynchronous inter-region block volume replication
Asynchronous cross-region block volume replication through snapshots and deltas addresses latency and recovery time issues in cloud-based data replication, ensuring efficient and reliable data protection.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- ORACLE INT CORP
- Filing Date
- 2026-01-14
- Publication Date
- 2026-05-26
AI Technical Summary
Existing cloud-based data replication methods, such as full volume backups, suffer from significant latency, increased computing requirements, and resource costs due to metadata overhead, leading to prolonged recovery times and potential data loss during failures.
Implementing asynchronous cross-region block volume replication by generating snapshots and deltas, allowing incremental data transfer and reducing the impact on I/O operations, with delta generation and transfer occurring independently of unchanged data.
This approach minimizes data loss and recovery time, enhances system performance, and improves disaster recovery capabilities by maintaining data availability and reducing latency during failures.
Smart Images

Figure 2026086422000001_ABST
Abstract
Description
Technical Field
[0001] Cross - reference to related applications This application claims the benefit and priority of U.S. Non - Provisional Application No. 17 / 091,635, filed on November 6, 2020, entitled "ASYNCHRONOUS CROSS - REGION BLOCK VOLUME REPLICATION". The entire content of U.S. Non - Provisional Application No. 17 / 091,635 is hereby incorporated by reference in its entirety for all purposes.
Background Art
[0002] Background Cloud - based platforms provide users with scalable and flexible computing resources. Such cloud - based platforms are also called infrastructure as a service (IaaS) and can provide an entire suite of cloud solutions based on customer data, for example, solutions for orchestrating transformations, loading data, and presenting data. In some cases, customer data may be stored in block - volume storage and / or object storage within a distributed storage system (e.g., cloud storage).
Summary of the Invention
Means for Solving the Problems
[0003] Summary Techniques for asynchronous cross - region block - volume replication are provided (e.g., a method, system, non - transient computer - readable medium for storing code or instructions executable by one or more processors).
[0004] In one embodiment, the method includes the step of a computer system creating a first snapshot of a block volume containing multiple partitions in a first geographical area at a first logical time. The method includes the computer system transmitting first snapshot data corresponding to the first snapshot to an object storage system in a second geographical area. The method includes the computer system creating a second snapshot of the block volume in the first geographical area at a second logical time. The method includes the computer system generating multiple deltas, each delta corresponding to one of the multiple partitions. The method includes the computer system transmitting multiple delta datasets corresponding to the multiple deltas to an object storage system in a second geographical area. The method includes the computer system generating a checkpoint by at least partially aggregating the multiple deltas and the object metadata associated with the first snapshot. The method includes the computer system receiving a restore request for generating a restore volume. The method also includes the computer system generating a restore volume from the checkpoint.
[0005] In one variation, the step of generating multiple deltas includes a step of generating a comparison between the second snapshot and the first snapshot. Based on this comparison, the step of generating multiple deltas uses the first snapshot data and the second snapshot data. The method may include a step of determining corrected data corresponding to changes between the first snapshot data and a second snapshot data corresponding to the first snapshot data. Multiple deltas may describe corrected data for multiple partitions. The step of creating the first snapshot may include, in logical time, a step of interrupting I / O operations for the multiple partitions, a step of generating multiple block images describing volume data in the multiple partitions, and a step of enabling I / O operations for the multiple partitions. A restore request may be or include a failover request. The method may further include a step of enabling the creation of a restore volume in a second geographical area, and a step of enabling I / O operations using the restore volume in the second geographical area. A restore request may be or include a failback request. The method may further include a step of creating a restore volume in a second geographical area, a step of enabling the creation of a failback volume in the first geographical area by at least partially cloning the restore volume, and a step of restoring the first snapshot data in the first geographical area. The step of sending multiple delta datasets may include the steps of generating multiple chunk objects from the multiple delta datasets, transferring the multiple deltas, and transferring the multiple chunk objects to an object storage system. The checkpoint may include an object metadata manifest. The object metadata may include chunk pointers corresponding to multiple chunk objects in the object storage system. The step of aggregating the object metadata may include updating the manifest to reflect multiple differences between the multiple delta datasets and the first snapshot data.
[0006] In some embodiments, the computer system includes one or more processors and a memory that communicates with the one or more processors, the memory being configured to store computer executable instructions, and by executing these computer executable instructions, the one or more processors are made to perform one or more of the steps of the method or modified versions described above.
[0007] In some embodiments, the computer-readable storage medium stores computer-executable instructions, which, when executed, cause one or more processors of a computer system to perform one or more steps of the method or variations thereof described above. [Brief explanation of the drawing]
[0008] [Figure 1] This figure shows an exemplary system for asynchronous inter-region block volume replication according to one or more embodiments. [Figure 2] This figure shows exemplary techniques for asynchronous block volume replication according to one or more embodiments. [Figure 3] This figure shows an exemplary technique for restoring a block volume system by asynchronous replication according to one or more embodiments. [Figure 4] This figure shows an exemplary technique for aggregating block volume replication metadata according to one or more embodiments. [Figure 5] This figure shows an exemplary technique for generating a failover volume from a standby volume according to one or more embodiments. [Figure 6] This figure shows exemplary techniques for resizing a standby volume according to one or more embodiments. [Figure 7] This figure shows an exemplary flow for generating a recovery volume according to one or more embodiments. [Figure 8] This figure shows an exemplary flow for generating a failover volume according to one or more embodiments. [Figure 9] This figure shows an exemplary flow for generating a failback volume according to one or more embodiments. [Figure 10] This figure shows an exemplary flow for restoring a block volume system using a standby system, according to one or more embodiments. [Figure 11] This figure shows an exemplary flow for resizing a block volume system and a standby system according to one or more embodiments. [Figure 12] This block diagram shows one pattern for implementing cloud infrastructure as a service system, according to at least one embodiment. [Figure 13] This block diagram shows another pattern for implementing cloud infrastructure as a service system, following at least one embodiment. [Figure 14] This block diagram shows another pattern for implementing cloud infrastructure as a service system, following at least one embodiment. [Figure 15] This block diagram shows another pattern for implementing cloud infrastructure as a service system, following at least one embodiment. [Figure 16] This is a block diagram illustrating an exemplary computer system according to at least one embodiment. [Modes for carrying out the invention]
[0009] In the attached drawings, similar components and / or features may have the same reference label. Furthermore, various components of the same type may be distinguished by a dash and a second label following the reference label to identify similar components. Where only the first reference label is used herein, the description is applicable to any one of the similar components having the same first reference label, notwithstanding the second reference label.
[0010] Detailed explanation The following description will explain various embodiments. For the purpose of explanation, specific configurations and details will be described to ensure that the embodiments are fully understood. However, it will be apparent to those skilled in the art that embodiments may be carried out without these specific details. Furthermore, well-known features may be omitted or simplified in order to avoid obscuring the embodiments described.
[0011] Cloud-based platforms provide users with scalable and flexible computing resources. Such cloud-based platforms, also known as Infrastructure as a Service (IaaS), can provide a complete suite of cloud solutions based on customer data, such as solutions for authoring transformations, loading data, and presenting data. In some cases, customer data may be stored in block volume storage and / or object storage within a distributed storage system (e.g., cloud storage). Customer data may also be stored in geographically located data centers, for example, as part of a wide-area distributed storage system. Data centers may be selected based at least partially on one or more performance metrics, including latency, input-output operations per second (IOPS), throughput, cost, and stability. In some cases, the optimal location of a data center may correspond to the customer's location (e.g., if the customer generates a significant amount of internal data). In other cases, the optimal location may correspond to the customer's client and / or user locations (e.g., if the customer operates a content delivery network).
[0012] Customer data is copied to a backup system in a different geographical area or data center (also known as an availability domain or "AD"), and backed up By designating the backup system as the primary system, it can be recovered after a failure (referred to as "failover"). The failover system can be characterized by multiple metrics including the recovery point objective (RPO), the recovery time objective (RT O), and disaster recovery (DR). Generally, the RPO describes how much data may be lost during a failover, so minimizing the RPO can be a goal of the backup system. Typically, generating a backup copy involves taking a backup in one area, waiting for the backup to be uploaded to the cloud system, then starting a copy operation between areas and waiting for the copy to complete in the destination area. Considerable metadata overhead may be associated with each backup, which can increase computing requirements and resource costs at a high repetition rate. Therefore, the RPO is typically available on the order of several hours for a backup system (where time indicates the amount of data considering the average data transfer rate). Similarly, since the RTO indicates the length of time related to restoring data after a failover, it can be an important metric for the failover system. When a backup is restored from another geographical area, the input / output operations may have a longer waiting time than typical conditions until the volume is fully restored. Additionally, the waiting time for restore time and on-demand data operations can increase significantly, for example, in response to spikes in system load corresponding to shifts in data traffic due to a failure across the entire area. The waiting time may be even longer while the distributed data storage system is implementing backup restoration from an inter-area backup copy.
[0013]
[0014] Since customer data can be beneficial, customers may expect that the data is protected by one or more approaches, including data redundancy and infrastructure implemented to enhance storage stability (e.g., power boost, surge protection, etc.). Data redundancy can take multiple forms, including but not limited to replicating customer data to a local storage medium (e.g., physical storage medium), a distributed storage system, and / or another distributed storage system. In some cases, customer data may be replicated in a batch, which is referred to as a backup. Such an approach may involve significant latency in system input / output operations during the execution of the backup. For example, the storage system may stop all read and write operations during the execution of the backup to ensure that a complete copy of the customer data is securely stored. In such cases, additional buffer capacity may be employed to temporarily store new data within the distributed storage system, and / or customer data requests may be delayed.
[0015] In contrast to a full volume backup, customer data can be replicated asynchronously, as will be explained in more detail with reference to some diagrams below. For example, customer data can be replicated persistently (e.g., at intervals of minutes, as opposed to intervals of several hours, which may be specific to a backup system) by generating objects (e.g., partitions in a block volume system) that describe incremental changes in a subunit of the data storage system. In the following paragraphs, the objects describing incremental changes will also be referred to as "deltas." A delta can describe one or more changes to customer data between two asynchronous records, as will be explained in more detail with reference to Figure 1. Records of the state of user data in a distributed storage system are also referred to as "snapshots" and can represent the state of data at a single logical time, as opposed to a single time series, as will be explained in more detail with reference to the diagrams below. In some cases, deltas may be generated, and such deltas may be stored and / or transferred / replicated once generated. Thus, asynchronous replication replicates customer data by transferring deltas rather than by backup. This could make it possible.
[0016] For at least these reasons, the technologies described herein offer one or more advantages over data redundancy approaches that rely on backup customer data. For example, by generating multiple deltas to replicate changes to data after a snapshot, it may be possible to subdivide the data replication process chronologically and reduce the impact of data replication on latency, IOPS, and / or other performance metrics.
[0017] An approach that relies on backing up the entire volume may carry the risk of data loss if a system failure occurs when available backups do not reflect beneficial changes to customer data (for example, if a failure occurs long after an existing backup has been performed, or if significant changes have occurred to customer data since the last backup). Implementing asynchronous replication can also improve overall system performance and, in some cases, improve the retention of customer data. Therefore, an asynchronous replication approach can improve RPO, RTO, and DR compared to a backup system.
[0018] As a concrete example, a block volume system can be implemented in a distributed storage system having multiple data centers in multiple geographical areas around the world. Customer data stored in a block volume system in one of the data centers can be replicated in a second data center to protect against data loss if the first data center suffers a catastrophic failure. Customer data may also be replicated by generating both snapshots and deltas, in which case snapshots may provide an overall replica of the customer data across multiple partitions, while deltas may enable incremental tracking of changes to the customer data for a single partition following a preceding snapshot. Once a delta is generated, it may be transferred to the second data center along with the customer data it describes, and may be used by the storage system in the second data center to restore the customer data in the event of a failure in the first data center. While a delta is being generated for one partition, the other partitions remain available for I / O operations, which may allow the storage system in the first data center to maintain the desired performance level. If a failure occurs in the first data center, the second data center may take over the role of the first data center (referred to as “failover”), and / or customer data may be returned to the first data center to resume operations (referred to as “failback”).
[0019] Figure 1 shows an exemplary system 100 for asynchronous inter-region block volume replication according to one or more embodiments. As described above, system 100 may facilitate data redundancy by replicating customer data from a source system to a destination system in order to provide customer data during failure recovery. In some embodiments, a first data center 120 (e.g., the source system) may store data in a block volume system 122. In some embodiments, the first data center 120 may store data in an object storage system instead of a block volume system. The data may be generated and / or provided by users of the first data center 120. The first data center 120 may be located in a first geographical region (e.g., region A) that is close to the location of users and / or may be responsive to one or more operational criteria such as performance metrics (e.g., latency, IOPS, throughput, cost, stability, etc.). In some cases, the normal operation of the first data center 120 may be interrupted (e.g., power outage, data corruption, distributed denial of service attack (DDOS), natural disaster, etc.), which may create a risk of data loss. To potentially mitigate the risk of data loss, the data is stored in the second data center 130 (e.g., destination data). It can be replicated in the system. In some embodiments, the second data center 130 may be located in a different geographical area from the location of the first data center 120, thereby reducing the risk of data loss caused by natural disasters or infrastructure failures. Similarly, storing data in a second data center 130 separate from the first data center 120 may reduce potential vulnerabilities to malicious actions (e.g., DDOS, data corruption, etc.) that could interrupt I / O operations and / or result in data loss by targeting the first data center. As will be described in more detail below, the second data center 130 may store data in the datastore 132 as standby volumes, object storage, and / or other storage formats (e.g., based on customer configuration and / or preference).
[0020] In some embodiments, data stored in a first data center 120 may be replicated in a second data center 130 via an asynchronous replication system 140. As described above, asynchronous replication may offer technical advantages over periodic backup replication in that it can reduce interruptions to normal I / O processes during operation recovery (e.g., RTO) and reduce the degree of data loss caused by outages (e.g., RPO) in the first data center 120. In some embodiments, the asynchronous replication 140 may include incrementally generating and transferring replicated data from the first data center 120 to the second data center 130, rather than as a coherent backup image generated in a specific time series that can be periodically replaced with a new backup image. Alternatively, as will be described in more detail below with reference to Figure 2, the asynchronous replication 140 may dynamically update the data stored in the second data center 130, for example, by generating data and transferring a record of the data stored in the first data center 120 and / or changes to the data. In some embodiments, customers and / or users of the block volume system 122 may configure asynchronous replication 140 as a subsequent option available when the configuration volume is created and / or after the block volume system 122 is already operating as a distributed data storage system. In some embodiments, configuring asynchronous replication 140 may include specifying a destination region (e.g., a second geographic region, i.e., "region B"). In some embodiments, for example, if the replicated data should be stored in a standby block volume system, configuring asynchronous replication 140 may include specifying a destination AD within the destination region.
[0021] In some embodiments, the asynchronous replication 140 may include a snapshot generation 142 subsystem that can generate snapshots of data stored in the block volume system 122. In some cases, the snapshot may describe the instantaneous state of user data stored in the block volume system 122 at a particular logical time, as will be described in more detail below. In contrast to the backup replication approach, the asynchronous replication 140 may also include a delta generation 144 subsystem that can identify changes made to the data in the block volume system 122 after a snapshot has been taken. These changes can then be represented as deltas that can be used to update the snapshot, rather than replacing the entire snapshot. The snapshot generation 142 may generate new snapshots periodically, for example, on an order of minutes (e.g., in ranges of 1-10 minutes, 10-20 minutes, 20-30 minutes, etc.) and / or on an order of hours (e.g., 1-10 hours).
[0022] In some embodiments, snapshot generation 142 may proceed by a two-phase commit protocol involving one or more partitions of the block volume system 122. Two-phase commit is a commitment that I / O operations on a partition have been interrupted. This may include generating a partition image after the snapshot generation 142 receives the data from the block volume system. After the partition image is complete, the partition can be released and I / O operations can be resumed. In some embodiments, a two-phase commit may be applied to the block volume system 122 as a whole, thereby blocking all read and write operations on the entire block volume system while the snapshot is being generated. In some embodiments, the snapshot may be generated on the order of milliseconds (e.g., 1 to 15 milliseconds, 5 to 10 milliseconds, etc.).
[0023] In some embodiments, snapshot generation 142 includes generating images of data stored in the block volume system 122 on a partition-by-partition basis. If the block volume system 122 stores data in one or more partitions, snapshot generation 142 may schedule the generation of partition images so as not to interfere with the I / O operations of the partitions. Snapshot generation 142 may generate and compile images of all partitions included in the block volume system, and may, in some cases, avoid interrupting the normal I / O operations of partitions other than those being imaged.
[0024] A snapshot can record an image of data at a specific logical time, where the logical time refers to the implementation of the snapshot generation 142. For example, a first snapshot generated by the snapshot generation 142 may be identified by a first logical time and may be generated over a period of time during which one or more partitions of the block volume system 122 are processed by the snapshot generation 142. Similarly, the snapshot generation 142 may generate a second snapshot of the block volume system 122 after a period (e.g., 1 to 10 minutes) has elapsed, which may be identified by a second logical time, corresponding to a second iteration of the snapshot generation 142 subsystem.
[0025] In some embodiments, the delta generator 144 can facilitate asynchronous replication 140 by enabling it to transfer new or modified data to the second data center 130 without transmitting any unchanged data. In some cases, the delta can describe changes to data stored in partitions of the block volume system 122 between a first snapshot and a second snapshot. For example, the delta generator 144 can compare the second snapshot with the first snapshot and determine which blocks have been added, removed, modified, etc. In this way, the delta can also describe data removed between two consecutive snapshots.
[0026] In some cases, the delta may be a logical structure containing unique identifiers and even a list of memory pointers. The memory pointers may describe memory locations in the first data center 120 and / or the second data center 130 for data that has changed between two consecutive snapshots. Implementing asynchronous replication 140, as will be explained in more detail below with reference to Figures 2-6, may involve transferring the delta for a partition to the second data center 130, rather than transferring the entire snapshot.
[0027] In some embodiments, the datastore 132 in the second data center 130 may store the replicated data as object storage. Therefore, the asynchronous replication 140 may include a data conversion subsystem 146 for converting data from the first data center 120 into chunk objects. In some embodiments, data from the block volume system 122 may be incorporated into 4MB chunk objects. In some embodiments, the chunk objects are used for system backup operations. The same format used may be implemented, which may allow the asynchronous replication 140 to be integrated into existing distributed storage systems that employ backup and / or tiered upload technologies.
[0028] In some embodiments, the asynchronous replication 140 may generate a checkpoint after deltas corresponding to partitions of the block volume system 122 have been generated. In some embodiments, the checkpoint generation subsystem 148 may generate a checkpoint by aggregating deltas and applying the aggregated changes to a pre-generated checkpoint, as will be described in more detail below with reference to Figure 4.
[0029] If the first data center suffers a failure (e.g., a catastrophic failure) and, for example, the system is desired to be restored by users and / or customers of the block volume system 122, a failover / failback request 150 may be provided to the asynchronous replica 140. The failover / failback request 150 may also be in the form of a restore request causing the asynchronous replica 140 to generate a restore volume using the replicated data stored in the second geographical area. In some embodiments, the asynchronous replica 140 may implement one or more restore operations in response to receiving the failover / failback request 150. In some embodiments, the failover / failback request 150 may include requests from external services of the distributed storage system and / or requests from users of the distributed storage system. In some cases, the failover request may indicate that the datastore 132 should be configured to take on the role of the block volume system 122, as is the case when the datastore 132 is a replica of the block volume system 122 (e.g., a standby volume). For example, the datastore may handle I / O operations involving user data. In some embodiments, the failover / failback request 150 may indicate that the first data center 120 should be configured to receive the replicated data as described above and to resume operations that were in progress before the failure caused by the failover request (e.g., failback restoration).
[0030] Figure 2 shows an exemplary technique 200 for asynchronous block volume replication according to one or more embodiments. As described above, asynchronous replication can enable user data to be transferred from the source system to the destination system in incremental increments, rather than as a single backup image. In some embodiments, the block volume system 122 may store user data that can be received from the user and provided to the user by input-output (I / O) operations 210. The data thus received may be stored in one or more configuration partitions 220 of the block volume system 122. An asynchronous replication system (e.g., the asynchronous replication 140 in Figure 1) can replicate and transfer user data from the block volume system 122 to the datastore 132 by one or more processes, as described below.
[0031] In some embodiments, user data may be duplicated by creating a first snapshot 230 (e.g., operation 250). As will be described in more detail with reference to Figure 1, the snapshot may be a record of data stored in the block volume system 122. The snapshot generation of the first snapshot 130 (e.g., snapshot generation 142 in Figure 1) may occur in a first logical time 260. As will be described in more detail with reference to Figure 1, the logical time may describe iterations of a snapshot generation operation (e.g., operation 250) that implement a two-phase commit protocol in each of the constituent partitions 220 of the block volume system 122, for example, if snapshot generation proceeds on a partition basis. In such a case, the first logical time 260 may correspond to the period during which the snapshot is being created.
[0032] In some cases, the first snapshot 230 is a first implementation example of asynchronous replication of user data stored in the block volume system 122. In such a case, the entire dataset represented by the first snapshot 230 can be transferred to the datastore 132. In this way, the first snapshot 130 can be similar to a backup in that at least the entire record of user data and metadata (e.g., manifests as described below) can be transferred from the source system to the destination system (e.g., operation 252).
[0033] In some embodiments, the first snapshot may be a disk image of the block volume system 122 and / or the configuration partition 220. In some embodiments, the first snapshot 230 may include multiple block images 232, so that the data described by the first snapshot 230 can be transferred asynchronously to the datastore 132 by transferring the block images 232 individually and / or in groups, at least partially. As will be described in more detail with reference to Figure 3, the block images 232 may be converted into chunk objects to be stored in the datastore 132, for example, if the datastore 132 is an object storage system.
[0034] In some embodiments, asynchronous replication may include creating a second snapshot 240 at a second logical time 262 (e.g., operation 254). In some embodiments, the second logical time 262 corresponds to a period following the first logical time 260, as described above in more detail with reference to Figure 1. For example, the second snapshot 240 may be created by the two-phase commit protocol described above a few minutes (e.g., 1 to 15 minutes, 5 to 10 minutes, etc.) after the creation of the first snapshot 230. In some embodiments, the second snapshot 240 may include one or more changes to the data stored in the block volume system 122 with respect to the first snapshot 230.
[0035] Technique 200 may include generating a delta 242 corresponding to the constituent partitions of the block volume system 122 (e.g., operation 256) rather than directly transferring the second snapshot 240. As described above in more detail with reference to Figure 1, delta generation (e.g., delta generation 144 in Figure 1) may include comparing the state of the data reflected in the first snapshot 230 with that of the second snapshot 240 to generate a list of modified blocks (identified by pointers) paired with reference chunk identifiers of the destination system (e.g., datastore 132), which may constitute at least a portion of the data contained in the delta 242 for each partition 220. In some embodiments, the delta may also include a unique delta identifier. In some embodiments, the unique delta identifier may correspond to an identifier of the second snapshot 240. In this way, the delta 242 may reflect changes to the data traceable up to a second logical time 262, for example, as an approach to provide an I / O history.
[0036] Once generated, delta 242 can be transferred to a destination system, for example, datastore 132 (e.g., operation 258). In some embodiments, delta 242 can be transferred along with any new data added to block volume system 122. Data removed from block volume system 122 between first logical time 260 and second logical time 262 can be reflected by metadata contained in delta 242 without transferring data (e.g., chunk objects) from the source system to the destination system.
[0037] Figure 3 shows an exemplary technique 300 for restoring a block volume system by asynchronous replication according to one or more embodiments. The asynchronous replication described above with reference to Figures 1 and 2. As part of the replication system, a destination system, for example, a second data center 130, may receive and maintain a copy of the data stored in the source system (for example, the first data center 120 in Figure 1). In some embodiments, the second data center 130 may be located in a geographic region different from the geographic region of the source system (for example, region B, opposite region A). The second data center 130 may include a data store 132 in which the replicated data can be stored in object storage as one or more data objects 310.
[0038] In some embodiments, the datastore 132 may operate as an object storage system based at least in part on the configuration of the asynchronous replication system, so that the asynchronous replication may be configured to store the replicated data as chunk objects rather than as blocks. That said, in some embodiments, the second data center 130 may include a standby volume, which may allow deltas to be applied directly to the standby volume. In embodiments in which the datastore 132 includes a standby volume, the asynchronous replication may be completed without converting blocks to objects (e.g., data conversion 146).
[0039] In some embodiments, asynchronous replication may include generating a first checkpoint 320 (e.g., operation 350). A checkpoint may differ from a snapshot in that it may include at least a manifest (e.g., metadata) describing the location and identifier of chunk objects stored in the datastore 132, as will be described in more detail with reference to Figure 4. For example, if a first checkpoint 320 corresponds to a first time in which it was generated as part of asynchronous replication, that first checkpoint may include a list of chunk pointers for each of the chunk objects that make up the replicated data.
[0040] As part of asynchronous replication, the second data center 130 receives a delta 330 and corresponding data from the source system (e.g., operation 352). As described above, the delta may include metadata describing changes to the data stored in the source system resulting from I / O operations between the first and second snapshots (I / O operation 210 in Figure 2). The corresponding data may be received as chunk objects (e.g., 4MB chunk objects).
[0041] Delta 330 may be applied to the first checkpoint 320 as part of generating a second checkpoint 340 (e.g., operation 354). Generating the second checkpoint may include identifying one or more chunk objects identified in Delta 330 and applying the modifications indicated by the Delta. For example, one or more transformations to chunk data may be indicated by a first Delta 330-1 of Delta 330, in which case the first Delta 330-1 indicates (e.g., by a chunk pointer) the location in the object storage of the datastore 132 of the referenced data with respect to a given partition. In some embodiments, rather than directly modifying the referenced data, asynchronous replication may include updating the first checkpoint 320 to reflect the changes indicated by the first Delta 330-1 as part of generating the second checkpoint 340, as will be described in more detail below with reference to Figure 4.
[0042] In some embodiments, the second data center 130 (e.g., the destination system) may be subject to a failover / failback request (e.g., failover / failback request 150 in Figure 1) that may be received from a user of the asynchronous replication system (e.g., the asynchronous replication system 140 in Figure 1 by operation 356). A check request may respond to a failure (e.g., a natural disaster) affecting the source system (e.g., the first data center 120 in Figure 1). Failover / failback requests may include parameters that guide how the block volume system should be restored from replicated data and a second checkpoint, as will be explained in more detail below with reference to Figures 7-10.
[0043] Implementing a failover / failback request may include generating a failover / failback volume, also referred to as a restore volume (e.g., operation 358). As described above, the failover system may be hosted in the destination system, while the failback system may be hosted in the source system. In some embodiments, when a failback request is received, a second checkpoint 340 may be used to map the data stored in object 310 to blocks in the source system (e.g., block volume system 122). The process of generating the failback volume will be described in more detail below with reference to Figure 5. Similarly, failover may include generating a block volume system in the second data center 130 to take on the role of the first data center 120 in terms of I / O operations until the first data center 120 recovers from a failure. Generating a failover volume in the second data center 130 may include mapping the data stored in object 310 to blocks, for example by using a second checkpoint 340, as will be described in more detail below with reference to Figure 4. In some embodiments, restoring block volume data may include, but is not limited to, one or more actions: creating a new block volume in the destination region, implementing failover in the destination region, enabling inter-region replication to the new volume in the source region, and performing failover in the source region.
[0044] In some embodiments, as in cases where the second data center 130 maintains a standby volume to store replicated data, the standby volume may already contain all changes from the source volume, except for changes from delta that were not transferred from the source system in the event of a failure. Thus, failover can be immediately available with relatively small amounts of lost data (e.g., RPO on the order of snapshot generation time of 1 to 15 minutes) (e.g., exhibiting a low RTO), as will be explained in more detail below with reference to Figure 5. In contrast, a backup system that implements synchronous replication of the entire backup image may result in an RPO of several hours.
[0045] In some embodiments, asynchronous replication may involve reverse replication rather than generating failover / failback volumes. In such cases, the destination system may be designated as the source system, but the previous source system may be re-designated as the destination system. In this way, I / O operations may be performed in a second data center 130, in which case data is written and read in datastore 132, and snapshot creation and delta generation also take place there. Correspondingly, the replicated data may be transferred to the first data center 120.
[0046] Figure 4 shows an exemplary technique 400 for aggregating block volume replication metadata according to one or more embodiments. Asynchronous replication may involve multiple iterations of one or more configuration operations (e.g., snapshot creation, delta generation, etc.) over the period during which data may be replicated. In such cases, updating the destination system (e.g., the second data center 130 in Figures 1-3) may involve maintaining an updated record of memory locations corresponding to the replicated data and reflecting the changes introduced by each subsequent iteration.
[0047] In some embodiments, asynchronous replication may include generating a second checkpoint 340 by one or more operations on a delta (e.g., delta 330 in Figure 3) and a first checkpoint 320. For example, if multiple deltas are received in the destination system from the source system after the generation of the first checkpoint 320, the accurate record of memory locations may be better represented by the changes indicated by the deltas, so that they apply to the first checkpoint 320 rather than referring only to the first checkpoint.
[0048] In some embodiments, the first checkpoint 320 may include a manifest 420. The manifest 420 may correspond to partitions of the source system (for example, to guide failover / failback operations), so if the source system has multiple partitions, multiple manifests may describe asynchronous replication of data from the source system. The manifest 420 may include a list of chunk pointers 422 (for example, chunk pointers 1 to N, where "N" is an integer pointing to the number of chunk pointers 422 included in the manifest 420). The chunk pointers 422 may describe memory locations associated with the replicated data (for example, stored in the datastore 132 as chunk objects). After generating the first checkpoint 320, the destination system may receive multiple deltas, such as those generated by iterating through snapshot creation and delta generation. Thus, generating the second checkpoint 340 may include aggregating the deltas (for example, operation 450). For example, aggregating deltas may involve referencing the identifier of the first checkpoint 320 to identify the snapshot from which the delta was generated. In this way, the checkpoint reference may allow determining whether the received delta (including the identifier referencing the snapshot) is already reflected in the manifest 420.
[0049] In some embodiments, deltas may be combined with a first checkpoint to provide an updated manifest 440. Combining deltas may involve modifying a first chunk pointer 422-1 in the manifest as indicated by the delta (for example, a pointer pointing to the location of data may change in response to a transformation on the replicated data described by the delta). In this way, by applying all transformations indicated in all received deltas to the manifest, the updated manifest 440 may enable the generation of failover / failback volumes with a potentially lower RPO than when the previous manifest (e.g., manifest 420) was used. Similarly, with each iteration of asynchronous replication, newly generated deltas may be generated and forwarded to the destination system. In this case, these deltas may be aggregated with the most recent manifest (e.g., the updated manifest).
[0050] Figure 5 shows an exemplary technique 500 for generating a failback volume from a standby volume according to one or more embodiments. As described above, a failback request (e.g., a failover / failback request 150 in Figure 1) may be received as part of a failure recovery in a source system, e.g., a first data center 120. The failback volume may be preferred for the same reasons that the first data center 120 is selected to host user data (e.g., performance metrics, stability, etc.) as a failover volume or reverse copy hosted in a destination system.
[0051] In some embodiments, a destination system, for example, a second data center 130, may implement a standby volume 560 to replicate data on the source volume 522. The standby volume 560 may contain all deltas received from the source volume so that the standby volume reflects the current state of the source volume 522. For example The first delta 540 may be generated from the source volume 522 (for example, by operation 512), and the first delta 540 may be applied directly to the standby volume 560 (for example, by operation 514). In some embodiments, applying the delta to the standby volume 560 may include copying data referenced from the source volume (e.g., modified blocks) and replicating data referenced in the standby volume in a manner similar to an input / output operation (e.g., a user write operation). Furthermore, applying the delta may include updating the manifest as described with reference to Figure 4. Figure 4 illustrates the manifest with respect to chunk pointers, but the manifest may similarly describe memory locations in the standby volume where the replicated data may be stored.
[0052] In some embodiments, updating the standby volume 560 may involve applying the delta in a separate operation so that the delta is not partially applied to the standby volume 560. For example, a first snapshot 530 of the standby volume 560 may be generated to provide a fallback position if a failover / failback request is received during an operation to apply the delta (e.g., operation 514). In this way, system recovery may implement a first snapshot 530 that can provide an improved RTO at the expense of an increasing RPO, rather than waiting for the delta to be applied. In some embodiments, the first snapshot 530 may be deleted when the first delta 540 is applied to the standby volume 560. Then, when the second delta 550 (or delta) is generated (e.g., operation 516), a second snapshot 532 may be generated on the standby volume 560 before the second delta is applied (e.g., operation 520) (e.g., operation 518). After the second delta has been fully applied, the second snapshot 550 may also be deleted.
[0053] In some embodiments, the standby volume 560 may be used to generate a failover or failback volume in response to receiving a failover / failback request (e.g., operation 522). In some embodiments, data recovery may include transferring the data back to the source region by enabling inter-region replication to the failover volume. In some embodiments, if the request is a failback request, the standby volume 560 may be cloned and restored in the first data center 120 (e.g., operation 524). Cloning the standby volume 560 may include creating a failback volume 562 in the first data center 120 and restoring the data (e.g., also referred to as "hydrating the clone"). In some embodiments, if the second delta was fully applied when the failback request was implemented, the restored data may include modifications indicated by the second delta 550. Implementing failover recovery using cloning may, in some cases, allow for continued inter-region replication to an existing standby volume. Advantageously, cloned volumes may be available with little or no latency, potentially making their impact on RTO negligible.
[0054] Figure 6 shows an exemplary technique 600 for resizing a standby volume according to one or more embodiments. Normal operation of a source system, such as a source volume 522 in a first data center 120, may include resizing the source volume 522 (for example, to add partitions, to remove partitions, to resize one or more partitions of the source volume 522, etc.). In some embodiments, a delta is generated for the partition, so resizing the source volume 522 may affect the mapping of the source volume delta 522 to the standby volume 560. Thus, the standby volume 560 is... - After implementing one or more approaches to limit the potential error in delta application introduced by resizing volume 522, it can be resized to take into account the resizing of source volume 522.
[0055] In some embodiments, one or more deltas 620 may be generated from the source volume 522 as described in more detail with reference to Figures 1 and 2 above (e.g., operation 610). The deltas 620 may be applied to the standby volume 560 as described above with reference to Figure 5 (e.g., operation 612). In some cases, resizing the source volume 522 may involve receiving a resize request from a user or another system (e.g., operation 614). To potentially limit the impact of resizing on the RPO, the final delta 630 may be generated after receiving the resize request but before resizing the source volume 522 (e.g., operation 616). In some embodiments, generating the final delta 630 may involve creating a new snapshot of the source volume 522, generating a delta based at least partially on the new snapshot, and applying the delta thus generated. In this way, the standby volume 560 may describe the state of the source volume immediately before the execution of the resize request by resizing the source volume 522 (e.g., operation 618). Applying the final delta (e.g., operation 620) may be decoupled from resizing the source volume 522, so that once the final delta 630 is generated, the source volume is resized, and then the final delta can be applied to the standby volume 560. In some embodiments, the standby volume 560 may be resized in a corresponding manner after the application of the final delta (e.g., operation 622).
[0056] Figure 7 shows an exemplary flow 700 for generating a restored volume according to one or more embodiments. The operation of flow 700 can be implemented as hardware circuitry and / or stored as computer-readable instructions on a non-temporary computer-readable medium of a computer system, such as the asynchronous replication system 140 in Figure 1. When implemented, such instructions represent a module containing circuitry or code that can be executed by the processor of the computer system. Execution of such instructions configures the computer system to perform specific operations described herein. Each circuitry or code, in combination with the processor, performs its respective operation. While these operations are shown in a specific order, it should be understood that this order is not required, and one or more operations may be omitted, skipped, and / or rearranged.
[0057] For example, flow 700 includes operation 702 in which the computer system creates a snapshot of a block volume in a first geographical area. As will be described in more detail with reference to Figures 1 and 2, the block volume (e.g., block volume system 122 in Figure 1) may be hosted in a first data center (e.g., first data center 120 in Figure 1) within the first geographical area, which may be determined at least in part on one or more performance metrics associated with I / O operations, stability, cost, etc. Creating a first snapshot (e.g., operation 250 in Figure 2) may include generating one or more block images (e.g., block image 232 in Figure 2) of one or more blocks that make up the block volume. Creating a snapshot may include implementing a two-phase commit protocol, which may interrupt I / O operations (e.g., I / O operation 210 in Figure 2) for a partition of the block volume, during which a block image of that partition is created. After creating a snapshot for a partition or block volume system, the computer system may resume I / O operations on the block volume.
[0058] In one example, flow 700 includes operation 704 in which the computer system sends snapshot data corresponding to a first snapshot to a second storage system in a second geographical area. In some embodiments, the snapshot data may include a block image, as in the case where the first snapshot is a first snapshot created from a block volume, or when a previously created snapshot is not available. Thus, the block image may be sent to the system (e.g., datastore 132 in Figure 1). In some embodiments, the block image may be converted into a chunk object and sent to the second storage system as snapshot data (e.g., operation 252 in Figure 2). In some cases, the chunk object is sent with metadata describing the snapshot data, including, for example, a manifest that enumerates chunk pointers that identify the location of the snapshot data in memory.
[0059] In one example, flow 700 includes operation 706 in which the computer system creates a second snapshot of a block volume in a first geographical area. As will be described in more detail with reference to Figure 2, the second snapshot (e.g., the second snapshot 240 in Figure 2) may describe the state of the block volume at a second logical time (e.g., the second logical time 262 in Figure 2) following the first logical time (e.g., the first logical time 260 in Figure 2) in which the first snapshot was created. The second snapshot may include information about modifications to the data stored in the block volume system that have occurred since the creation of the first snapshot (e.g., read and write operations on the data).
[0060] For example, flow 700 includes operation 708 in which the computer system generates multiple deltas. These multiple deltas may describe one or more changes to the data stored in the block volume between the creation of the second snapshot (e.g., second logical time) and the creation of the first snapshot (e.g., first logical time), rather than directly sending the second snapshot to the second storage system. Generating a delta may involve comparing the block image of the partition in the second snapshot with the corresponding image of the partition in the first snapshot, and identifying one or more modifications to the data in the partition. As will be illustrated in more detail with reference to Figures 1 and 2, a delta may include metadata including, but not limited to, a delta identifier corresponding to the snapshot, one or more block identifiers describing the data in the modified block volume, and / or a corresponding number of object identifiers describing the location of the data identified by the block identifiers.
[0061] In one example, flow 700 includes operation 710, in which the computer system transmits data about multiple deltas. Transmitting data may involve creating a copy of the data from the block volume system to a second storage system, as will be explained in more detail with reference to Figure 2. This may involve generating chunk objects and copying the chunk objects to the second storage system (e.g., the destination system), similar to how the replicated data would be stored in object storage. The data may be accompanied by multiple deltas, which may provide metadata describing the correspondence between the chunk objects and blocks in the block volume system on a partition-by-partition basis.
[0062] In one example, flow 700 includes operation 712 in which the computer system generates a checkpoint in a second geographical area. As will be described in more detail with reference to Figures 3 and 4, the checkpoint (e.g., the first checkpoint 320 in Figure 3) may include a manifest (e.g., the manifest 420 in Figure 4) that describes a plurality of chunk pointers. In some embodiments, generating a checkpoint is performed by the checkpoint This may include aggregating the delta with a previously generated checkpoint (e.g., operation 450 in Figure 4) to generate an updated checkpoint (e.g., second checkpoint 340 in Figure 3), which may include an updated manifest (e.g., updated manifest 420 in Figure 4) describing a second set of chunk pointers that reflect the modifications contained in multiple deltas, such as those applied to the updated delta.
[0063] In one example, flow 700 includes operation 714 in which the computer system receives a restore request. The restore request (e.g., failover / failback request 150 in Figure 1) may include a user request to generate a restore volume at a second geographic location, also referred to as a failover volume. In some embodiments, the restore request may include a user request to generate a restore volume at a first geographic location, also referred to as a failback volume. In some embodiments, the restore request may include a request to reverse asynchronous replication, thereby designating a second storage system (e.g., datastore 132 in Figure 1) as the source system and a block volume system (e.g., block volume system 122 in Figure 1) as the destination system for asynchronous replication. In some embodiments, the restore request may be generated automatically (e.g., without user interaction) as part of the asynchronous replication configuration. For example, the asynchronous replication system may be configured to automatically generate a restore request in response to a failure occurring in the source system.
[0064] In one example, flow 700 includes operation 716 in which the computer system generates a restore volume. As illustrated with reference to Figures 3 and 5, generating a restore volume may include generating a block volume by mapping replica data from chunk objects stored in object storage to blocks in a block volume, for example, by referencing an updated manifest (e.g., updated manifest 440 in Figure 4). In this way, the restore volume can be generated with a lower RPO than can be achieved by the backup system. Similarly, the RTO of a system restore operation, which may reflect the time between receiving a restore request and resuming normal I / O operations, may depend on whether the restore request is a failover request or a failback request, which will be explained in more detail below with reference to Figures 8 and 9.
[0065] Figure 8 shows an exemplary flow 800 for generating a failover volume according to one or more embodiments. The operation of flow 800 can be implemented as hardware circuitry and / or stored as computer-readable instructions on a non-temporary computer-readable medium of a computer system, such as the asynchronous replication system 140 in Figure 1. When implemented, such instructions represent a module containing circuitry or code that can be executed by the processor of the computer system. Execution of such instructions configures the computer system to perform specific operations described herein. Each circuitry or code, in combination with the processor, performs its respective operation. While these operations are shown in a specific order, it should be understood that this order is not required, and one or more operations may be omitted, skipped, and / or rearranged.
[0066] In one example, flow 800 includes operation 802 in which a computer system receives a restore request, which is a failover request. As described above in more detail with reference to Figures 1 to 6, a failover request may be received by an asynchronous replica (e.g., an asynchronous replica 140 in Figure 1) after a block volume system (e.g., block volume system 122 in Figure 1), which may form part of a distributed storage system (e.g., cloud storage), has been affected by a disaster or other failure. In some embodiments, the failover request may include a request to generate a failover volume that regenerates the block volume system in a second geographical area that may not be affected by the failure affecting the block volume system.
[0067] For example, flow 800 includes operation 804 in which the computer system generates a failover volume in a second geographical area. The failover volume may, in some cases, allow the resumption of I / O operations (e.g., I / O operation 210 in Figure 2) before a failure affecting the block volume system is resolved. In such cases, the RTO may be improved by generating the failover volume compared to a failback volume. As described above with reference to Figure 7, generating a failover volume using replica data stored in object storage may include generating the failover volume by referencing an updated manifest (e.g., updated manifest 440 in Figure 4) that describes the location of the replica data in memory, which is updated with the most recently aggregated delta generated during asynchronous replication.
[0068] In one example, flow 800 includes operation 806 in which the computer system hydrates a failover volume. As described in more detail above, hydrating a failover volume may include restoring the data described in the updated manifest from object storage (similar to the chunk objects described above with reference to Figure 1, for example) into blocks that can be handled by read / write operations of the block volume system (e.g., I / O operation 210 in Figure 2). In some embodiments, partitions of the block volume system may be stored within the failover volume.
[0069] In one example, flow 800 includes operation 808, in which the computer system initiates I / O operations. After "hydration," in which replicated data from the block volume system is restored and mapped to the failover volume, the failover volume may initiate I / O operations and take on the role of the block volume system while the failure can be resolved.
[0070] Figure 9 shows an exemplary flow 900 for generating a failback volume according to one or more embodiments. The operation of flow 900 can be implemented as hardware circuitry and / or stored as computer-readable instructions on a non-temporary computer-readable medium of a computer system, such as the asynchronous replication system 140 in Figure 1. When implemented, such instructions represent a module containing circuitry or code that can be executed by the processor of the computer system. Execution of such instructions configures the computer system to perform specific operations described herein. Each circuitry or code, in combination with the processor, performs its respective operation. While these operations are shown in a specific order, it should be understood that this order is not required, and one or more operations may be omitted, skipped, and / or rearranged.
[0071] In one example, flow 900 includes operation 902 in which the computer system receives a restore request, which is a failback request. In some embodiments, the failure causing the restore request may not be for an indefinite period. For example, the failure may be solvable within a predictable length of time, as is the case when the failure is caused by a temporary power or network interruption that can be reliably overcome. Thus, the failback request may represent a preferred alternative (for example, in terms of distributed storage performance) for generating a failover volume and a potentially limited difference in terms of RTO for the failover volume.
[0072] In one example, flow 900 includes operation 904 in which the computer system generates a standby volume in a second geographical area. In some embodiments, the standby volume is inbound and outbound in the second geographical area (e.g., the second data center 130 in Figure 1). This may include a block volume system similar to a failover volume, except that it is not designed to perform force operations. For example, a standby volume may not be visible to users of the distributed storage system and / or may be configured to be unsuitable for I / O operations. In some embodiments, a standby volume may replicate the structure of a block volume system (e.g., in terms of partitioning, size, etc.) and the mapping of blocks to replicated data stored as chunk objects.
[0073] For example, flow 900 includes operation 906 in which the computer system clones a standby volume in a first geographical area. Cloning a standby volume (e.g., operation 524 in Figure 5), as described above in more detail with reference to Figure 5, may include restoring the block volume system to the first geographical area by generating a failback volume using at least partially the structure of the standby volume. For example, the standby volume may describe the structure and mapping of the replicated data (e.g., by reference to an updated manifest) so that when the replicated data is copied back from a second geographical area to the first geographical area, the failback volume is restored with a new mapping.
[0074] In one example, flow 900 includes operation 908 in which the computer system hydrates a failback volume in a first geographical area. In some embodiments, hydrating a failback volume may include restoring replicated data to the first geographical area (e.g., the first data center 120 in Figure 1) so that the failback volume can resume I / O operations. Restoring replicated data may include copying data from a second geographical area to the first geographical area. Restoring replicated data may also include remapping data from mappings from asynchronous replication iterations (e.g., those generated by delta application) to new mappings corresponding to new data locations in the failback block volume system.
[0075] In one example, flow 900 includes operation 910, in which the computer system initiates I / O operations in a first geographical area. The I / O operations describe accessing data using a failback volume, storing the data, and modifying it. As part of system recovery, a low RTO and low RPO may improve system performance. For at least this reason, in some embodiments, the I / O operations may be initiated while the data recovery of operation 908 is in progress. For example, after a read / write operation has been initiated, the replicated data may be fully copied back from the second geographical area to the first geographical area. In some embodiments, the I / O operations on the failback volume may be initiated after the copyover and remapping are fully completed.
[0076] Figure 10 shows an exemplary flow 1000 for generating a failback volume according to one or more embodiments. The operation of flow 1000 can be implemented as a hardware circuit and / or stored as a computer-readable instruction on a non-temporary computer-readable medium of a computer system, such as the asynchronous replication system 140 in Figure 1. When implemented, the instruction represents a module containing a circuit or code that can be executed by the processor of the computer system. The execution of such an instruction configures the computer system to perform the specific operations described herein. Each circuit or code, in combination with the processor, performs its respective operation. While these operations are shown in a specific order, it should be understood that the specific order is not required, and one or more operations may be omitted, skipped, and / or rearranged.
[0077] As will be explained in more detail with reference to Figure 6, in some embodiments, asynchronous replication may include a standby volume in the object storage of the replicated data at a second geographic location, additionally or alternatively. Rather than converting the data described by delta into chunk objects (as described, for example, with reference to data transformation 146 in Figure 1), the standby volume may include a direct transfer of block data from the block volume system to the standby volume. In some embodiments, the standby volume may be invisible to users of the block volume system (e.g., unavailable for I / O operations), but may maintain updated block data to potentially minimize RTO and RPO in the event of an outage affecting the block volume system. In some embodiments, the standby volume may interact with object storage in the destination region, for example, by caching data from delta, while implementing checkpoint operations on the object storage data.
[0078] In one example, flow 1000 includes operation 1002 in which the computer system generates a delta from a block volume system in a first geographical region. The delta may represent modifications to the data stored in the block volume system between a first snapshot in a first logical time and a second snapshot in a second logical time, as will be explained in more detail with reference to Figure 2.
[0079] In one example, flow 1000 includes operation 1004 in which the computer system creates a snapshot of the standby volume system in a second geographical area. In some embodiments, the creation of a snapshot of the standby volume may occur before updating the standby volume. As will be described in more detail with reference to Figure 1, creating a snapshot may involve creating multiple block images of the blocks that make up the standby volume on a partition basis. When combined, these images can describe the state of the data in the standby volume at the logical time corresponding to the creation of the snapshot (for example, as described above in relation to a two-phase commit protocol). The snapshot may provide a fallback position that the asynchronous replication system should use when restoring the block volume system. For example, if a restore request is received while the standby volume is being updated with a new delta, the asynchronous replication may use the snapshot to generate a restored volume.
[0080] In one example, flow 1000 includes operation 1006 in which the computer system applies a delta to the standby volume system. In contrast to the approach described in relation to data replication using object storage, the standby volume can be updated by directly applying a delta. For example, as described in more detail with reference to Figure 2, the delta may include metadata (e.g., block identifiers) that describes the location in memory of the block volume system for the data being replicated as a result of the modification. Thus, modifications may be made to the standby volume by applying the indicated modifications from the delta.
[0081] In one example, flow 1000 includes operation 1008 in which the computer system receives a failback request. As described above, a failback request may be received after an operational failure or other interruption in the block volume system, so that a standby system may be used to generate a failback system after the interruption is resolved.
[0082] For example, flow 1000 includes operation 1010 in which the computer system clones a standby volume. Cloning a standby volume (e.g., operation 524 in Figure 5), as will be explained in more detail with reference to Figure 5, involves (e.g., size This may include creating a failback volume (e.g., failback volume 562 in Figure 5) in a first geographical area (e.g., first data center 120 in Figure 1) that can replicate the structure of the standby volume (in terms of partitions, etc.).
[0083] In one example, flow 1000 includes operation 1012 in which the computer system restores replicated data in a first geographical area to a failback volume. As described above, restoring replicated data may include copying block data from a second geographical area to the first geographical area (also referred to as "hydrating") and mapping the block data to a failback volume. In some embodiments, input / output operations may be initiated on the failback volume once the replicated data is restored and / or while the replicated data is being restored.
[0084] Figure 11 shows an exemplary flow 1100 for generating a failback volume according to one or more embodiments. The operation of flow 1100 can be implemented as hardware circuitry and / or stored as computer-readable instructions on a non-temporary computer-readable medium of a computer system, such as the asynchronous replication system 140 in Figure 1. When implemented, such instructions represent a module containing circuitry or code that can be executed by the processor of the computer system. Execution of such instructions configures the computer system to perform specific operations described herein. Each circuitry or code, in combination with the processor, performs its respective operation. While these operations are shown in a specific order, it should be understood that this order is not required, and one or more operations may be omitted, skipped, and / or rearranged.
[0085] For example, flow 1100 includes operation 1102 in which the computer system receives a resize request in a first geographical region. During the operation of an asynchronous replication system, for example, during snapshot creation iterations or delta generation, the block volume system in the first geographical region may receive a resize request. Since the resize request may involve adding and / or removing one or more partitions, or resizing one or more partitions of the block volume system, the delta generated before the resize may not be compatible with the standby volume after the resize.
[0086] In one example, flow 1100 includes operation 1104 in which the computer system generates a final delta in a first geographical area. In some embodiments, the asynchronous replication system may implement a modified replication protocol in response to receiving a resize request from the block volume system. For example, instead of resizing the block volume system immediately after receiving the resize request, another snapshot may be created to generate one or more final deltas. Such an approach may allow the replicated data to reflect the updated state of the block volume system before resizing, and therefore may further reduce the RPO of the restore volume generated from the replicated data.
[0087] For example, flow 1100 includes operation 1106 in which the computer system applies the last delta to the standby volume in a second geographical area. Applying the last delta may include modifying the standby volume to reflect the modifications to the block volume system described by one or more last deltas generated in operation 1104, as described above.
[0088] In one example, flow 1100 includes operation 1108 in which the computer system resizes a block volume in a first geographical area. After operation 1104, the resize request may be applied to the block volume system.
[0089] For example, flow 1100 includes operation 1110 in which the computer system resizes the standby volume in response to a resize request. Similarly, after operation 1106, the standby volume may also be resized in a manner that responds to a resize request. Since asynchronous replication can generate partition-specific deltas, the standby volume may be configured to replicate the structure of a block volume system. For example, the standby volume may contain the same number of partitions of the same size as the block volume system.
[0090] As mentioned above, Infrastructure as a Service (IaaS) is a specific type of cloud computing. IaaS can be configured to provide virtualized computing resources over a public network (e.g., the internet). In the IaaS model, the cloud computing provider may host infrastructure components (e.g., servers, storage devices, network nodes (e.g., hardware), deployment software, platform virtualization (e.g., hypervisor layer)). In some cases, the IaaS provider may also supply various services (e.g., billing, monitoring, logging, security, load balancing, and clustering) to accompany those infrastructure components. Since these services can be policy-driven, IaaS users may be able to implement policies to drive load balancing to maintain application availability and performance.
[0091] In some cases, IaaS customers may access resources and services over a wide area network (WAN), such as the internet, and install the remaining elements of their application stack using the cloud provider's services. For example, a user might log into an IaaS platform, create virtual machines (VMs), install an operating system (OS) on each VM, deploy middleware such as databases, create storage buckets for workloads and backups, and even install enterprise software on those VMs. The customer can then use the provider's services to perform various functions, including balancing network traffic, troubleshooting application issues, monitoring performance, and managing disaster recovery.
[0092] In most cases, the cloud computing model will require the participation of a cloud provider. This cloud provider may be, but does not have to be, a third-party service specializing in providing IaaS (e.g., giving away, renting, or selling). Entities may also choose to deploy a private cloud and become their own provider of infrastructure services.
[0093] In some cases, IaaS deployment is the process of placing a new application or a new version of an application onto a prepared application server, and may also include the process of preparing the server (e.g., installing libraries, daemons, etc.). This is often managed by the cloud provider under the hypervisor layer (e.g., servers, storage, network hardware, and virtualization). Therefore, the customer may be responsible for handling (OS), middleware, and / or application deployment (e.g., on self-service virtual machines that can be started on demand).
[0094] In some examples, IaaS provisioning can refer to acquiring computers or virtual hosts for use and even installing the necessary libraries or services on them. In most cases, deployment includes provisioning. First, provisioning may need to be performed first.
[0095] In some cases, IaaS provisioning presents two distinct challenges. First, there's the initial challenge of provisioning an early set of infrastructure before anything is operational. Second, once everything is provisioned, there's the challenge of evolving the existing infrastructure (e.g., adding new services, modifying services, removing services, etc.). In some cases, these two challenges can be addressed by allowing the infrastructure configuration to be defined declaratively. In other words, the infrastructure (e.g., what components are needed and how they interact) can be defined by one or more configuration files. Thus, the overall topology of the infrastructure (e.g., which resources depend on which and how each of them works together) can be described declaratively. In some examples, once the topology is defined, it can generate workflows to create and / or manage the various components described in the configuration files.
[0096] In some examples, infrastructure can have many interconnected elements. For example, there may be one or more virtual private clouds (VPCs) (e.g., potentially on-demand pools of configurable and / or shared computing resources) that are also known as the core network. In some examples, there may also be one or more security group rules and one or more virtual machines (VMs) provisioned to define how network security can be set up. Other infrastructure elements such as load balancers and databases may also be provisioned. Infrastructure can evolve incrementally as more and more infrastructure elements are desired and / or added.
[0097] In some cases, sequential deployment techniques may be employed to enable the deployment of infrastructure code across various virtual computing environments. In addition, the techniques described may enable infrastructure management within these environments. In some cases, a service team may write code that is to be deployed to one or more, but often many, different production environments (e.g., across various different geographical locations, sometimes spanning the entire world). However, in some cases, the infrastructure to which the code can be deployed must be set up first. In some cases, provisioning may be done manually, resources may be provisioned using provisioning tools, and / or the code may be deployed using deployment tools once the infrastructure is provisioned.
[0098] Figure 12 is a block diagram 1200 showing an exemplary pattern of an IaaS architecture according to at least one embodiment. The service operator 1202 may be communicatively coupled to a secure host tenancy 1204 which may include a virtual cloud network (VCN) 1206 and a secure host subnet 1208. In some examples, the service operator 1202 may use one or more client computing devices, which may be handheld portable devices (e.g., iPhone®, cellular phone, iPad®, computing tablet, personal digital assistant (PDA)) or wearable devices (e.g., Google® Glass® head-mounted display), running software such as Microsoft Windows Mobile® and / or various mobile operating systems such as iOS, Windows Phone, Android, BlackBerry 8, Palm OS, and may use the Internet, email, short message service (SMS), Blackberry®, or other enabled communication protocols. Alternatively, the client computing devices may be For example, various versions of Microsoft Windows®, Apple Macintosh®, and / or Linux® operating systems can run on this system. This may include a single computer and / or a general-purpose personal computer, including a laptop computer. Client computing devices include, but are not limited to, various GNU / Linux operating systems, such as Google® Chrome OS. This could be a workstation computer running any of the various commercially available UNIX® or UNIX-like operating systems. Alternatively, or in addition, the client computing device may be any other electronic device, such as a thin client computer, an internet-enabled gaming system (e.g., a Microsoft Xbox game console with or without a Kinect® gesture input device), and / or a personal messaging device, that can communicate over a network that can access the VCN1206 and / or the Internet.
[0099] VCN1206 may include a local peering gateway (LPG) 1210 that can be communicatively coupled to Secure Shell (SSH) VCN1212 via LPG1210 included in SSH VCN1212. SSH VCN1212 may include an SSH subnet 1214, and SSH VCN1212 may be communicatively coupled to control plane VCN1216 via LPG1210 included in control plane VCN1216. Furthermore, SSH VCN1212 may be communicatively coupled to data plane VCN1218 via LPG1210. Control plane VCN1216 and data plane VCN1218 may be included in a service tenancy 1219 that may be owned and / or operated by an IaaS provider.
[0100] The control plane VCN 1216 may include a control plane demilitary zone (DMZ) layer 1220 that functions as a peripheral network (e.g., a portion of the corporate network between the corporate intranet and the external network). DMZ-based servers have limited responsibility and may be more susceptible to security breaches. In addition, the DMZ layer 1220 may include a control plane application layer 1224 that may include one or more load balancer (LB) subnets 1222 and application subnets 1226, and a control plane data layer 1228 that may include database (DB) subnets 1230 (e.g., a frontend DB subnet and / or a backend DB subnet). The LB subnet 1222 included in the control plane DMZ layer 1220 can be communicatively coupled to the application subnet 1226 included in the control plane application layer 1224 and to the internet gateway 1234 which may be included in the control plane VCN 1216. The application subnet 1226 can be communicatively coupled to the DB subnet 1230 included in the control plane data layer 1228, to the service gateway 1236 and to the network address translation (NAT) gateway 1238. The control plane VCN 1216 may include the service gateway 1236 and the NAT gateway 1238.
[0101] The control plane VCN 1216 may include a data plane mirror application layer 1240 which may include an application subnet 1226. The application subnet 1226 included in the data plane mirror application layer 1240 may include a virtual network interface controller (VNIC) 1242 which may run a compute instance 1244. The compute instance 1244 can communicately connect the application subnet 1226 of the data plane mirror application layer 1240 to the application subnet 1226 which may be included in the data plane application layer 1246.
[0102] The data plane VCN1218 may include the data plane application layer 1246, the data plane DMZ layer 1248, and the data plane data layer 1250. The data plane DMZ layer 1248 includes the application subnet 1226 of the data plane application layer 1246. The data plane data layer 1250 may also include an LB subnet 1222 that can be communicatively coupled to the internet gateway 1234 of the data plane VCN 1218. The application subnet 1226 may also be communicatively coupled to the service gateway 1236 and the NAT gateway 1238 of the data plane VCN 1218. The data plane data layer 1250 may also include a DB subnet 1230 that can be communicatively coupled to the application subnet 1226 of the data plane application layer 1246.
[0103] The Internet gateway 1234 of the control plane VCN1216 and data plane VCN1218 can be communicatively coupled to a metadata management service 1252 which can be communicatively coupled to the public internet 1254. The public internet 1254 can be communicatively coupled to the NAT gateway 1238 of the control plane VCN1216 and data plane VCN1218. The service gateway 1236 of the control plane VCN1216 and data plane VCN1218 can be communicatively coupled to a cloud service 1256.
[0104] In some examples, a service gateway 1236 of the control plane VCN1216 or data plan VCN1218 may make application programming interface (API) calls to a cloud service 1256 without traversing the public internet 1254. API calls from the service gateway 1236 to the cloud service 1256 can be one-way, where the service gateway 1236 makes an API call to the cloud service 1256, and the cloud service 1256 sends the requested data to the service gateway 1236. However, the cloud service 1256 does not have to initiate an API call to the service gateway 1236.
[0105] In some examples, a secure host tenancy 1204 may be directly connected to a service tenancy 1219, which may otherwise be isolated. A secure host subnet 1208 may communicate with an SSH subnet 1214 via an LPG 1210, which may otherwise enable bidirectional communication through an isolated system. Connecting the secure host subnet 1208 to the SSH subnet 1214 may give the secure host subnet 1208 access to other entities within the service tenancy 1219.
[0106] The control plane VCN1216 may allow users of service tenancy 1219 to set up or provision desired resources. Desired resources provisioned within the control plane VCN1216 may be deployed or used within the data plane VCN1218. In some examples, the control plane VCN1216 may be isolated from the data plane VCN1218, and the data plane mirror application layer 1240 of the control plane VCN1216 may communicate with the data plane application layer 1246 of the data plane VCN1218 via a VNIC 1242 which may be included in the data plane mirror application layer 1240 and the data plane application layer 1246.
[0107] In some examples, a system user or customer may perform requests, such as create, read, update, or delete (CRUD) operations, via the public internet 1254, which can communicate requests to the metadata management service 1252. The metadata management service 1252 may communicate requests to the control plane VCN 1216 via the internet gateway 1234. This request may be received by the LB subnet 1222, which is included in the control plane DMZ layer 1220. The LB subnet 1222 may determine that the request is valid, and in response to this determination, the LB subnet 1222 may send the request to the application subnet 1226, which is included in the control plane application layer 1224. If the request is validated and requires a call to the public internet 1254, the call to the public internet 1254 is made public It may be sent to a NAT gateway 1238 that can make calls to the public internet 1254. Memory that may be desired to be stored by request may be stored in DB subnet 1230.
[0108] In some examples, the data plane mirror application layer 1240 can facilitate direct communication between the control plane VCN 1216 and the data plane VCN 1218. For example, it may be desirable that configuration changes, updates, or other appropriate modifications be applied to resources contained in the data plane VCN 1218. Through VNIC 1242, the control plane VCN 1216 can communicate directly with the resources contained in the data plane VCN 1218, thereby enabling it to perform changes, updates, or other preferred modifications to the resources contained in the data plane VCN 1218.
[0109] In some embodiments, the control plane VCN1216 and data plane VCN1218 may be included in the service tenancy 1219. In this case, the system user or customer does not have to own or operate either the control plane VCN1216 or the data plane VCN1218. Instead, the IaaS provider may own or operate the control plane VCN1216 and the data plane VCN1218, and these may both be included in the service tenancy 1219. This embodiment may enable network isolation that can prevent a user or customer from interacting with resources of other users or other customers. This embodiment may also enable the system user or customer to store databases privately without having to rely on the public internet 1254 for storage, which may not have the desired level of security.
[0110] In another embodiment, the LB subnet 1222 included in the control plane VCN 1216 may be configured to receive signals from the service gateway 1236. In this embodiment, the control plane VCN 1216 and the data plane VCN 1218 may be configured to be invoked by the IaaS provider's customer without calling the public internet 1254. The IaaS provider's customer may desire this embodiment because the database used by the customer may be stored on a service tenancy 1219 that can be controlled by the IaaS provider and isolated from the public internet 1254.
[0111] Figure 13 is a block diagram 1300 showing another exemplary pattern of an IaaS architecture according to at least one embodiment. A service operator 1302 (e.g., service operator 1202 in Figure 12) may be communicatively coupled to a secure host tenancy 1304 (e.g., secure host tenancy 1204 in Figure 12), which may include a virtual cloud network (VCN) 1306 (e.g., VCN1206 in Figure 12) and a secure host subnet 1308 (e.g., secure host subnet 1208 in Figure 12). VCN 1306 may include a local peering gateway (LPG) 1310 (e.g., LPG1210 in Figure 12), which may be communicatively coupled to a secure shell (SSH) VCN 1312 (e.g., SSH VCN1212 in Figure 12) via an LPG 1210 contained in an SSH VCN 1312. SSH VCN1312 may include SSH subnet 1314 (e.g., SSH subnet 1214 in Figure 12), and SSH VCN1312 may be communicably coupled to control plane VCN1316 (e.g., control plane VCN1216 in Figure 12) via LPG1310 included in control plane VCN1316. Control plane VCN1316 may be included in service tenancy 1319 (e.g., service tenancy 1219 in Figure 12), and data plane VCN1318 (e.g., data plane VCN1218 in Figure 12) may be included in customer tenancy 1321, which may be owned or operated by a user or customer of the system.
[0112] The control plane VCN1316 may include a control plane DMZ layer 1320 (e.g., control plane DMZ layer 1220 in Figure 12) which may include an LB subnet 1322 (e.g., LB subnet 1222 in Figure 12), a control plane application layer 1324 (e.g., control plane application layer 1224 in Figure 12) which may include an application subnet 1326 (e.g., application subnet 1226 in Figure 12), and a control plane data layer 1328 (e.g., control plane data layer 1228 in Figure 12) which may include a database (DB) subnet 1330 (e.g., similar to DB subnet 1230 in Figure 12). The LB subnet 1322 included in the control plane DMZ layer 1320 can be communicatively coupled to the application subnet 1326 included in the control plane application layer 1324 and to an internet gateway 1334 (e.g., internet gateway 1234 in Figure 12) which may be included in the control plane VCN 1316. The application subnet 1326 can also be communicatively coupled to the DB subnet 1330 included in the control plane data layer 1328, to a service gateway 1336 (e.g., service gateway in Figure 12) and to a network address translation (NAT) gateway 1338 (e.g., NAT gateway 1238 in Figure 12). The control plane VCN 1316 may include the service gateway 1336 and the NAT gateway 1338.
[0113] The control plane VCN 1316 may include a data plane mirror application layer 1340 (e.g., the data plane mirror application layer 1240 in Figure 12) which may include an application subnet 1326. The application subnet 1326 included in the data plane mirror application layer 1340 may include a virtual network interface controller (VNIC) 1342 (e.g., VNIC 1242) which may run a compute instance 1344 (e.g., similar to compute instance 1244 in Figure 12). The compute instance 1344 may facilitate communication between the application subnet 1326 of the data plane mirror application layer 1340 and the application subnet 1326 included in the data plane application layer 1346 (e.g., the data plane application layer 1246 in Figure 12) via the VNIC 1342 included in the data plane mirror application layer 1340 and the VNIC 1342 included in the data plane application layer 1346.
[0114] The Internet gateway 1334 included in the control plane VCN 1316 can be communicably coupled to a metadata management service 1352 (e.g., metadata management service 1252 in Figure 12), which can be communicably coupled to the public internet 1354 (e.g., public internet 1254 in Figure 12). The public internet 1354 can be communicably coupled to a NAT gateway 1338 included in the control plane VCN 1316. The service gateway 1336 included in the control plane VCN 1316 can be communicably coupled to a cloud service 1356 (e.g., cloud service 1256 in Figure 12).
[0115] In some examples, the data plane VCN1318 may be included in customer tenancy 1321. In this case, the IaaS provider may provide a control plane VCN1316 per customer, and the IaaS provider may set up a unique compute instance 1344 included in service tenancy 1319 for each customer. Each compute instance 1344 may enable communication between the control plane VCN1316 included in service tenancy 1319 and the data plane VCN1318 included in customer tenancy 1321. The compute instance 1344 may enable resources provisioned in the control plane VCN1316 included in service tenancy 1319 to be deployed or used in the data plane VCN1318 included in customer tenancy 1321.
[0116] In another example, an IaaS provider's customer may have a database residing in customer tenancy 1321. In this example, the control plane VCN1316 is the application subnet. It may include a dataplane mirror application layer 1340 which may include 1326. The dataplane mirror application layer 1340 may reside within the dataplane VCN 1318, but does not have to reside within the dataplane VCN 1318. That is, the dataplane mirror application layer 1340 may have access to customer tenancy 1321, but does not have to reside within the dataplane VCN 1318, nor does it have to be owned or operated by the IaaS provider's customer. The dataplane mirror application layer 1340 may be configured to make calls to the dataplane VCN 1318, but does not have to be configured to make calls to any entity contained within the control plane VCN 1316. The customer may also want to deploy or use resources in the dataplane VCN 1318 that are provisioned within the control plane VCN 1316, and the dataplane mirror application layer 1340 can facilitate the customer's desired deployment or other use of resources.
[0117] In some embodiments, a customer of the IaaS provider may apply filters to the data plane VCN 1318. In this embodiment, the customer may determine what the data plane VCN 1318 can access, and may also restrict access from the data plane VCN 1318 to the public internet 1354. The IaaS provider may not be able to filter or control access from the data plane VCN 1318 to any external network or database. By applying filters and controls to the data plane VCN 1318 included in customer tenancy 1321, the customer can help isolate the data plane VCN 1318 from other customers and the public internet 1354.
[0118] In some embodiments, the cloud service 1356 may be invoked by the service gateway 1336 to access services that may not reside on the public internet 1354, the control plane VCN 1316, or the data plane VCN 1318. The connection between the cloud service 1356 and the control plane VCN 1316 or data plane VCN 1318 may not be live or continuous. The cloud service 1356 may reside on another network owned or operated by the IaaS provider. The cloud service 1356 may be configured to receive calls from the service gateway 1336 and not to receive calls from the public internet 1354. Some cloud services 1356 may be isolated from other cloud services 1356, and the control plane VCN 1316 may be isolated from cloud services 1356 that may not be in the same region as the control plane VCN 1316. For example, the control plane VCN 1316 may be located in "Region 1", and the cloud service "Deployment 12" may be located in Region 1 and "Region 2". If a call to deployment 12 is made by a service gateway 1336 included in the control plane VCN 1316 located in region 1, the call may be transmitted to deployment 12 in region 1. In this example, the control plane VCN 1316, or deployment 12 in region 1, does not have to be communicatively coupled to, or communicate with, deployment 12 in region 2.
[0119] Figure 14 is a block diagram 1400 showing another exemplary pattern of an IaaS architecture according to at least one embodiment. A service operator 1402 (e.g., service operator 1202 in Figure 12) may be communicatively coupled to a secure host tenancy 1204 (e.g., secure host tenancy 1204 in Figure 12), which may include a virtual cloud network (VCN) 1406 (e.g., VCN1206 in Figure 12) and a secure host subnet 1408 (e.g., secure host subnet 1208 in Figure 12). VCN1406 may include an LPG1410 (e.g., LPG1210 in Figure 12), which may be communicatively coupled to SSH VCN1412 (e.g., SSH VCN1212 in Figure 12) via an LPG1410 contained in SSH VCN1412. 412 may include SSH subnet 1214 (e.g., SSH subnet 1214 in Figure 12), and SSH VCN 1412 may be communicatively coupled to control plane VCN 1416 (e.g., control plane VCN 1216 in Figure 12) via LPG 1410 included in control plane VCN 1416, and may be communicatively coupled to data plane VCN 1418 (e.g., data plane 1218 in Figure 12) via LPG 1410 included in data plane VCN 1418. Control plane VCN 1416 and data plane VCN 1418 may be included in service tenancy 1419 (e.g., service tenancy 1219 in Figure 12).
[0120] The control plane VCN1416 may include a control plane DMZ layer 1420 (e.g., control plane DMZ layer 1220 in Figure 12) which may include a load balancer (LB) subnet 1422 (e.g., LB subnet 1222 in Figure 12), a control plane application layer 1424 (e.g., control plane application layer 1224 in Figure 12) which may include an application subnet 1426 (e.g., similar to application subnet 1226 in Figure 12), and a control plane data layer 1428 (e.g., control plane data layer 1228 in Figure 12) which may include a DB subnet 1430. The LB subnet 1422 included in the control plane DMZ layer 1420 can be communicatively coupled to the application subnet 1426 included in the control plane application layer 1424 and to an internet gateway 1234 (e.g., internet gateway 1234 in Figure 12) which may be included in the control plane VCN 1416. The application subnet 1426 can also be communicatively coupled to the DB subnet 1430 included in the control plane data layer 1428 and to a service gateway 1436 (e.g., service gateway in Figure 12) and a network address translation (NAT) gateway 1438 (e.g., NAT gateway 1238 in Figure 12). The control plane VCN 1416 may include the service gateway 1436 and the NAT gateway 1438.
[0121] The data plane VCN1418 may include a data plane application layer 1446 (e.g., data plane application layer 1246 in Figure 12), a data plane DMZ layer 1448 (e.g., data plane DMZ layer 1248 in Figure 12), and a data plane data layer 1450 (e.g., data plane data layer 1250 in Figure 12). The data plane DMZ layer 1448 may include a trusted application subnet 1460 and an untrusted application subnet 1462 of the data plane application layer 1446, and an LB subnet 1422 that can be communicatively coupled to the internet gateway 1234 included in the data plane VCN1418. The trusted application subnet 1460 may be communicatively coupled to the service gateway 1436 included in the data plane VCN1418, the NAT gateway 1438 included in the data plane VCN1418, and the DB subnet 1430 included in the data plane data layer 1450. An untrusted application subnet 1462 may be communicatively coupled to a service gateway 1436 included in the data plane VCN 1418 and to a DB subnet 1430 included in the data plane data layer 1450. The data plane data layer 1450 may include a DB subnet 1430 that can be communicatively coupled to a service gateway 1436 included in the data plane VCN 1418.
[0122] An untrusted application subnet 1462 may contain one or more primary VNICs 1464(1)-(N) that can be communicatively coupled to tenant virtual machines (VMs) 1466(1)-(N). Each tenant VM 1466(1)-(N) may be communicatively coupled to each application subnet 1467(1)-(N) that may be included in each container egress VCN 1468(1)-(N) that may be included in each customer tenancy 1470(1)-(N). Each secondary VNIC 1472(1)-(N) can facilitate communication between the untrusted application subnet 1462 included in the data plane VCN 1418 and the application subnets included in the container egress VCN 1468(1)-(N). Each container egress VCN 1468(1)-(N) is connected to the public internet 1454( For example, it may include a NAT gateway 1438 that can be communicatively coupled to the public internet 1254) in Figure 12.
[0123] The Internet gateway 1234, included in the control plane VCN1416 and the data plane VCN1418, can be communicatively coupled to a metadata management service 1452 (e.g., the metadata management system 1252 in Figure 12), which can be communicatively coupled to the public internet 1454. The public internet 1454 can be communicatively coupled to a NAT gateway 1438, included in the control plane VCN1416 and the data plane VCN1418. The service gateway 1436, included in the control plane VCN1416 and the data plane VCN1418, can be communicatively coupled to a cloud service 1456.
[0124] In some embodiments, the data plane VCN1418 may be integrated with customer tenancy 1470. This integration may be useful or desirable for the IaaS provider's customer in several cases, such as when they may want support when executing code. The customer may provide code to be executed that is disruptive, communicates with other customer resources, or may cause undesirable effects. In response, the IaaS provider may decide whether to execute the code provided to the IaaS provider by the customer.
[0125] In some examples, an IaaS provider's customer might request temporary network access to the IaaS provider and that a certain function be attached to a dataplane tier application 1446. The code for performing that function may run in VMs 1466(1) to (N) and may not be configured to run anywhere else on the dataplane VCN 1418. Each VM 1466(1) to (N) may be connected to a single customer tenancy 1470. Each container 1471(1) to (N) contained within VMs 1466(1) to (N) may be configured to run code. In this case, there may be double isolation (for example, containers 1471(1)-(N) may execute code and containers 1471(1)-(N) may be contained in at least VM1466(1)-(N) contained in the untrusted app subnet 1462), which may help prevent incorrect or unwanted code from corrupting the IaaS provider's network or another customer's network. Containers 1471(1)-(N) may be communicably coupled to customer tenancy 1470 and may be configured to send or receive data to or from customer tenancy 1470. Containers 1471(1)-(N) may not be configured to send or receive data to or from any other entity in the data plane VCN1418. Once code execution is complete, the IaaS provider may disable or discard containers 1471(1)-(N).
[0126] In some embodiments, a trusted application subnet 1460 may execute code that may be owned or operated by the IaaS provider. In this embodiment, the trusted application subnet 1460 may be communicatively coupled to the DB subnet 1430 and configured to perform CRUD operations within the DB subnet 1430. An untrusted application subnet 1462 may be communicatively coupled to the DB subnet 1430, but in this embodiment, the untrusted application subnet may be configured to perform read operations within the DB subnet 1430. Containers 1471(1)~(N) that may be included in each customer's VM 1466(1)~(N) and may execute code from that customer do not have to be communicatively coupled to the DB subnet 1430.
[0127] In other embodiments, the control plane VCN1416 and the data plane VCN1418 They do not have to be directly communicatively coupled. In this embodiment, there may be no direct communication between the control plane VCN1416 and the data plane VCN1418. However, communication can be performed indirectly by at least one method. An LPG1410 that facilitates communication between the control plane VCN1416 and the data plane VCN1418 may be established by the IaaS provider. In another example, the control plane VCN1416 or the data plane VCN1418 may make a call to the cloud service 1456 via the service gateway 1436. For example, a call from the control plane VCN1416 to the cloud service 1456 may include a request for a service that can communicate with the data plane VCN1418.
[0128] Figure 15 is a block diagram 1500 showing another exemplary pattern of an IaaS architecture according to at least one embodiment. A service operator 1502 (e.g., service operator 1202 in Figure 12) may be communicatively coupled to a secure host tenancy 1504 (e.g., secure host tenancy 1204 in Figure 12), which may include a virtual cloud network (VCN) 1506 (e.g., VCN 1206 in Figure 12) and a secure host subnet 1508 (e.g., secure host subnet 1208 in Figure 12). VCN 1506 may include an LPG 1510 (e.g., LPG 1210 in Figure 12), which may be communicatively coupled to SSH VCN 1512 (e.g., SSH VCN 1212 in Figure 12) via an LPG 1510 contained in SSH VCN 1512. SSH VCN1512 may include SSH subnet 1514 (e.g., SSH subnet 1214 in Figure 12), and SSH VCN1512 may be communicatively coupled to control plane VCN1516 (e.g., control plane VCN1216 in Figure 12) via LPG1510 included in control plane VCN1516, and may be communicatively coupled to data plane VCN1518 (e.g., data plane 1218 in Figure 12) via LPG1510 included in data plane VCN1518. Control plane VCN1516 and data plane VCN1518 may be included in service tenancy 1519 (e.g., service tenancy 1219 in Figure 12).
[0129] The control plane VCN1516 may include a control plane DMZ layer 1520 (e.g., control plane DMZ layer 1220 in Figure 12) which may include an LB subnet 1522 (e.g., LB subnet 1222 in Figure 12), a control plane application layer 1524 (e.g., control plane application layer 1224 in Figure 12) which may include an application subnet 1526 (e.g., application subnet 1226 in Figure 12), and a control plane data layer 1528 (e.g., control plane data layer 1228 in Figure 12) which may include a DB subnet 1530 (e.g., DB subnet 1430 in Figure 14). The LB subnet 1522 included in the control plane DMZ layer 1520 can be communicatively coupled to the application subnet 1526 included in the control plane application layer 1524 and to an internet gateway 1534 (e.g., internet gateway 1234 in Figure 12) which may be included in the control plane VCN 1516. The application subnet 1526 can also be communicatively coupled to the DB subnet 1530 included in the control plane data layer 1528 and to a service gateway 1536 (e.g., service gateway in Figure 12) and a network address translation (NAT) gateway 1538 (e.g., NAT gateway 1238 in Figure 12). The control plane VCN 1516 may include the service gateway 1536 and the NAT gateway 1538.
[0130] The data plane VCN1518 may include the data plane application layer 1546 (e.g., data plane application layer 1246 in Figure 12), the data plane DMZ layer 1548 (e.g., data plane DMZ layer 1248 in Figure 12), and the data plane data layer 1550 (e.g., data plane data layer 1250 in Figure 12). The data plane DMZ layer 1548 includes trusted application subnets 1560 (e.g., trusted application subnet 1460 in Figure 14) and untrusted application subnets of the data plane application layer 1546. The trusted application subnet 1560 may include an LB subnet 1522 that can be communicatively coupled to an internet gateway 1534 included in the data plane VCN 1518, and a service gateway 1536 included in the data plane VCN 1518, and a DB subnet 1530 included in the data plane data layer 1550. The untrusted application subnet 1562 may be communicatively coupled to a service gateway 1536 included in the data plane VCN 1518, and a DB subnet 1530 included in the data plane data layer 1550. The data plane data layer 1550 may include a DB subnet 1530 that can be communicatively coupled to a service gateway 1536 included in the data plane VCN 1518.
[0131] An untrusted application subnet 1562 may include primary VNICs 1564(1)-(N) that can be communicatively coupled to tenant virtual machines (VMs) 1566(1)-(N) residing within the untrusted application subnet 1562. Each tenant VM 1566(1)-(N) may execute code in its respective container 1567(1)-(N) and may be communicatively coupled to an application subnet 1526 that may be included in a dataplane application layer 1546 that may be included in a container egress VCN 1568. Each secondary VNIC 1572(1)-(N) may facilitate communication between the untrusted application subnet 1562 included in the dataplane VCN 1518 and the application subnet included in the container egress VCN 1568. The container egress VCN may include a NAT gateway 1538 that can be communicatively coupled to the public internet 1554 (e.g., the public internet 1254 in Figure 12).
[0132] The Internet gateway 1534, which is included in both the control plane VCN1516 and the data plane VCN1518, can be communicatively coupled to a metadata management service 1552 (for example, the metadata management system 1252 in Figure 12), which can be communicatively coupled to the public internet 1554. The public internet 1554 can be communicatively coupled to a NAT gateway 1538, which is included in both the control plane VCN1516 and the data plane VCN1518. The service gateway 1536, which is included in both the control plane VCN1516 and the data plane VCN1518, can be communicatively coupled to a cloud service 1556.
[0133] In some examples, the pattern shown by the architecture of block diagram 1500 in Figure 15 may be considered an exception to the pattern shown by the architecture of block diagram 1400 in Figure 14, which may be desirable for the IaaS provider's customers when the IaaS provider cannot communicate directly with the customers (e.g., in a disconnected area). Each container 1567(1)-(N) contained within VM 1566(1)-(N) for each customer can be accessed by the customer in real time. Each container 1567(1)-(N) may be configured to make calls to each secondary VNIC 1572(1)-(N) contained within the application subnet 1526 of the data plane application layer 1546, which may be contained within the container egress VCN 1568. The secondary VNICs 1572(1)-(N) can send calls to the NAT gateway 1538, which may send calls to the public internet 1554. In this example, containers 1567(1)-(N), which can be accessed by customers in real time, may be isolated from the control plane VCN1516 and from other entities included in the data plane VCN1518. Containers 1567(1)-(N) may also be isolated from resources from other customers.
[0134] In another example, a customer uses containers 1567(1) to (N) to access cloud service 15 56 can be called. In this example, the customer may execute code within containers 1567(1)~(N) to request a service from cloud service 1556. Containers 1567(1)~(N) may send this request to secondary VNICs 1572(1)~(N), secondary VNICs 1572(1)~(N) may send the request to the NAT gateway, and the NAT gateway may send the request to the public internet 1554. The public internet 1554 may send the request to LB subnet 1522, which is included in control plane VCN 1516, via internet gateway 1534. In response to the determination that the request is valid, the LB subnet may send the request to application subnet 1526, and application subnet 1526 may send the request to cloud service 1556 via service gateway 1536.
[0135] Please understand that the illustrated IaaS architectures 1200, 1300, 1400, and 1500 may have components other than those shown. Furthermore, the illustrated embodiments are only some examples of cloud infrastructure systems that may incorporate embodiments of this disclosure. In some other embodiments, the IaaS system may have more or fewer components than those shown, may combine two or more components, or may have different configurations or arrangements of components.
[0136] In some embodiments, the IaaS system described herein may include a set of application, middleware, and database service offerings that are self-service, subscription-based, flexibly scalable, reliable, highly available, and delivered to customers in a secure manner. An example of such an IaaS system is Oracle Cloud Infrastructure (OCI) provided by the Assignee.
[0137] Figure 16 shows an exemplary computer system 1600 in which various embodiments of the present disclosure can be realized. System 1600 can be used to realize any of the computer systems described above. As shown, computer system 1600 includes a processing unit 1604 that communicates with several peripheral subsystems via a bus subsystem 1602. These peripheral subsystems may include a processing acceleration unit 1606, an I / O subsystem 1608, a storage subsystem 1618, and a communication subsystem 1624. The storage subsystem 1618 includes a tangible computer-readable storage medium 1622 and system memory 1610.
[0138] The bus subsystem 1602 provides a mechanism for various components and subsystems of the computer system 1600 to communicate with each other as intended. While the bus subsystem 1602 is schematically shown as a single bus, alternative embodiments of the bus subsystem may utilize multiple buses. The bus subsystem 1602 may be one of several types of bus structures, including a memory bus or memory controller, peripheral bus, and local bus, using one of various bus architectures. For example, such architectures may include industry standard architecture (ISA) buses, microchannel architecture (MCA) buses, extended ISA (EISA) buses, video electronics standards association (VESA) local buses, and peripheral component interconnect (PCI) buses, which can be implemented as mezzanine buses manufactured according to the IEEE P1386.1 standard.
[0139] The processing unit 1604 can be implemented as one or more integrated circuits (e.g., conventional microprocessors or microcontrollers) and controls the operation of the computer system 1600. One or more processors may be included in the processing unit 1604. These processors may include single-core processors or multi-core processors. In certain embodiments, the processing unit 1604 may include single-core processors or multi-core processors. The multicore processor may be implemented as one or more independent processing units 1632 and / or 1634, each containing a processing unit. In other embodiments, processing unit 1604 may also be implemented as a quadcore processing unit formed by integrating two dual-core processors onto a single chip.
[0140] In various embodiments, the processing unit 1604 can execute various programs in response to program code and can maintain multiple programs or processes running simultaneously. At any given time, some or all of the program code to be executed may reside in the processor 1604 and / or the storage subsystem 1618. With appropriate programming, the processor 1604 can provide the various functionalities described above. The computer system 1600 may also include a processing acceleration unit 1606, which may include a digital signal processor (DSP), a special-purpose processor, and the like.
[0141] The I / O subsystem 1608 may include user interface input devices and user interface output devices. User interface input devices may include pointing devices such as keyboards, mice or trackballs, touchpads or touchscreens integrated into displays, scroll wheels, click wheels, dials, buttons, switches, keypads, voice input devices with voice command recognition systems, microphones, and other types of input devices. User interface input devices may also include motion sensing and / or gesture recognition devices such as Microsoft Kinect® motion sensors, which enable users to control and interact with input devices such as Microsoft Xbox® 360 game controllers through a natural user interface using gestures and speech commands. User interface input devices may also include eye gesture recognition devices such as Google Glass® blink detectors, which detect eye movements from the user (e.g., blinks while taking pictures and / or making menu selections) and translate eye gestures into input to an input device (e.g., Google Glass®). In addition, the user interface input device may include a voice recognition sensing device that enables the user to interact with a voice recognition system (e.g., Siri® Navigator) via voice commands.
[0142] User interface input devices may include, but are not limited to, three-dimensional (3D) mice, joysticks or pointing sticks, gamepads and graphic tablets, as well as auditory / visual devices such as speakers, digital cameras, digital camcorders, portable media players, webcams, image scanners, fingerprint scanners, barcode readers, 3D scanners, 3D printers, laser rangefinders, and eye-tracking devices. In addition, user interface input devices may also include, for example, medical imaging input devices such as computed tomography, magnetic resonance imaging, positron emission tomography, and medical ultrasound equipment. User interface input devices may also include, for example, audio input devices such as MIDI keyboards and digital musical instruments.
[0143] User interface output devices may include non-visual displays such as display subsystems, indicator lights, or audio output devices. Display subsystems may include flat panel devices such as those using cathode ray tubes (CRTs), liquid crystal displays (LCDs), or plasma displays, projection devices, touchscreens, etc. Generally, the use of the term “output device” is intended to include all conceivable types of devices and mechanisms for outputting information from the computer system 1600 to a user or another computer. For example, user interface The output devices may include, but are not limited to, various display devices that visually convey text, graphics, and audio / video information, such as monitors, printers, speakers, headphones, car navigation systems, plotters, audio output devices, and modems.
[0144] The computer system 1600 may include a storage subsystem 1618 containing software elements, which are currently shown to be located in the system memory 1610. The system memory 1610 may store program instructions that can be loaded and executed on the processing unit 1604, as well as data generated during the execution of these programs.
[0145] Depending on the configuration and type of the computer system 1600, the system memory 1610 may be volatile (such as random access memory (RAM)) and / or non-volatile (such as read-only memory (ROM), flash memory, etc.). RAM typically includes data and / or program modules that are immediately accessible to the processing unit 1604 and / or currently being operated and executed by the processing unit 1604. In some implementations, the system memory 1610 may include several different types of memory, such as static random access memory (SRAM) or dynamic random access memory (DRAM). In some implementations, a basic input / output system (BIOS) containing basic routines that help transfer information between elements within the computer system 1600, such as during startup, may be stored typically in ROM. As an example, but not an limitation, the system memory 1610 also includes application programs 1612, program data 1614, and an operating system 1616, which may include client applications, web browsers, middle-tier applications, relational database management systems (RDBMS), etc. For example, operating system 1616 is compatible with various versions of Microsoft Windows®, Apple Macintosh®, and / or Linux®. This may include operating systems, various commercially available UNIX® or UNIX-like operating systems (including, but not limited to, various GNU / Linux operating systems, Google Chrome® OS, etc.), and / or mobile operating systems such as iOS, Windows® Phone, Android® OS, BlackBerry® 16 OS, and Palm® OS.
[0146] The storage subsystem 1618 may also provide a tangible computer-readable storage medium for storing basic programming and data structures that provide the functionality of several embodiments. Software (programs, code modules, instructions) that, when executed by a processor, provides the functionality described above may be stored in the storage subsystem 1618. These software modules or instructions may be executed by the processing unit 1604. The storage subsystem 1618 may also provide a repository for storing data used in accordance with this disclosure.
[0147] The storage subsystem 1600 may also include a computer-readable storage medium reader 3220 which may be further connected to the computer-readable storage medium 1622. Together with the system memory 1610, and optionally in combination with the system memory 1610, the computer-readable storage medium 1622 may comprehensively represent a combination of a storage medium, remote, local, fixed, and / or removable storage device, for temporarily and / or more permanently accommodating, storing, transmitting, and retrieving computer-readable information.
[0148] The computer-readable storage medium 1622 containing code or a portion of code may also include any suitable media known or used in the art, including, but not limited to, volatile and non-volatile, removable and non-removable media, as well as storage and communication media, which are implemented by any method or technique for storing and / or transmitting information. This may include tangible computer-readable storage media such as RAM, ROM, electronically erasable programmable ROM (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical storage devices, magnetic cassettes, magnetic tapes, magnetic disk storage devices or other magnetic storage devices, or other tangible computer-readable media. This may also include intangible computer-readable media such as any other media that can be used to transmit data signals, data transmissions, or desired information and are accessible by the computing system 1600.
[0149] For example, the computer-readable storage medium 1622 is a hard disk drive that reads and writes to a non-removable non-volatile magnetic medium, a magnetic disk drive that reads and writes to a removable non-volatile magnetic disk, a CD-ROM, a DVD and a Blu-ray (registered The computer-readable storage medium 1622 may include optical disc drives that read from and write to removable non-volatile optical discs such as (trademark) discs, or other optical media. The computer-readable storage medium 1622 may include, but is not limited to, Zip® drives, flash memory cards, Universal Serial Bus (USB) flash drives, Secure Digital (SD) cards, DVD discs, digital videotapes, etc. The computer-readable storage medium 1622 may also include flash memory-based SSDs, enterprise flash drives, solid-state drives (SSDs) based on non-volatile memory such as solid-state ROM, SSDs based on volatile memory such as solid-state RAM, dynamic RAM, static RAM, DRAM-based SSDs, magnetoresistive RAM (MRAM) SSDs, and hybrid SSDs that use a combination of DRAM and flash memory-based SSDs. The disk drives and the computer-readable media associated therewith may provide the computer system 1600 with non-volatile storage of computer-readable instructions, data structures, program modules, and other data.
[0150] The communication subsystem 1624 provides interfaces to other computer systems and networks. The communication subsystem 1624 functions as an interface for sending and receiving data between other systems and the computer system 1600. For example, the communication subsystem 1624 may enable the computer system 1600 to connect to one or more devices via the Internet. In some embodiments, the communication subsystem 1624 may include radio frequency (RF) transceiver components for accessing wireless voice and / or data networks (using, for example, cellular telephone technology, 3G, 4G, or EDGE (High Speed Data Rate for Global Evolution)), Global Positioning System (GPS) receiver components, and / or other components. In some embodiments, the communication subsystem 1624 may provide, in addition to or instead of a wireless interface, a wired network connection (e.g., Ethernet®).
[0151] In some embodiments, the communication subsystem 1624 may also receive input communications in the form of structured and / or unstructured data feeds 1626, event streams 1628, event updates 1630, etc., on behalf of one or more users who may be using the computer system 1600.
[0152] For example, the communication subsystem 1624 handles Twitter® feeds, Facebook ( (Registered Trademark) Updates, Rich Site Summary (RSS) feeds and other web feeds, and / Alternatively, it may be configured to receive data feeds 1626 in real time from users of social networks and / or other communication services, such as real-time updates from one or more third-party sources.
[0153] In addition, the communication subsystem 1624 may also be configured to receive data in the form of a continuous data stream, which may include an event stream 1628 and / or event update 1630 of real-time events that are essentially continuous or infinite and have no explicit termination. Examples of applications that generate continuous data may include, for example, sensor data applications, financial stock market boards, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, and automotive traffic monitoring.
[0154] The communication subsystem 1624 may also be configured to output structured and / or unstructured data feeds 1626, event streams 1628, event updates 1630, etc., to one or more databases that can communicate with one or more streaming data source computers coupled to the computer system 1600.
[0155] The computer system 1600 may be one of various types, including handheld portable devices (e.g., iPhone® mobile phones, iPad® computing tablets, PDAs), wearable devices (e.g., Google Glass® head-mounted displays), PCs, workstations, mainframes, kiosks, server racks, or any other data processing systems.
[0156] Because the nature of computers and networks is constantly changing, the description of the computer system 1600 shown in the figure is intended only as an example. Many other configurations with more or fewer components than the system shown in the figure are possible. For example, customized hardware may be used, and / or certain elements may be implemented in hardware, firmware, software (including applets), or a combination thereof. Furthermore, connections to other computing devices, such as network input / output devices, may be employed. Those skilled in the art will understand other aspects and / or methods for implementing various embodiments based on the disclosures and teachings contained herein.
[0157] While specific embodiments of the Disclosure have been described, various modifications, changes, alternative configurations, and equivalent examples are also included within the scope of the Disclosure. The embodiments of the Disclosure are not limited to operation within a particular data processing environment, but are freely operable in multiple data processing environments. Furthermore, while the embodiments of the Disclosure have been described using a specific set of transactions and steps, it should be apparent to those skilled in the art that the scope of the Disclosure is not limited to the described set of transactions and steps. Various features and aspects of the embodiments described above may be used individually or in combination.
[0158] Furthermore, while embodiments of this disclosure have been described using specific combinations of hardware and software, it should be recognized that other combinations of hardware and software are also within the scope of this disclosure. Embodiments of this disclosure may be implemented using hardware alone, software alone, or a combination thereof. Various processes described herein may be implemented on the same processor or different processors of any combination thereof. Therefore, where a component or module is described as being configured to perform a particular operation, such a configuration may, for example, include electronic circuits for performing the operation. This can be achieved by designing, programming programmable electronic circuits (such as microprocessors) to perform operations, or any combination thereof. Processes can communicate using a variety of techniques, including but not limited to conventional techniques for inter-process communication, and different pairs of processes may use different techniques, or the same pair of processes may use different techniques at different times.
[0159] Therefore, the specification and drawings should be considered illustrative rather than restrictive. However, it will be clear that additions, reductions, deletions, and other modifications and changes may be made without departing from the broader spirit and scope described in the claims. Thus, while specific embodiments of the disclosure have been described, these are not intended to be limiting. Various examples of modifications and equivalents are within the claims.
[0160] The use of the words “a,” “an,” and “the,” as well as similar referents, in the context of describing the disclosed embodiments (particularly in the context of the claims), should be interpreted as encompassing both singular and plural forms unless otherwise indicated herein or unless the context clearly contradicts this interpretation. The words “comprising,” “having,” “including,” and “containing” should be interpreted as non-restrictive unless otherwise specified (i.e., “including The term “connected” should be interpreted as meaning “but not limited to.” The term “connected” should be interpreted as being partially or entirely included, attached, or combined with something, even if something is intervening. Unless otherwise indicated herein, the descriptions of value ranges are intended merely as a concise way of referring individually to each distinct value that falls within that range, and each distinct value is incorporated herein as if it were individually described herein. All methods described herein may be carried out in any preferred order unless otherwise indicated herein or unless it is clearly inconsistent with the context. Any and all examples or exemplary words provided herein (e.g., “etc.”) are intended merely to better illustrate embodiments of the disclosure and, unless otherwise requested, do not limit the scope of the disclosure. Nothing in this specification should be interpreted as indicating that any unrequested element is essential to the implementation of the disclosure.
[0161] Disjunctive phrases such as "at least one of X, Y, or Z" are intended to be understood, unless otherwise specified, in contexts where they are used schematically to indicate that an item, term, etc., can be X, Y, Z, or any combination thereof (e.g., X, Y, and / or Z). Therefore, such disjunctive phrases are not, and should not, be intended to imply that a particular embodiment requires the presence of at least one X, at least one Y, or at least one Z, respectively.
[0162] Preferred embodiments of the Disclosure, including the best known form for carrying out the Disclosure, are described herein. Modifications of these preferred embodiments will become apparent to those skilled in the art by reading the preceding description. Those skilled in the art should be able to adopt such modifications as appropriate, and the Disclosure may be practiced in ways other than those specifically described herein. Accordingly, the Disclosure includes all modifications and equivalents of the subject matter described in the claims, as permitted by applicable law. Furthermore, any combination of the elements described above in all possible modifications is incorporated herein unless otherwise indicated herein.
[0163] All documents, including publications, patent applications, and patents, cited herein are as if they were... Each reference is indicated to be cited individually and specifically, and is cited here to the same extent as if it were presented as a whole.
[0164] While the above-described specification illustrates aspects of the disclosure with reference to specific embodiments, those skilled in the art will recognize that the disclosure is not limited thereto. The various features and aspects of the above-described disclosure may be used individually or together. Furthermore, embodiments may be used in any number of environments and applications beyond those described herein, without departing from the broader spirit and scope of this specification. Accordingly, the specification and drawings should be considered illustrative rather than restrictive.
[0165] The following sections describe embodiments of the disclosed implementation examples. Item 1: A method, A computer system creates a first snapshot of a block volume containing multiple partitions in a first geographical area at a first logical time; The computer system transmits the first snapshot data corresponding to the first snapshot to an object storage system in a second geographical area; The computer system performs the steps of creating a second snapshot of the block volume in the first geographical area at a second logical time, The method includes the step of generating a plurality of deltas by the computer system, wherein each delta corresponds to one of the plurality of partitions, and the method further The computer system transmits multiple delta datasets corresponding to the multiple deltas to the object storage system in the second geographical area; The computer system generates a checkpoint by aggregating, at least partially, the object metadata associated with the multiple deltas and the first snapshot; The computer system receives a restore request to generate a restore volume, A method comprising the step of generating the recovery volume from the checkpoint using the computer system.
[0166] Section 2: The step of generating the multiple deltas is: A step of generating a comparison between the second snapshot and the first snapshot, Based on the comparison, the steps include determining corrected data corresponding to the changes between the first snapshot data and the second snapshot data corresponding to the second snapshot, The method according to item 1, further comprising the step of generating a plurality of deltas that describe the modified data for the plurality of partitions.
[0167] Section 3: The step of creating the first snapshot is: A step of interrupting input / output operations for the multiple partitions in accordance with logical time, The steps include generating multiple block images that describe the volume data within the multiple partitions, The method according to item 1, further comprising the step of enabling input / output operations for the plurality of partitions.
[0168] Section 4: The restoration request is a failover request, and the method further, A step that enables the generation of the restored volume in the second geographical area, The method according to item 1, further comprising the step of enabling input / output operations using the restored volume in the second geographical area.
[0169] Section 5: The restoration request is a failback request, and the method further, The steps include generating the restored volume in the second geographical area, A step that enables the generation of a failback volume in the first geographical area, The method according to item 1, further comprising the step of restoring the first snapshot data in the first geographical area.
[0170] Section 6: The step of sending the multiple delta datasets is: The steps include generating multiple chunk objects from the multiple delta datasets, The steps include transferring the multiple deltas, The method according to item 1, further comprising the step of transferring the plurality of chunk objects to the object storage system.
[0171] Section 7: The checkpoint includes the manifest of the object metadata, The method according to item 6, wherein the object metadata includes chunk pointers corresponding to the multiple chunk objects in the object storage system.
[0172] Section 18: The method according to Section 7, wherein the step of aggregating the object metadata includes updating the manifest to reflect multiple differences between the multiple delta datasets and the first snapshot data.
[0173] Item 9: A computer system, One or more processors, The system includes a memory that communicates with one or more processors, the memory being configured to store computer executable instructions, and by executing these computer executable instructions, the system causes one or more processors to perform the following steps, the following steps being: A computer system creates a first snapshot of a block volume containing multiple partitions in a first geographical area at a first logical time; The computer system transmits the first snapshot data corresponding to the first snapshot to an object storage system in a second geographical area; The computer system performs the steps of creating a second snapshot of the block volume in the first geographical area at a second logical time, The computer system includes the step of generating a plurality of deltas, each delta of which corresponds to one of the plurality of partitions, and the following steps further include: The computer system transmits multiple delta datasets corresponding to the multiple deltas to the object storage system in the second geographical area; The computer system will aggregate, at least partially, the object metadata associated with the multiple deltas and the first snapshot. The step of generating checkpoints, The computer system receives a restore request to generate a restore volume, A computer system comprising the step of generating the restore volume from the checkpoint.
[0174] Item 10: The step of generating the multiple deltas is: A step of generating a comparison between the second snapshot and the first snapshot, Based on the comparison, the steps include determining corrected data corresponding to the changes between the first snapshot data and the second snapshot data corresponding to the second snapshot, The computer system according to paragraph 9, comprising the step of generating a plurality of deltas that describe the modified data for the plurality of partitions.
[0175] Section 11: The step of creating the first snapshot is: A step of interrupting input / output operations for the multiple partitions in accordance with logical time, The steps include generating multiple block images that describe the volume data within the multiple partitions, The computer system according to paragraph 9, further comprising the step of enabling input / output operations for the multiple partitions.
[0176] Section 12: The restoration request is a failover request, and the method further, A step that enables the generation of the restored volume in the second geographical area, The computer system according to paragraph 9, further comprising the step of enabling input / output operations using the restored volume in the second geographical area.
[0177] Section 13: The restoration request is a failback request, and the method further, The steps include generating the restored volume in the second geographical area, A step that enables the generation of a failback volume in the first geographical area, The computer system according to paragraph 9, comprising the step of restoring the first snapshot data in the first geographical area.
[0178] Section 14: The step of sending the multiple delta datasets is: The steps include generating multiple chunk objects from the multiple delta datasets, The steps include transferring the multiple deltas, The computer system according to paragraph 9, further comprising the step of transferring the plurality of chunk objects to the object storage system.
[0179] Section 15: The checkpoint includes the manifest of the object metadata, The computer system described in paragraph 14, wherein the object metadata includes chunk pointers corresponding to the multiple chunk objects in the object storage system.
[0180] Section 16: The step of aggregating the object metadata is performed by the multiple delta data The computer system described in paragraph 15, which includes the step of updating the manifest to reflect multiple differences between the set and the first snapshot data.
[0181] Item 17: A computer-readable storage medium for storing computer executable instructions, wherein, when executed, the computer executable instructions cause one or more processors of a computer system to perform the following steps, the following steps: A computer system creates a first snapshot of a block volume containing multiple partitions in a first geographical area at a first logical time; The computer system transmits the first snapshot data corresponding to the first snapshot to an object storage system in a second geographical area; The computer system performs the steps of creating a second snapshot of the block volume in the first geographical area at a second logical time, The computer system includes the step of generating a plurality of deltas, each delta of which corresponds to one of the plurality of partitions, and the following steps further include: The computer system transmits multiple delta datasets corresponding to the multiple deltas to the object storage system in the second geographical area; The computer system generates a checkpoint by aggregating, at least partially, the object metadata associated with the multiple deltas and the first snapshot; The computer system receives a restore request to generate a restore volume, A computer-readable storage medium, comprising the step of generating the recovery volume from the checkpoint by the computer system.
[0182] Section 18: The step of generating the multiple deltas is: A step of generating a comparison between the second snapshot and the first snapshot, Based on the comparison, the steps include determining corrected data corresponding to the changes between the first snapshot data and the second snapshot data corresponding to the second snapshot, A computer-readable storage medium according to paragraph 17, comprising the step of generating a plurality of deltas that describe the modified data for the plurality of partitions.
[0183] Section 19: The step of sending the multiple delta datasets is: The steps include generating multiple chunk objects from the multiple delta datasets, The steps include transferring the multiple deltas, A computer-readable storage medium according to paragraph 17, comprising the step of transferring the plurality of chunk objects to the object storage system.
[0184] Item 20: The checkpoint includes the manifest of the object metadata, The computer-readable storage medium described in paragraph 19 includes the object metadata, which includes chunk pointers corresponding to the multiple chunk objects in the object storage system.
Claims
1. It is a method, A computer system creates a first snapshot of a block volume containing multiple partitions in a first geographical area at a first logical time; The computer system transmits the first snapshot data corresponding to the first snapshot to an object storage system in a second geographical area. The computer system performs the steps of creating a second snapshot of the block volume in the first geographical area at a second logical time, The method further includes the step of generating a plurality of deltas by the computer system, wherein each delta of the plurality of deltas corresponds to one of the plurality of partitions, and the method further includes the step of generating a plurality of deltas by the computer system, wherein each delta of the plurality of deltas corresponds to one of the plurality of partitions, The computer system transmits a plurality of delta datasets corresponding to the plurality of deltas to the object storage system in the second geographical area. The steps include generating a checkpoint by aggregating, at least partially, the object metadata associated with the plurality of deltas and the first snapshot using the computer system, The computer system receives a restore request to generate a restore volume, A method comprising the step of generating the recovery volume from the checkpoint using the computer system.
2. The step of generating the aforementioned multiple deltas is: A step of generating a comparison between the second snapshot and the first snapshot, Based on the above comparison, the step of determining corrected data corresponding to the changes between the first snapshot data and the second snapshot data corresponding to the second snapshot, The method according to claim 1, further comprising the step of generating a plurality of deltas that describe the modified data for the plurality of partitions.
3. The step of creating the first snapshot is: A step of interrupting input / output operations for the multiple partitions in accordance with logical time, The steps include generating multiple block images that describe the volume data within the multiple partitions, The method according to claim 1, further comprising the step of enabling input / output operations for the plurality of partitions.
4. The restoration request is a failover request, and the method further, A step that enables the generation of the restored volume in the second geographical area, The method according to claim 1, further comprising the step of enabling input / output operations using the recovery volume in the second geographical area.
5. The restoration request is a failback request, and the method further, The steps include generating the recovery volume in the second geographical area, A step that enables the generation of a failback volume in the first geographical region, The method according to claim 1, further comprising the step of restoring the first snapshot data in the first geographical area.
6. The step of sending the aforementioned multiple delta datasets is, The steps include generating multiple chunk objects from the multiple delta datasets, The steps include transferring the plurality of deltas, The method according to claim 1, further comprising the step of transferring the plurality of chunk objects to the object storage system.
7. The aforementioned checkpoint includes the manifest of the object metadata, The method according to claim 6, wherein the object metadata includes chunk pointers corresponding to the plurality of chunk objects in the object storage system.
8. The method according to claim 7, wherein the step of aggregating the object metadata includes updating the manifest to reflect a plurality of differences between the plurality of delta datasets and the first snapshot data.
9. A computer system, One or more processors, The system includes a memory that communicates with one or more processors, the memory being configured to store computer executable instructions, and by executing the computer executable instructions, the system causes one or more processors to perform the following steps, the following steps are: A computer system creates a first snapshot of a block volume containing multiple partitions in a first geographical area at a first logical time; The computer system transmits the first snapshot data corresponding to the first snapshot to an object storage system in a second geographical area. The computer system performs the steps of creating a second snapshot of the block volume in the first geographical area at a second logical time, The steps include generating a plurality of deltas by the computer system, wherein each delta corresponds to one of the plurality of partitions, and the following steps further include: The computer system transmits a plurality of delta datasets corresponding to the plurality of deltas to the object storage system in the second geographical area. The steps include generating a checkpoint by aggregating, at least partially, the object metadata associated with the plurality of deltas and the first snapshot using the computer system, The computer system receives a restore request to generate a restore volume, A computer system comprising the step of generating the recovery volume from the checkpoint.
10. The step of generating the aforementioned multiple deltas is: A step of generating a comparison between the second snapshot and the first snapshot, Based on the above comparison, the step of determining corrected data corresponding to the changes between the first snapshot data and the second snapshot data corresponding to the second snapshot, The computer system according to claim 9, comprising the step of generating a plurality of deltas that describe the modified data for the plurality of partitions.
11. The step of creating the first snapshot is: A step of interrupting input / output operations for the multiple partitions in accordance with logical time, The steps include generating multiple block images that describe the volume data within the multiple partitions, The computer system according to claim 9, further comprising the step of enabling input / output operations for the plurality of partitions.
12. The restoration request is a failover request, and the method further, A step that enables the generation of the restored volume in the second geographical area, The computer system according to claim 9, further comprising the step of enabling input / output operations using the recovery volume in the second geographical area.
13. The restoration request is a failback request, and the method further, The steps include generating the recovery volume in the second geographical area, A step that enables the generation of a failback volume in the first geographical region, The computer system according to claim 9, further comprising the step of restoring the first snapshot data in the first geographical area.
14. The step of sending the aforementioned multiple delta datasets is, The steps include generating multiple chunk objects from the multiple delta datasets, The steps include transferring the plurality of deltas, The computer system according to claim 9, further comprising the step of transferring the plurality of chunk objects to the object storage system.
15. The aforementioned checkpoint includes the manifest of the object metadata, The computer system according to claim 14, wherein the object metadata includes chunk pointers corresponding to the plurality of chunk objects in the object storage system.
16. The computer system according to claim 15, wherein the step of aggregating the object metadata includes updating the manifest to reflect a plurality of differences between the plurality of delta datasets and the first snapshot data.
17. A computer-readable storage medium for storing computer executable instructions, wherein, when executed, the computer executable instructions cause one or more processors of a computer system to perform the following steps, and the following steps are: A computer system creates a first snapshot of a block volume containing multiple partitions in a first geographical area at a first logical time; The computer system provides the first snapshot corresponding to the first The steps include sending snapshot data to an object storage system in a second geographical area, The computer system performs the steps of creating a second snapshot of the block volume in the first geographical area at a second logical time, The steps include generating a plurality of deltas by the computer system, wherein each delta corresponds to one of the plurality of partitions, and the following steps further include: The computer system transmits a plurality of delta datasets corresponding to the plurality of deltas to the object storage system in the second geographical area. The steps include generating a checkpoint by aggregating, at least partially, the object metadata associated with the plurality of deltas and the first snapshot using the computer system, The computer system receives a restore request to generate a restore volume, A computer-readable storage medium comprising the steps of generating the recovery volume from the checkpoint by the computer system.
18. The step of generating the aforementioned multiple deltas is: A step of generating a comparison between the second snapshot and the first snapshot, Based on the above comparison, the step of determining corrected data corresponding to the changes between the first snapshot data and the second snapshot data corresponding to the second snapshot, The computer-readable storage medium according to claim 17, comprising the step of generating a plurality of deltas that describe the modified data for the plurality of partitions.
19. The step of sending the aforementioned multiple delta datasets is, The steps include generating multiple chunk objects from the multiple delta datasets, The steps include transferring the plurality of deltas, The computer-readable storage medium according to claim 17, further comprising the step of transferring the plurality of chunk objects to the object storage system.
20. The aforementioned checkpoint includes the manifest of the object metadata, The computer-readable storage medium according to claim 19, wherein the object metadata includes chunk pointers corresponding to the plurality of chunk objects in the object storage system.