Hierarchical Key Management for Inter-Region Replication
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- ORACLE INT CORP
- Filing Date
- 2023-06-02
- Publication Date
- 2026-05-22
AI Technical Summary
Current disaster recovery solutions for cloud infrastructure regions lack efficient and secure mechanisms for end-to-end file storage replication, particularly during failover and failback processes, leading to management challenges and potential data loss.
A hierarchical key management system is implemented, using three distinct keys (source file system key, session key, and target file system key) to encrypt and decrypt file data across different cloud infrastructure regions, ensuring secure and efficient replication through parallel processing and asynchronous operations.
The system provides scalable, reliable, and secure end-to-end file storage replication with minimal management effort, maintaining data consistency and reducing recovery time objectives by utilizing high-throughput object storage and parallel processing threads.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 352,992, filed on June 16, 2022, U.S. Provisional Patent Application No. 63 / 357,526, filed on June 30, 2022, U.S. Provisional Patent Application No. 63 / 412,243, filed on September 30, 2022, and U.S. Provisional Patent Application No. 63 / 378,486, filed on October 5, 2022, under 35 U.S.C. 119(e), and claims priority to U.S. Non - Provisional Patent Application No. 18 / 094,302, filed on January 6, 2023, entitled "HIERARCHICAL KEY MANAGEMENT FOR CROSS - REGION REPLICATION", the disclosures of which are hereby incorporated by reference in their entirety for all purposes.
[0002] Field This disclosure generally relates to file systems. More particularly, but not by way of limitation, techniques for key management including end - to - end file storage replication between different cloud infrastructure regions are described.
Background Art
[0003] Background Enterprise businesses contain extremely important data. While excellent disaster recovery solutions are important, security is an absolute essential aspect. The security of extremely important data must be protected both when stored within a data center and during transmission between data centers during failover and failback processes. Therefore, there is a need for security during disaster recovery.
Summary of the Invention
[0004] Brief Summary This disclosure generally relates to file systems. More particularly, but not by way of limitation, techniques for key management including end-to-end file storage replication between different cloud infrastructure regions are described. **Means for Solving the Problems**
[0005] In one embodiment, a computing system generates a first security key associated with a source file system for encrypting and decrypting a plurality of file keys in the source file system, at least partially based on a first master key, wherein the source file system is configured to transmit snapshot deltas during a replication process and a snapshot delta is identified between two snapshots of the source file system; the computing system generates a second security key associated with a target file system for encrypting and decrypting a plurality of file keys in the target file system, at least partially based on a second master key, wherein the target file system is configured to receive snapshot deltas during the replication process; and the computing system generates a session key for encrypting and decrypting snapshot deltas transferred between the source file system and the target file system during the replication process, at least partially based on a third master key, wherein the session key is valid during a session and the first master key, the second master key, and the third master key are different keys. A technique is provided that includes a method including these operations.
[0006] In yet another embodiment, the session is a period between the start of a replication process in the source file system and the end of the replication process in the target file system.
[0007] In yet another embodiment, the snapshot difference transferred between the source file system and the target file system further includes object storage configured to receive the snapshot difference from the source file system and transfer the snapshot difference to the target file system.
[0008] In yet another embodiment, the method further includes encrypting the snapshot difference using a session key before transferring the snapshot difference to the object storage and decrypting the snapshot difference using the session key after transferring the snapshot difference to the target file system.
[0009] In yet another embodiment, the method further includes transferring the session key from the control plane of the source file system to the control plane of the target file system.
[0010] In yet another embodiment, the session key is associated with globally unique resource identification information.
[0011] In yet another embodiment, each file key among the plurality of file keys in the source file system is associated with a specific file in the source file system for encrypting and decrypting the file data of the specific file in the source file system, and each file key among the plurality of file keys in the target file system is associated with a specific file in the target file system for encrypting and decrypting the file data of the specific file in the target file system.
[0012] In yet another embodiment, the method further includes authenticating a key requester in the source file system to require the use of a first security key and authenticating a key requester in the target file system to require the use of a second security key.
[0013] In yet another embodiment, authenticating a key requester within a source file system includes checking an identification number of a replication process and an identification number of the source file system.
[0014] In various embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more of the methods disclosed herein.
[0015] In various embodiments, the non-transitory computer-readable medium stores computer-executable instructions that, when executed by one or more processors, cause one or more processors of a computer system to perform one or more of the methods disclosed herein.
[0016] In various embodiments, a computer program product includes a computer program / instructions that, when executed by a processor, cause the processor to perform any of the methods disclosed herein.
[0017] The techniques described above and below can be implemented in a variety of methods and in a variety of situations. Referring to the following figures, which are described in more detail below, a number of exemplary implementations and situations are provided. However, the following implementations and situations are only a part of many implementations and situations.
[0018] The features, embodiments, and advantages of the present disclosure will be better understood when the following detailed description is read with reference to the accompanying drawings.
Brief Description of the Drawings
[0019]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6A
Figure 6B
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Best Mode for Carrying Out the Invention
[0020] Detailed Description In the following description, for purposes of explanation, specific details are set forth in order to provide a thorough understanding of an embodiment. It will be apparent, however, that various embodiments may be practiced without these specific details. The figures and the description are not intended to be restrictive. The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment or design described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments or designs.
[0021] The present disclosure generally relates to file systems. More particularly, but not by way of limitation, techniques for key management including end-to-end file storage replication between different cloud infrastructure regions are described.
[0022] In one embodiment, the file storage service (FSS) disclosed in the present disclosure utilizes inter-region replication of a three-layer key architecture. During the replication process, three different keys may be included: a source file system key (FSKa), a session key, and a target file system key (FSKb). Within each source region and target region, the file system key is also in a layered structure such that it encrypts local files using a master key from the customer to securely store the file system key.
[0023] In some embodiments, the source file system has three keys: a source file key (FKa) for each file, a source file system key (FSKa), and a master key (based on a key managed by the customer). The source file system includes a plurality of files each having its own file key FKa for encryption and decryption. The target file system has three keys: a target file key (FKb) for each file, a target file system key (FSKb), and a master key (based on a key managed by the customer). In certain embodiments, the master key of the source file system is different from the master key of the target file system.
[0024] Explanation of terms in certain embodiments The "Recovery Time Objective" (RTO) refers to, in some embodiments, the period during which a user needs to make replication available within a secondary (or target) region, regardless of whether the failure is planned or unplanned, after a failure occurs within the availability domain (AD) of the primary (or source) region.
[0025] The "Recovery Point Objective" (RPO) refers to, in some embodiments, the maximum allowable range with respect to the time of data loss between a failure in the primary region (usually due to an unplanned failure) and the availability of the secondary region.
[0026] In certain embodiments, a "replicator" can refer to a component (e.g., a virtual machine (VM)) within the data plane of a file system that uploads differences to a remote object store (i.e., an object storage service) when the component is located in a source region, or downloads differences from the object storage for applying the differences when the component is located in a target region. The replicator is formed as a fleet (i.e., a plurality of VMs or replicator threads) called a replicator fleet and can execute inter-region (or cross-region) replication processes (e.g., uploading differences to a target region) in parallel.
[0027] In certain embodiments, a "delta generator" (DG) can refer to a component within the data plane of a file system that extracts differences (i.e., changes) between the keys and values of two snapshots when the component is located in a source region, or applies differences to the latest snapshot within the B-tree of the file system when the component is located in a target region. The delta generator within the source region can use multiple threads (referred to as delta generator threads or range threads for multiple divided B-tree key ranges) to execute the extraction of differences (or B-tree traversal) in parallel. The delta generator within the target region can use multiple threads to apply the downloaded differences to the latest snapshot in parallel.
[0028] For the purposes of the present disclosure, in certain embodiments, a "shared database" (SDB) can refer to a key-value store that components (e.g., a replicator fleet) within both the control plane and the data plane of a file system can read from and write to in order to communicate with each other. In certain embodiments, the SDB can be part of a B-tree.
[0029] In certain embodiments, a "file system communicator" (FSC) may refer to a file manager layer that runs on a storage node within the data plane of a file system. This service helps with file creation, deletion, read, and write requests, and works with an FNS server (e.g., Orca) to service I / O to clients. The replicator fleet can communicate with multiple storage nodes, thereby distributing file system data read / write operations across the storage nodes.
[0030] In certain embodiments, a "blob" may refer to a data type for storing information (e.g., a formatted binary file) in a database. Blobs are generated during replication by a source region and uploaded to an object store (i.e., object storage) within a target region. A blob may contain binary tree (B-tree) keys and values as well as file data. A blob within an object store is called an object. The pairs of B-tree keys and values and the data associated with them are packed together into the blob that is uploaded to the object store within the target region.
[0031] In one embodiment, a "manifest" may refer to information transmitted by a file system in a source region (referred to herein as the source file system) to a file system in a target region (referred to herein as the target file system) to facilitate the inter-region replication process. There are two types of manifest files: the master manifest and the checkpoint manifest. A range manifest file (or master manifest file) is created by the source file system at the start of the replication process and describes information (e.g., B-tree key ranges) required by the target file system. A checkpoint manifest file is created after a checkpoint in the source file system and notifies the target file system of the number of blobs included in the checkpoint and uploaded to the object store. In response, the target file system can then download that number of blobs.
[0032] In one embodiment, a "difference" may refer to the differences identified between two specific snapshots after a replicator has recursively visited all the nodes of a B-tree (also referred to herein as scanning the B-tree). A difference generator identifies pairs of keys and values of the B-tree with respect to the differences, traverses the B-tree nodes, and retrieves the file data associated with the B-tree keys. The difference between two snapshots may include multiple blobs. The term "difference" may include blobs and manifests when used in the context of uploading information by the source file system to the object store and downloading by the target file system from the object store.
[0033] In certain embodiments, an "object" may refer to a partial set of information representing the entire difference during an inter-region replication cycle, and is stored in an object store. An object may be several megabytes in size and stored at a specific location within a bucket of the object store. An object may contain many differences (i.e., blobs and manifests). A blob uploaded and stored in an object store is called an object.
[0034] In certain embodiments, a "bucket" may refer to a container that stores objects in compartments within an object storage namespace (tenancy). In the present disclosure, a bucket is used by a source replicator to store differences protected using server-side encryption (SSE), and is also used by a target replicator to download changes and apply them to a snapshot.
[0035] In certain embodiments, "difference application" may refer to the process by which a target file system applies the downloaded differences to the latest snapshot to create a new snapshot. Difference application may include analyzing a manifest file, applying snapshot metadata, inserting B-tree keys and values into a B-tree, and storing the data associated with the B-tree keys (i.e., the file data or the data portion of a blob) in local storage. Snapshot metadata is created and applied at the start of a replication cycle.
[0036] In certain embodiments, a "region" may refer to a logical abstraction corresponding to a geographical area. Each region can include one or more connected data centers. A region is independent of other regions and can be separated by a vast distance.
[0037] End-to-End Inter-Region Replication Architecture Disclosed herein is an end-to-end inter-region replication architecture that provides new technologies for end-to-end file storage replication and security between file systems within different cloud infrastructure regions. In certain embodiments, a file storage service generates differences between snapshots within a source file system and, upon disaster recovery, transfers the differences and associated data via high-throughput object storage to recreate new snapshots in a target file system located in a different region. The file storage service utilizes new technologies to achieve scalable, reliable, and restartable end-to-end replication. New technologies for ensuring secure transfer and consistency of information during end-to-end replication are also described.
[0038] In the context of the cloud, a region refers to a local geographic area that includes one or more connected data centers. A region is independent from other regions and can be separated by vast distances, even across countries or continents. A realm refers to a logical collection of one or more regions. Realms are typically separated from each other and do not share data. Within a region, the data centers within the region can be organized into one or more availability domains (ADs). An availability domain is separated from others, is fault tolerant, and has a very low probability of failing simultaneously. An AD is configured such that a failure in one AD within a region is unlikely to affect the availability of other ADs within the same region.
[0039] Current practices for disaster recovery can include taking periodic snapshots and resynchronizing those snapshots to a different file system within a different availability domain (AD) or region. Resynchronization is manageable and maintained by the customer, but lacks a user interface to show progress, is a slow serialized process, and is not easy to manage as data grows over time.
[0040] Accordingly, different approaches are needed to address these and other challenges. The file storage replication of a cloud service provider (e.g., Oracle Cloud Infrastructure (OCI)) disclosed in this disclosure is based on incremental snapshots and provides a consistent point-in-time view of the entire file system by propagating the differences in the changed data from the primary AD within a region to a secondary AD within the same or a different region. As used herein, the primary site (or source side) may refer to the location where the file system is located and where the replication process for disaster recovery is initiated (e.g., an AD or a region). The secondary site (or target side) may refer to the location where the file system receives information from the file system within the primary site during the replication process and becomes the new operational file system after disaster recovery (e.g., an AD or a region). The file system located at the primary site is called the source file system, and the file system located at the secondary site is called the target file system. Accordingly, the primary site, the source side, the source region, the primary file system, or the source file system (referring to one of the file systems on the source side) may be used interchangeably. Similarly, the secondary site, the target side, the target region, the secondary file system, or the target file system (referring to one of the file systems on the target side) may be used interchangeably.
[0041] The file storage service (FSS) of the present disclosure supports complete disaster recovery for failover or failback with minimal management effort. Failover is a series of actions to make the secondary site / target site the primary / source (i.e., start providing services for the workload), which may include planned failover and / or unplanned failover. A planned failover (sometimes called a planned migration) is initiated by the user to perform a planned failover from the source side (e.g., source region) to the target side (e.g., target region) without data loss. An unplanned failover is the case where, for example, due to a disaster, the source side stops unexpectedly and the source side is lost, so the user needs to start using the target side. Failback is to restore the primary side / source side to become the primary / source again before the failover. Failback may occur when the user wants to reuse the source side as the primary AD by reversing the failover process after a planned failover or an unplanned failover and a trigger event (e.g., power outage) has ended. The user can resume either from the last point in time on the source side before the trigger event or from the latest changes on the target side. The replication process described in the present disclosure can maintain the identity of the file system after round-trip replication. In other words, the source file system can resume providing services for the workload again after performing a failover and then a failback.
[0042] The technologies disclosed in this disclosure (e.g., methods, computer-readable media, and systems) use consistent snapshot information to replicate the differences between snapshots from a source region to multiple remote (or target) regions, and then scan (or recursively visit) all keys and values within one or more file trees (e.g., B-trees) of a source file system (referred to herein as "scanning the B-tree" or "scanning the keys") to construct consistent information (e.g., the differences or discrepancies between the keys and values of two snapshots created at different times), including region-to-region replication of file system data and / or metadata. The constructed consistent information is in blob format and is transferred to the remote side (e.g., the target region) using an object interface, such as an object store (described later), so that the target file system on the remote side can immediately detect the information transferred through the object interface and start downloading and applying it. This process is implemented using a control plane and can be extended to thousands of file systems and hundreds of replication machines. Both the source file system and the target file system can operate simultaneously and asynchronously. Operating simultaneously means that the data upload process by the source file system and the data download process by the target file system can occur simultaneously. Operating asynchronously means that the source file system and the target file system can each operate at their own pace without waiting for each other at all stages, e.g., with different start times, end times, processing speeds, etc.
[0043] In one embodiment, multiple file systems may exist in the same region and be represented by the same B-tree. Each of these file systems within the same region can be replicated independently across regions. For example, file system A may have a set of parallel execution replicator threads that scan the B-tree to perform replication of file system A. File system B, represented by the same B-tree, may have another set of such parallel execution replicator threads that scan the same B-tree to perform replication of file system B.
[0044] Regarding security, cross-region replication is completely secure. Information is transferred securely and applied securely. The disclosed technology provides separation between the source region and the target region so that keys are not shared between the two without being encrypted. Thus, if the source key is involved, the target is not affected. Further, the disclosed technology includes ways to read keys, convert those keys into a certain format, and upload and download those keys securely. Since different keys are created and used in different regions, separate keys are created at the target and applied to the information with a target-centric security mechanism. For example, FSS generates a session key that is only valid during one replication cycle or session to encrypt data uploaded from the source region to the object store and decrypt data downloaded from the object store to the target region. Separate keys are used locally within the source region and the target region.
[0045] In the disclosed technology, each upload process and download process via the object store during replication has different pipeline stages. For example, the upload process has multiple pipeline stages including scanning a B-tree to generate differences, accessing storage I / O, and uploading data (or blobs) to the object store. The download process has multiple pipeline stages including downloading data, applying differences to a snapshot, and storing the data in storage. Each of these pipelines also includes parallel processing threads to improve the throughput and performance of the replication process. Further, the parallel processing threads can take over a failed processing thread and resume the replication process from the point of failure without restarting from the beginning. Thus, the replication process is highly scalable and reliable.
[0046] Figure 1 shows an exemplary concept of the target recovery point in time (RPO) and target recovery time (RTO) for an unplanned failover according to an embodiment. The RPO is the maximum allowable range of data loss between the failure of the primary site and the availability of the secondary site (usually specified in minutes). As shown in Figure 1, the primary site A102 encounters an unplanned incident at time 110 and triggers the failover replication process by copying the latest snapshot and its delta to the secondary site B104. The information first copied reaches the secondary site B104 at time 112. The primary site A102 completes the copy of the information to the secondary site B104 at time 114, and the secondary site B104 completes the replication process at time 116. Thus, the secondary site B104 becomes fully operational at time 116. As a result, the user's data is not accessible within the primary site A110 from point 110 until the point 116 where the data becomes available again. Thus, the RPO is the time between point 110 and point 116. For example, if there is data equivalent to 10 minutes that the user is not interested in, the RPO is 10 minutes. If the data loss exceeds 10 minutes, the RPO is not met. An RPO of 0 means synchronous replication.
[0047] RTO is the time it takes for the secondary to become fully operational after a failure, so that the user can access the data again (usually specified in minutes). RTO is considered from the perspective of the secondary site. Referring back to Figure 1, the primary site A102 starts the failover replication process at time 120. However, the secondary site B104 remains operational until time 122 when it recognizes the incident (or power outage) at the primary site A102. Therefore, the secondary site B104 stops its service at time 122. The secondary site B104 becomes fully operational at time 126 using the same failover replication process as described for RPO. Therefore, RTO is the time between 122 and 126. Here, the secondary site B104 can take over the role of the primary site. However, for customers using the primary site A102, the service loss is between times 120 and 126.
[0048] The primary (or source) site is where the action is taking place, and the secondary (or target) site is inactive and cannot be used until a disaster occurs. However, the customer may be provided with a point in time to continue using for test-related activities at the secondary site. This relates to how the customer sets up the replication, how the customer can start using the target if any problems occur, and how the customer can return to the source after the source fails over.
[0049] FIG. 2 is a simplified block diagram showing an architecture for inter-region remote replication according to an embodiment. In FIG. 2, the end-to-end replication architecture shown includes two regions: a source region 290 and a target region 292. Each region may include one or more file systems. In one embodiment, the end-to-end replication architecture includes data planes 202 and 212, a control plane (only control APIs 208a-n and 218a-n are shown), local storages 204 and 214, an object store 260, and a key management service (KMS) 250 for both the source region 290 and the target region 292. FIG. 2 shows only one file system 280 in the source region 290 and one file system 282 in the target region 292 for simplicity. If there are two or more file systems in one region, the same replication architecture is applied to each pair of source and target file systems. File systems within a region may share resources. For example, certain resources within the KMS 250, the object store 260, and the data plane may be shared by multiple file systems within the same region depending on the implementation.
[0050] The data plane within the architecture includes local storage nodes 204a - n and 214a - n and replicators (or replicator fleets) 206a - n and 216a - n. The control API hosts within each region perform all orchestration between different regions. The FSS receives a request from a customer to set up replication between the source file system 280 and the target file system 282 where the customer's data will be moved. The control plane 208 obtains the request, performs resource allocation, and notifies the replicator fleet 206a - n within the source data plane 202 to start uploading data 230a from different snapshots to the object storage 260 (or sometimes called that only the differences are uploaded). An API is available to assist the customer in setting the target time and the recovery time objective (RTO) of the replication. The replication model disclosed in this disclosure is a "push - based" model based on snapshot differences, that is, the source region initiates the replication.
[0051] As used herein, the data 230a and 230b transferred between the source file system 280 and the target file system 282 are general terms and may include an initial snapshot, keys and values of different B - trees between two snapshots, file data (e.g., fmap), snapshot metadata (i.e., a set of B - tree keys of the snapshots reflecting different snapshots obtained within the source file system), and other information (e.g., manifest files) that helps facilitate the replication process.
[0052] Regarding the data plane of the inter-region replication architecture, the replicator is a component within the data plane of the file system. The replicator performs differential generation or differential application on the file system according to the region where the file system is located. For example, the replicator fleet 206 within the file system 280 of the source region performs the generation and replication of the difference 230a. The replicator fleet 216 within the file system 282 of the target region downloads the differences 230b and applies them to the latest snapshot within the file system 282 of the target region. The file system 282 of the target region can also use the control plane and workflow to ensure end-to-end transfer.
[0053] All incremental operations are based on snapshots, which are existing resources within file storage as a service. A snapshot is a point in time, data point, or image of what is happening within the file system and is executed periodically within the file system 280 of the source region. In the very first replication (for example, where replication has not been obtained before), FSS obtains a base snapshot that is a snapshot of all the contents of the source file system and transfers all that content to the target system. In other words, the replicator reads from the storage layer of that particular file system and stores all the data in the object storage bucket.
[0054] After the data plane 202 of the source file system 280 uploads all data 230a to the object storage (or object store) 260, the source-side control plane 208 notifies the target-side control plane 218 that there is new work to be done on the target side, and then this notification is relayed to the target-side replicator. Thereafter, the target-side replicators 216a~n begin to download objects (e.g., initial snapshots and deltas) from the object storage bucket 260 and apply the deltas captured on the source side.
[0055] For a base copy (e.g., the entire contents of the file system up to a point in time ranging from the past 5 days to 5 years), the upload process may take time. To assist in meeting service-level goals regarding time and performance, the source system 280 can take replication snapshots at specific intervals such as one hour. The source side 280 can then transfer all data within that one hour to the target side 282 and take new snapshots every hour. If there is any cache with many changes, the replication can be set to a shorter replication interval.
[0056] To illustrate the above, consider a situation where a first snapshot is created on a file system within a source region (referred to as the source file system). Replication is performed periodically, and thus the first snapshot is replicated to a file system within a target region (referred to as the target file system). Thereafter, when some update is performed within the source file system, a second snapshot is created. If an unplanned power outage occurs after the second snapshot is created, the source file system attempts to replicate the second snapshot to the target file system. During failover, the source file system may well identify the difference (i.e., the delta) between the first snapshot and the second snapshot, which includes the keys and values of the B-tree and the file data associated therewith within the B-tree representing both the first snapshot and the second snapshot. Next, deltas 230a and 230b are transferred from the source file system to the target file system via an object store 260 within the target region, and the target file system recreates the second snapshot by applying the deltas to the first snapshot previously established within the target region. When the second snapshot is created on the target file system, the failover replication process is complete and the target file system is ready to operate.
[0057] Regarding the control plane and its application programming interfaces (APIs), the control plane provides instructions for the data plane that includes replicators as executors for executing instructions. Storage (204 and 214) and replicator fleets (206 and 216) are both within the data plane. The control plane is not shown in Figure 2. As used herein, a "cycle" may refer to a period that starts when the source file system 280 begins to transfer data 230a to the target file system 282 and ends when the target file system 282 has received all the data 230b and completed the application of the received data. The data 230a - b is captured on the source side and then applied on the target side. When all changes on the target side are applied to the cycle, the source file system 280 takes another snapshot and starts another cycle.
[0058] The control APIs (208a - n and 218a - n) are a set of hosts within the overall architecture of the control plane and execute the configuration of the file system. The control APIs are responsible for communicating state information between different regions. State machines that track various state activities within a region, such as the progress of a job, the location of keys, and future tasks to be executed, are distributed across multiple regions. All this information is stored in the control plane of each region and communicated between regions via the control APIs. In other words, the state information relates to the details of the life cycle, the details of the differences, and the life cycle of the resources. The state machine can also be useful for tracking the progress of replication and cooperating with the data plane to estimate the time taken for replication. Thus, the state machine can provide the user with a status regarding whether the replication is proceeding as expected and the health of the job.
[0059] Furthermore, communication between the control APIs (208a - n) of the source file system 280 and the control APIs (218a - n) of the target file system 218 in a different region includes the transfer of snapshots and metadata for creating an exact copy from the source to the target. For example, when a customer regularly obtains snapshots within the source file system, the control plane can ensure that snapshots of the same user, including metadata tracking, transfer, and recreation, are created in the target file system.
[0060] The object store 260 in FIG. 2 (also referred to as an "object" herein) is an object storage service (e.g., Oracle's object storage service) that enables reading blobs and writing files for archival purposes. The advantages of using an object store are, firstly, ease of configuration, secondly, ease of streaming data to the object store, and thirdly, having the advantages of security streaming as a reliable repository for maintaining information, all of which are because there is no network loss, the data can be immediately downloaded, and it exists permanently. Direct communication between replicators within the source region and the target region is possible, but direct communication requires the configuration of an inter-region network, which is not scalable and difficult to manage.
[0061] For example, when there is a large amount of data to be moved from a source to a target, the source can upload the data to the object store 260, and the target 282 does not need to wait for all the information uploaded to the object store 260 to start downloading. Thus, both the source 280 and the target 282 can operate continuously and simultaneously. The use of the object store enables the system to scale and achieve higher throughput. Further, the Key Management Service (KMS) 250 can control access to the object store 260 to ensure security. In other words, the source tries to move the data out of the source region as fast as possible and hold the data somewhere so that the data is not lost before it can be applied to the target.
[0062] Compared to using a network pipe with packet loss and recovery issues, the use of the object store 260 between the source region and the target region enables continuous data streaming where hundreds of file systems can be written from the source region to the object store, and at the same time, the target region can apply hundreds of files simultaneously. Thus, data streaming via the object store can achieve high throughput. Further, both the source region and the target region can operate at their own speeds for uploading and downloading.
[0063] Whenever a user changes some data in the source file system 280, a snapshot is taken and the difference before and after the change is updated. These changes are accumulated in the source file system 280 and can be streamed to the object store 260. The target file system 282 can detect that the data is available in the object store 260 and immediately download the changes and apply them to that file system. In some embodiments, only the differences are uploaded to the object storage after the base snapshot.
[0064] In some embodiments, the replicator can communicate with many different regions (e.g., from Phoenix to Ashburn and further to other remote regions), and the file system can manage many different endpoints on the replicator. Each replicator 206 within the source file system 280 can maintain a cache of these object storage endpoints, and further, in cooperation with the KMS 250, generate a transfer key (e.g., a session key) for encrypting the data address of the data in the object storage 260 (e.g., server-side encryption or SSE) to protect the data stored in the bucket. There is one master bucket for each AD within the target region. A bucket is a container that stores objects in a compartment within the object storage namespace (tenancy). Since all remote clients can communicate with the bucket and write information in a specific format, the information of each file system can be uniquely identified, preventing the mixing of data from different customers or file systems.
[0065] The object store 260 is a high-throughput system, and the techniques disclosed in this disclosure can utilize the object store. In one embodiment, the replication process includes multiple pipeline stages, a B-tree scan within the source file system 280, storage IO access, data upload to the object store 260, data download from the object store 260, and differential application within the target file system 282. Each stage includes parallel processing threads that participate in improving the performance of data streaming from the source region 290 to the target region 292 via the object store 260.
[0066] In one embodiment, each file system within the source region may include a set of replicator threads 206a-n that are executed in parallel to upload the differences to the object store 260. Each file system within the target region may also include a set of replicator threads 216a-n that are executed in parallel to download the differences from the object store 260. Since both the source side and the target side operate asynchronously at the same time, the source can upload as fast as possible, while the target can start downloading after detecting that the differences are available in the object store. Thereafter, the target file system applies the differences to the latest snapshot and deletes the differences in the object store after the application. Thus, the FSS consumes very little space in the object store, and the object store has a very high throughput (e.g., gigabyte-scale transfers).
[0067] In one embodiment, multiple threads are also executed in parallel for storage IO access (e.g., DASD) 204a-n and 214a-n. Thus, all processes related to the replication process, including accessing storage, uploading the snapshot and data 230a from the source file system 280 to the object store 260, and downloading the snapshot and data 230b to the target file system 282, include multiple threads that are executed in parallel to perform data streaming.
[0068] File storage is a local service of the AD. When a file system is created, that file system is within a specific AD. When a customer transfers or replicates data from one file system to another file system within the same region or a different region, artifact (also called manifest) transfer may need to be used.
[0069] As an alternative to using an object store to transfer data, a network connection between remote machines (e.g., between source and target replicator nodes) can be set up and VCN peering can be used to use Classless Inter-Domain Routing (CIDR) for each region.
[0070] Referring again to Figure 2, the Key Management System (KMS) 250 provides security for replication and provides storage services to a cloud service provider (e.g., OCI). In some embodiments, the file systems 280 on the source (or primary) side and the target (or secondary) side use separate KMS keys and key management is hierarchical. The reason for using separate keys is that if the source is compromised, an unauthorized actor cannot decrypt the target using the same key. The FSS has a three-tier key architecture. Since the source and target use different keys during data transfer, the source first decrypts the data, re-encrypts it using an intermediate key, and then re-encrypts the data on the target side. The FSS defines a session, and each session is one data cycle. A key for transferring data in that session is created. In other words, a new key is used for each new session. In other embodiments, a key can be used for two or more sessions (e.g., two or more data transfers) before creating another key. The key is not transferred via the object store 260, the key is only available on the source side and is not visible from outside the source for security reasons.
[0071] The replication cycle (also called a session) is periodic and adjustable. For example, the replicators (206a~n and 216a~n) execute replication once every hour. The cycle starts when a new snapshot is created on the source side 280 and ends when all the differences 230b have been applied to the target side 282 (i.e., the target has reached the DONE state). Each session is completed before another session starts. Therefore, there is always only one session and no overlap between sessions.
[0072] Secret management (i.e., replication using KMS) processes the transfer of confidential materials between the source (primary) file system 290 and the target (or secondary) file system 292 using the KMS250. The source file system 280 calculates the differences, reads the file data, and then decrypts the file data in cooperation with the key management service using the encryption key of the local file system. Next, the source file system 280 generates a session key (referred to as a delta encryption key (DEK)), encrypts it to become an encrypted session key (referred to as a delta transfer key (DTK)), and transfers the DTK to the target file system 282 via the respective control planes 208 and 218. The source file system 280 further encrypts the data 230a using the DEK and uploads the encrypted data 230a to the object store 260 via the Transport Layer Security (TLS) protocol. Next, the object store 260 uses server-side encryption (SSE) to ensure the security for the storage of the data (e.g., differences, manifests, and metadata) 230a.
[0073] The target file system 282 securely obtains the encrypted session key DTK via the control plane 218 (using HTTPS via inter-region API communication), decrypts the session key DTK via the KMS 250 to obtain the DEK, and places the DEK at a location within the target region 292. When a replication job is scheduled within the target file system 282, the DEK is provided to a replicator (one of the replication fleets 216a - n), and the replicator uses this key to decrypt the data (e.g., the delta including file data) 230b downloaded from the object store 260 for application and re-encrypts the file data using the local file system key.
[0074] Replication between the source file system 280 and the target file system 282 is a parallel process, and both the source file system 280 and the target file system 282 operate at their own paces. When the source side completes the upload (which can occur before the target download process), the source side cleans up the memory and removes all keys. When the target completes the application of the delta to the latest snapshot, it similarly cleans up the memory and removes all keys. The FSS service also releases the KMS key. In other words, there are two copies of the session key, one within the source file system 280 and another within the target file system 282. Both copies are deleted at the end of each session, and a new session key is generated in the next replication cycle. This process ensures that the same key is not used for different purposes. Further, the session key is encrypted by the file system key, creating a two-fold protection. This is to ensure that only a specific file system can use this session key.
[0075] Figure 3 is a simplified schematic diagram of components involved in inter-region remote replication according to an embodiment. In one embodiment, components called the differential generator (DG) 310 in the source region A 302 and 330 in the target region B 304 are part of the replicator fleet 318 and operate on thousands of storage nodes within the fleet. The replicator 318 in the source region A makes remote procedural calls (RPCs) to the differential generator 310 (e.g., obtaining a set of keys and values, locking a block, etc.), and collects the keys, values, and data pages of the B-tree from a direct-access storage device (DASD) 314, which is a replication storage service for accessing storage and is regarded as a data server. The DG 310 in the source region A is a helper for the replicator 318, divides the key range of the difference, and packs all the keys / values in a specific range into a blob to be returned to the replicator 318. There are multiple storage nodes 322 and 342 connected to the DASDs 314 and 334 in both regions, and each node contains a large number of disks (e.g., 10TB or more).
[0076] In one embodiment, the file system communicators (FSCs) 312 and 332 in both regions are metadata servers that help update the source file system for user updates to the system. The FSCs 312 and 332 are used for file system communication, and the differential generator 310 is used for replication. Both the DG 310 and 330 and the FSCs 312 and 332 are metadata servers. User traffic passes through the FSCs 312 and 332 and the DASDs 314 and 334, while replication traffic passes through the DG. In an alternative embodiment, the functions of the FSC can be merged with the functions of the DG.
[0077] In one embodiment, the shared databases (SDBs) 316 and 336 of both regions are key-value stores, and through these components, both the control plane and the data plane (e.g., the replicator fleet) can read and write for each other to communicate. The control planes 320 and 340 of both regions can put new jobs into the queues in their respective shared databases 316 and 336, and the replicator fleets 318 and 338 continuously read the queues in the shared databases 316 and 336, and when the replicator fleets 318 and 338 detect a job request, they can initiate file system replication. In other words, the shared databases 316 and 336 are conduits between the replicator fleet and the control plane. Further, the shared databases 316 and 336 are resources distributed across different regions, and the IO traffic between the shared databases 316 and 336 should be minimized. Similarly, the IO traffic between the DASD should be minimized so as not to affect the user's performance. However, the replication process may be adjusted because it is a secondary service compared to the primary service.
[0078] The replicator fleet 318 within the source region A, in cooperation with the DG310, can start scanning the B-tree in the file system within the source region A, collect keys and values, and convert those keys and values into flat files or blobs to be uploaded to the object store. Once the data blobs (including keys and values and the actual data) are uploaded, the target can apply those data blobs immediately without waiting for a large number of blobs to be present in the object store 360. The object store 360 is located in the target region B for disaster recovery reasons. The goal is to push from the source to the target region B as quickly as possible and keep the data safe.
[0079] Optimize space by using lower-cost machines with smaller footprints, schedule as many replications as possible while ensuring fair bandwidth allocation among those machines, and there are multiple replicators to replicate thousands of file systems. The replicator fleets 318 and 338 in both regions are run on virtual machines that can be automatically scaled up and down to build the entire fleet for running replications. The replicators and replication services can dynamically adapt based on capacity to support each job. If the load on one replicator is high, another replicator can be selected to share the load. Different replicators in the fleet can balance the load with each other to ensure that jobs can continue and are not stopped due to overloading individual replicators.
[0080] FIG. 4 is a simplified flowchart showing steps executed during inter-region remote replication, according to an embodiment.
[0081] Step S1: When the customer sets up replication, the customer provides a source (or primary) file system (A) 402, a target (or secondary) file system (B) 404, and an RPO. The file systems are uniquely identified by file system identification information (e.g., Oracle Cloud ID or OCID), which is globally unique to the file system. The data is stored in the file storage service ("FSS") control plane database.
[0082] Step S2: The source (A) control plane (CP-A) 410 coordinates to periodically create system snapshots at regular intervals (less than the RPO), and notifies the data plane (including replicator / uploader 412) of the latest snapshot and the last snapshot successfully copied to the target (B) file system 404.
[0083] Step S3: CP-A410 notifies the replicator 412 (or uploader), which is a component within the data plane, to copy the latest snapshot.
[0084] S3a: The replicator 412 of the source (A) scans the B-tree to calculate the difference between two specific snapshots. The existing key infrastructure is used to decrypt the file system data.
[0085] S3b: These differences 414 are uploaded to the object store 430 within the target (B) region (the data can be compressed and / or deduplicated during copying). This upload can be executed in parallel by multiple replicator threads 412.
[0086] Step S4: CP-A410 notifies the target (B) control plane (CP-B) 450 of the completion of the upload.
[0087] Step S5: CP-B450 calls the target replicator B452 (or downloader) to apply the differences.
[0088] S5a: The replicator B452 downloads the data 454 from the object store 430.
[0089] S5b: The replicator B452 applies these differences to the target file system (B).
[0090] Step S6: After the difference application is completed, CP-A410 is notified of the new snapshot currently available at the target (B).
[0091] Step 7: The inter-region remote replication process repeats from Step S2 to Step S6.
[0092] Figure 5 is a simplified diagram showing a high-level concept of B-tree scanning according to an embodiment. The B-tree structure can be used within a file system. The difference generator scans the B-tree and ensures the consistency of the scan. In other words, the scan confirms that at the end of the scan, the keys and values are as expected so that data corruption cannot occur, and captures all information between any two snapshots. The file system is a transactional file system that may be modified, and since another user may update the same transaction or data, the user needs to be aware of the changes and retry the transaction.
[0093] The keys, values, and snapshots are immutable (i.e., cannot be changed except that the garbage collector may remove them). As shown in Figure 5, there are many snapshots (Snapshot 1 to Snapshot N) in the file system. When the difference generator is scanning the B-tree keys (510 to 560) in the source file system, the garbage collector 580 may come in and clean up the keys of the snapshots that it considers garbage, so the snapshots may be removed. When the difference generator scans the B-tree keys, the difference generator needs to ensure that the keys associated with the remaining snapshots (e.g., keys not removed by the garbage collector) are copied. When keys, such as 540 and 550, are removed by the garbage collector 580, the B-tree page can be shrunk, for example, from 2 pages before garbage collection to 1 page after garbage collection. A way for the difference generator to ensure consistency when scanning the B-tree keys is for the garbage collector 580 to confirm that it has not changed or deleted any of the keys in the page (or section between two snapshots) that the difference generator has just scanned (e.g., between two keys). Once consistency is confirmed, the difference generator collects the keys and sends them to the replicator for processing and uploading.
[0094] The B-tree key can indicate what has changed. The technology disclosed in this disclosure can determine which B-tree keys are new and what has been updated between two snapshots. The diff generator can collect the metadata part, keys and values, and related data, and then send it to the target. The target can understand that the received information is within the range of two snapshots and applies to the target file system. The diff generator (or a thread of the diff generator) scans the section between two keys, confirms its consistency, and then uses the last end key as the next start key for the next scan. This process is repeated until all keys are checked, and the diff generator collects related data each time the consistency is confirmed.
[0095] For example, when a file is changed (e.g., created, deleted, and then re-created) within a file system, this process creates multiple versions of the corresponding file directory entry. During the replication process, the garbage collector may clean up (or remove) the version of the file directory entry corresponding to the deleted file, which may cause a consistency problem called a whiteout. A whiteout occurs when there is a mismatch between the source file system and the target file system, because the target file system may fail to reconstruct the original snapshot chain containing the modified file. The disclosed technology can detect whiteout files (i.e., modified files affected by the garbage collector) during the B-tree scan, extract the version of the modified file that is not affected, and provide related information to the target file system within the same replication cycle to ensure the consistency between the source file system and the target file system by properly reconstructing the correct snapshot chain.
[0096] Figures 6A and 6B are diagrams showing the pipeline stages of inter-region replication according to an embodiment. The inter-region replication of the source file system disclosed in the present disclosure includes four pipeline stages, namely, the start of inter-region replication, the B-tree scan within the source file system (i.e., the differential generation pipeline stage), the storage IO access for retrieving data (i.e., the data read pipeline stage), and the data upload to the object store (i.e., the data upload pipeline stage), which are included within the source file system. The target file system includes four pipeline stages in a similar but reverse order, namely, the preparation for inter-region replication, the download of data from the object store, the application of differences within the target file system, and the storage IO access for storing data. Figure 6A shows the four pipeline stages within the source file system, and the same concept applies to the target file system. Figure 6B shows the processes and interactions between the components involved in the pipeline stages. These pipeline stages can all operate in parallel. Each pipeline stage operates independently and can pass information to the next pipeline stage when the processing at the current stage is completed. Each pipeline stage receives a portion of the total bandwidth and is guaranteed not to use more than necessary. In other words, resources are fairly allocated among all jobs. When no other jobs are operating within the system, the operating job can acquire as many resources as possible.
[0097] Threads within each pipeline stage also execute tasks independently of each other in parallel (or simultaneously) within the same pipeline stage (i.e., if a thread fails, it does not affect other threads). Further, the tasks (or replication jobs) executed by threads at each pipeline stage are restartable, i.e., if a thread fails, a new thread (also called a replacement thread) can take over from the failed thread and continue the original task from the last successful point.
[0098] In some embodiments, the B-tree scan can be performed using parallel processing threads within the source file system 280. The B-tree can be divided into a plurality of key ranges between the first key and the last key in the file system. The number of key ranges can be determined by the customer. For each file system, a plurality of (e.g., about 8 to 16) range threads can be used for the B-tree scan. One range thread can perform a B-tree scan of one key range, and all range threads operate in parallel simultaneously. The number of threads used varies depending on factors such as the size of the file system, the availability of resources, and the bandwidth for balancing resource and traffic congestion. Usually, the number of key ranges is more than the number of available range threads for fully utilizing the range threads. Therefore, the B-tree scan is scalable and can be processed by simultaneous parallel scans (e.g., using multiple threads).
[0099] After the differential generator scans the pages, if some keys are missing and thus some keys are inconsistent, the system can remove the ongoing uncommitted transactions and return to the starting point for scanning again. During the repetition of the B-tree scan due to the inconsistency, the differential generator can ignore the missing keys and the related data in order to minimize the amount of information to be processed or uploaded to the target side, because these related data are regarded as garbage and thus not collected. Therefore, the B-tree scan and data transfer can be made more efficient. Further, the differential generator does not need to wait for the garbage collector to remove the information to be deleted before scanning the B-tree keys. For example, keys have dependencies on each other. If a key or iNode points to a block that has been deleted by the garbage collector or should be deleted, the system (or the differential generator) can itself understand that a particular block is garbage and that the differential generator does not need to carry it.
[0100] The differential generator usually makes no changes on the source side (e.g., does not delete the keys or blocks of data considered as garbage), and simply does not copy them to the target side. The B-tree scanning process and garbage collection are asynchronous processes. For example, when the block of data pointed to by a key no longer exists, the file system can flag the key as garbage, notify that the key should not be changed (e.g., is immutable), and only the garbage collector can remove the key. The differential generator can continue scanning the next key without waiting for the garbage collector. In other words, the differential generator and the garbage collector can proceed at their own paces.
[0101] In FIG. 6A, when the source region starts an inter-region replication process that can include multiple file systems, the main threads 610a - n select a replication job (one job per file system). The main thread of the file system (e.g., 610a or 610 for later use) within the source region (i.e., the source file system) communicates with the differential generator 620 (shown in FIG. 6B) to obtain the number of key ranges requested by the customer and updates the corresponding records in the SDB 622. Once the main thread 610 of the source file system knows the number of key ranges required, it further creates a set of range threads 612a - n based on the number of key ranges required. These range threads 612a - n are executed by the differential generator 620. These range threads 612a - n initialize the GETKEYVAL buffer 640 (shown in FIG. 6B), update the checkpoint record 642 in the SDB 622 (shown in FIG. 6B), and perform storage I / O access 644 by interacting with the DASD I / O threads 614a - n.
[0102] In one embodiment, each main thread 610 is responsible for monitoring all range threads 612a - n that it creates. During replication, the main thread 610 may generate a master manifest file that outlines the entire replication. The range threads 612a - n generate a range manifest file that includes the number of key ranges (i.e., the subdivision of the entire replication), and then generate a checkpoint manifest (CM) file for each range to provide updates to the target file system regarding the number of blobs per checkpoint, where the checkpoints are created during a B - tree scan. One checkpoint is created by a range thread 612. When the main thread 610 determines that all range threads 612a - n are complete, it creates a final checkpoint manifest (CM) file that includes an end - of - file mark, and then uploads the CM file to the object store so that the target file system can understand the progress within the source file system. The CM file includes an overview of all individual ranges, such as the number of ranges, the final state of the checkpoint records, and other information.
[0103] Range threads 612a~n are used for parallel processing to significantly shorten the time for B-tree traversal of a large source file system. In one embodiment, the B-tree keys are divided into ranges of approximately equal size. One range thread can perform a B-tree traversal of one key range. The number of range threads 612a~n used varies depending on factors such as the size of the file system, resource availability, and bandwidth to balance resources, the amount of data generated, and traffic congestion. Usually, the number of key ranges is about two to four times more than the number of available range threads 612a~n to fully utilize the range threads. Each of the range threads 612a~n has a dedicated buffer (GETKEYVAL) 640 that contains jobs available for work. Each range thread 612 operates independently of other range threads and periodically updates checkpoint records 642 in the SDB622.
[0104] Range threads 612a~n may need to collect file data (e.g., FMAP) associated with the B-tree keys and request IO access 644 to storage when traversing the B-tree (i.e., when recursively visiting all nodes of the B-tree). These IO requests are queued by each range thread 612 so that the DASD IO threads 614a~n (i.e., the data read pipeline stage) can handle those IO requests. These DASD IO threads 614a~n are common threads shared by all range threads 612a~n. After the DASD IO threads 614a~n obtain the requested data, the data is placed in the output buffer 646 to serialize the data into a blob so that the replica object threads 616a~n (i.e., the data upload pipeline stage) upload it to the object store located in the target region. Each object thread selects an upload job that can include a portion of all the data to be uploaded, and all object threads execute the uploads in parallel.
[0105] FIG. 7 is a diagram showing a hierarchical structure in the FSS data plane according to an embodiment. In FIG. 7, the replicator fleet 710 includes four layers: a job layer 712, a differential generator client 714, an encryption / DASD IO 716, and an object 718. The replicator fleet 710 is a single process responsible for communicating information with the storage fleet 720, the KMS 730, and the object storage 740. In one embodiment, the job layer 712 polls the SDB 704 for a job 706 that is queued as either an upload job or a download job. The replicator fleet 710 includes VMs (or threads) that select enqueue replication jobs up to the maximum capacity. A replicator thread may own a part of a replication job, but coordinates with another replicator thread that owns the remaining part of the same replication job to complete the entire replication job simultaneously. The replication jobs executed by the replicator fleet 710 are restartable in that if a replicator thread fails during replication, another replicator thread can take over and continue from the last successful point to complete the job that the failed replicator thread originally owned. If a strayed replicator thread (e.g., a replicator thread that fails and restarts) competes with another replicator thread, the FSS can avoid the conflict by using a mechanism called a generation number to cause both replicator threads to update different records.
[0106] The differential generator client layer 714 performs a B-tree scan by accessing the differential generator server 724 in which the B-tree in the storage fleet 720 exists. The encryption / DASD IO layer 716 assumes the roles of security and storage access. After the B-tree scan, the replicator fleet 710 may request IO access via the encryption / DASD IO layer 716 to access the DASD range 722 of the file data associated with the differences identified during the B-tree scan. The replicator fleet 710 and the storage fleet 720 both periodically update the status of the control API 702 (e.g., checkpoint and lease of the replicator fleet 710) via the SDB 704, enabling the control API 702 to trigger an alarm or execute an action if necessary.
[0107] During the inter-region replication process, the encryption / DASD IO layer 716 exchanges information with the KMS and the FSK fleet 730 on the target side to create a session key (or snapshot encryption key), and uses the FSK for encryption and decryption of the session key. Finally, the object layer 718 is responsible for uploading differential and file data from the source file system to the object store 740 and downloading them from the object store 740 to the target file system.
[0108] The data plane of the FSS is responsible for differential generation. The data plane stores FSS data using a B-tree, which includes various types of key-value pairs including, but not limited to, a leader block, a superblock, an iNode, a file name key, a cookie map (cookies associated with directory entries), and a block map (also called FMAP in the case of file content data).
[0109] These B-tree keys are processed together by replicators and difference generators within the data plane. An algorithm for calculating the pairs of keys and values (i.e., part of the difference) changed between two specific snapshots within the file system continuously reads the keys, returns the keys to the replicator using the transaction budget, and finally ensures that the transaction is committed to obtain a consistent pair of keys and values for processing.
[0110] In other embodiments, difference generation and calculation can be scalable. A scalable approach can calculate the difference (i.e., the change in the pairs of keys and values) between two snapshots by utilizing multiple threads to divide the B-tree into many key ranges. A pool of threads (i.e., difference generators) can execute a scan of the B-tree (i.e., traverse the B-tree) and calculate the differences in parallel.
[0111] FIG. 8 shows a simplified exemplary binary large object (BLOB) format according to an embodiment. A blob is a data type for storing information (e.g., binary data) in a database. Blobs are generated by a source region during replication and uploaded to an object store. The target region needs to download and apply the blobs. Blobs and objects can be used interchangeably depending on the context.
[0112] During the B-tree scan, when the differential generator encounters the iNode of a specific file (i.e., data content) and its block map (also called FMAP, data associated with the B-tree key), the differential generator cooperates with the replicator to traverse all pages within the blocks (FMAP blocks) within the DASD range pointed to by the FMAP, reads them into the data buffer, decrypts the data using the local encryption file key, puts it into the output buffer, and serializes it into a blob for the replicator to upload to the object store. In other words, the differential generator needs to collect all FMAPs of the identified differences in order to obtain all data related to the differences between two snapshots.
[0113] Snapshot differences stored in the object store may span multiple blobs (or objects if stored in the object store). The blob format of these blobs includes a key, a value, and, if present, data associated with the key. For example, in Figure 8, the snapshot difference 800 includes at least three blobs 802, 804, and 806. The first blob 802 includes a prefix 810 indicating the types of the key and value, the length of the key, and the length of the value, followed by a key 812 (key 1) and a value 814 (value 1). The second blob 804 includes a prefix 820 (types of the key and value, length of the key, and length of the value), a key 822 (key 2), a value 824 (value 2), a data length 826, and data 828 (data 2). In the prefix 820 of this second blob 804, since this blob includes additional data 828 associated with the key 822, the types of the key and value are fmap. The third blob 830 includes a format similar to that of the first blob 810, for example, a prefix 830, a key 832 (key 3), and a value 834 (value 3).
[0114] The data is decrypted, collected, and then written to a blob. All processes are executed in parallel. Multiple blobs can be processed and updated simultaneously. When all processes are complete, the data is written in blob format (shown in Figure 8), and then can be uploaded to an object store in the format (shown in Figure 9) or path name.
[0115] Figure 9 shows an exemplary replication bucket format according to an embodiment. A "bucket" can refer to a container that stores objects in a compartment within an object storage namespace. In one embodiment, a bucket is used by a source replicator to store data protected using server-side encryption (SSE) technology, and is also used by a target replicator to download changes and apply them to a snapshot. Replication data for all file systems in a target region can share a bucket within that region.
[0116] The data layout of a bucket in the object store has a directory structure that includes, but is not limited to, a file system ID (e.g., Oracle Cloud ID), a delta including a start snapshot number and an end snapshot number, a manifest that describes the content of the information within the object's layout, and blobs. For example, the bucket in FIG. 9 includes two objects 910 and 930. The first object 910 includes two deltas 912 and 920. This object starts with a path name 911 (e.g., ocid1.filesystem.oc1.iad...) that uses the source file system ID as a prefix, followed by a first delta 912 generated from snapshot 1 and snapshot 2, and a second snapshot 920 generated from snapshot 2 and snapshot 3. Each delta includes one or more blobs that represent the content of that delta. The first delta 912 stores two blobs 914 and 916 in the order of generation. The second delta 920 includes only one blob 922. Each delta also includes a manifest that describes the content of the information within the layout of this delta, e.g., manifest 918 of the first delta 912 and manifest 924 of the second delta 920. The manifest within the bucket is content that describes the delta, such as the file system number and the snapshot range. The manifest can be a master manifest, a range manifest, or a checkpoint manifest depending on the stage of the replication process.
[0117] The second object 930 also includes two deltas 932 and 940 in a similar format starting with path name 931. The two objects 910 and 930 within the bucket come from different source regions, namely IAD for object 910 and PHX for object 930. After the blobs are applied, the corresponding information within the layout can be removed to reduce space utilization.
[0118] The final manifest object (i.e., the checkpoint manifest, CM file) is uploaded from the source region to the object store, and the source file system indicates to the target region that it has completed uploading the snapshot delta of a specific object. The source CP transmits this event to the target CP, and the target CP can notify the target DP via the SDB to trigger the download process of that object by the target replicator.
[0119] The control plane within the source region or the target region coordinates all of the replication workflow and drives the replication of data. The control plane performs the following functions: (1) creates the underlying system snapshot for creating deltas, (2) determines when such snapshots need to be created, (3) initiates replication based on the snapshots, (4) monitors the replication, (5) triggers the download of deltas by the secondary (or target side), and (6) indicates to the primary (or source) side that the snapshot has reached the secondary.
[0120] The file system has several operations for processing its resources, including but not limited to creating, reading, updating, and deleting (CRUD). These operations are typically synchronized within the same region, starting a workflow when the file system receives an HTTPS request from the API server, making changes in the backend to create a resource, and returning a response to the customer. Resources are split into a source region and a target region. The state is maintained for the same resource between the source region and the target region. Therefore, there is asynchronous communication between the source region and the target region. Customers can interact with the source region to create or update resources, and these creations or updates can be automatically reflected in secondary or auxiliary resources within the target region. The state machine in the control plane also targets recovery in many aspects, including but not limited to fleet failures, key management failures, disk failures, and object failures.
[0121] Regarding the application programming interfaces (APIs) within the control plane, there are various APIs for users to configure replication. The control API for any new resource only functions within the region where the object was created. In the target file system, a field named "IsTargetable" can be set in the API to ensure that the target file system receiving replication cannot be accidentally used by consumers. In other words, setting this field to false means that consumers can view the target file system, but no one can export the target file system or access any data within the live system. Export is not a read-only permission but a read / write permission for export, so any export has the potential to change the data. Therefore, during the replication process, exports are not permitted to prevent any changes to the target file system. Consumers can only access data within old snapshots that have already been replicated. Any newly created or replicated file system can have this field set to true. The reason is that the target can only obtain data from a single source. Otherwise, conflicts may occur when data is written or deleted. The system needs to know whether the target file system in use is already part of some replication. Setting the "IsTargetable" field to "true" means that replication is not in progress, and setting it to "false" means that the target file system cannot be used.
[0122] Regarding inter-region communication between components of the control plane, the primary resource on the source file system is called an application, and the auxiliary (or secondary) source on the target file system is called an application target. When source and target objects are created, they have a single replication relationship. Both objects can be updated only from the source side, including changes to compartments, editing of details, or deletion. If the user wants to delete the target side, the replication itself can be deleted. In the case of a planned failover, it is possible to delete the source side, and both the source side and the target replication are deleted. In the case of an unplanned failover, the source side is unavailable, and thus only the target replication can be deleted. In other words, there are two resources for a single replication, and those resources should be kept in a synchronized state. There are various workflows for updating metadata on both the source and target sides. Additionally, inter-region APIs for retries, fault handling, and failovers are also part of the inter-region communication process.
[0123] When creating the necessary security and other related artifacts, the source uploads the security and artifacts to the object store, starts a job at the target (i.e., notifies the target that the job is available), and the target can start downloading the artifacts (e.g., snapshots or deltas). Subsequently, the target continues to search for an end-of-file marker (also referred to herein as a checkpoint manifest (CM) file) within the object store. The CM file is used as a mechanism for the source side and the target side to communicate the completion of the upload of the object during the replication process. At every checkpoint, the source side uploads this CM file containing information such as the number of blobs uploaded up to this checkpoint, enabling the target side to download this number of blobs and apply them to the current snapshot. This CM file is a mechanism for the source side to communicate to the target side that the upload of the object to the object store is complete, allowing the target to start working on that object. In other words, the target continues to download until there are no more objects in the object storage. Therefore, this approach enables simultaneous processing on both the source side and the target side.
[0124] FIG. 10 is a flowchart showing a state machine for simultaneous source upload and target download according to an embodiment. As described previously, both the source file system and the target file system can execute replication simultaneously and thus can have their own state machines. In one embodiment, each file system can have its own state machine while sharing some common job-level states. In FIG. 10, the source file system has states 1030-1034 for session key generation and transfer in addition to states 1002-1018 for performing data upload. The target file system has states 1050-1068 related to data download. The session key can be generated at any time within the source file system while the differences are being uploaded to the object storage. Thus, the session key transfer has its own state sequence 1030-1034. In FIG. 10, the target file system cannot start the replication download process (i.e., Ready_to_Reconcile state 1050) until it receives an indication that at least an object has been uploaded to the object storage by the source file system (i.e., Mainfest_Copied state 1014) and that it is ready to download the session key (i.e., Copied_DTK state 1034).
[0125] In the source file system, multiple functional blocks such as a snapshot generator, a control API, and a delta monitor are part of the CP. The replicator fleet is part of the DP. The snapshot generator is responsible for periodically generating snapshots. The delta monitor periodically monitors the progress of the replicator in replication-related tasks, including the creation of snapshots and the replication schedule. When the delta monitor detects that the replicator has completed a replication job, it transitions the state to a copied state on the source side (e.g., the Manifest_Copied state 1014) or a replicated state on the target side (e.g., the Replicated state 1058). In certain embodiments, multiple file systems can simultaneously perform replication from a source region to a target region.
[0126] Referring to FIG. 10, in certain embodiments, in the source file system, in the simultaneous mode state machine, after creating a snapshot signal to the delta monitor indicating that a snapshot has been generated, the snapshot generator. The delta monitor that executes the CP replication state (CpRpSt) workflow is responsible for initiating the upload of snapshot metadata to the object store on the target side. The snapshot metadata may include the type of snapshot, snapshot identification information, the time of the snapshot, etc. The CpRpSt workflow sets the Ready_to_Copy_Metadata state 1002 for the replicator fleet to start copying the metadata. When the replicator obtains a replication job, it creates a copy of the snapshot metadata (i.e., the Snapshot_Metadata_Copying state 1004) and uploads those copies to the object store. When all replicators have completed the upload of the snapshot metadata, the state is set to the Snapshot_Metadata_Copied state 1006. Thereafter, the CpRpSt workflow continues to poll the source SDB for the session key.
[0127] Here, the CpRtSt workflow returns control to the differential monitor to monitor the differential upload process that transitions to the Ready_to_Copy state 1008 indicating that the differential calculation is scheduled. Next, the source CP API sends a request to the replicator to start the next stage of replication by uploading the differential and creating a copy of the manifest. The replicator that selects the replication job can start creating a copy of the manifest (i.e., the Mainfest_Copying state 1010). When the source file system completes the copy of the manifest, it transitions to the Manifest_Copied state 1014 and at the same time notifies the target file system that it can start the internal state (the Ready_to_Reconcile state 1050).
[0128] As described above, the session key can be generated by the source file system during the upload of data. The replicator of the source file system communicates with the target KMS vault to obtain the master key that can be provided by the customer and creates a session key (referred to herein as the differential encryption key or DEK). Next, the replicator encrypts the session key using the local file system key (FSK: file system key) (here, it becomes the encrypted DEK, also referred to herein as the differential transfer key (DTK)). Then, the DTK is stored in the SDB within the source region and reused by the replicator thread during the replication cycle. The state machine transitions to the Ready_to_Copy_DTK state 1030.
[0129] The source file system transfers the resource identification information of DTK and KMS to the target API. Then, the target API stores those resource identification information in the SDB within the target region. During this transfer process, the state machine is set to the Copying_DTK state 1032. When the CpRpSt workflow in the source file system finishes polling the source SDB for the session key, the target file system downloads the session key (DTK) and sends a notification to the target side, informing that it is ready to decrypt the downloaded differences for application using that session key. Then, the state machine migrates to the Copied_DTK state 1034. The replicator on the target side retrieves the DTK from the SDB and requests the KMS API to decrypt the DTK into the plaintext DEK (i.e., the decrypted session key).
[0130] When the source file system completes the upload of data for a specific replication cycle including session key transfer, the difference monitor notifies the target control API of the status such as validation information and migrates to the X-region_Copied_Done state 1016. This can occur before the target file system completes the download and application of the data. The source file system further cleans up the memory and removes all keys. Then, the source file system migrates to the Awaiting_Target_Response state 1018, waits for a response from the target file system, and starts a new replication cycle.
[0131] As described above, the target file system cannot start the replication download process until it receives an indication that at least the object has been uploaded to the object storage by the source file system (i.e., the Mainfest_Copied state 1014), and that it is ready to download the session key (i.e., the Copied_DTK state 1034). When these two conditions are met, the state machine transitions to the Ready_To_Reconcile state 1050. Next, in the Reconciling state 1052, the target file system starts an adjustment process with the source side, such as synchronizing snapshots of the source file system and the target file system, obtaining snapshots, and generating statistical values, and also performs some internal CP management operations, including communication within the target file system between the delta monitor and the CP API.
[0132] After the adjustment process is completed, the replication job is passed to the target replicator (i.e., the Ready_to_Replicate state 1054). The target replicator monitors the checkpoint manifest (CM) file uploaded by the source file system. The CM file is marked by the target. Then, the target replicator thread starts to download the manifest and apply the downloaded and decrypted deltas (i.e., the Replicating state 1056). The target replicator thread also reads the FMAP data blocks from the blobs downloaded from the object store, communicates with the local FSK service to obtain the file system key FSK, and the FSK is used to re-encrypt each FMAP data block and store it in local storage.
[0133] When the source file system finishes uploading data, it updates the final CM file by setting the end-of-file (eof) field to true and uploads it to the object store. As soon as the target file system detects this final CM file, it ends the download of the blobs and applies them, and the state machine transitions to the Replicated state 1058.
[0134] After the target file system applies all the differences (or blobs), it continues to download the snapshot metadata from the object store and inputs the information of the source file system's snapshot into the target file system's snapshot (i.e., the Snapshot_metadata_Populating state 1060). When the target file system's snapshot is input, the state machine transitions to the Snapshot_Metadata_Populated state 1062.
[0135] In the Snapshot_Deleting state 1064, the target file system deletes all the blobs in the object store for the blobs that were downloaded and applied to the latest snapshot. Then, the target control API notifies the target difference monitor when the blobs in the object store are deleted and proceeds to the Snapshot_Deleted state 1066. The target file system further cleans up the memory and removes all the keys. The FSS service also releases the KMS key.
[0136] When the target DP finishes applying the differences and cleaning up, it uses the target control API to verify the validity regarding the status of the source file system and whether it has received an X-region_Copied_Done notification from the source file system. If the notification has been received, the target difference monitor transitions to the X-region DONE state 1068 and sends an X-region DONE notification to the source file system. In some embodiments, the target file system checks whether the end of file exists for all key ranges and all upload processing threads because all objects uploaded to the object store have special markers such as end-of-file markers in the CM file, so as to detect whether the source file system has completed the upload.
[0137] Referring again to the state machine of the source file system, while the source file system is in the Awaiting_Target_Response state 1018, it checks whether the status of the target CP has changed to completed, indicating that all differences downloaded by the target have been applied and the file data has been stored locally. If the status of the target CP has changed to completed, this ends the replication cycle.
[0138] The source side and the target side operate asynchronously. When the source file system completes the replication upload, it notifies the target control API of the X-region_Copied_Done notification. Then, when the target file system completes the replication process, the difference monitor target communicates in the reverse direction with the source control API using the X-region DONE notification. The source file system returns to the Ready_to_Copy_Metadata state 1002 and starts another replication cycle.
[0139] FIG. 11 is an exemplary flow diagram showing the exchange of information between the data plane and the control plane within a source region according to an embodiment. The data plane components and the control plane components communicate with each other using a shared database (SDB), such as 1106. The SDB is a key-value store that both the control plane components and the data plane components can read from and write to. The data plane components include a replicator and a delta generator. The exchange of information between the components within source region A1101 and target region B1102 is also shown.
[0140] In FIG. 11, at step S1, the source control plane (CPa) 1103 requests the object store within the target region B (OSb) 1112 to create a bucket. At step S2, the source replicator (REPLICATORa) 1108 periodically updates the heartbeat status to the source SDB (SDBa) 1106. The heartbeat is a concept used to track the progress of replication executed by the replicator. The heartbeat can use a mechanism called lease, where the heartbeat is continuously updated each time the replicator works on a job, enabling the control plane to recognize the overall release information. For example, the byte count moves continuously on the job. If the replicator cannot function properly, the heartbeat may become stale, and then another replicator can detect and take over to continue working on the remaining jobs. Therefore, if the system crashes midway, the system can start exactly from the last point based on the checkpoint mechanism. The checkpoint helps the system know where the last point of progress is and enables it to continue from that point without re-executing the entire job.
[0141] In step S3, CPa1103 further requests the File System Service Workflow (FSW_CPa) 1104 to create snapshots periodically. In step S4, FSW_CPa1104 notifies CPa1103 about the new snapshot. In step S5, next, CPa1103 stores the snapshot information in SDBa1106. In step S6, REPLICATORa1108 polls SDB1106 for any changes to the existing snapshots. If a change is detected, in step S7, it retrieves the job specification. In step S8, when REPLICATORa1108 detects a change to the snapshot, this initiates the replication process. In step S8, REPLICATORa1108 provides information about two snapshots (SNa and SNb) including the changes between the snapshots to the Differencer (DGa) 1110. In step S9, REPLICATORa1108 puts work item information such as the number of key ranges into SDBa1106. In step 10, REPLICATORa1108 checks the replication job queue in SDBa1106 to obtain work items. In step S11, it assigns those work items to the Differencer (DGa) 1110 to scan the B-tree keys of the snapshot (i.e., traverse the B-tree) and calculate the differences and the corresponding key-value pairs. In step 12, REPLICATORa1108 decrypts the file data associated with the identified B-tree keys and packs them into blobs together with the key-value pairs. In step 13, REPLICATORa1108 encrypts the blobs using the session key and uploads them as objects to OSb1112. In step S14, REPLICATORa executes a checkpoint and stores the checkpoint record in SDBa1106. This replication process repeats (as a loop) until all differences are identified and the data is uploaded to OSb1112.In step S15, REPLICATORa1108 then notifies SDBa1106 of the details of the replication job, and this detail is then passed to CPa1103 in step S16 and further relayed to CPb1114 as the final CM file in step S17. In step S18, CPb1114 stores the job details in SDBb1116.
[0142] The exchange of information between the data plane and the control plane within target region B is similar. At the end of applying the delta to the target file system, the control plane within target region B notifies the control plane within source region A that the snapshot has been successfully applied. Thereby, the control plane within source region A can start over using the new snapshot.
[0143] Authentication is performed on all components. There is an authentication mechanism that uses the replication ID and the file system number, from the replicator to the file system key (FSK). The key can be given to the replicator only if the replicator provides appropriate content. Thus, the authentication mechanism can prevent fraudsters from obtaining the decryption key. Other security mechanisms include blocking network ports. A component called the file system key server (FSKS) is a gatekeeper for properly checking the requester by checking metadata such as the job the requester is running and other information. For example, assume that the replicator is attempting to request the key to the file system. In that case, FSKS can check whether the replicator is associated with a specific job (e.g., whether the replication is actually associated with that file system) to confirm the validity of the requester.
[0144] Availability addresses situations where a machine can automatically restart after going down or where services remain available while software deployment is in progress. For example, since all replicators are stateless, losing a replicator is transparent to the customer because another replicator can pick up and continue the job's work. The job's state is maintained not locally but in a shared database and other reliable locations. The shared database is a service such as the database used by the control plane to maintain information about the file system and is based on a B-tree.
[0145] The system has thousands of storage nodes that enable any storage node to perform differential replication, so storage availability in the FSS of the present disclosure is high. By utilizing many machines that can take over from each other in case of some failure, the availability of the control plane is increased. For example, the progress of replication is not simply hindered by the failure of a single control plane. Thus, there is no single point of failure. Network access availability uses congestion management, including various types of throttling, to prevent source nodes from becoming overloaded.
[0146] Replication is durable by writing the replication state to the shared database and by utilizing checkpointing where the replicator is stateless. The replication process should be idempotent. Idempotency can refer to deterministic reapplications where, in case of an operation failure, retrying the same operation, for example, using the same key, upload process, or scan process, should result in the same outcome.
[0147] Operations within multiple regions should be equivalent. In the control plane, the actions taken should be stored. For example, in the case of repeated HTTP requests, an equivalence cache can be useful for remembering that a particular operation has been performed and that it is the same operation. For example, in the data plane, when a block is allocated, the block and the file map key of the file system are written together. Thus, if the block is allocated again, the block can be identified. If the block is sealed, the write operation fails. The equivalence mechanism can know that the block has been sealed in the past and the write operation does not need to be redone. In yet another example, the equivalence mechanism stores a chain of steps that need to be performed for the processing of a particular key and value. In other words, the equivalence mechanism enables all operations to be checked to ensure they are in the correct state. Thus, the system can simply proceed to the next step without repetition.
[0148] Atomic replay enables the application of differences to start as soon as the first difference object reaches the object store when a snapshot is rolled back, for example, when going back from snapshot 10 to snapshot 5. To make the replay atomic, the entire difference needs to be maintained in the object store before the differences can be applied.
[0149] Regarding the expansion of replicators, the FSS of the present disclosure enables adding the number of replication machines (e.g., replicator virtual machines (“VMs”)) required to support many file systems. The number of replicators can be dynamically increased or decreased by considering the bandwidth requirements and availability of resources. Regarding the expansion of storage, thousands of storages can be used to parallelize the process and improve the working speed. Regarding the inter-region bandwidth, the bandwidth allocation is automatically adjusted, such as by adjusting all the inter-region bandwidths by grasping the increase in latency and reducing the speed of requests, to ensure that each workload is not overused or does not exceed the predefined throughput limit. All replicator processors (or threads) have this function.
[0150] In the expansion of checkpoint storage, the uploader and downloader checkpoint the progress to persistent storage, and the shared storage is used as a work queue for dividing key ranges. If the checkpoint workload overly burdens the shared database, the checkpoint storage function can be added to the differential generator for the purpose of expansion. The current workload of the shared database may consume less than 10 IOPs.
[0151] FIG. 12 is a schematic diagram showing a failback mode according to an embodiment. The failback mode enables restoring the primary side / source side so as to become the primary / source again before failover. As shown in FIG. 12, the primary AD 1202 includes the source file system 1206, and the secondary AD 1204 includes the target file system 1208. The secondary AD 1204 may exist in the same region or a different region from the region of the primary AD 1202.
[0152] In FIG. 12, snapshot 1 1220 and snapshot 2 1222 in the source file system 1206 exist before the failover due to the power outage event. Similarly, snapshot 1 1240 and snapshot 2 1242 of the target file system 1208 exist before the failover. When a power outage occurs in snapshot 3 1224 in the primary AD 1202, the FSS performs an unplanned failover 1250, and snapshot 3 1224 in the source file system 1206 is replicated to the target file system 1208 and becomes the new snapshot 3 1224. After the target file system 1208 starts operating, the customer can make changes to create snapshot 4 1246 for the target file system 1208.
[0153] If the customer decides to reuse the source file system again, the FSS service may execute a failback. When executing the failback, the user has two options: (1) the last point in time in the source file system before the trigger event 1251, or (2) the latest changes in the target file system 1252.
[0154] In the case of the first option, the user can resume from the last point in the source file system 1206 prior to the trigger event (i.e., snapshot 3 1224). In other words, since snapshot 3 1224 has previously successfully failed over to the target file system 1208, it becomes the snapshot for use after failback. To execute the failback 1251, the state of the source file system 1206 is changed to inaccessible. Next, the FSS service identifies snapshot 3 1224, the last point in the source file system 1206 before the failover was successful. The FSS may execute a clone of snapshot 3 1224 (i.e., a replica within the same region) in the primary AD 1202. Now, the primary AD 1202 returns to its initial settings prior to the power outage, and the user can reuse the source file system 1206 again. Since snapshot 3 1224 already exists in the file system being used, no data transfer from the secondary AD 1204 to the primary AD 1202 is required.
[0155] In the case of the second option, the user wants to reuse the source file system with the latest changes in the target file system 1208. In other words, since snapshot 4 1246 in the target file system 1208 was the latest change in the target file system 1208, it becomes the snapshot for use after failback. The failback process 1252 for this option includes reverse replication (i.e., reversing the roles of the source file system and the target file system for the replication process), and the FSS performs the following steps.
[0156] Step 1. The state of the source file system 1206 is changed to inaccessible. Step 2. Next, the FSS service identifies the latest snapshot in the successfully replicated target file system 1208, e.g., snapshot 3 1244.
[0157] In step 3, the FSS service also detects the corresponding snapshot 3 1224 in the source file system 1206 and performs a clone (i.e., a replication within the same region).
[0158] In step 4, the FSS service initiates reverse replication 1252 in a process similar to that described in relation to FIG. 4, but in the reverse direction. In other words, both the source file system 1206 and the target file system 1208 need to be synchronized, after which the target file system 1208 can upload the differences to the object store in the primary AD 1202. The source file system 1206 can download the differences from the object store, complete the application to snapshot 3 1224, and create a new snapshot 4 1226.
[0159] Here, the primary AD 1202 returns to its initial settings before the power outage, and the user can reuse the source file system 1206 again without transferring the data that already exists in both the source file system 1206 and the target file system 1208, such as snapshots 1 to 3 (1220 to 1224) in the source file system 1206. This saves time and prevents unnecessary bandwidth.
[0160] Hierarchical Key Management Architecture For the purposes of this application, in some examples, the KMS Bolt is a key management service that stores and manages master encryption keys and secrets for secure access to resources. The KMS Bolt enables the secure storage of master encryption keys (or master keys) and secrets, which otherwise might be stored in configuration files or code. The Bolt service is integrated with the FSS and supports the use of keys managed by the customer. Master encryption keys obtained from the KMS Bolt can be used to generate other encryption keys, such as file system keys (FSK) and / or differential encryption keys (DEK, also referred to herein as session keys), for encrypting / decrypting other keys or data. The FSK and DEK may be referred to herein as security keys for security purposes. A key version may be automatically assigned to each Bolt's master key. The Bolt service supports keys as a cloud infrastructure and assigns globally unique resource identification information (e.g., Oracle cloud ID (OCID)) to each key.
[0161] In one embodiment, the file storage service (FSS) disclosed in this disclosure utilizes cross-region replication of a three-tier key architecture. During the replication process, three different keys may be included: a source file system key (FSKa), a session key, and a target file system key (FSKb). Within each source region and target region, the file system keys are also structured in layers to encrypt local files using the master key from the customer to securely store the file system keys.
[0162] In some embodiments, the source file system has three keys: a source file key (FKa) per file, a source file system key (FSKa), and a master key (based on a key managed by the customer). The source file system includes a plurality of files each having its own file key FKa for encryption and decryption. Each FKa is encrypted by the FSKa and stored in the iNode of the file B tree. The FSKa, which is unique for each file system, is then generated and encrypted / decrypted by a file master key (sometimes referred to herein as the source-side master key) provided by the customer and stored in the metadata of the data plane. A similar scheme is used within the target region. The target file system has three keys: a target file key (FKb) per file, a target file system key (FSKb), and a master key (based on a key managed by the customer), which may sometimes be referred to herein as the target-side master key. In some embodiments, the master key of the source file system is different from the master key of the target file system.
[0163] During inter-region replication, the three keys mentioned with respect to inter-region replication can be used to decrypt and encrypt data during transfer between regions. The FSS defines a session, and each session is one data cycle (or one replication cycle). A session key is created to transfer data only within that session. First, the source region decrypts the file data using FSKa (via FKa) during difference generation. Second, the data (e.g., blobs and manifest files) is encrypted using the session key within the source region and uploaded for storage in the object store. Third, the session key is transferred from the CP of the source region to the CP of the target region using the Key Management Service (KMS). Fourth, the target region downloads the data from the object store and decrypts it using the session key. Finally, the target region re-encrypts the data using FSKb (via FKb) and stores it locally in the B-tree.
[0164] Existing key management for file system replication uses one encryption key for two or more regions and may transfer data between different regions. This increases the risk of compromising the integrity of the data across regions when the security of the local region is breached. Therefore, different approaches are needed to address these issues related to security risks and the like.
[0165] The three - layer key architecture disclosed in this disclosure prevents the disclosure of FSKa or customer - managed keys (or master keys) outside the source region by encrypting data (e.g., blobs and manifests) using a session key before uploading to the object store. Further, the key infrastructure of FSS has a unique key (e.g., FSKa) for each file system instead of sharing file - system - level keys among multiple file systems within one or more regions to add protection within each file system. Finally, the session key is transferred via the CP API (with the support of KMS before being reused) instead of being transferred with the data via the object store, which is the third level of protection.
[0166] In some embodiments, all keys involved in the replication process are stored in an encrypted form. For example, customer - managed keys (or service keys) are securely stored in the KMS vault. In some embodiments, FSKa and FKa are encrypted and stored locally in the source B - tree, while FSKb and FKB are encrypted and stored locally in the target B - tree. To further ensure security, an authentication process is performed in both the source region and the target region to check the identity of the key requester. For example, the control plane (including the file - system key server (FSK service)) can check the replication ID (i.e., the replication job executed by the key requester) and the file - system number (i.e., the file system to which the key requester belongs) provided by the key requester before granting the file - system key (FSK).
[0167] When creating a session key in the source region for replication, for the purpose of placing the KMS bolt in the target region, if an unexpected disaster occurs in the source region, the target file system can still download the data (or objects) in the object store and decrypt the application's data. Therefore, there is an advantage in holding the session key within the target region. In certain embodiments, the session key is temporarily stored in the source region to encrypt the data uploaded to the object store for security purposes.
[0168] Figure 13 is a flowchart showing key management for cross-region replication according to an embodiment. In Figure 13, the components within the source region 1301 involved in key management include a snapshot monitor 1310, a shared database (SDB) 1312, a replicator 1314, and a file system key (FSK) service 1316. The components within the target region 1302 involved in key management include a key management service (KMS) 1350, an application programming interface (API) 1330, an SDB 1332, an object storage 1360, a replicator 1334, and an FSK service 1336. When the replication process is initiated, the control plane of the source file system can start the key management process shown in Figure 13 in cooperation with the data plane.
[0169] In step S1 of FIG. 13, the source replicator 1314 communicates with the target KMS bolt 1350 to obtain a master key (which may be referred to herein as an inter-region master key and is different from the master key of the aforementioned source FSK), and generates a new session key (i.e., DEK, plain text key) for this replication cycle. In certain situations, the replicator (or replicator thread, also referred to herein as an alternative replicator) can take over a replication job from another failed replicator after the failed replicator previously generated a session key and resume replication. In certain embodiments, the alternative replicator first checks the source SDB 1312 (see step S3 described below) and obtains the locally stored session key (i.e., decrypts the existing DTK). In some rare situations, if the source SDB loses the record of the session key, the alternative replicator needs to call the KMS API using the CGO code, communicate with the target KMS bolt by providing the KMS resource ID, obtain the master key to generate a new DEK, and resume the replication job (i.e., encrypt the remaining part of the previously incomplete difference).
[0170] In step S2, the source replicator 1314 communicates with the FSK service 1316 to obtain the source file system key (FSKa). The FSK service 1316 also performs authentication on the source replicator 1314, which is the key requester, by checking the replication ID and the file system number (FSNum) to confirm that the key requester is actually associated with a specific file system and a specific replication process. In step S3, the source replicator 1314 encrypts the session key DEK using FSKa (the encrypted session key is referred to as the differential transfer key (DTK) in this specification), stores the encrypted session key DEK in the source SDB 1312 for later use by other replicator threads (including alternative threads) during the replication cycle, and avoids the need to communicate with the KMS again via an inter-region call. In step S4, the source replicator 1314 locally caches the FSKa. In step S5, in addition to storing the DTK in the source SDB 1312, the source replicator 1314 also locally caches the DTK. In step S6, during the replication cycle, each replicator thread that processes a replication job (i.e., a part of the entire replication) checks the source SDB 1312 and re-uses the session key (i.e., DEK) by decrypting the encrypted session key (i.e., DTK) using the locally cached FSKa. The source replicator 1314 further decrypts the local FKa using the FSKa, and this local FKa is then used to decrypt the file data collected during the B-tree scan. So far, the above steps have been executed within the source region 1301.
[0171] From step S7 to S9, the key management process enters the inter-region phase. In step S7, in one embodiment, the source replicator 1314 then decrypts the data using FSKa (as described above), re-encrypts the data (e.g., keys and values of the B-tree, file data, manifest files, and metadata) using the session key, and uploads it to the object store 1360 as an object within the target region 1302. In some embodiments, some of the data (e.g., internal metadata for housekeeping purposes) need not be encrypted.
[0172] In step S8, the source replicator 1314 notifies the state machine to change to the Manifest_Copied state by marking the records in the source SDB 1312. In step S9, the snapshot monitor in the source region 1301 transfers the encrypted session key (i.e., DTK) and the KMS key resource ID to the target API 1330. The flow diagram shows that the step of transferring the data object to the object store (S7) occurs before the step of transferring the DTK to the target API (S9), but the source region need not wait for the transfer of all data objects to complete before starting step S9. The earlier the source file system can transfer the DTK to the target file system, the earlier the target file system can download and decrypt the data.
[0173] In step S10, the target API 1330 stores the received encrypted session key (i.e., DTK) and the KMS key resource ID in the target SDB 1332 for reuse during the replication cycle. In step S11, the target replicator 1334 (or replicator thread) retrieves the encrypted session key and the KMS key resource ID from the target SDB 1332. The target replicator 1334 further checks whether the expiration date of the DTK has passed. If the expiration date has passed, the target replicator 1334 triggers an alarm and reports a hard failure. Otherwise, the target replicator 1334 proceeds to the next step. In step S12, the target replicator 1334 sends the encrypted session key and the KMS key resource ID to the KMS vault 1350 and requests that the DTK be decrypted (i.e., made into a DEK) for use. The KMS vault 1350 also authenticates the target replicator 1334, which is the key requester, and verifies that the DTK is for the session associated with this particular replication cycle. In step S13, the target replicator 1334 locally caches the session key (i.e., DEK). In step S14, the target replicator 1334 downloads data from the object store 1360 and decrypts the data using the session key. In step S15, the target replicator 1334 communicates with the target FSK service 1336 to obtain the target file system key (FSKb). The target FSK service 1336 also performs authentication on the source replicator 1334, which is the key requester, by checking the replication ID and the file system number to verify that the key requester is actually associated with a specific file system and a specific replication process. In step S16, the target replicator 1334 locally caches the FSKb for reuse.In step S17, the target replicator 1334 uses FSKb to decrypt the target file key (FKb), and the target file key (FKb) is used to re-encrypt the data after applying the data to the latest snapshot, and then the data is used to be locally stored in the B-tree. In step S18, the target replicator 1334 notifies the state machine to change to the Replicated state by marking the records in the target SDB 1332.
[0174] When the source side completes the upload (which can occur before the target download process), the source side cleans up the memory and removes all keys. When the target completes the application of the differences to the latest snapshot, it similarly cleans up the memory and removes all keys. The FSS service also releases the KMS master key. Thus, the keys are only valid for the current replication cycle.
[0175] FIG. 14 is a flowchart showing the steps of session key generation according to an embodiment. In FIG. 14, the components within the source region 1401 involved in session key (i.e., DEK) generation can be the source SDB 1410, the source control plane (CP) API 1412, the source replicator 1414, and the source FSK service 1416. The components within the target region 1402 include the target KMS bolt 1418.
[0176] In one embodiment, in step S1, the source replicator 1414 receives a request from the CP, prepares for session key generation for inter-region replication, and performs initialization for communicating with the KMS and the object store within the target region. In step S2, the source replicator 1414 communicates with the target KMS 1418 to list all the bolts within a specific compartment. In step S3, since there is a prefix name unique to the key and the bolt, the source replicator 1414 obtains all the keys of the specific bolts whose bolt prefix matches the desired prefix. In step S4, the source replicator 1414 caches the master key resource ID for all future AD uses. Since the master key bolt and the key are unique for each AD, caching the resource IDs of other ADs within the same source region while processing the keys of this AD can help reduce future communications with the KMS. In step S5, the source replicator 1414 generates a session key (i.e., DEK) using additional authentication data (AAD) as the snapshot resource ID. The AAD is metadata provided to the KMS during encryption / decryption and helps the FSS confirm the validity that a specific session key is associated with the snapshot resource. This provides an additional layer of protection. In step S6, as described above, all keys can be stored in an encrypted form. Therefore, the source replicator 1414 encrypts the session key using the source file system key FSK1316 (not shown) and stores the encrypted session key (i.e., DTK) and the KMS source ID in the source SDB 1410 for reuse by other replicator threads within the current replication cycle. The key and its KMS resource ID need to be stored together for presentation to the target KMS bolt for later decryption.In step S7, the source replicator (working replicator thread) 1414 caches the DEK for use, and in step S8, it re-encrypts the data to be uploaded to the object store using the session key DEK.
[0177] As previously explained in relation to the pipeline stage of FIG. 6, multiple replicator threads within the replicator fleet can each process a part of the entire replication process. Steps 9 to 11 can be repeated for each replicator thread that executes the replication process. In step S9, a new source replicator thread 1414 obtains the encrypted session key DTK and the KMS resource ID from the source SDB 1410. In step S10, the new replicator thread 1414 communicates with the KMS vault by calling the KMS API using the CGO together with the snapshot resource ID of the source region to obtain the decrypted session key DEK. In step S11, a new replicator thread 1414 re-encrypts the data for uploading to the object store.
[0178] When all replicator threads have completed the replication job for the current replication cycle, in step S12, the source CP API 1412 deletes the encrypted session key DTK and the KMS resource ID stored in the source SDB 1410. In step 13, all replicator threads 1414 erase the cached session key DEK from memory. Thus, this can make the session key valid only for one replication cycle (i.e., one session) and prevent it from being reused in another replication cycle again.
[0179] Figure 15 shows a simplified table format (or schema) of a shared database (SDB) for storing an encrypted session key DTK according to an embodiment. As described above, each region includes an SDB for storing an encrypted session key DTK. In one embodiment, the SDB can be a key-value store. Accordingly, this schema includes two parts: a key part and a value part. The information in this table can be used for communication between DPs and CPs within each region. In Figure 15, the key part 1510 includes a key type 1511, an object type 1512, an object number (i.e., a differential transfer ID or a replication number) 1514, and a snapshot number 1516. The value part 1520 includes a version 1522, a KMS key resource ID 1524, an encrypted key 1526, a target file system number 1528, a target fleet 1530, and a target availability domain (AD) 1532.
[0180] In one embodiment, the key part 1510 of the schema 1500 begins with a key type 1511 that indicates the type of metadata within the B-tree to assist in efficient parsing of the B-tree key. For example, the key types of metadata can include iNode, file map (i.e., fmap), etc. The object type 1512 can be used to distinguish the purpose of the key entry. For example, setting the object type 1512 to 10 can indicate that the entry is written by a source replicator within the source region and read by the CP of the source region for transferring the encrypted session key 1526 to the target region. On the other hand, setting the object type 1512 to 11 can indicate that the entry is written by the target CP and read by the target replicator after receiving the encrypted session key 1526 from the source region for decrypting the differential / data downloaded from the target object store.
[0181] In some embodiments, object number 1514 is used to identify which replication process (e.g., a particular replication cycle) is associated with encrypted session key DTK1526. Snapshot number 1516 is used to identify the particular snapshot involved in this replication.
[0182] In one embodiment, the value portion 1520 of schema 1500 begins with a version 1522 field indicating the version of the resource ID. KMS key resource ID 1524 is the resource ID of the KMS master key within the target region. In an alternative embodiment, KMS resource ID 1524 can be a key managed by the customer. Encrypted key 1526 includes both the ciphertext, the vault prefix, and a counted octet string indicating the length of the byte array of this key. Target file system number 1528 is used for communication with the FSK service and determines which target file system requires FSK for authentication / validation purposes. Target fleet 1530 is for identifying the replicator fleet within the target DP. Target AD 1532 serves the purpose of authentication / validation for verifying the region and AD of the target file system.
[0183] FIG. 16 is a flowchart showing an end-to-end key management process for inter-region replication according to an embodiment. In one embodiment, the three-layer key architecture of the FSS can have three different system-level security keys: a source file system key (FSKa), a target file system key (FSKb), and a session key (DEK). A security key may refer to a key used for security purposes and can encrypt / decrypt another key or data. FSKa is for the source file system. FSKb is for the target file system. The session key is used between the source file system and the target file system. Each of these three system-level security keys can be generated by three different master keys respectively. The first version of the master key from the KMS vault can be used to generate FSKa within the source file system. The second version of the master key from the KMS vault can be used to generate FSKb within the target file system. The third version of the master key from the KMS vault can be used to generate DEK within the source file system. Further, there are two different file-level security keys: a source file key (FKa) and a target file key (FKb).
[0184] In some embodiments, FSKa and the first version of the master key are held within the source region. FSKb and the second version of the master key are held within the target region. The third version of the master key can also be held within the source region. Only the DEK is transferred between the source region and the target region and can be used for encrypting / decrypting data transferred between the source region and the target region. Thus, the exposure of keys outside the region where the keys are generated can be minimized.
[0185] In step 1610, the FSS may generate a source file system key (FSKa) within the source file system based on a source-side master key for encrypting / decrypting a file key (e.g., FKa) within the source file system. In step 1612, the source file system may decrypt the file data of the corresponding file previously encrypted by FKa using the source file key (FKa). In step 1614, the source replicator (e.g., 1314 in FIG. 13) may communicate with the target KMS bolt to obtain an inter-region master key and generate a new session key (i.e., DEK, plaintext key) for this replication cycle (or session). In step 1616, during the delta generation process within the source file system, the source file system may encrypt the generated snapshot delta (including file data and manifest files) using the DEK.
[0186] In step 1630, the source file system may upload the encrypted snapshot delta to the object store for the target file system to download. In step 1632, the source file system may encrypt the session key using FSKa to generate an encrypted session key (i.e., DTK) and then transfer it from the source CP to the target CP.
[0187] In step 1650, the target file system may decrypt the DTK with the assistance of the KMS to obtain the DEK and then decrypt the downloaded snapshot delta using the DEK for delta application. In step 1652, the FSS may generate a target file system key (FSKb) within the target file system based on a target-side master key for encrypting / decrypting a file key (e.g., FKB) within the target file system. In step 1654, during delta application, the target file system may encrypt the file data from the downloaded snapshot delta using the target file key (FKa) and store it in the corresponding file.
[0188] Exemplary cloud architecture As described above, infrastructure as a service (IaaS) is a specific type of cloud computing. IaaS can be configured to provide virtualized computing resources via a public network (e.g., the Internet). In the IaaS model, a cloud computing provider can host infrastructure components (e.g., servers, storage devices, network nodes (e.g., hardware), deployment software, platform virtualization (e.g., hypervisor layer), etc.). In some cases, the IaaS provider can provide various services that arise in connection with those infrastructure components (examples of services include billing software, monitoring software, logging software, load balancing software, clustering software, etc.). Thus, since these services can be policy-driven, IaaS users may be able to implement policies to drive load balancing to maintain application availability and performance.
[0189] In some cases, IaaS customers may access resources and services via a wide area network (WAN), such as the Internet, and use the cloud provider's services to install the remaining elements of the application stack. For example, a user can log in to an IaaS platform, create virtual machines (VMs), install an operating system (OS) on each VM, deploy middleware such as a database, create storage buckets for workloads and backups, and install enterprise software on the VM as well. Next, the customer can use the provider's services to perform various functions, including load balancing network traffic, troubleshooting application problems, monitoring performance, and managing disaster recovery.
[0190] In most cases, the cloud computing model requires the participation of a cloud provider. The cloud provider can be a third-party service that specializes in providing (e.g., offering, lending, selling) IaaS, but it doesn't have to be. An entity can choose to deploy a private cloud and become its own provider of infrastructure services.
[0191] In some examples, the deployment of IaaS is the process of placing a new application or a new version of an application on a prepared application server, etc. This process can include the process of preparing the server (e.g., installing libraries, daemons, etc.). This process is often managed by the cloud provider under the hypervisor layer (e.g., servers, storage, network hardware, and virtualization). Thus, the customer can play a role in handling the deployment of (e.g., an operating system (OS), middleware, and / or applications) on top of (e.g., self-service virtual machines that can be spun up on demand).
[0192] In some cases, IaaS provisioning can also refer to obtaining computers or virtual hosts for use and installing the required libraries or services on those computers or virtual hosts. In most cases, deployment does not include provisioning, and provisioning may need to be done first.
[0193] In some cases, there are two different issues with IaaS provisioning. First, there is the initial issue of provisioning the initial set of infrastructure before anything is run. Second, after everything is provisioned, there is the issue of evolving the existing infrastructure (e.g., adding new services, changing services, removing services, etc.). In some cases, these two issues can be addressed by enabling the infrastructure configuration to be defined declaratively. In other words, the infrastructure (e.g., which components are required and how those components exchange information) can be defined by one or more configuration files. In this way, the entire infrastructure topology (e.g., which resources depend on which resources and how each of those resources cooperate) can be described declaratively. In some cases, after the topology is defined, a workflow can be generated to create and / or manage the various components described in the configuration files.
[0194] In some examples, the infrastructure can include many interconnected elements. For example, there can be one or more virtual private clouds (VPCs), also known as core networks (e.g., a configurable and / or shared pool of computing resources, sometimes on-demand), of configurable and / or shared computing resources. In some examples, there can be one or more inbound traffic / outbound traffic group rules provisioned to define how inbound traffic and / or outbound traffic of the network is set, and one or more virtual machines (VMs). Other infrastructure elements, such as load balancers, databases, etc., can be provisioned. As more infrastructure elements are desired and / or added, the infrastructure can grow incrementally.
[0195] In some cases, continuous deployment techniques can be employed to enable the deployment of infrastructure code across various virtual computing environments. Additionally, the techniques described can enable infrastructure management within these environments. In some examples, a service team may write code that is desired to be deployed to one or more, but often many, different production environments (e.g., across various geographical locations, sometimes worldwide). However, in some examples, the infrastructure to which the code is deployed must be set up first. In some cases, provisioning can be done manually, provisioning tools can be used to provision resources, and / or deployment tools can be used to deploy the code after the infrastructure has been provisioned.
[0196] FIG. 17 is a block diagram 1700 showing an exemplary pattern of an IaaS architecture according to at least one embodiment. A service operator 1702 can be communicatively coupled to a secure host tenancy 1704 that can include a virtual cloud network (VCN) 1706 and a secure host subnet 1708. In some examples, the service operator 1702 may use one or more client computing devices, which can be portable handheld devices (e.g., iPhone®, mobile phone, iPad®, computing tablet, personal digital assistant (PDA)) or wearable devices (e.g., Google® Glass head-mounted display) that run software such as Microsoft Windows Mobile® and / or various mobile operating systems such as iOS, Windows Phone, Android, BlackBerry 8, Palm OS, and have Internet, email, short message service (SMS), Blackberry®, or other communication protocols enabled. Alternatively, the client computing device can be a general-purpose personal computer, including, for example, personal computers and / or laptop computers that run various versions of Microsoft Windows®, Apple Macintosh®, and / or Linux® operating systems. The client computing device can be a workstation computer that runs any of various commercially available UNIX® or UNIX-like operating systems, including, but not limited to, various GNU / Linux® operating systems such as Google Chrome OS.Alternatively or in addition, the client computing device can be any other electronic device such as a thin client computer, an Internet-enabled gaming system (e.g., a Microsoft Xbox gaming console with or without a Kinect (registered trademark) gesture input device), and / or a personal messaging device that can communicate via a network and / or the Internet that has access to the VCN 1706.
[0197] The VCN 1706 can include an LPG 1710 that can be communicatively coupled to an SSH VCN 1712 via a local peering gateway (LPG) 1710 included in a secure shell (SSH) VCN 1712. The SSH VCN 1712 can include an SSH subnet 1714 and can be communicatively coupled to a control plane VCN 1716 via an LPG 1710 included in the control plane VCN 1716. Also, the SSH VCN 1712 can be communicatively coupled to a data plane VCN 1718 via an LPG 1710. The control plane VCN 1716 and the data plane VCN 1718 can be included in a service tenancy 1719 that can be owned and / or operated by an IaaS provider.
[0198] The control plane VCN 1716 can include a control plane demilitarized zone (DMZ) layer 1720 that functions as a boundary network (e.g., a part of the enterprise network between the enterprise intranet and the external network). Servers based on the DMZ can have limited responsibilities and can help contain intrusions. Further, the DMZ layer 1720 can include a control plane application layer 1724 that can include one or more load balancer (LB) subnets 1722, an app subnet 1726, and a control plane data layer 1728 that can include a database (DB) subnet 1730 (e.g., a front-end DB subnet and / or a back-end DB subnet). The LB subnet 1722 included in the control plane DMZ layer 1720 can be communicatively coupled to the app subnet 1726 included in the control plane application layer 1724 that can be included in the control plane VCN 1716 and the Internet gateway 1734, and the app subnet 1726 can be communicatively coupled to the DB subnet 1730 included in the control plane data layer 1728 as well as the service gateway 1736 and the network address translation (NAT) gateway 1738. The control plane VCN 1716 can include the service gateway 1736 and the NAT gateway 1738.
[0199] The control plane VCN 1716 can include a data plane mirror application layer 1740 that can include an application subnet 1726. The application subnet 1726 included in the data plane mirror application layer 1740 can include a virtual network interface controller (VNIC) 1742 that can execute a compute instance 1744. The compute instance 1744 can communicatively couple the application subnet 1726 of the data plane mirror application layer 1740 to the application subnet 1726 that can be included in the data plane application layer 1746.
[0200] The data plane VCN 1718 can include a data plane application layer 1746, a data plane DMZ layer 1748, and a data plane data layer 1750. The data plane DMZ layer 1748 can include an LB subnet 1722 that can be communicatively coupled to the application subnet 1726 of the data plane application layer 1746 and the internet gateway 1734 of the data plane VCN 1718. The application subnet 1726 can be communicatively coupled to the service gateway 1736 and the NAT gateway 1738 of the data plane VCN 1718. The data plane data layer 1750 can also include a DB subnet 1730 that can be communicatively coupled to the application subnet 1726 of the data plane application layer 1746.
[0201] The internet gateways 1734 of the control plane VCN 1716 and the data plane VCN 1718 can be communicatively coupled to a metadata management service 1752 that can be communicatively coupled to the public internet 1754. The public internet 1754 can be communicatively coupled to the NAT gateways 1738 of the control plane VCN 1716 and the data plane VCN 1718. The service gateways 1736 of the control plane VCN 1716 and the data plane VCN 1718 can be communicatively coupled to a cloud service 1756.
[0202] In some examples, a service gateway 1736 of the control plane VCN 1716 or the data plane VCN 1718 can make application programming interface (API) calls to a cloud service 1756 without going through the public Internet 1754. An API call from the service gateway 1736 to the cloud service 1756 can be one-way, and the service gateway 1736 can make an API call to the cloud service 1756, and the cloud service 1756 can send the requested data to the service gateway 1736. However, the cloud service 1756 does not need to initiate an API call to the service gateway 1736.
[0203] In some examples, the secure host tenancy 1704 can be directly connected to the service tenancy 1719 or, otherwise, can be separated. The secure host subnet 1708 can communicate with the SSH subnet 1714 via the LPG 1710, and the LPG 1710 can enable two-way communication on a separated system if not. Connecting the secure host subnet 1708 to the SSH subnet 1714 can provide the secure host subnet 1708 with access to other entities within the service tenancy 1719.
[0204] The control plane VCN 1716 may enable users of service tenancy 1719 to set or otherwise provision desired resources. Desired resources provisioned within the control plane VCN 1716 may be deployed or otherwise used in the data plane VCN 1718. In some examples, the control plane VCN 1716 may be separable from the data plane VCN 1718, and the data plane mirror app layer 1740 of the control plane VCN 1716 may communicate with the data plane app layer 1746 of the data plane VCN 1718 via VNICs 1742 that may be included in the data plane mirror app layer 1740 and the data plane app layer 1746.
[0205] In some examples, a user or customer of the system may perform requests, such as create, read, update, or delete (CRUD) operations, via the public internet 1754 that can communicate requests to the metadata management service 1752. The metadata management service 1752 may communicate the requests to the control plane VCN 1716 via the internet gateway 1734. The requests may be received by the LB subnet 1722 included in the control plane DMZ layer 1720. The LB subnet 1722 may determine that the requests are valid, and in response, the LB subnet 1722 may send the requests to the app subnet 1726 included in the control plane app layer 1724. If the validity of the requests is confirmed and the requests require calls to the public internet 1754, the calls to the public internet 1754 may be sent to the NAT gateway 1738 that can make calls to the public internet 1754. Metadata that may desirably be stored by the requests may be stored within the DB subnet 1730.
[0206] In some examples, the data plane mirror application layer 1740 can facilitate direct communication between the control plane VCN 1716 and the data plane VCN 1718. For example, it may be desirable for changes, updates, or other appropriate modifications to the configuration to be applied to the resources included in the data plane VCN 1718. Through the VNIC 1742, the control plane VCN 1716 can communicate directly with the resources included in the data plane VCN 1718, thereby enabling changes, updates, or other appropriate modifications to the configuration of the resources.
[0207] In some embodiments, the control plane VCN 1716 and the data plane VCN 1718 may be included in the service tenant 1719. In this case, the user or customer of the system does not have to own or operate either the control plane VCN 1716 or the data plane VCN 1718. Instead, the IaaS provider may own or operate both the control plane VCN 1716 and the data plane VCN 1718, which can both be included in the service tenancy 1719. This embodiment can enable network isolation that can prevent a user or customer from exchanging information with the resources of other users or other customers. Also, this embodiment can enable a user or customer of the system to privately store a database without relying on the public Internet 1754, which may not have the desired level of threat prevention for storage.
[0208] In other embodiments, the LB subnet 1722 included in the control plane VCN 1716 may be configured to receive signals from the service gateway 1736. In this embodiment, the control plane VCN 1716 and the data plane VCN 1718 may be configured to be invoked by a customer of the IaaS provider without invoking the public Internet 1754. A customer of the IaaS provider may desire this embodiment because the databases used by the customer may be controlled by the IaaS provider and may be stored in a service tenancy 1719 that may be isolated from the public Internet 1754.
[0209] FIG. 18 is a block diagram 1800 showing another exemplary pattern of an IaaS architecture according to at least one embodiment. A service operator 1802 (e.g., the service operator 1702 of FIG. 17) can be communicatively coupled to a secure host tenancy 1804 (e.g., the secure host tenancy 1704 of FIG. 17) that can include a virtual cloud network (VCN) 1806 (e.g., the VCN 1706 of FIG. 17) and a secure host subnet 1808 (e.g., the secure host subnet 1708 of FIG. 17). The VCN 1806 can include a local peering gateway (LPG) 1810 (e.g., the LPG 1710 of FIG. 17) and can be communicatively coupled to a secure shell (SSH) VCN 1812 (e.g., the SSH VCN 1712 of FIG. 17) via the LPG 1710 included in the SSH VCN 1812. The SSH VCN 1812 can include an SSH subnet 1814 (e.g., the SSH subnet 1714 of FIG. 17), and the SSH VCN 1812 can be communicatively coupled to a control plane VCN 1816 (e.g., the control plane VCN 1716 of FIG. 17) via the LPG 1810 included in the control plane VCN 1816. The control plane VCN 1816 can be included in a service tenancy 1819 (e.g., the service tenancy 1719 of FIG. 17), and the data plane VCN 1818 (e.g., the data plane VCN 1718 of FIG. 17) can be included in a customer tenancy 1821 that can be owned or operated by a user or customer of the system.
[0210] The control plane VCN 1816 can include a control plane DMZ layer 1820 (e.g., the control plane DMZ layer 1720 in FIG. 17) that can include an LB subnet 1822 (e.g., the LB subnet 1722 in FIG. 17), a control plane application layer 1824 (e.g., the control plane application layer 1724 in FIG. 17) that can include an application subnet 1826 (e.g., the application subnet 1726 in FIG. 17), and a control plane data layer 1828 (e.g., the control plane data layer 1728 in FIG. 17) that can include a database (DB) subnet 1830 (e.g., similar to the DB subnet 1730 in FIG. 17). The LB subnet 1822 included in the control plane DMZ layer 1820 can be communicatively coupled to the application subnet 1826 included in the control plane application layer 1824 that can be included in the control plane VCN 1816, and an Internet gateway 1834 (e.g., the Internet gateway 1734 in FIG. 17), and the application subnet 1826 can be communicatively coupled to the DB subnet 1830 included in the control plane data layer 1828, as well as a service gateway 1836 (e.g., the service gateway 1736 in FIG. 17) and a network address translation (NAT) gateway 1838 (e.g., the NAT gateway 1738 in FIG. 17). The control plane VCN 1816 can include a service gateway 1836 and a NAT gateway 1838.
[0211] The control plane VCN 1816 can include a data plane mirror app layer 1840 (e.g., the data plane mirror app layer 1740 of FIG. 17) that can include an app subnet 1826. The app subnet 1826 included in the data plane mirror app layer 1840 can include a virtual network interface controller (VNIC) 1842 (e.g., the VNIC 1742) that can execute a compute instance 1844 (e.g., similar to the compute instance 1744 of FIG. 17). The compute instance 1844 can facilitate communication between the app subnet 1826 of the data plane mirror app layer 1840 and an app subnet 1826 that can be included in the data plane app layer 1846 (e.g., the data plane app layer 1746 of FIG. 17) via the VNIC 1842 included in the data plane mirror app layer 1840 and the VNIC 1842 included in the data plane app layer 1846.
[0212] The internet gateway 1834 included in the control plane VCN 1816 can be communicatively coupled to a metadata management service 1852 (e.g., the metadata management service 1752 of FIG. 17) that can be communicatively coupled to the public internet 1854 (e.g., the public internet 1754 of FIG. 17). The public internet 1854 can be communicatively coupled to the NAT gateway 1838 included in the control plane VCN 1816. The service gateway 1836 included in the control plane VCN 1816 can be communicatively coupled to a cloud service 1856 (e.g., the cloud service 1756 of FIG. 17).
[0213] In some examples, the data plane VCN 1818 can be included in the customer's tenancy 1821. In this case, the IaaS provider may provide a control plane VCN 1816 for each customer, and the IaaS provider may configure the specific compute instances 1844 included in the service tenancy 1819 for each customer. Each compute instance 1844 may enable communication between the control plane VCN 1816 included in the service tenancy 1819 and the data plane VCN 1818 included in the customer's tenancy 1821. The compute instance 1844 may enable the resources provisioned within the control plane VCN 1816 included in the service tenancy 1819 to be deployed or otherwise used in the data plane VCN 1818 included in the customer's tenancy 1821.
[0214] In another example, a customer of an IaaS provider may have a database that persists in the customer's tenancy 1821. In this example, the control plane VCN 1816 can include a data plane mirror app layer 1840 that can include an app subnet 1826. The data plane mirror app layer 1840 can exist in the data plane VCN 1818, but the data plane mirror app layer 1840 does not have to persist in the data plane VCN 1818. That is, the data plane mirror app layer 1840 can have access rights to the customer's tenancy 1821, but the data plane mirror app layer 1840 does not have to exist in the data plane VCN 1818 and does not have to be owned or operated by the customer of the IaaS provider. The data plane mirror app layer 1840 can be configured to make calls to the data plane VCN 1818, but does not have to be configured to make calls to any entity included in the control plane VCN 1816. The customer may wish to deploy or otherwise use resources within the data plane VCN 1818 that are provisioned within the control plane VCN 1816, and the data plane mirror app layer 1840 can facilitate the desired deployment or other use of the customer's resources.
[0215] In some embodiments, a customer of an IaaS provider can apply a filter to the data plane VCN 1818. In this embodiment, the customer can determine which data plane VCN 1818s are accessible, and the customer can restrict access from the data plane VCN 1818 to the public Internet 1854. The IaaS provider does not have to be able to apply a filter or otherwise control access of the data plane VCN 1818 to any external network or database. Applying the filter and control by the customer to the data plane VCN 1818 included in the customer's tenancy 1821 can help to isolate the data plane VCN 1818 from other customers and from the public Internet 1854.
[0216] In some embodiments, cloud service 1856 may be invoked by service gateway 1836 to access services that may not exist on any of public Internet 1854, control plane VCN 1816, or data plane VCN 1818. The connection between cloud service 1856 and control plane VCN 1816 or data plane VCN 1818 may not be operational or continuous. Cloud service 1856 may exist on a different network owned or operated by an IaaS provider. Cloud service 1856 may be configured to receive calls from service gateway 1836 and may be configured not to receive calls from public Internet 1854. Some cloud services 1856 may be isolated from other cloud services 1856, and control plane VCN 1816 may be isolated from cloud services 1856 that may not exist in the same region as control plane VCN 1816. For example, control plane VCN 1816 may be located in "Region 1," and "Deployment 17" of the cloud service may be located in Region 1 and "Region 2." When a call to Deployment 17 is made by service gateway 1836 included in control plane VCN 1816 located in Region 1, this call may be sent to Deployment 17 within Region 1. In this example, control plane VCN 1816, or Deployment 17 within Region 1, may or may not be communicatively coupled to Deployment 17 within Region 2.
[0217] FIG. 19 is a block diagram 1900 showing another exemplary pattern of an IaaS architecture according to at least one embodiment. A service operator 1902 (e.g., the service operator 1702 of FIG. 17) can be communicatively coupled to a secure host tenancy 1904 (e.g., the secure host tenancy 1704 of FIG. 17) that can include a virtual cloud network (VCN) 1906 (e.g., the VCN 1706 of FIG. 17) and a secure host subnet 1908 (e.g., the secure host subnet 1708 of FIG. 17). The VCN 1906 can be communicatively coupled to an SSH VCN 1912 (e.g., the SSH VCN 1712 of FIG. 17) via an LPG 1910 (e.g., the LPG 1710 of FIG. 17) included in the SSH VCN 1912. The SSH VCN 1912 can include an SSH subnet 1914 (e.g., the SSH subnet 1714 of FIG. 17), and the SSH VCN 1912 can be communicatively coupled to a control plane VCN 1916 (e.g., the control plane VCN 1716 of FIG. 17) via an LPG 1910 included in the control plane VCN 1916 and to a data plane VCN 1918 (e.g., the data plane 1718 of FIG. 17) via an LPG 1910 included in the data plane VCN 1918. The control plane VCN 1916 and the data plane VCN 1918 can be included in a service tenancy 1919 (e.g., the service tenancy 1719 of FIG. 17).
[0218] The control plane VCN 1916 can include a control plane DMZ layer 1920 (e.g., the control plane DMZ layer 1720 of FIG. 17) that can include a load balancer (LB) subnet 1922 (e.g., the LB subnet 1722 of FIG. 17), a control plane application layer 1924 (e.g., the control plane application layer 1724 of FIG. 17) that can include an application subnet 1926 (similar to the application subnet 1726 of FIG. 17), and a control plane data layer 1928 (e.g., the control plane data layer 1728 of FIG. 17) that can include a DB subnet 1930. The LB subnet 1922 included in the control plane DMZ layer 1920 can be communicatively coupled to the application subnet 1926 included in the control plane application layer 1924 that can be included in the control plane VCN 1916, and to an Internet gateway 1934 (e.g., the Internet gateway 1734 of FIG. 17). The application subnet 1926 can be communicatively coupled to the DB subnet 1930 included in the control plane data layer 1928, as well as to a service gateway 1936 (e.g., the service gateway of FIG. 17) and a network address translation (NAT) gateway 1938 (e.g., the NAT gateway 1738 of FIG. 17). The control plane VCN 1916 can include the service gateway 1936 and the NAT gateway 1938.
[0219] The data plane VCN 1918 can include a data plane application layer 1946 (e.g., the data plane application layer 1746 of FIG. 17), a data plane DMZ layer 1948 (e.g., the data plane DMZ layer 1748 of FIG. 17), and a data plane data layer 1950 (e.g., the data plane data layer 1750 of FIG. 17). The data plane DMZ layer 1948 can include a reliable application subnet 1960 and an unreliable application subnet 1962 of the data plane application layer 1946, and an LB subnet 1922 communicatively coupled to an Internet gateway 1934 included in the data plane VCN 1918. The reliable application subnet 1960 can be communicatively coupled to a service gateway 1936 included in the data plane VCN 1918, a NAT gateway 1938 included in the data plane VCN 1918, and a DB subnet 1930 included in the data plane data layer 1950. The unreliable application subnet 1962 can be communicatively coupled to a service gateway 1936 included in the data plane VCN 1918 and a DB subnet 1930 included in the data plane data layer 1950. The data plane data layer 1950 can include a DB subnet 1930 communicatively coupled to a service gateway 1936 included in the data plane VCN 1918.
[0220] The untrusted application subnet 1962 can include one or more primary VNICs 1964(1) to (N) communicatively coupled to tenant virtual machines (VMs) 1966(1) to (N). Each tenant VM 1966(1) to (N) can be communicatively coupled to respective application subnets 1967(1) to (N) that can be included in respective container egress VCNs 1968(1) to (N) that can be included in respective customer tenancies 1970(1) to (N). Each secondary VNIC 1972(1) to (N) can facilitate communication between the untrusted application subnet 1962 included in the data plane VCN 1918 and the application subnets included in the container egress VCNs 1968(1) to (N). Each container egress VCN 1968(1) to (N) can include a NAT gateway 1938 communicatively coupled to the public Internet 1954 (e.g., the public Internet 1754 of FIG. 17).
[0221] The Internet gateway 1934 included in the control plane VCN 1916 and in the data plane VCN 1918 can be communicatively coupled to a metadata management service 1952 (e.g., the metadata management system 1752 of FIG. 17) communicatively coupled to the public Internet 1954. The public Internet 1954 can be communicatively coupled to the NAT gateway 1938 included in the control plane VCN 1916 and in the data plane VCN 1918. The service gateway 1936 included in the control plane VCN 1916 and in the data plane VCN 1918 can be communicatively coupled to a cloud service 1956.
[0222] In some embodiments, the data plane VCN 1918 can be integrated with the customer's tenancy 1970. This integration can, in some cases, be useful or desirable for the customers of the IaaS provider, such as when they may want support when running code. The customer may provide code to run that can be disruptive, communicate with the resources of other customers, or otherwise cause undesirable effects. In response, the IaaS provider can determine whether to execute the code provided to the IaaS provider by the customer.
[0223] In some examples, a customer of an IaaS provider may grant the IaaS provider temporary network access rights and request a function that connects to the data plane application layer 1946. The code for executing this function may be executed in the VMs 1966(1) to (N), and this code need not be configured to execute elsewhere on the data plane VCN 1918. Each of the VMs 1966(1) to (N) may be connected to the tenancy 1970 of one customer. Each of the containers 1971(1) to (N) included in the VMs 1966(1) to (N) may be configured to execute the code. In this case, there can be a two-fold separation (for example, the containers 1971(1) to (N) that execute the code, the containers 1971(1) to (N) may be included in at least the VMs 1966(1) to (N) that are included in the untrusted application subnet 1962), which can help prevent incorrect or otherwise undesirable code from damaging the IaaS provider's network or damaging the networks of different customers. The containers 1971(1) to (N) may be communicatively coupled to the customer's tenancy 1970 and may be configured to send or receive data with the customer's tenancy 1970. The containers 1971(1) to (N) need not be configured to send or receive data with any other entity within the data plane VCN 1918. Upon completion of the execution of the code, the IaaS provider may force the containers 1971(1) to (N) to terminate or otherwise discard them.
[0224] In some embodiments, the trusted application subnet 1960 can execute code that can be owned or operated by an IaaS provider. In this embodiment, the trusted application subnet 1960 may be communicatively coupled to the DB subnet 1930 and may be configured to perform CRUD operations within the DB subnet 1930. The untrusted application subnet 1962 may be communicatively coupled to the DB subnet 1930, but in this embodiment, the untrusted application subnet may be configured to perform read operations within the DB subnet 1930. The containers 1971(1) to (N) that can execute customer code, which may be included in each customer's VMs 1966(1) to (N), need not be communicatively coupled to the DB subnet 1930.
[0225] In other embodiments, the control plane VCN 1916 and the data plane VCN 1918 need not be directly communicatively coupled. In this embodiment, there need not be a direct communication between the control plane VCN 1916 and the data plane VCN 1918. However, the communication can occur indirectly by at least one method. The LPG 1910 may be established by the IaaS provider, thereby facilitating communication between the control plane VCN 1916 and the data plane VCN 1918. In another example, the control plane VCN 1916 or the data plane VCN 1918 can make calls to the cloud service 1956 via the service gateway 1936. For example, a call from the control plane VCN 1916 to the cloud service 1956 can include a request for a service that can communicate with the data plane VCN 1918.
[0226] FIG. 20 is a block diagram 2000 showing another exemplary pattern of an IaaS architecture according to at least one embodiment. A service operator 2002 (e.g., service operator 1702 of FIG. 17) may be communicatively coupled to a secure host tenancy 2004 (e.g., secure host tenancy 1704 of FIG. 17) that can include a virtual cloud network (VCN) 2006 (e.g., VCN 1706 of FIG. 17) and a secure host subnet 2008 (e.g., secure host subnet 1708 of FIG. 17). The VCN 2006 can include an LPG 2010 (e.g., LPG 1710 of FIG. 17) and can be communicatively coupled to an SSH VCN 2012 (e.g., SSH VCN 1712 of FIG. 17) via the LPG 2010 included in the SSH VCN 2012. The SSH VCN 2012 can include an SSH subnet 2014 (e.g., SSH subnet 1714 of FIG. 17), and the SSH VCN 2012 can be communicatively coupled to a control plane VCN 2016 (e.g., control plane VCN 1716 of FIG. 17) via the LPG 2010 included in the control plane VCN 2016 and to a data plane VCN 2018 (e.g., data plane 1718 of FIG. 17) via the LPG 2010 included in the data plane VCN 2018. The control plane VCN 2016 and the data plane VCN 2018 can be included in a service tenancy 2019 (e.g., service tenancy 1719 of FIG. 17).
[0227] The control plane VCN 2016 can include a control plane DMZ layer 2020 (e.g., the control plane DMZ layer 1720 in FIG. 17) that can include an LB subnet 2022 (e.g., the LB subnet 1722 in FIG. 17), a control plane application layer 2024 (e.g., the control plane application layer 1724 in FIG. 17) that can include an application subnet 2026 (e.g., the application subnet 1726 in FIG. 17), and a control plane data layer 2028 (e.g., the control plane data layer 1728 in FIG. 17) that can include a DB subnet 2030 (e.g., the DB subnet 1930 in FIG. 19). The LB subnet 2022 included in the control plane DMZ layer 2020 can be communicatively coupled to the application subnet 2026 included in the control plane application layer 2024 that can be included in the control plane VCN 2016, and to an Internet gateway 2034 (e.g., the Internet gateway 1734 in FIG. 17). The application subnet 2026 can be communicatively coupled to the DB subnet 2030 included in the control plane data layer 2028, as well as to a service gateway 2036 (e.g., the service gateway in FIG. 17) and a network address translation (NAT) gateway 2038 (e.g., the NAT gateway 1738 in FIG. 17). The control plane VCN 2016 can include the service gateway 2036 and the NAT gateway 2038.
[0228] The data plane VCN 2018 can include a data plane application layer 2046 (e.g., the data plane application layer 1746 of FIG. 17), a data plane DMZ layer 2048 (e.g., the data plane DMZ layer 1748 of FIG. 17), and a data plane data layer 2050 (e.g., the data plane data layer 1750 of FIG. 17). The data plane DMZ layer 2048 can include a trusted application subnet 2060 (e.g., the trusted application subnet 1960 of FIG. 19) and an untrusted application subnet 2062 (e.g., the untrusted application subnet 1962 of FIG. 19) of the data plane application layer 2046, and an LB subnet 2022 communicatively coupled to an Internet gateway 2034 included in the data plane VCN 2018. The trusted application subnet 2060 can be communicatively coupled to a service gateway 2036 included in the data plane VCN 2018, a NAT gateway 2038 included in the data plane VCN 2018, and a DB subnet 2030 included in the data plane data layer 2050. The untrusted application subnet 2062 can be communicatively coupled to a service gateway 2036 included in the data plane VCN 2018, and a DB subnet 2030 included in the data plane data layer 2050. The data plane data layer 2050 can include a DB subnet 2030 communicatively coupled to a service gateway 2036 included in the data plane VCN 2018.
[0229] The untrusted application subnet 2062 can include primary VNICs 2064(1) to (N) communicatively coupled to tenant virtual machines (VMs) 2066(1) to (N) existing within the untrusted application subnet 2062. Each tenant VM 2066(1) to (N) can execute code within respective containers 2067(1) to (N) and can be communicatively coupled to an application subnet 2026 that can be included in a data plane application layer 2046 that can be included in a container egress VCN 2068. Each secondary VNIC 2072(1) to (N) can facilitate communication between the untrusted application subnet 2062 included in the data plane VCN 2018 and the application subnet included in the container egress VCN 2068. The container egress VCN can include a NAT gateway 2038 communicatively coupled to a public internet 2054 (e.g., the public internet 1754 of FIG. 17).
[0230] The internet gateway 2034 included in the control plane VCN 2016 and the data plane VCN 2018 can be communicatively coupled to a metadata management service 2052 (e.g., the metadata management system 1752 of FIG. 17) communicatively coupled to the public internet 2054. The public internet 2054 can be communicatively coupled to a NAT gateway 2038 included in the control plane VCN 2016 and the data plane VCN 2018. The service gateway 2036 included in the control plane VCN 2016 and the data plane VCN 2018 can be communicatively coupled to a cloud service 2056.
[0231] In some examples, the pattern shown by the architecture of block diagram 2000 in FIG. 20 may be regarded as an exception to the pattern shown by the architecture of block diagram 1900 in FIG. 19, which may be desirable for customers of the IaaS provider when the IaaS provider cannot communicate directly with the customer (e.g., a disconnected region). Each of the containers 2067(1) to (N) included in the VMs 2066(1) to (N) for each customer can be accessed in real time by the customer. The containers 2067(1) to (N) can be configured to make calls to the respective secondary VNICs 2072(1) to (N) included in the application subnet 2026 of the data plane application layer 2046 that may be included in the container egress VCN 2068. The secondary VNICs 2072(1) to (N) can send the calls to the NAT gateway 2038, and the NAT gateway 2038 can send the calls to the public Internet 2054. In this example, the containers 2067(1) to (N) that can be accessed in real time by the customer can be separated from the control plane VCN 2016 and can be separated from other entities included in the data plane VCN 2018. The containers 2067(1) to (N) can be separated from the resources of other customers.
[0232] In other examples, a customer can call cloud service 2056 using containers 2067(1) to (N). In this example, the customer may execute the code within containers 2067(1) to (N) that requests a service from cloud service 2056. Containers 2067(1) to (N) can send this request to secondary VNICs 2072(1) to (N), and secondary VNICs 2072(1) to (N) can send this request to a NAT gateway, and the NAT gateway can send this request to public internet 2054. Public internet 2054 can send this request to LB subnet 2022 included in control plane VCN 2016 via internet gateway 2034. In response to determining that this request is valid, the LB subnet can send this request to app subnet 2026, and app subnet 2026 can send this request to cloud service 2056 via service gateway 2036.
[0233] It should be understood that the IaaS architectures 1700, 1800, 1900, 2000 shown in the figures may include components other than the components shown. Further, the embodiments shown in the figures are merely some examples of cloud infrastructure systems that can incorporate embodiments of the present disclosure. In some other embodiments, the IaaS system may include more or fewer components than the components shown in the figures, combine two or more components, or may have different configurations or arrangements of components.
[0234] In one embodiment, the IaaS system described herein can include the provision of a series of application, middleware, and database services that are delivered to customers in a self-service, subscription-based, elastically scalable, reliable, highly available, and secure manner. An example of such an IaaS system is Oracle Cloud Infrastructure (OCI) provided by the present assignee.
[0235] FIG. 21 shows an exemplary computer system 2100 in which various embodiments can be implemented. System 2100 can be used to implement any of the computer systems described above. As shown in the figure, computer system 2100 includes a processing unit 2104 that communicates with a plurality of peripheral subsystems via a bus subsystem 2102. These peripheral subsystems can include a processing acceleration unit 2106, an I / O subsystem 2108, a storage subsystem 2118, and a communication subsystem 2124. Storage subsystem 2118 includes tangible computer-readable storage media 2122 and system memory 2110.
[0236] The bus subsystem 2102 provides a mechanism for the various components and subsystems of the computer system 2100 to communicate with each other as intended. Although the bus subsystem 2102 is schematically shown as a single bus, alternative embodiments of the bus subsystem may utilize multiple buses. The bus subsystem 2102 can be any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, and a local bus that uses any of various bus architectures. For example, such architectures can include an ISA (Industry Standard Architecture) bus, an MCA (Micro Channel Architecture) bus, an EISA (Enhanced ISA) bus, a VESA (Video Electronics Standards Association) local bus, and a PCI (Peripheral Component Interconnect) bus implemented as a mezzanine bus manufactured to the IEEE P1386.1 standard.
[0237] The processing unit 2104, which may be implemented as one or more integrated circuits (e.g., conventional microprocessors or microcontrollers), controls the operation of the computer system 2100. One or more processors may be included in the processing unit 2104. These processors can include single-core processors or multi-core processors. In certain embodiments, the processing unit 2104 may be implemented as one or more independent processing units 2132 and / or 2134, with a single-core processor or multi-core processor included in each processing unit. In other embodiments, the processing unit 2104 may be implemented as a quad-core processing unit formed by integrating two dual-core processors on a single chip.
[0238] In various embodiments, the processing unit 2104 can execute various programs according to program code and can maintain a plurality of programs or processes running simultaneously. At any given time, some or all of the program code being executed can be present in the processor 2104 and / or in the storage subsystem 2118. With appropriate programming, the processor 2104 can provide the various functions described above. The computer system 2100 may further include a processing acceleration unit 2106 that can include a digital signal processor (DSP), an application specific processor, and / or the like.
[0239] The I / O subsystem 2108 may include user interface input devices and user interface output devices. User interface input devices may include a keyboard, a pointing device such as a mouse or trackball, a touchpad or touch screen incorporated into a display, a scroll wheel, a click wheel, a dial, a button, a switch, a keypad, a voice input device with a voice command recognition system, a microphone, and other types of input devices. The user interface input devices may enable a user to interact with information by controlling input devices such as a Microsoft Xbox (registered trademark) 360 game controller via a natural user interface using gestures and spoken commands, and may include motion detection devices and / or gesture recognition devices such as a Microsoft Kinect (registered trademark) motion sensor. The user interface input devices may include gesture recognition devices such as a Google Glass (registered trademark) blink detector that detects a user's eye activity (e.g., a "blink" when taking a photo and / or selecting a menu) and converts the eye gesture into an input to the input device (e.g., Google Glass (registered trademark)). Further, the user interface input devices may include a voice recognition detection device that enables a user to interact with a voice recognition system (e.g., a Siri (registered trademark) navigator) via voice commands.
[0240] The user interface input device may include, but is not limited to, 3D mice, joysticks or pointing sticks, game pads, and graphic tablets, as well as audio / visual devices such as speakers, digital cameras, digital video cameras, portable media players, web cameras, image scanners, fingerprint scanners, barcode readers, 3D scanners, 3D printers, laser distance meters, and eye tracking devices. Further, the user interface input device may include medical image input devices such as, for example, computed tomography, magnetic resonance imaging, positron emission tomography, and medical ultrasonic examination devices. The user interface input device may include audio input devices such as, for example, MIDI keyboards, digital musical instruments, and the like.
[0241] The user interface output device may include, among others, visual displays other than display subsystems, indicator lights, or audio output devices. The display subsystem may be a flat panel device such as a flat panel device using a cathode ray tube (CRT), a liquid crystal display (LCD), or a plasma display, a projection device, a touch screen, or the like. Generally, the use of the term "output device" is intended to include all possible types of devices and mechanisms for outputting information from the computer system 2100 to the user or another computer. For example, the user interface output device may include, but is not limited to, various display devices for visually communicating text information, graphics information, and audio / video information, such as monitors, printers, speakers, headphones, car navigation systems, plotters, audio output devices, and modems.
[0242] The computer system 2100 may include a storage subsystem 2118 that provides a tangible non-transitory computer-readable storage medium for storing software and data structures that provide the functionality of the embodiments described in this disclosure. The software can include programs, code modules, instructions, scripts, etc., and when executed by one or more cores or processors of the processing unit 2104, provides the aforementioned functionality. The storage subsystem 2118 may provide a repository for storing data used in accordance with this disclosure.
[0243] As shown in the example of FIG. 21, the storage subsystem 2118 can include various components including a system memory 2110, a computer-readable storage medium 2122, and a computer-readable storage medium reader 2120. The system memory 2110 may store program instructions that are readable and executable by the processing unit 2104. The system memory 2110 may also store data used during the execution of the instructions and / or data generated during the execution of the program instructions. Various different types of programs may be loaded into the system memory 2110 including, but not limited to, client applications, web browsers, middle-tier applications, relational database management systems (RDBMS), virtual machines, containers, etc.
[0244] System memory 2110 may store an operating system 2116. Examples of operating systems 2116 include Microsoft Windows®, Apple Macintosh®, and / or Linux® operating systems, various commercially available UNIX® or UNIX-like operating systems (including, but not limited to, various GNU / Linux® operating systems, Google Chrome® OS, etc.), and / or various versions of mobile operating systems such as iOS, Windows® Phone, Android® OS, BlackBerry® OS, and Palm® OS. In some implementations where computer system 2100 runs one or more virtual machines, the virtual machines may be loaded into system memory 2110 along with guest operating systems (GOS) and executed by one or more processors or cores of processing unit 2104.
[0245] System memory 2110 can be provided in different configurations according to the type of computer system 2100. For example, system memory 2110 can be volatile memory (such as random access memory (RAM)) and / or non-volatile memory (such as read-only memory (ROM), flash memory, etc.). Various types of RAM configurations can be provided, including static random access memory (SRAM), dynamic random access memory (DRAM), and the like. In some implementations, system memory 2110 can include a basic input / output system (BIOS) that contains basic routines useful for transferring information between elements within computer system 2100, such as during startup.
[0246] Computer-readable storage medium 2122 represents a storage medium for temporarily and / or more persistently containing and storing computer-readable information for use by computer system 2100, including instructions executable by processing unit 2104 of computer system 2100, in addition to a remote storage device, a local storage device, a fixed storage device, and / or a removable storage device.
[0247] The computer-readable storage medium 2122 can include any suitable medium known in or used in the art, including storage media and communication media such as volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing and / or transmitting information, but not limited to these. The computer-readable storage medium 2122 can include tangible computer-readable storage media such as RAM, ROM, electronically erasable programmable ROM (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD), or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage, or other magnetic storage devices, or other tangible computer-readable media.
[0248] As an example, computer-readable storage medium 2122 can include a hard disk drive that reads from or writes to a removable non-volatile magnetic medium, a magnetic disk drive that reads from or writes to a removable non-volatile magnetic disk, and an optical disk drive that reads from or writes to a removable non-volatile optical disk such as a CD ROM, a DVD, and a Blu-ray (registered trademark) disk, or other optical media. The computer-readable storage medium 2122 can include, but is not limited to, a Zip (registered trademark) drive, a flash memory card, a universal serial bus (USB) flash drive, a secure digital (SD) card, a DVD disk, a digital video tape, etc. The computer-readable storage medium 2122 can include an SSD (solid-state drives) based on non-volatile memory such as a flash memory-based semiconductor drive, an enterprise flash drive, a semiconductor ROM, an SSD based on volatile memory such as a semiconductor RAM, a dynamic RAM, a static RAM, a DRAM-based SSD, a magnetoresistive RAM (MRAM) SSD, and a hybrid SSD that uses a combination of DRAM and a flash memory-based SSD. Disk drives and associated computer-readable media can provide non-volatile storage of computer-readable instructions, data structures, program modules, and other data of computer system 2100.
[0249] Machine-readable instructions executable by one or more processors or cores of processing unit 2104 may be stored on a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium can include a physically tangible memory or storage device, including a volatile memory storage device and / or a non-volatile storage device. Examples of non-transitory computer-readable storage media include magnetic storage media (e.g., disks or tapes), optical storage media (e.g., DVDs, CDs), various types of RAM, ROM, or flash memory, hard drives, floppy (registered trademark) drives, removable memory drives (e.g., USB drives), or other types of storage devices.
[0250] The communication subsystem 2124 provides an interface to other computer systems and networks. The communication subsystem 2124 functions as an interface for receiving data from other systems of the computer system 2100 and for transmitting data to other systems. For example, the communication subsystem 2124 may enable the computer system 2100 to connect to one or more devices via the Internet. In some embodiments, the communication subsystem 2124 can include components of a radio frequency (RF) transceiver for accessing wireless voice and / or data networks (such as cellular phone technology, advanced data network technologies such as 3G, 4G, or EDGE (enhanced data rates for global evolution), WiFi (IEEE 802.11 family of standards, or other mobile communication technologies, or any combination thereof), components of a global positioning system (GPS) receiver, and / or other components. In some embodiments, the communication subsystem 2124 can provide a wired network connection (such as Ethernet (registered trademark)) in addition to, or instead of, the wireless interface.
[0251] In some embodiments, the communication subsystem 2124 may receive input communications in the form of structured and / or unstructured data feeds 2126, event streams 2128, event updates 2130, etc., on behalf of one or more users who may use the computer system 2100.
[0252] As an example, the communication subsystem 2124 can be configured to receive in real time a data feed 2126 from social networks such as Twitter (registered trademark) feeds, Facebook (registered trademark) updates, and / or other communication services, web feeds such as Rich Site Summary (RSS) feeds, and / or real-time updates from one or more third-party information sources.
[0253] Furthermore, the communication subsystem 2124 may be configured to receive data in the form of a continuous data stream, which can include an event stream 2128 and / or event updates 2130 of real-time events that have no explicit end and are essentially continuous or boundaryless. Examples of applications that generate continuous data can include, for example, sensor data applications, financial tickers, network performance measurement tools (such as network monitoring and traffic management applications), clickstream analysis tools, automotive traffic monitoring, and the like.
[0254] The communication subsystem 2124 can be configured to output structured and / or unstructured data feeds 2126, event streams 2128, event updates 2130, etc. to one or more databases that can communicate with one or more streaming data source computers coupled to the computer system 2100.
[0255] The computer system 2100 can be one of various types, including a handheld portable device (such as an iPhone (registered trademark) mobile phone, an iPad (registered trademark) computing tablet, a PDA), a wearable device (such as a Google Glass (registered trademark) head-mounted display), a PC, a workstation, a mainframe, a ticket vending machine, a server rack, or any other data processing system.
[0256] Due to the constantly changing nature of computers and networks, the description of the computer system 2100 shown in the figures is merely intended to be a specific example. Many other configurations are possible that include more or fewer components than the system shown in the figures. For example, customized hardware may be used and / or certain elements may be implemented in hardware, firmware, software (including applets), or combinations thereof. Additionally, connections to other computing devices such as network input / output devices may be employed. Based on the disclosure and teachings provided herein, one of ordinary skill in the art will understand other methods and / or ways to implement various embodiments.
[0257] Although specific embodiments have been described, various modifications, changes, alternative structures, and equivalents are also encompassed within the scope of the present disclosure. Embodiments are not limited to operating within a particular data processing environment and can operate freely within multiple data processing environments. Further, although embodiments have been described using a specific series of transactions and steps, it should be apparent to one of ordinary skill in the art that the scope of the present disclosure is not limited to the series of transactions and steps described. The various features and aspects of the foregoing embodiments may be used individually or together.
[0258] Furthermore, although embodiments have been described using specific combinations of hardware and software, it should be recognized that other combinations of hardware and software are within the scope of the present disclosure. Embodiments may be implemented using only hardware, or only software, or combinations thereof. The various processes described herein may be implemented on the same processor or on different processors in any combination. Thus, when a component or service is described as being configured to perform an operation, such a configuration may be realized, for example, by designing an electronic circuit to perform this operation, by programming a programmable electronic circuit (such as a microprocessor) to perform this operation, or by any combination thereof. Processes can communicate using a variety of techniques, including but not limited to prior art techniques for inter-process communication, and different pairs of processes may use different techniques, or the same pair of processes may use different techniques at different times.
[0259] Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive. However, it is apparent that additions, deletions, omissions, and other modifications and changes may be made without departing from the broader ideas and scope as set forth in the claims. Thus, although specific embodiments of the disclosure have been described, these are not intended to be limiting. Various changes and equivalents are within the scope of the appended claims.
[0260] The terms "a," "an," and "the" and similar referents used in the context of describing the disclosed embodiments (in particular, in the context of the appended claims) are to be construed to cover both the singular and the plural unless otherwise indicated herein or clearly contradicted by the context. The terms "comprising," "having," "including," and "containing" are to be construed as open-ended terms (i.e., meaning "including, but not limited to") unless otherwise noted. The term "connected" is to be construed as being either internally or partially contained in, connected to, or joined together with, even if there is something intervening. The recitation of a range of values herein is merely intended to provide a convenient way of referring individually to each separate value falling within the range, and each separate value is incorporated herein as if it were individually recited herein. All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by the context. The use of any examples, or exemplary language (e.g., "such as") provided herein is merely intended to better illuminate the embodiments and does not impose a limitation on the scope of the disclosure unless otherwise claimed. No language in this specification should be construed as indicating any non-claimed element as essential to the practice of the disclosure.
[0261] Disjunctive language, such as the phrase "at least one of X, Y, or Z," is generally intended, unless specifically stated otherwise, to convey that items, conditions, etc. can be any one of X, Y, or Z, or any combination thereof (e.g., X, Y, and / or Z) within the context in which it is used. Thus, such disjunctive language is not generally intended to, and should not, mean that a particular embodiment requires the presence of at least one of each of at least one of X, at least one of Y, or at least one of Z.
[0262] In this specification, preferred embodiments of the present disclosure are described, including the best mode known to the applicant for carrying out the present disclosure. Variations of such preferred embodiments may become apparent to those skilled in the art upon reading the foregoing description. Those skilled in the art should be able to adopt such variations as needed, and the present disclosure may be practiced otherwise than as specifically described herein. Accordingly, the present disclosure includes all modifications and equivalents of the subject matter recited in the claims appended hereto as permitted by applicable law. Further, any combination of the foregoing elements in all possible variations of the embodiments is included in the present disclosure unless otherwise specifically indicated herein.
[0263] All references, including publications, patent applications, and patents, cited herein are hereby incorporated by reference in their entirety, to the same extent as if each reference were individually and specifically indicated to be incorporated by reference and were set forth in its entirety herein.
[0264] In the foregoing specification, aspects of the present disclosure have been described with reference to specific embodiments of the present specification. Those skilled in the art will recognize that the present disclosure is not limited thereto. The various features and aspects of the foregoing disclosure may be used individually or together. Further, embodiments may be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of the present specification. Accordingly, the present specification and drawings are to be regarded as illustrative rather than restrictive.
[0265] Embodiments may be implemented by using a computer program product including computer programs / instructions, which, when executed by a processor, cause the processor to execute any of the methods described in the present disclosure.
Claims
1. It is a method, The computing system includes generating a first security key associated with the source file system to encrypt and decrypt multiple file keys in the source file system, at least in part, based on a first master key, the source file system is configured to send snapshot differences during a replication process, the snapshot differences between two snapshots of the source file system are identified, and the method further includes: The computing system includes generating a second security key associated with the target file system for encrypting and decrypting a plurality of file keys in the target file system, at least in part, based on a second master key, the target file system is configured to receive the snapshot difference during the replication process, and the method further includes: The computing system includes generating a session key, at least in part, based on a third master key, for encrypting and decrypting the snapshot differences transferred between the source file system and the target file system during the replication process, wherein the session key is valid during the session. A method wherein the first master key, the second master key, and the third master key are different keys.
2. The method according to claim 1, wherein the session is the period between the start of the replication process in the source file system and the end of the replication process in the target file system.
3. The method according to claim 1, further comprising object storage configured to receive the snapshot difference from the source file system and transfer the snapshot difference to the target file system, wherein the snapshot difference transferred between the source file system and the target file system is transferred to the target file system.
4. The method according to claim 3, further comprising: encrypting the snapshot difference using the session key before transferring the snapshot difference to the object storage; and decrypting the snapshot difference using the session key after transferring the snapshot difference to the target file system.
5. The method according to claim 1, further comprising transferring the session key from the control plane of the source file system to the control plane of the target file system.
6. The method according to claim 1, wherein the session key is associated with globally unique resource identification information.
7. The method according to claim 1, wherein each of the plurality of file keys in the source file system is associated with a specific file in the source file system and is used to encrypt and decrypt the file data of the specific file in the source file system, and each of the plurality of file keys in the target file system is associated with a specific file in the target file system and is used to encrypt and decrypt the file data of the specific file in the target file system.
8. In order to request the use of the first security key, the key requester in the source file system is authenticated, The method according to claim 1, further comprising authenticating a key requester in the target file system in order to request the use of the second security key.
9. The method according to claim 8, wherein authenticating the key requester in the source file system includes checking the identification number of the replication process and the identification number of the source file system.
10. A program for causing one or more processors to execute the method according to any one of claims 1 to 9.
11. It is a system, One or more processors, The system comprises one or more computer-readable media for storing computer-executable instructions, and when the computer-executable instructions are executed by the one or more processors, the system... The computing system is configured to generate a first security key associated with the source file system in order to encrypt and decrypt multiple file keys in the source file system, at least in part on a first master key, the source file system is configured to send snapshot differences during the replication process, the snapshot differences between two snapshots of the source file system are identified, and the computer executable instructions further to the system The computing system is configured to generate a second security key associated with the target file system in order to encrypt and decrypt a plurality of file keys in the target file system, at least in part on a second master key, the target file system is configured to receive the snapshot difference during the replication process, and the computer executable instructions further enable the system to The computing system is made to generate a session key, at least partially based on a third master key, in order to encrypt and decrypt the snapshot differences transferred between the source file system and the target file system during the replication process, and the session key is valid during the session. A system in which the first master key, the second master key, and the third master key are different keys.
12. The system according to claim 11, wherein the session is the period between the start of the replication process in the source file system and the end of the replication process in the target file system.
13. The system according to claim 11 or 12, wherein the session key is associated with globally unique resource identification information.
14. In order to request the use of the first security key, the key requester in the source file system is authenticated, The system according to claim 11 or 12, further comprising authenticating a key requester in the target file system in order to request the use of the second security key.
15. The system according to claim 14, wherein authenticating the key requester in the source file system includes checking the identification number of the replication process and the identification number of the source file system.