Snapshot hardware security module and disk metadata store
The snapshot management system addresses the challenge of recreating key management data in cloud infrastructure by capturing and storing snapshots, ensuring secure and efficient access to encrypted client data across nodes, even in failure scenarios.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- ORACLE INT CORP
- Filing Date
- 2022-04-27
- Publication Date
- 2026-04-23
AI Technical Summary
Cloud infrastructure systems face challenges in efficiently recreating key management data and metadata across multiple nodes in the event of failures, such as node outages or data corruption, leading to loss of access to encrypted client data.
A snapshot management system captures and stores snapshots of key management data and metadata across nodes, using a snapshot orchestrator to validate and store these snapshots efficiently, enabling recreation of key management data in case of failures.
Ensures efficient recreation of key management data and metadata, maintaining access to encrypted client data even in the event of node failures or data loss, with secure and efficient inter-regional storage and validation processes.
Smart Images

Figure 0007850747000001 
Figure 0007850747000002 
Figure 0007850747000003
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 63 / 194,023, filed on May 27, 2021, entitled "SNAPSHOTTING HARDWARE SECURITY MODULES AND DISK METADATA STORES", and U.S. Non - Provisional Patent Application No. 17 / 719,010, filed on April 12, 2022. The entire contents of the foregoing applications are hereby incorporated by reference in their entirety for all purposes.
[0002] Field The disclosed technology relates to key management in a cloud environment. More specifically, the disclosed technology relates to capturing snapshots of key management data and corresponding metadata across a series of nodes within a cloud infrastructure system and storing the snapshots to efficiently recreate the key management data in the event of a failure in one or more nodes.
Background Art
[0003] Background A cloud infrastructure (CI) system can perform multiple functions, such as storing data across nodes (e.g., servers) within the CI system and enabling queries of the stored data. Further, a CI system can store and maintain large amounts of client data. Often, client data can be encrypted using keys such that it can only be accessed and / or modified upon the provision of an appropriate key. Client keys can be used to encrypt / decrypt portions of client data across the CI system.
[0004] Furthermore, when client data is added or modified in the CI system, multiple keys may be modified. For example, a new key can be added to encrypt a new dataset, or a key can be removed from multiple keys in response to another dataset being removed from the CI system. Changes to multiple keys associated with a client can be recorded and maintained by one or more nodes within the CI system. For example, when keys are added, deleted, modified, etc., a key management node can record and remember all changes to the multiple keys and the associated metadata on one or more security modules within the CI system. The key management node can be used to retrieve keys across the CI system and decrypt encrypted client data. [Overview of the project]
[0005] overview This embodiment relates to taking snapshots of key management data and storing these snapshots for efficient recreation of the key management data in the event of an outage on one or more nodes. A first exemplary embodiment relates to a method for capturing snapshots of key management data across a set of nodes. This method may include a snapshot orchestrator requesting snapshot instances from each of a set of nodes across one or more regions within a cloud infrastructure service. Each snapshot instance may provide multiple changes to multiple client keys maintained by each of the nodes in the set. Furthermore, each change may correspond to an entry in the log.
[0006] This method may also include the snapshot orchestrator retrieving snapshot instances and their corresponding metadata from each of a set of nodes. This method may also include validating the snapshot instances received from the set of nodes. This method may also include storing the snapshot instances and their corresponding metadata on a storage node in response to the validation of the snapshot instances. This allows for subsequent acquisition of snapshot instances, recreation of logs, and modification of multiple keys on any of the nodes in the set.
[0007] Another exemplary embodiment relates to a snapshot management system. The snapshot management system may comprise a processor and a non-temporary computer-readable medium. The non-temporary computer-readable medium, when executed by the processor, contains instructions that cause the processor to request a snapshot instance for each of a set of nodes. Each snapshot instance may provide multiple changes to multiple client keys maintained by each of the set of nodes. Each change may correspond to a log sequence record in an append-only log.
[0008] The instruction can further cause the processor to retrieve snapshot instances and their corresponding metadata from each of the set of nodes. This instruction can further cause the processor to validate the snapshot instances received from the set of nodes based on the identified entropy value of each snapshot instance. This instruction can further cause the processor to store the snapshot instances and their corresponding metadata in a storage node in response to the validation of the snapshot instances. The snapshot instances and their corresponding metadata can be configured to be used on any of the nodes in the set for append-only logging and to recreate changes to multiple keys.
[0009] Another exemplary embodiment relates to a non-temporary computer-readable medium. This non-temporary computer-readable medium can store a set of instructions that, when executed by a processor, cause the processor to perform a process. This process may include requesting snapshot instances from each of a set of nodes across one or more regions within a cloud infrastructure service. Each snapshot instance may specify multiple changes to multiple client keys maintained by each of the set of nodes. Furthermore, each change may correspond to an entry in a log.
[0010] This process may further include retrieving snapshot instances and their corresponding metadata from each of a set of nodes. This process may further include validating the snapshot instances received from the set of nodes by determining that the entropy value of each received snapshot instance falls within a threshold similarity. In response to the validation of the snapshot instances, this process may further include storing the snapshot instances and their corresponding metadata in a storage node. This enables subsequent acquisition of snapshot instances, recreation of logging, and modification of multiple keys on any of the nodes in the set.
[0011] Furthermore, embodiments may be implemented by using a computer program product which includes a computer program / instruction that, when executed by the processor, causes the processor to execute any of the methods / technologies described herein. [Brief explanation of the drawing]
[0012] [Figure 1] This is a block diagram of an exemplary snapshot management system according to at least one embodiment. [Figure 2] A block diagram of an example system for capturing a snapshot of key management data, according to at least one embodiment, is shown. [Figure 3] A block diagram of an exemplary snapshot orchestrator according to at least one embodiment is shown. [Figure 4] This is an exemplary bitmap of a snapshot instance according to at least one embodiment. [Figure 5] A block diagram of an exemplary snapshot file format bitmap, according to at least one embodiment, is shown. [Figure 6] This is a block diagram of an exemplary method for capturing a snapshot of key management data across a series of nodes, according to at least one embodiment. [Figure 7] This block diagram shows one pattern for implementing a cloud infrastructure as a service system, according to at least one embodiment. [Figure 8] This block diagram shows another pattern for implementing cloud infrastructure as a service system, according to at least one embodiment. [Figure 9] This block diagram shows another pattern for implementing cloud infrastructure as a service system, according to at least one embodiment. [Figure 10] This block diagram shows another pattern for implementing cloud infrastructure as a service system, according to at least one embodiment. [Figure 11] A block diagram illustrating an exemplary computer system according to at least one embodiment. [Modes for carrying out the invention]
[0013] Detailed explanation Cloud infrastructure (CI) services can include one or more storage modules (or "vaults") for securely managing and storing cryptographic keys across various CI services and applications. CI vaults can include secure and resilient management services that enable secure data encryption while maximizing efficiency in performing highly available hardware provisioning and / or software patching.
[0014] In CI services, security-sensitive workloads can utilize cryptographic keys stored in a Hardware Security Module (HSM) for encryption. CI services can provide key management services using HSMs. Client keys can be stored in a vault, and cross-domain key replication allows clients to recover from domain failures by replicating keys across CI service domains. Depending on whether replication is enabled for the client, existing and new keys can be replicated along with associated metadata that accurately reflects the replication status.
[0015] The CI key management service can store customer keys using HSMs. Each HSM can contain multiple partitions, each providing a logical and security boundary for the keys stored within that HSM. Depending on the client creating a vault, the vault can contain associated partitions containing all keys within the vault created within that partition. The contents of one partition may not be shareable with another partition unless they share the same trust chain and maintain a security explosion radius local to a single partition.
[0016] The HSM can be modeled into a replicated state machine (RSM). Each bolt for customers can include an associated write-ahead log (WAL). The WAL can record operations on the bolt in an ordered manner, including but not limited to key creation and deletion, metadata updates, etc. The WAL can be stored in a persistent key and value store of an internal CI that supports transaction reading and writing. Similar to the aforementioned computer, the bolt can be fully recreated by applying the WAL entries associated with the bolt.
[0017] Furthermore, when a bolt is created for a client, by default, the associated split within the HSM can be replicated across all hosts within the data center in the region. Enabling inter-region replication for the bolt allows the WAL to be relayed to the destination region and also applied to all hosts within that region. By default, client keys can be replicated across hosts distributed across data centers. The KMS service within a region can run within a firewall without accessing an external network (e.g., the Internet). To enable inter-region communication, a high-performance RESTful service (e.g., a WAL service) that takes WAL traffic from the KMS service to other regions within the realm can be run in each region. An inter-region call from one region to the WAL service of another region can go through an access-controlled proxy at the edge of the source region. Any service can incorporate Representational State Transfer (REST) technology and can be run geographically dispersed and on one or more load balancers. Each write-ahead log (WAL) can include an additional dedicated log that includes key changes over time. Each WAL can include a series of entries, and each entry includes an associated log sequence number (LSN). The LSN can be monotonically dense and can increase for each WAL entry.
[0018] Furthermore, CIKMS can include a plurality of distributed services that operate in cooperation across multiple regions. However, processing that spans regions can cause various problems (e.g., problems arising from Byzantine failures). For example, in a distributed system that includes such multiple regions, when an incident occurs, it can lead to partial or complete loss of hosts, disks, data stores, HSMs, or network problems, resulting in data corruption during transfer, disruption of the message delivery order, or the occurrence of splits. Additionally, this can cause bugs or incidents that manipulate data in the WAL, or other data center failures. Accordingly, since the WAL can include the entire various system states related to its customer vault, the service can be made stateless. Each WAL entry can encode both client data and metadata regarding the CI service.
[0019] For example, for any of various reasons, a node that hosts a key management service (e.g., a hardware security module (HSM)) can malfunction or lose its functionality. For example, when a large-scale outage event occurs, the power or functionality of the device hosting the HSM is lost, and as a result, the data in the HSM can be lost. When the key management service data is lost, multiple keys (and the key tracking logs) can be lost, and as a result, access to the encrypted client data becomes impossible.
[0020] In some cases, CIKMS can be a sharded system with multiple region shards. The region shard identifier can encode a combination of the CI realm, region, and KMS shard as an integer. The region shard identifier can include a 32-bit integer representing each shard across the region / realm. This enables efficient calculation for generating the WAL membership.
[0021] Furthermore, multiple WALs can be compared to determine if they are identical up to a specified point in time. This can be done by calculating an incremental checksum for each log entry and comparing the actual entropy value with the calculated entropy value during recalculation. Performing anti-entropy checks can help identify bit flips or resolve inter-WAL contention to converge WALs across regions.
[0022] Furthermore, each WAL can contain a large number of entries (e.g., millions). Therefore, applying multiple WALs every time a host needs to be bootstrapped or every time cross-region replication is configured for a client can be inefficient. To replicate WAL data across regions, snapshots of the WAL can be taken periodically and replicated between regions. Each snapshot can include an entropy value. Snapshots can be taken at the same time point in all single state machines, and a centralized snapshot orchestrator can compare the entropy values and promote the snapshots as valid after consensus. Additionally, by performing a checksum of the entire snapshot, the data can be processed to determine if bit flips have occurred.
[0023] In many cases, various snapshot approaches can be used to capture snapshots of an RSM. The first snapshot approach includes a stop-the-world snapshot, where no changes to the RSM state are permitted while the snapshot is being taken. Another example of a snapshot approach is a concurrent snapshot, where a snapshot is taken while new behavior is being processed and WAL records are being processed while the snapshot is being taken. Furthermore, a complete snapshot may include the entire replicated state associated with the LSN. Alternatively, an incremental snapshot may include a series of snapshots where one snapshot depends on previous snapshots. Using incremental snapshots can reduce the size of snapshots and shorten the time it takes to create them.
[0024] This embodiment relates to capturing snapshots of key management service data (e.g., WAL for multiple keys). The snapshot orchestrator can initiate snapshots across multiple hosts on a specific instance (e.g., by a specified log sequence number (LSN) of the WAL) and synchronize the captured snapshots. The LSN can include a monotonically dense and increasing sequence number that can be used to order WAL entries. The stored snapshots can include resilient copies of the log data that allow hosts to efficiently restore in the event of a CI system failure.
[0025] This embodiment can implement a snapshot orchestrator that can periodically adjust data snapshots based on various thresholds. The snapshot orchestrator can implement a verification (anti-entropy) process to check and validate snapshot content across hosts in a cluster of host devices. Furthermore, the snapshot orchestrator can retrieve content from one or more HSMs, including metadata, to ensure that client data does not leave the boundaries of a device or system according to one or more standards / protocols. To ensure the security of snapshot data, the snapshot data can be encrypted and checksummed. For example, a portion of the snapshot data can be derived from the snapshot data and compared with the snapshot data to detect errors in the snapshot data or to verify the integrity of the snapshot data. Snapshot data can be stored across multiple devices (e.g., inter-region storage). Furthermore, the payload structure of the snapshot data and the truncation of the snapshot data allow for efficient restoration of snapshot data without a sudden increase in memory resources.
[0026] A snapshot orchestrator may include logging handlers that handle the application of individual logging records of a specific type. A WAL processor can process WAL and dispatch it to one or more logging handlers. A WAL processor can maintain its current state (e.g., current LSN) when WAL logging is applied. A WAL manager can manage WAL processors based on WAL added to / removed from the system.
[0027] In some embodiments, a method for taking snapshots of key management data is provided. A snapshot orchestrator service running on a device within a CI system can perform the methods described herein. The snapshot orchestrator can receive requests for snapshot data from multiple host devices. Snapshots captured by host devices may include key management data (e.g., a WAL providing a set of changes to multiple keys that encrypt client data) and metadata associated with the snapshot data (e.g., the host device, an identification of the LSN instance from which the snapshot data was taken). The multiple host devices may include HSM nodes and / or database nodes located across one or more regions within a CI system.
[0028] A request for snapshot data can be sent to a set of host devices. The snapshot service can take a snapshot instance on each host device and provide the data (e.g., the entropy value of the snapshot data) to the snapshot orchestrator. The snapshot service can take snapshot data in response to snapshot data in a specified state (e.g., a specified LSN instance in the write append log (WAL)). In some cases, the snapshot orchestrator can compare the obtained entropy values received from the snapshot service, and if successful, upload the set of snapshot data to an external persistent store.
[0029] A snapshot orchestrator device can acquire snapshot instances from multiple host devices and validate the snapshot data. This may include performing a validation or anti-entropy process to verify the snapshot data. In some cases, the WAL may be truncated at the time the snapshot is taken to remove a portion of the WAL prior to a specified LSN instance.
[0030] A snapshot orchestrator can store snapshot data on storage nodes. In some cases, the primary storage node can act as the truth point for snapshot data instances to coordinate / synchronize snapshot data between nodes in inter-region storage. In the event of a downtime event or loss of key management data, snapshot data can be taken and the key data can be efficiently provided to clients.
[0031] A. Overview of the Snapshot Management System The system described herein can store client keys in one or more hardware security modules (HSMs) and associated metadata on on-disk stores across one or more regions. This can be done by an internal replication system that maintains data synchronization between these hosts located in different regions. Furthermore, the system described herein can periodically take snapshots of these HSMs and associated metadata stores to back up client key data and compress replication logs. This allows for efficient retrieval of key data in the event of key data failure or loss in one or more regions.
[0032] Figure 1 is a block diagram of an exemplary snapshot management system 100. The snapshot management system 100 can provide a snapshot infrastructure that efficiently captures snapshot instances and associated metadata. The snapshot management system 100 can periodically adjust snapshots based on various configurable thresholds. The snapshot management system 100 can further validate captured snapshot instances (for example, by performing anti-entropy checks).
[0033] In some cases, the snapshot management system 100 can encrypt or perform checksums on snapshot instances to enhance their security. Furthermore, the snapshot management system 100 can enable inter-regional storage of snapshot instances. The snapshot payload structure allows for efficient restoration of snapshot instances without causing memory ballooning.
[0034] As shown in Figure 1, the snapshot management system 100 may include a snapshot orchestrator 102. The snapshot orchestrator 102 can request snapshot instances from several host nodes (e.g., 106A-B) across one or more regions. The snapshot orchestrator 102 can further process the acquired snapshot instances (e.g., WAL logging) to validate them. Depending on the validation of the snapshot instances, the instances can be stored for later acquisition.
[0035] The snapshot orchestrator 102 may include a snapshot acquisition subsystem 104. The snapshot acquisition subsystem 104 can initiate a request for a snapshot instance from host nodes 106A-B. For example, the snapshot acquisition subsystem 104 can add an entry to the WAL (e.g., on the LSN instance) requesting that a snapshot be taken on a specified LSN instance. Alternatively, the snapshot acquisition subsystem 104 can send a message to host nodes 106A-B requesting that a snapshot be taken by each of them.
[0036] The snapshot orchestrator 102 can run within the cluster manager and can initiate snapshots on host nodes by adding new entries to the WAL. The snapshot orchestrator 102 can perform anti-entropy checks. For example, if the anti-entropy check fails and some host nodes share a common anti-entropy value, the node where the snapshot failed is removed and allowed to catch up. Alternatively, it can issue an alarm.
[0037] The snapshot management system 100 may include several host nodes (e.g., 106A-B). Host nodes 106A-B may include HSMs deployed across multiple regions within the CI service. Each host node 106A-B may maintain separate logging of all keys created / deleted for clients. Furthermore, each host node 106A-B may include snapshot generation subsystems 108A-B for detecting requests to ingest snapshots, interacting with snapshot shotters 110A-B to ingest snapshots, and providing snapshot instances to the snapshot orchestrator 102. The snapshot shotters may be worker nodes running within each replica / service instance configured to perform snapshot-related tasks (e.g., executing snapshots, storing snapshots locally, pulling / pushing remote snapshots, etc.).
[0038] Each snapshotter 110A-B can run within each host node instance 106A-B. Snapshotters 110A-B can be part of a write-ahead log manager and are responsible for creating and storing local snapshots. Snapshotters 110A-B can expose a set of application programming interfaces (APIs) that can be called by the snapshot orchestrator 102. In some cases, several previous snapshots can be remotely stored in each snapshotter 110A-B.
[0039] The snapshot acquisition subsystem 104 can acquire snapshot instances 112A to B provided by each host node 106A to B. Snapshot instances 112A to B can be provided to the snapshot verification subsystem 114 for verification. For example, the snapshot verification subsystem 114 can compare snapshot instances 112A to B, identify any inconsistencies within instances 112A to B, truncate snapshot instances 112A to B, perform anti-entropy checks, and so on.
[0040] The verified snapshot instance 116 can be stored in the storage module 118. The storage module 118 can maintain several snapshot instances for subsequent acquisition in the event of a failure in any of the host nodes 106A to B. For example, in response to a failure of host node 106a, the latest snapshot instance can be provided to host node 106a to replicate the client's WAL.
[0041] Figure 2 shows a block diagram 200 of an example system for capturing snapshots of key management data. As shown in Figure 2, the snapshot orchestrator 102 can request snapshots from a set of snapshot services running on the host device. As shown in Figure 2, the write-ahead log 202 can include an append-only log which can include several instances (e.g., 204A-B). Furthermore, the snapshot orchestrator 102 can add an entry (e.g., 206) to the WAL 202 that contains an LSN instance instructing the host nodes 106A-B to capture the snapshot. The host nodes 106A-B can then identify the WAL instance 206 in the WAL 202 and begin capturing the snapshot instance as described herein.
[0042] The snapshot orchestrator 102 can coordinate the snapshot creation process across multiple host nodes and is responsible for pushing snapshots to the storage module 118 for long-term storage. Each region can include a pair region for inter-region redundancy. The snapshot orchestrator 102 can push snapshots to the storage module 118 for each region. The storage module 118 can delete snapshots associated with the WAL after a pre-configured tombstone interval (e.g., 30 days).
[0043] The snapshot orchestrator 102 can track WAL snapshots and associated metadata within a table (e.g., a database such as a key-value store database). The snapshot orchestrator 102 can run as a lease-based daemon and can periodically request snapshots based on time intervals. The snapshot orchestrator 102 can expose an API for performing snapshots on demand by writing to the WAL 202. Furthermore, the snapshot orchestrator 102 daemon can perform WAL truncation based on the latest snapshot LSN. A snapshot can be initiated when one or more conditions are met. Examples of conditions may be based on the latest snapshot creation time being greater than the snapshot interval and the state of the latest snapshot in progress. Another example of a condition may be based on the latest LSN in the WAL being greater than the maximum allow record in the WAL. In response to the conditions being met, a log entry may be written to the WAL, and a snapshot metadata record may be added to the snapshot's ongoing metadata table.
[0044] In some cases, backups of multiple key partitions can capture snapshots of HSM partitions. This allows for encrypted snapshots of each partition on the HSM. Any backup functionality provided by the HSM (configuration, keys, users) can be included in the encrypted blob. To ensure the accuracy of the snapshots, the backups can be validated and the accuracy of the data checked (e.g., using cyclic redundancy checks (CRC) such as the CRC32 error detection algorithm). Since partition data may be read from a card, this validation may include performing checksums. The backups can be stored on a designated storage device, and when the restore function is initialized, the restore function can be run on a new temporary partition that has passed the integrity check.
[0045] In some cases, when a snapshot entry is detected in the WAL, all host nodes can perform backups on their HSM cards. However, because the backups are encrypted, comparing them to each other can be difficult. Therefore, all data for each backup type can be downloaded and returned along with a blob of the encrypted data. The orchestrator can then sort the data using the metadata and compare the backups to determine if they point to the same object.
[0046] Each host node 106A-B can implement a snapshotter 110A-B. Snapshotters 110A-B can run as part of a write-ahead log manager. Snapshotters 110A-B can capture snapshots and save them locally. Furthermore, snapshotters 110A-B can expose an API that allows returning the status of a snapshotter associated with a particular LSN, listing snapshots that exist locally for the LSN, and returning snapshots that can be written directly to local / remote storage modules.
[0047] Snapshotters 110A-B can capture new local snapshots. Snapshot requests can flow through the WAL on each host node 106A-B using a specific logging type. The WAL manager then invokes the WAL processor to take the snapshot using a snapshot writer, which is responsible for writing the snapshot and its associated metadata.
[0048] In some cases, a snapshot may be stored as a collection of "mini" snapshots within an overall WAL snapshot. Each logging handler processor can be responsible for versioning the data within the snapshot associated with a particular logging handler. A snapshot writer can perform integrity checks and also manage the versioning of the entire snapshot file. Integrity checks can be signed with the encryption key value of the snapshot via a private key available on each individual machine.
[0049] Downloaded snapshots and / or locally saved snapshots are stored in subdirectory directories within the WAL Manager. Snapshots and the corresponding device status can be managed by the Snapshot Manager. Each snapshot can reside in a storage device subdirectory named with a specified LSN, along with a marker file indicating whether the snapshot is complete and verified.
[0050] B. Snapshot Orchestrator As described above, a snapshot orchestrator can take snapshots from one or more host nodes, maintain those snapshots (and their corresponding metadata), and recreate the WAL if any host nodes experience loss. For example, in response to a failure in a region containing host nodes, the snapshot orchestrator can provide the host nodes with the snapshots and associated metadata and recreate the WAL that provides client key management data.
[0051] Figure 3 shows a block diagram 300 of an exemplary snapshot orchestrator 302. The snapshot orchestrator 302 can acquire snapshot data (WAL) from one or more host nodes as described herein. The orchestrator 302 can poll the acquired WAL and perform additional processing (e.g., tombstonening). Snapshots received from host nodes and associated metadata can be maintained in a table. The metadata may include the host node name, LSN instance, snapshot status, and the time the record was created. Snapshots and their corresponding metadata maintained in the table can be stored locally in the WAL repository 304.
[0052] The snapshot orchestrator 302 can truncate the WAL via the WAL truncator 306. For example, snapshots ingested by each host node can be truncated by a specified LSN instance. Truncating snapshots reduces the size of each snapshot while simultaneously providing a more accurate representation of the WAL for recreating it on a particular host node.
[0053] Snapshots and their corresponding metadata can be stored in the storage module 310. The storage module 310 may include one or more interconnected storage nodes that enable efficient storage and retrieval of snapshot data in response to snapshot requests from the orchestrator 302.
[0054] The inter-region snapshot copier 308 may include components responsible for transporting backups and existing snapshots between regions. The inter-region snapshot copier 308 can copy snapshot data across regions by passing through the snapshot metadata table if there are any mismatches between nodes in different regions. The inter-region snapshot copier 308 may have a lease-based daemon that can connect to different regions individually. For each snapshot whose status indicates successful snapshot capture, the snapshot can be copied to the inter-region snapshot copier 308. The data replication handler of the inter-region snapshot copier 308 may be responsible for applying data replication WAL records to local databases. Each local database may contain fewer than a threshold number of entries. In some cases, only writes to data may be blocked during snapshot creation, and the database may be read by the snapshot orchestrator's data plane while snapshot creation is in progress. Local databases may contain arrays of key-value pairs. Key-value pairs can be ordered within the database by LSN values and calculated CRC32 values. The CRC32 can be incrementally changed as data records are updated. The CRC32 function converts a variable-length string into an 8-character string, which is the text representation of the hexadecimal value of a 32-bit binary sequence. Therefore, as the precision of the data increases, the CRC32 value can also improve.
[0055] C. Bitmap of snapshot data Figure 4 shows an example of a bitmap 400 of snapshot instances. Snapshot data can include snapshot instances taken from a host node, as described herein. Snapshot data can include multiple handler snapshots of variable length. Handler snapshots can include individual snapshot instances taken by each snapshot service.
[0056] The bitmap of snapshot data 400 may include any of the handler snapshot fields 402, 404, snapshot header field 406, CRC32 field 408, and version field 410. For example, a first snapshot instance (e.g., an instance from an HSM) may be included in the first handler snapshot field 402. A second snapshot instance (e.g., from a Berkeley database (BDB) instance) may be included in the second handler snapshot field 404. The snapshot instances can be combined by the snapshot orchestrator to form a snapshot payload that can later be parsed and restored by nodes (e.g., HSM, BDB nodes). The handler snapshot fields 402, 404 may include a variable length (vlen).
[0057] Snapshot data may also include a snapshot header 406 that identifies the snapshot size in bytes. Snapshot data can be encrypted using an asymmetric key with Public Key Cryptography Standard (PKCS) envelope encryption. Snapshot data may also include a cyclic redundancy check (CRC32) field 408 and a version field 410 that specifies the version of the node providing the snapshot data. The CRC32 field 408 and the version field 410 may not be encrypted so that verification of each snapshot instance can be confirmed without needing to decrypt the snapshot instance or access the public key.
[0058] A snapshot instance can be associated with a metadata file that specifies the handler snapshot metadata (e.g., name and offset). For example, this might include the LSN associated with the snapshot, the write-ahead log name, the handler snapshot metadata for each handler snapshot, the number of handler snapshots, the length of the snapshot (in bytes) including the header, the snapshot version, the logging type, and the entropy value.
[0059] Figure 5 shows a block diagram of an exemplary snapshot file format bitmap 500. The snapshot file format bitmap 500 may include a version field 502, a header CRC32 field 504, a header field 506, and several table fields (e.g., 508, 510, 512).
[0060] The version field 502 can provide the version of the snapshot. The snapshot length can include a signed integer (e.g., 4 bytes) that limits the snapshot size to 2GB. The header CRC32 field 504 can include 4 bytes specifying the CRC of the remaining bytes, including the header. The header field 506 can include metadata necessary to index each table. The file format can contain multiple tables of variable length (e.g., Table 1, Table 2, Table 3) (e.g., within table fields 508, 510, 512).
[0061] D. Flow process for taking and managing snapshots of key management data As described above, a snapshot management system can periodically capture snapshots of key management data across a set of nodes and store these snapshots for subsequent recreation of the key management data. For example, in response to an outage at a first node, the latest snapshot can be provided to the first node, allowing the key management data to be recreated at that node. This enables the efficient recreation of key management data (e.g., changes to client keys over time) in response to an outage event (e.g., failure) at one or more nodes. Figure 6 is a block diagram of an exemplary method 600 for capturing snapshots of key management data across a set of nodes. For example, the method described herein can be performed by a snapshot orchestrator (e.g., 102) interacting with host nodes (e.g., 106A-B) as part of a snapshot management system (e.g., 100).
[0062] In 602, the method may include requesting snapshot instances from a set of nodes. Each of the set of nodes may have a snapshotter capable of capturing snapshots of key management data on each node. The key management data may include changes (e.g., additions / deletions) to any of several keys specific to a client. Each change to any of the keys may be associated with an entry in the logging (e.g., a specified LSN instance in the write append log (WAL)). In some cases, the set of nodes may be located across one or more regions within the cloud infrastructure service, for example, in different data centers in geographically different regions. In this example, each change to any of several client keys identified on the first node in the first region is synchronized between other nodes in other regions of the cloud infrastructure service by an inter-region snapshot copier.
[0063] In some cases, requesting a snapshot involves adding an entry to the log that specifies the request to ingest a snapshot instance. In response, each node in the set can ingest the snapshot instance depending on the identification of the entry added to the log. In some cases, a snapshot instance can be requested in response to either the expiration of a threshold period (e.g., every minute, every 10 minutes) or the detection of the addition of a threshold number of entries to the log (e.g., every 100 entries, every 1000 entries).
[0064] In 604, the snapshot orchestrator can retrieve snapshot instances and corresponding metadata from a set of nodes. A snapshot instance can include snapshots of logs (e.g., WAL) on each of the nodes in the set. The metadata for each snapshot instance can include various aspects of the snapshot and / or the node that ingested the snapshot. For example, the metadata may specify the log sequence number in the WAL that triggered the snapshot, the snapshot status, the identifier of the host node ingesting the snapshot, the entropy value, etc. In some cases, the metadata may include a key for accessing a specified partition of a first node, each partition of the first node independently maintaining client logging. The snapshot orchestrator can use the key to access the partition and provide the stored snapshot instances and corresponding metadata to the partition of the first node.
[0065] In 606, the snapshot orchestrator can validate the set of snapshot data. This may include identifying that the set of snapshot data contains human-readable data that can be used to restore the client's key state on a given instance. Furthermore, this may include performing a validation or anti-entropy process to validate the snapshot data.
[0066] As an example, a snapshot orchestrator can identify an entropy value specific to each snapshot instance and each corresponding node. As used herein, an entropy value may include a value (or set of values) derived from one or more characteristics, such as the LSN from which the snapshot was requested, a checksum value specifying the number of bytes in the snapshot instance's header or the snapshot instance itself, or the time the snapshot was captured. In other words, the entropy values described and used herein do not need to be numerical measures of uncertainty of results, as commonly used in the industry. Instead, according to the techniques described herein, entropy values can be generated based on the aforementioned characteristics and may be specific to each snapshot instance and / or each corresponding node. The snapshot orchestrator can determine whether the identified entropy values of each received snapshot instance are within a threshold similarity. Snapshot instances can be validated depending on whether their entropy values are within the threshold similarity. For example, a snapshot instance can be validated if all entropy values are the same or within the threshold similarity.
[0067] In some cases, the snapshot orchestrator can truncate each snapshot instance to remove changes to multiple client keys that have corresponding entries in the log prior to a specified entry in the log. For example, if each snapshot instance contains 1000 entries, the snapshot instance can be truncated to only the most recent 100 entries to improve the storage efficiency of the snapshot instance while simultaneously allowing all associated client keys to be recreated.
[0068] In some cases, depending on the validation of the snapshot instance, the table can be updated to include metadata corresponding to the snapshot instance. The metadata can be associated with a specified log sequence number in the append-only log, the status of the snapshot instance at the time of ingestion, or the timestamp when the request for the snapshot instance was validated.
[0069] In version 608, the snapshot orchestrator can store snapshot data on a storage node and a set of host nodes (e.g., inter-region storage). In some cases, the primary storage node can act as the truth point for snapshot data instances to coordinate / synchronize snapshot data among nodes that store WAL separately. In the event of a downtime event or loss of key management data, snapshot data can be taken and the key data can be efficiently provided to clients.
[0070] As an example, the snapshot orchestrator may receive an outage notification at a first node. This may result from an outage at the first node (e.g., a power outage) or another failure at the first node. The snapshot orchestrator may periodically check for outages at any node. In another example, a node (or corresponding alarm system) may send an outage notification in response to the detection of an outage at any node. The snapshot orchestrator may retrieve stored snapshot instances and corresponding metadata from the storage node. Furthermore, the snapshot orchestrator may provide the stored snapshot instances and corresponding metadata to the first node. The first node may use the stored snapshot instances and corresponding metadata to recreate a log specifying changes to multiple client keys. Furthermore, various modifications and equivalents include relevant and appropriate combinations of the features disclosed in the embodiments.
[0071] E. Overview of IaaS As mentioned above, Infrastructure as a Service (IaaS) is a specific type of cloud computing. IaaS can be configured to provide computing resources that are virtualized over a public network (such as the internet). In the IaaS model, a cloud computing provider can host infrastructure components (e.g., servers, storage, network nodes (e.g., hardware), deployment software, platform virtualization (e.g., hypervisor layer)). In some cases, the IaaS provider can also provide various services that accompany these infrastructure components (e.g., billing, monitoring, logging, load balancing, and clustering). Therefore, since these services can be policy-driven, IaaS users may be able to implement policies that drive load balancing to maintain application availability and performance.
[0072] In some cases, IaaS customers can access resources and services over a wide area network (WAN), such as the internet, and install the rest of their application stack using the cloud provider's services. For example, a user can log into an IaaS platform and create virtual machines (VMs), install operating systems (OS) on each VM, deploy middleware such as databases, create storage buckets for workloads and backups, and even install enterprise software on those VMs. The customer can then use the provider's services to perform various functions such as balancing network traffic, troubleshooting application issues, monitoring performance, and managing disaster recovery.
[0073] In most cases, the cloud computing model may require the participation of a cloud provider. This cloud provider may, but does not have to be, a third-party service specializing in providing IaaS (e.g., offering, renting, or selling). Entities can also choose to deploy a private cloud and become their own infrastructure service provider.
[0074] In some cases, IaaS deployment is the process of deploying a new application, or a new version of an application, to a prepared application server, etc. This may also include the process of preparing the server (e.g., installing libraries, daemons, etc.). This is often managed by the cloud provider under the hypervisor layer (e.g., servers, storage, network hardware, and virtualization). Therefore, the customer may be responsible for handling things like the OS, middleware, and / or application deployment (e.g., self-service virtual machines, e.g., that can be spun up on demand).
[0075] In some examples, IaaS provisioning can even refer to acquiring the computers or virtual hosts to be used and installing the necessary libraries or services on them. In most cases, provisioning is not included in the deployment, so you may need to perform provisioning first.
[0076] In some cases, IaaS provisioning presents two distinct challenges. First, there's the initial challenge of provisioning an initial set of infrastructure before doing anything. Second, there's the challenge of evolving existing infrastructure after everything has been provisioned (e.g., adding new services, modifying services, removing services, etc.). In some cases, these two challenges can be addressed by declaratively defining the configuration of the infrastructure. In other words, the infrastructure (e.g., what components are needed and how they interact) can be defined by one or more configuration files. Thus, the overall topology of the infrastructure (e.g., which resources depend on which resources and how they all work together) can be described declaratively. In some cases, once the topology is defined, a workflow can be generated to create and / or manage the various components described in the configuration files.
[0077] In some examples, infrastructure can have many interconnected elements. For example, there may be one or more virtual private clouds (VPCs), also known as core networks (e.g., a potentially on-demand pool of configurable and / or shared computing resources). In some examples, there may also be one or more inbound / outbound traffic group rules and one or more virtual machines (VMs) provisioned to define how inbound and / or outbound network traffic is configured. Other infrastructure elements such as load balancers and databases can also be provisioned. Infrastructure can evolve incrementally as more infrastructure elements are desired or added.
[0078] In some cases, continuous deployment techniques may be used to enable the deployment of infrastructure code across various virtual computing environments. Furthermore, the techniques described enable infrastructure management within these environments. In some examples, a service team may write code that they wish to deploy to one or more, but often many, different production environments (e.g., across various different geographical locations, and possibly worldwide). However, in some examples, the infrastructure to which the code will be deployed must first be set up. In some cases, provisioning can be done manually, resources can be provisioned using provisioning tools, and / or code can be deployed using deployment tools after the infrastructure has been provisioned.
[0079] Figure 7 is a block diagram 700 showing an example pattern of an IaaS architecture according to at least one embodiment. A service operator 702 can be communicatively coupled to a secure host tenant 704 which may include a virtual cloud network (VCN) 706 and a secure host subnet 708. In some examples, the service operator 702 may use one or more client computing devices, which may be portable handheld devices (e.g., iPhone®, mobile phones, iPad®, computing tablets, personal digital assistants (PDAs)) or wearable devices (e.g., Google Glass® head-mounted displays), running software such as Microsoft Windows Mobile®, and / or various mobile operating systems such as iOS, WindowsPhone, Android, BlackBerry 9, PalmOS, as well as the internet, email, short message service (SMS), Blackberry®, or other valid communication protocols. Alternatively, the client computing device may be a general-purpose personal computer, including, for example, personal computers and / or laptop computers running various versions of Microsoft Windows®, Apple Macintosh®, and / or Linux® operating systems. The client computing device may be a workstation computer running one of a variety of commercially available UNIX® or UNIX-like operating systems, including, but not limited to, various GNU / Linux operating systems such as Google Chrome OS.Alternatively, or in addition, the client computing device may be any other electronic device, such as a thin client computer, an internet-enabled game system (e.g., a Microsoft Xbox game console with or without Kinect® gesture input), and / or a personal messaging device that can communicate via a network that has access to the VCN706 and / or the Internet.
[0080] VCN706 may include a local peering gateway (LPG) 710 that can communicately connect to Secure Shell (SSH) VCN712 via LPG710 included in SSHVCN712. SSHVCN712 may include an SSH subnet 714, and SSHVCN712 may communicately connect to control plane VCN716 via LPG710 included in control plane VCN716. Furthermore, SSHVCN712 may communicately connect to data plane VCN718 via LPG710. Control plane VCN716 and data plane VCN718 may be included in a service tenant 719 owned and / or operated by an IaaS provider.
[0081] The control plane VCN 716 may include a control plane demilitarized zone (DMZ) layer 720 that functions as a perimeter network (e.g., part of the corporate network between the corporate intranet and the external network). DMZ-based servers have limited liability and can help deter breaches. Furthermore, the DMZ layer 720 may include a control plane application layer 724 that may include one or more load balancer (LB) subnets 722, an application subnet 726, and a control plane data layer 728, which may include a database (DB) subnet 730 (e.g., a front-end DB subnet and / or a back-end DB subnet). The LB subnet 722 included in the control plane DMZ layer 720 may be communicatively coupled to the application subnet 726 included in the control plane application layer 724 and an internet gateway 734 that may be included in the control plane VCN 716, and the application subnet 726 may be communicatively coupled to the DB subnet 730 included in the control plane data layer 728, as well as a service gateway 736 and a network address translation (NAT) gateway 738. The control plane VCN716 may include a service gateway 736 and a NAT gateway 738.
[0082] The control plane VCN 716 may include a data plane mirror application layer 740 which may include an application subnet 726. The application subnet 726 included in the data plane mirror application layer 740 may include a virtual network interface controller (VNIC) 742 which can run a compute instance 744. The compute instance 744 may communicatively combine the application subnet 726 of the data plane mirror application layer 740 with the application subnet 726 which may be included in the data plane application layer 746.
[0083] The data plane VCN718 may include a data plane application layer 746, a data plane DMZ layer 748, and a data plane data layer 750. The data plane DMZ layer 748 may include an LB subnet 722 that can be communicatively coupled to the application subnet 726 of the data plane application layer 746 and the internet gateway 734 of the data plane VCN718. The application subnet 726 may be communicatively coupled to the service gateway 736 of the data plane VCN718 and the NAT gateway 738 of the data plane VCN718. The data plane data layer 750 may also include a DB subnet 730 that can be communicatively coupled to the application subnet 726 of the data plane application layer 746.
[0084] The Internet gateway 734 of the control plane VCN716 and data plane VCN718 can be communicatively coupled to a metadata management service 752, which can be communicatively coupled to the public internet 754. The public internet 754 can be communicatively connected to the NAT gateway 738 of the control plane VCN716 and data plane VCN718. The service gateway 736 of the control plane VCN716 and data plane VCN718 can be communicatively coupled to a cloud service 756.
[0085] In some cases, a service gateway 736 on the control plane VCN716 or data plane VCN718 can make application programming interface (API) calls to a cloud service 756 without going through the public internet 754. API calls from the service gateway 736 to the cloud service 756 can be one-way: the service gateway 736 can make an API call to the cloud service 756, and the cloud service 756 can send the requested data to the service gateway 736. However, the cloud service 756 may not be able to initiate an API call to the service gateway 736.
[0086] In some examples, a secure host tenant 704 can connect directly to a service tenant 719, or it may be isolated otherwise. A secure host subnet 708 can communicate with an SSH subnet 714 via an LPG 710, which can enable bidirectional communication through systems that would otherwise be isolated. Connecting the secure host subnet 708 to the SSH subnet 714 allows the secure host subnet 708 to access other entities within the service tenant 719.
[0087] The control plane VCN716 allows users of service tenant 719 to set up or provision desired resources. Desired resources provisioned within the control plane VCN716 can be deployed or used within the data plane VCN718. In some examples, the control plane VCN716 can be separated from the data plane VCN718, and the data plane mirror application layer 740 of the control plane VCN716 can communicate with the data plane application layer 746 of the data plane VCN718 via a VNIC 742, which can be included in the data plane mirror application layer 740 and the data plane application layer 746.
[0088] In some examples, a user or customer of the system may make requests, such as create, read, update, or delete (CRUD) operations, via the public internet 754, which can communicate the requests to the metadata management service 752. The metadata management service 752 can communicate the requests to the control plane VCN 716 via the internet gateway 734. This request may be received by the LB subnet 722, which is included in the control plane DMZ layer 720. The LB subnet 722 may determine that the request is valid, and in response to this determination, the LB subnet 722 may send the request to the application subnet 726, which is included in the control plane application layer 724. If the request is validated and a call to the public internet 754 is required, the call to the public internet 754 may be sent to the NAT gateway 738, which can make calls to the public internet 754. Memory that may be desirable to be stored by the request can be stored in the DB subnet 730.
[0089] In some cases, the data plane mirror application layer 740 can facilitate direct communication between the control plane VCN716 and the data plane VCN718. For example, it may be desirable to apply configuration changes, updates, or other appropriate modifications to resources contained in the data plane VCN718. Through VNIC742, the control plane VCN716 can communicate directly with the resources contained in the data plane VCN718, thereby enabling it to perform configuration changes, updates, or other appropriate modifications to the resources contained in the data plane VCN718.
[0090] In some embodiments, the control plane VCN716 and data plane VCN718 can be included in the service tenant 719. In this case, the system's user or customer cannot own or operate either the control plane VCN716 or the data plane VCN718. Instead, the IaaS provider can own or operate the control plane VCN716 and the data plane VCN718, both of which may be included in the service tenant 719. This embodiment can enable network isolation that can prevent a user or customer from interacting with resources of other users or other customers. This embodiment also allows the system's user or customer to store databases privately without having to rely on the public internet 754, which may not have the desired level of threat protection for storage.
[0091] In another embodiment, the LB subnet 722 included in the control plane VCN 716 may be configured to receive signals from the service gateway 736. In this embodiment, the control plane VCN 716 and the data plane VCN 718 may be configured to be invoked by the IaaS provider's customers without calling the public internet 754. The IaaS provider's customers may prefer this embodiment because the database used by the customer may be controlled by the IaaS provider and stored in a service tenant 719 which can be isolated from the public internet 754.
[0092] Figure 8 is a block diagram 800 illustrating another pattern example of an IaaS architecture according to at least one embodiment. A service operator 802 (e.g., service operator 702 in Figure 7) can be communicatively coupled to a secure host tenant 804 (e.g., secure host tenant 704 in Figure 7), which may include a virtual cloud network (VCN) 806 (e.g., VCN706 in Figure 7) and a secure host subnet 808 (e.g., secure host subnet 708 in Figure 7). The VCN 806 may include a local peering gateway (LPG) 810 (e.g., LPG710 in Figure 7), which may be communicatively coupled to a secure shell (SSH) VCN 812 (e.g., SSHVCN712 in Figure 7) via the LPG710 contained in SSHVCN 812. SSHVCN812 may include SSH subnet 814 (e.g., SSH subnet 714 in Figure 7), and SSHVCN812 may be communicably coupled to control plane VCN816 (e.g., control plane VCN716 in Figure 7) via LPG810 included in control plane VCN816. Control plane VCN816 may include service tenant 819 (e.g., service tenant 719 in Figure 7), and data plane VCN818 (e.g., data plane VCN718 in Figure 7) may include customer tenant 821, which may be owned or operated by a user or customer of the system.
[0093] The control plane VCN816 may include a control plane DMZ layer 820 (e.g., control plane DMZ layer 720 in Figure 7) which may include an LB subnet 822 (e.g., LB subnet 722 in Figure 7), a control plane application layer 824 (e.g., control plane application layer 724 in Figure 7) which may include an application subnet 826 (e.g., application subnet 726 in Figure 7), and a control plane data layer 828 (e.g., control plane data layer 728 in Figure 7) which may include a database (DB) subnet 830 (e.g., similar to DB subnet 730 in Figure 7). The LB subnet 822 included in the control plane DMZ layer 820 can be communicatively coupled to the application subnet 826 included in the control plane application layer 824 and to an internet gateway 834 (e.g., internet gateway 734 in Figure 7) which may be included in the control plane VCN 816. The application subnet 826 can be communicatively coupled to the DB subnet 830 included in the control plane data layer 828, as well as to a service gateway 836 (e.g., service gateway in Figure 7) and a network address translation (NAT) gateway 838 (e.g., NAT gateway 738 in Figure 7). The control plane VCN 816 may include the service gateway 836 and the NAT gateway 838.
[0094] The control plane VCN 816 may include a data plane mirror application layer 840 (e.g., data plane mirror application layer 740 in Figure 7) which may include an application subnet 826. The application subnet 826 included in the data plane mirror application layer 840 may include a virtual network interface controller (VNIC) 842 (e.g., VNIC 742) which may run a compute instance 844 (e.g., similar to compute instance 744 in Figure 7). The compute instance 844 may facilitate communication between the application subnet 826 of the data plane mirror application layer 840 and the application subnet 826, which may be included in the data plane application layer 846 (e.g., data plane application layer 746 in Figure 7) via the VNIC 842 included in the data plane mirror application layer 840 and the VNIC 842 included in the data plane application layer 846.
[0095] The Internet gateway 834 included in the control plane VCN816 can be communicatively coupled to the metadata management service 852 (e.g., the metadata management service 752 in Figure 7), which can be communicatively coupled to the public internet 854 (e.g., the public internet 754 in Figure 7). The public internet 854 can be communicatively coupled to the NAT gateway 838 included in the control plane VCN816. The service gateway 836 included in the control plane VCN816 can be communicatively coupled to the cloud service 856 (e.g., the cloud service 756 in Figure 7).
[0096] In some examples, the data plane VCN818 may be contained within the customer tenant 821. In this case, the IaaS provider may provide a control plane VCN816 for each customer, and the IaaS provider may set up a unique compute instance 844 contained within the service tenant 819 for each customer. Each compute instance 844 may enable communication between the control plane VCN816 contained within the service tenant 819 and the data plane VCN818 contained within the customer tenant 821. The compute instance 844 may enable resources provisioned within the control plane VCN816 contained within the service tenant 819 to be deployed or otherwise used within the data plane VCN818 contained within the customer tenant 821.
[0097] In another example, an IaaS provider's customer may have a database residing within customer tenant 821. In this example, control plane VCN 816 may include a data plane mirror app tier 840 that can include app subnet 826. The data plane mirror app tier 840 may reside within data plane VCN 818, but does not have to. That is, the data plane mirror app tier 840 can access customer tenant 821, but does not have to reside within data plane VCN 818 and may be owned or operated by the IaaS provider's customer. The data plane mirror app tier 840 may be configured to make calls to data plane VCN 818, but does not have to be configured to make calls to any entity contained within control plane VCN 816. A customer may want to deploy or otherwise use resources in data plane VCN 818 that are provisioned within control plane VCN 816, and the data plane mirror app tier 840 can facilitate the customer's desired deployment or other use of resources.
[0098] In some embodiments, a customer of the IaaS provider can apply filters to the data plane VCN818. In this embodiment, the customer can determine what the data plane VCN818 can access and can restrict access from the data plane VCN818 to the public internet 854. The IaaS provider may not be able to apply filters or control the data plane VCN818's access to external networks or databases. Applying customer filters and controls to the data plane VCN818 contained in a customer tenant 821 can help isolate the data plane VCN818 from other customers and the public internet 854.
[0099] In some embodiments, the cloud service 856 can be invoked by the service gateway 836 to access services that may not reside on the public internet 854, the control plane VCN 816, or the data plane VCN 818. The connection between the cloud service 856 and the control plane VCN 816 or data plane VCN 818 may not be live or continuous. The cloud service 856 may reside on a different network owned or operated by the IaaS provider. The cloud service 856 may be configured to receive calls from the service gateway 836, or it may be configured not to receive calls from the public internet 854. Some cloud services 856 may be isolated from other cloud services 856, and the control plane VCN 816 may be isolated from cloud services 856 that do not have to be in the same region as the control plane VCN 816. For example, the control plane VCN 816 may be located in "region 1", and the cloud service "deployment 7" may be located in regions 1 and "region 2". If a call to deployment 7 is made by a service gateway 836 included in the control plane VCN816 in region 1, that call may be sent to deployment 7 in region 1. In this example, the control plane VCN816, or deployment 7 in region 1, may not be communicatively coupled to or communicating with deployment 7 in region 2.
[0100] Figure 9 is a block diagram 900 illustrating another pattern example of an IaaS architecture according to at least one embodiment. A service operator 902 (e.g., service operator 702 in Figure 7) can be communicatively coupled to a secure host tenant 904 (e.g., secure host tenant 704 in Figure 7), which may include a virtual cloud network (VCN) 906 (e.g., VCN706 in Figure 7) and a secure host subnet 908 (e.g., secure host subnet 708 in Figure 7). The VCN 906 may include an LPG 910 (e.g., LPG710 in Figure 7) which can be communicatively coupled to the SSHVCN 912 (e.g., SSHVCN712 in Figure 7) via an LPG 910 contained within the SSHVCN 912. SSHVCN912 may include SSH subnet 914 (e.g., SSH subnet 714 in Figure 7), and SSHVCN912 may be communicatively coupled to control plane VCN916 (e.g., control plane VCN716 in Figure 7) via LPG910 included in control plane VCN916, and may be communicatively coupled to data plane VCN918 (e.g., data plane 718 in Figure 7) via LPG910 included in data plane VCN918. Control plane VCN916 and data plane VCN918 may be included in service tenant 919 (e.g., service tenant 719 in Figure 7).
[0101] The control plane VCN916 may include a control plane DMZ layer 920 (e.g., control plane DMZ layer 720 in Figure 7) which may include a load balancer (LB) subnet 922 (e.g., LB subnet 722 in Figure 7), a control plane application layer 924 (e.g., control plane application layer 724 in Figure 7) which may include an application subnet 926 (e.g., similar to application subnet 726 in Figure 7), and a control plane data layer 928 (e.g., control plane data layer 728 in Figure 7) which may include a DB subnet 930. The LB subnet 922 included in the control plane DMZ layer 920 may be communicatively coupled to the application subnet 926 included in the control plane application layer 924 and to an internet gateway 934 (e.g., internet gateway 734 in Figure 7) which may be included in the control plane VCN 916. The application subnet 926 may be communicatively coupled to the DB subnet 930, service gateway 936 (e.g., service gateway in Figure 7), and network address translation (NAT) gateway 938 (e.g., NAT gateway 738 in Figure 7) included in the control plane data layer 928. The control plane VCN 916 may include the service gateway 936 and the NAT gateway 938.
[0102] The data plane VCN918 can include a data plane application layer 946 (e.g., data plane application layer 746 in Figure 7), a data plane DMZ layer 948 (e.g., data plane DMZ layer 748 in Figure 7), and a data plane data layer 950 (e.g., data plane data layer 750 in Figure 7). The data plane DMZ layer 948 can include an LB subnet 922, which can be communicatively coupled to the trusted application subnet 960 and untrusted application subnet 962 of the data plane application layer 946, and the internet gateway 934 included in the data plane VCN918. The trusted application subnet 960 can be communicatively coupled to the service gateway 936 included in the data plane VCN918, the NAT gateway 938 included in the data plane VCN918, and the DB subnet 930 included in the data plane data layer 950. The untrusted application subnet 962 can be communicatively coupled to the service gateway 936 included in the data plane VCN918 and the DB subnet 930 included in the data plane data layer 950. The data plane data layer 950 may include a DB subnet 930 that can be communicatively coupled to a service gateway 936 included in the data plane VCN 918.
[0103] An untrusted application subnet 962 may include one or more primary VNICs 964(1)-(N) that can be communicatively connected to tenant virtual machines (VMs) 966(1)-(N). Each tenant VM 966(1)-(N) may be communicatively connected to its respective application subnet 967(1)-(N), which may be included in its respective container exit VCN 968(1)-(N), which may be included in its respective customer tenant 970(1)-(N). Each secondary VNIC 972(1)-(N) can facilitate communication between the untrusted application subnet 962, which is included in the data plane VCN 918, and the application subnets included in the container exit VCN 968(1)-(N). Each container exit VCN 968(1)-(N) may include a NAT gateway 938 that can be communicatively connected to the public internet 954 (e.g., public internet 754 in Figure 7).
[0104] The Internet gateway 934, included in the control plane VCN916 and the data plane VCN918, can communicate with a metadata management service 952 (for example, the metadata management system 752 in Figure 7), which can communicate with the public internet 954. The public internet 954 can communicate with a NAT gateway 938, included in the control plane VCN916 and the data plane VCN918. The service gateway 936, included in the control plane VCN916 and the data plane VCN918, can communicate with a cloud service 956.
[0105] In some embodiments, the data plane VCN918 can be integrated with the customer tenant 970. This integration may be beneficial or desirable for the IaaS provider's customer, such as when support during code execution is required. The customer may provide execution of potentially destructive code, code that may communicate with other customer resources, or code that may cause other undesirable effects. Accordingly, the IaaS provider can decide whether to execute the code provided to the IaaS provider by the customer.
[0106] In some examples, an IaaS provider's customer may request temporary network access from the IaaS provider to add functionality to a dataplane tier application 946. The code that performs the functionality can run on VMs 966(1)-(N), and the code cannot be configured to run anywhere else on the dataplane VCN 918. Each VM 966(1)-(N) can connect to one customer tenant 970. Each container 971(1)-(N) contained within VMs 966(1)-(N) may be configured to run code. In this case, a double isolation may exist (e.g., code execution in containers 971(1)-(N), which may be contained within at least VMs 966(1)-(N) that are in an untrusted application subnet 962), which may help prevent incorrect or undesirable code from damaging the IaaS provider's network or another customer's network. Containers 971(1)-(N) may be communicatively coupled to customer tenant 970 and may be configured to send or receive data from customer tenant 970. Containers 971(1)-(N) may not be configured to send or receive data from any other entities in the data plane VCN918. Once code execution is complete, the IaaS provider may terminate or otherwise destroy containers 971(1)-(N).
[0107] In some embodiments, a trusted application subnet 960 may execute code owned or operated by the IaaS provider. In this embodiment, the trusted application subnet 960 may be communicatively coupled to a DB subnet 930 and configured to perform CRUD operations within the DB subnet 930. An untrusted application subnet 962 may be communicatively coupled to the DB subnet 930, but in this embodiment, the untrusted application subnet may be configured to perform read operations within the DB subnet 930. Containers 971(1)-(N), which may be included in each customer's VM966(1)-(N) and capable of executing code from the customer, do not necessarily have to be communicatively coupled to the DB subnet 930.
[0108] In other embodiments, the control plane VCN916 and the data plane VCN918 do not have to be directly communicatively coupled. In this embodiment, direct communication between the control plane VCN916 and the data plane VCN918 is not required. However, communication can be performed indirectly through at least one method. The LPG910 may be established by an IaaS provider that can facilitate communication between the control plane VCN916 and the data plane VCN918. In another example, the control plane VCN916 or the data plane VCN918 can make a call to a cloud service 956 via a service gateway 936. For example, a call from the control plane VCN916 to the cloud service 956 may include a request for a service that can communicate with the data plane VCN918.
[0109] Figure 10 is a block diagram 1000 illustrating another pattern example of an IaaS architecture according to at least one embodiment. A service operator 1002 (e.g., service operator 702 in Figure 7) can be communicatively coupled to a secure host tenant 1004 (e.g., secure host tenant 704 in Figure 7), which may include a virtual cloud network (VCN) 1006 (e.g., VCN706 in Figure 7) and a secure host subnet 1008 (e.g., secure host subnet 708 in Figure 7). VCN 1006 may include an LPG 1010 (e.g., LPG710 in Figure 7) which can be communicatively coupled to SSHVCN 1012 (e.g., SSHVCN712 in Figure 7) via an LPG 1010 contained within SSHVCN 1012. SSHVCN1012 may include SSH subnet 1014 (e.g., SSH subnet 714 in Figure 7), and SSHVCN1012 may be communicatively coupled to control plane VCN1016 (e.g., control plane VCN716 in Figure 7) via LPG1010 included in control plane VCN1016, and may be communicatively coupled to data plane VCN1018 (e.g., data plane 718 in Figure 7) via LPG1010 included in data plane VCN1018. Control plane VCN1016 and data plane VCN1018 may be included in service tenant 1019 (e.g., service tenant 719 in Figure 7).
[0110] The control plane VCN1016 may include a control plane DMZ layer 1020 (e.g., control plane DMZ layer 720 in Figure 7) which may include an LB subnet 1022 (e.g., LB subnet 722 in Figure 7), a control plane application layer 1024 (e.g., control plane application layer 724 in Figure 7) which may include an application subnet 1026 (e.g., application subnet 726 in Figure 7), and a control plane data layer 1028 (e.g., control plane data layer 728 in Figure 7) which may include a DB subnet 1030 (e.g., DB subnet 930 in Figure 9). The LB subnet 1022 included in the control plane DMZ layer 1020 can be communicatively coupled to the application subnet 1026 included in the control plane application layer 1024, and can be communicatively coupled to the internet gateway 1034 (e.g., internet gateway 734 in Figure 7), which can be included in the control plane VCN 1016. The application subnet 1026 can be communicatively coupled to the DB subnet 1030 included in the control plane data layer 1028, and can be communicatively coupled to the service gateway 1036 (e.g., the service gateway in Figure 7) and the network address translation (NAT) gateway 1038 (e.g., NAT gateway 738 in Figure 7). The control plane VCN 1016 can include the service gateway 1036 and the NAT gateway 1038.
[0111] The data plane VCN 1018 may include a data plane application layer 1046 (e.g., data plane application layer 746 in Figure 7), a data plane DMZ layer 1048 (e.g., data plane DMZ layer 748 in Figure 7), and a data plane data layer 1050 (e.g., data plane data layer 750 in Figure 7). The data plane DMZ layer 1048 may include an LB subnet 1022, which may be communicatively coupled to a trusted application subnet 1060 (e.g., trusted application subnet 960 in Figure 9), and may be communicatively coupled to an untrusted application subnet 1062 (e.g., untrusted application subnet 962 in Figure 9) of the data plane application layer 1046 and an internet gateway 1034 included in the data plane VCN 1018. A trusted application subnet 1060 can be communicatively coupled to a service gateway 1036 included in the data plane VCN 1018, and can be communicatively coupled to a NAT gateway 1038 included in the data plane VCN 1018, and to a DB subnet 1030 included in the data plane data layer 1050. An untrusted application subnet 1062 can be communicatively connected to a service gateway 1036 included in the data plane VCN 1018 and to a DB subnet 1030 included in the data plane data layer 1050. The data plane data layer 1050 may include a DB subnet 1030 that can be communicatively coupled to a service gateway 1036 included in the data plane VCN 1018.
[0112] An untrusted application subnet 1062 may contain primary VNICs 1064(1)-(N), which can be communicatively coupled to tenant virtual machines (VMs) 1066(1)-(N) residing within the untrusted application subnet 1062. Each tenant VM 1066(1)-(N) can execute code within its respective container 1067(1)-(N) and can be communicatively coupled to an application subnet 1026, which can be included in a dataplane application layer 1046, which can be included in a container exit VCN 1068. Each secondary VNIC 1072(1)-(N) can facilitate communication between the untrusted application subnet 1062, which is included in the dataplane VCN 1018, and the application subnet, which is included in the container exit VCN 1068. The container exit VCN may contain a NAT gateway 1038, which can be communicatively coupled to the public internet 1054 (e.g., public internet 754 in Figure 7).
[0113] The Internet gateway 1034, included in the control plane VCN1016 and the data plane VCN1018, can communicate with the metadata management service 1052 (for example, the metadata management system 752 in Figure 7), which can communicate with the public internet 1054. The public internet 1054 can communicate with the NAT gateway 1038, included in the control plane VCN1016 and the data plane VCN1018. The service gateway 1036, included in the control plane VCN1016 and the data plane VCN1018, can communicate with the cloud service 1056.
[0114] In some examples, the pattern shown by the architecture in block diagram 1000 of Figure 10 can be considered an exception to the pattern shown by the architecture in block diagram 900 of Figure 9, which may be desirable for the IaaS provider's customers when the IaaS provider cannot communicate directly with the customer (e.g., in a disconnected area). Each container 1067(1)-(N) contained within each customer's VM 1066(1)-(N) is accessible by the customer in real time. Each container 1067(1)-(N) may be configured to make calls to each secondary VNIC 1072(1)-(N) contained within the application subnet 1026 of the data plane application layer 1046, which may be contained within the container exit VCN 1068. The secondary VNICs 1072(1)-(N) may send calls to a NAT gateway 1038, which can send calls to the public internet 1054. In this example, the containers 1067(1)-(N), which customers can access in real time, can be isolated from the control plane VCN1016 and from other entities included in the data plane VCN1018. The containers 1067(1)-(N) may also be isolated from resources from other customers.
[0115] In another example, a customer can use containers 1067(1)-(N) to invoke cloud service 1056. In this example, the customer can execute code within containers 1067(1)-(N) to request a service from cloud service 1056. Containers 1067(1)-(N) can send this request to secondary VNICs 1072(1)-(N), which can then send the request to a NAT gateway that can send the request to the public internet 1054. The public internet 1054 can then send the request to LB subnet 1022, which is included in control plane VCN 1016, via internet gateway 1034. In response to the determination that the request is valid, the LB subnet can send the request to application subnet 1026, which can then send the request to cloud service 1056 via service gateway 1036.
[0116] It should be understood that the IaaS architectures 700, 800, 900, and 1000 shown in the figures may have components other than those shown. Furthermore, the embodiments shown in the figures are only some examples of cloud infrastructure systems that may incorporate embodiments of this disclosure. In some other embodiments, the IaaS system may have more or fewer components than those shown, may combine two or more components, or may have different configurations or arrangements of components.
[0117] In certain embodiments, the IaaS system described herein may include a suite of applications, middleware, and database service products delivered to customers in a self-service, subscription-based, elastically scalable, reliable, highly available, and secure manner. An example of such an IaaS system is Oracle Cloud Infrastructure (OCI), offered by the assignee.
[0118] Figure 11 shows an exemplary computer system 1100 in which various embodiments can be implemented. System 1100 can be used to implement any of the computer systems described above. As shown in the figure, computer system 1100 includes a processing unit 1104 that communicates with several peripheral subsystems via a bus subsystem 1102. These peripheral subsystems may include a processing accelerator 1106, an I / O subsystem 1108, a storage subsystem 1118, and a communication subsystem 1124. The storage subsystem 1118 includes a tangible computer-readable storage medium 1122 and system memory 1110.
[0119] The bus subsystem 1102 provides a mechanism that enables various components and subsystems of the computer system 1100 to communicate with each other as intended. Although the bus subsystem 1102 is schematically shown as a single bus, multiple buses can be utilized in alternative embodiments of the bus subsystem. The bus subsystem 1102 may be any of several types of bus structures, including a memory bus or memory controller, peripheral bus, and local bus, using any of the various bus architectures. For example, such architectures may include the Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus. This can be implemented as a mezzanine bus manufactured according to the IEEEP1386.1 standard.
[0120] The processing unit 1104 can be implemented as one or more integrated circuits (e.g., conventional microprocessors or microcontrollers) and controls the operation of the computer system 1100. One or more processors may be included in the processing unit 1104. These processors may include single-core processors or multi-core processors. In certain embodiments, the processing unit 1104 may be implemented as one or more independent processing units 1132 and / or 1134, each containing a single-core processor or a multi-core processor. In other embodiments, the processing unit 1104 may be implemented as a quad-core processing unit formed by integrating two dual-core processors onto a single chip.
[0121] In various embodiments, the processing unit 1104 can execute various programs in response to program code and can maintain multiple concurrently running programs or processes. At any given time, some or all of the program code to be executed may reside in the processor 1104 and / or the storage subsystem 1118. Through appropriate programming, the processor 1104 can provide the various functions described above. The computer system 1100 may further include a processing accelerator 1106 which may include a digital signal processor (DSP), a dedicated processor, etc.
[0122] The I / O subsystem 1108 may include user interface input devices and user interface output devices. User interface input devices may include keyboards, pointing devices such as mice and trackballs, touchpads and touchscreens integrated into displays, scroll wheels, click wheels, dials, buttons, switches, keypads, audio input devices with voice command recognition systems, microphones, and other types of input devices. User interface input devices may also include motion sensing and / or gesture recognition devices, such as the Microsoft Kinect® motion sensor, which allow the user to control and interact with input devices, such as the Microsoft Xbox® 360 game controller, through a natural user interface using gestures and voice commands. User interface input devices may also include eye gesture recognition devices, such as the Google Glass® blink detector, which detects eye activity from the user (e.g., blinking while taking photos and / or selecting menus) and translates eye gestures as input to an input device (e.g., Google Glass®). Furthermore, the user interface input device may include a voice recognition sensing device that enables the user to interact with a voice recognition system (e.g., Siri® Navigator) through voice commands.
[0123] User interface input devices may include, but are not limited to, three-dimensional (3D) mice, joysticks or pointing sticks, gamepads and graphic tablets, as well as audio / visual devices such as speakers, digital cameras, digital video cameras, portable media players, webcams, image scanners, fingerprint scanners, barcode readers, 3D scanners, 3D printers, laser rangefinders, and eye-tracking devices. Furthermore, user interface input devices may also include medical imaging input devices such as computed tomography, magnetic resonance imaging, positional radiography, and medical ultrasound equipment. User interface input devices may also include audio input devices such as MIDI keyboards and digital musical instruments.
[0124] User interface output devices may include non-visual displays such as display subsystems, indicator lights, or audio output devices. Display subsystems may include flat panel devices such as those using cathode ray tubes (CRTs), liquid crystal displays (LCDs), or plasma displays, projection devices, touchscreens, etc. Generally, the use of the term “output device” is intended to include all possible types of devices and mechanisms for outputting information from the computer system 1100 to a user or another computer. For example, user interface output devices include, but are not limited to, a variety of display devices that visually convey text, graphics, and audio / video information, such as monitors, printers, speakers, headphones, car navigation systems, plotters, audio output devices, and modems.
[0125] The computer system 1100 may include a storage subsystem 1118 having software elements that are shown to be currently located in the system memory 1110. The system memory 1110 can store program instructions that can be loaded and executed on the processing unit 1104, as well as data generated during the execution of these programs.
[0126] Depending on the configuration and type of the computer system 1100, the system memory 1110 may be volatile (such as random access memory (RAM)) and / or non-volatile (such as read-only memory (ROM) or flash memory). RAM typically contains data and / or program modules that are immediately accessible to the processing unit 1104 and / or currently operating and executing by the processing unit 1104. In some implementations, the system memory 1110 may contain several different types of memory, such as static random access memory (SRAM) or dynamic random access memory (DRAM). In some implementations, a basic input / output system (BIOS), which contains basic routines that help transfer information between elements within the computer system 1100, such as during startup, may typically be stored in ROM. As an example, and not an limitation, the system memory 1110 also refers to application programs 1112, program data 1114, and the operating system 1116, which may include client applications, web browsers, middle-tier applications, relational database management systems (RDBMS), etc. For example, operating systems 1116 may include various versions of Microsoft Windows®, Apple Macintosh®, and / or Linux operating systems, various commercially available UNIX® or UNIX-like operating systems (including, but not limited to, various GNU / Linux operating systems, Google Chrome® OS, etc.), and / or mobile operating systems such as iOS, Windows® Phone, Android® OS, BlackBerry® OS, and Palm® OS.
[0127] The storage subsystem 1118 may also provide a tangible, computer-readable storage medium for storing basic programming and data structures that provide functionality in some embodiments. When executed by a processor, software (programs, code modules, instructions) that provides the aforementioned functionality may be stored in the storage subsystem 1118. These software modules or instructions may be executed by the processing unit 1104. The storage subsystem 1118 may also provide a repository for storing data used in accordance with this disclosure.
[0128] The storage subsystem 1100 may also include a computer-readable storage medium reader 1120 that can be further connected to the computer-readable storage medium 1122. Together, and optionally in combination with the system memory 1110, the computer-readable storage medium 1122 can comprehensively represent a storage medium for temporarily and / or more permanently storing, storing, transmitting, and retrieving computer-readable information, in addition to remote, local, fixed, and / or removable storage devices.
[0129] The computer-readable storage medium 1122 containing code or a portion of code may also include any suitable medium known or used in the art, including, but not limited to, storage and communication media such as volatile and non-volatile, removable and non-removable media, implemented in any way or technique for storing and / or transmitting information. This may include tangible computer-readable storage media such as RAM, ROM, electronically erasable programmable ROM (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD), or other optical storage devices, magnetic cassettes, magnetic tapes, magnetic disk storage devices or other magnetic storage devices, or other tangible computer-readable media. This may also include intangible computer-readable media such as any other medium that can be used to transmit data signals, data transmissions, or desired information and is accessible by the computing system 1100.
[0130] As an example, the computer-readable storage medium 1122 may include a hard disk drive that reads or writes to a non-removable non-volatile magnetic medium, a magnetic disk drive that reads or writes to a removable non-volatile magnetic disk, and an optical disk drive that reads or writes to a removable non-volatile optical disk such as a CD-ROM, DVD, Blu-ray® disc, or other optical medium. The computer-readable storage medium 1122 may also include, but is not limited to, Zip® drives, flash memory cards, Universal Serial Bus (USB) flash drives, Secure Digital (SD) cards, DVD discs, digital videotapes, etc. The computer-readable storage medium 1122 may also include solid-state drives (SSDs) based on non-volatile memory such as flash memory-based SSDs, enterprise flash drives, solid-state ROMs, SSDs based on volatile memory such as solid-state RAM, dynamic RAM, and static RAM, DRAM-based SSDs, magnetoresistive RAM (MRAM) SSDs, and hybrid SSDs that use a combination of DRAM and flash memory-based SSDs. Disk drives and associated computer-readable media can provide non-volatile storage for computer-readable instructions, data structures, program modules, and other data for the computer system 1100.
[0131] The communication subsystem 1124 provides interfaces to other computer systems and networks. The communication subsystem 1124 functions as an interface for sending and receiving data between the computer system 1100 and other systems. For example, the communication subsystem 1124 can enable the computer system 1100 to connect to one or more devices via the Internet. In some embodiments, the communication subsystem 1124 may include radio frequency (RF) transceiver components for accessing wireless voice and / or data networks (e.g., using advanced data network technologies such as cellular technology, 3G, 4G, or EDGE (Enhanced Data Rate for Global Evolution)), WiFi (IEEE 902.11 family standards, or other mobile communication technologies, or any combination thereof), Global Positioning System (GPS) receiver components, and / or other components. In some embodiments, the communication subsystem 1124 may provide, in addition to or instead of a wireless interface, a wired network connection (e.g., Ethernet).
[0132] In some embodiments, the communication subsystem 1124 may also receive input communications in the form of structured and / or unstructured data feeds 1126, event streams 1128, event updates 1130, etc., on behalf of one or more users who can use the computer system 1100.
[0133] As an example, the communication subsystem 1124 may be configured to receive data feeds 1126 in real time from users of social networks and / or other communication services such as Twitter® feeds, Facebook® updates, and web feeds such as Rich Site Summary (RSS) feeds, as well as / or real-time updates from one or more third-party information sources.
[0134] Furthermore, the communication subsystem 1124 may be configured to receive data in the form of a continuous data stream, which may include an event stream 1128 of real-time events and / or event updates 1130, which may be continuous or have no explicit end and may be essentially unlimited. Examples of applications that generate continuous data may include, for example, sensor data applications, financial tickers, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, and automotive traffic monitoring.
[0135] The communication subsystem 1124 may also be configured to output structured and / or unstructured data feeds 1126, event streams 1128, event updates 1130, etc., to one or more databases that can communicate with one or more streaming data source computers coupled to the computer system 1100.
[0136] The computer system 1100 may be one of various types, including handheld portable devices (e.g., iPhone® mobile phones, iPad® computing tablets, PDAs), wearable devices (e.g., Google Glass® head-mounted displays), PCs, workstations, mainframes, kiosks, server racks, or other data processing systems.
[0137] Due to the constantly changing nature of computers and networks, the description of the computer system 1100 shown in the figure is intended only as a specific example. Many other configurations are possible with more or fewer components than the system shown in the figure. For example, customized hardware may also be used, or certain elements may be implemented in hardware, firmware, software (including applets), or a combination thereof. Furthermore, connections to other computing devices such as network input / output devices may be used. Based on the disclosures and teachings provided herein, those skilled in the art will understand other techniques and / or methods for implementing various embodiments.
[0138] While specific embodiments have been described, various modifications, changes, alternative structures, and equivalents are also included within the scope of this disclosure. The embodiments are not limited to operation within a specific, particular data processing environment, but can freely operate within multiple data processing environments. Furthermore, while the embodiments have been described using a specific set of transactions and steps, it will be apparent to those skilled in the art that the scope of this disclosure is not limited to the described set of transactions and steps. The various features and aspects of the embodiments described above can be used individually or in combination.
[0139] Furthermore, while embodiments have been described using specific combinations of hardware and software, it should be recognized that other combinations of hardware and software are also within the scope of this disclosure. Embodiments can be implemented using hardware alone, software alone, or a combination thereof. The various processes described herein can be implemented on the same processor or on any combination of different processors. Thus, where a component or module is described as being configured to perform a particular operation, such configuration can be achieved, for example, by designing an electronic circuit to perform the operation, by programming a programmable electronic circuit (such as a microprocessor) to perform the operation, or by any combination thereof. Processes can communicate using a variety of techniques, including but not limited to conventional techniques for inter-process communication, and different pairs of processes may use different techniques, or pairs of the same process may use different techniques at different times.
[0140] Therefore, the specification and drawings should be considered illustrative, not restrictive. However, it is clear that additions, subtractions, deletions, and other modifications and changes can be made without departing from the broader intent and scope set forth in the claims. Thus, while specific embodiments of disclosure have been described, they are not intended to be limiting. Various modifications and equivalents are included in the following claims.
[0141] In the context describing the disclosed embodiments (particularly in the context of the following claims), the use of the terms “a,” “an,” “the,” and similar reference subjects should be construed to cover both singular and plural forms unless otherwise indicated herein or otherwise clearly inconsistent with the context. The terms “include,” “have,” “contain,” and “contain” should be construed as unrestricted terms (i.e., “include but not limited to”) unless otherwise specified herein. The term “connected” should be construed to be included, attached, or combined with, even if something is intervening within it. The descriptions of value ranges herein are merely intended to serve as a simplified way of referring individually to each individual value within that range unless otherwise indicated herein, and each individual value is incorporated into the specification as if it were individually described herein. All methods described herein may be performed in any appropriate order unless otherwise indicated herein or otherwise clearly inconsistent with the context. Any examples or exemplary language provided herein (e.g., "etc.") are for illustrative purposes only and, unless otherwise requested, do not limit the scope of this disclosure. Nothing in this specification should be construed as indicating that any unclaimed element is essential for the practice of this disclosure.
[0142] Disjunctive expressions, such as the phrase "at least one of X, Y, or Z," are intended to be understood in context to be generally used to indicate that an item, term, etc., can be any one of X, Y, or Z, or any combination thereof (e.g., X, Y, and / or Z), unless otherwise specified. Therefore, such disjunctive expressions are not, and should not, be intended to mean that a particular embodiment requires the presence of at least one X, at least one Y, or at least one Z, respectively.
[0143] Preferred embodiments of the present disclosure, including the best known modes for carrying out the present disclosure, are described herein. Variations of these preferred embodiments will be apparent to those skilled in the art by reading the preceding description. Those skilled in the art should be able to adopt such variations as needed, and the present disclosure can also be carried out in ways other than those specifically described herein. Accordingly, the present disclosure includes all modifications and equivalents of the subject matter described in the claims appended herein, as permitted by applicable law. Furthermore, unless otherwise indicated herein, any combination of the elements described above in all possible modifications is incorporated herein.
[0144] All references cited herein, including publications, patent applications, and patents, are incorporated herein by reference to the same extent that each reference is incorporated by reference to the same extent that it is incorporated in whole herein, as is indicated individually and specifically.
[0145] While aspects of the disclosure are described in the aforementioned specification with reference to specific embodiments, those skilled in the art will recognize that the disclosure is not limited thereto. The various features and aspects of the aforementioned disclosure can be used individually or in combination. Furthermore, embodiments can be used in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of this specification. Accordingly, the specification and drawings should be considered illustrative and not restrictive.
Claims
1. A method for capturing a snapshot of key management data across a series of nodes, the method being: The snapshot orchestrator includes requesting snapshot instances from each of the set of nodes across one or more regions within the cloud infrastructure service, each snapshot instance providing multiple changes to multiple client keys maintained by each of the set of nodes, each change corresponding to an entry in the log. The aforementioned method, The snapshot orchestrator retrieves the snapshot instance and corresponding metadata from each of the set of nodes, The snapshot orchestrator verifies the snapshot instances taken from each of the series of nodes, A method comprising: in response to the validation of the snapshot instance, the snapshot orchestrator storing the snapshot instance and its corresponding metadata in a storage node, thereby enabling subsequent acquisition of the snapshot instance, recreation of the log, and modification of the multiple client keys on any of the set of nodes.
2. The method according to claim 1, wherein requesting the snapshot includes adding an entry to the log specifying a request to capture the snapshot instance, and each of the set of nodes captures the snapshot instance in the log in response to identification of the added entry.
3. The method according to claim 1, wherein the snapshot instance is requested in response to either the expiration of a threshold period or the detection of the addition of a threshold number of entries to the log.
4. The method according to claim 1, wherein each change to any of the plurality of client keys identified at a first node in a first region is synchronized between other nodes in other regions of the cloud infrastructure service by an inter-region snapshot copier.
5. Verifying the aforementioned snapshot instance is For each of the aforementioned acquired snapshot instances, the entropy value unique to each snapshot instance and each corresponding node is identified, The method according to claim 1, comprising determining whether the identified entropy values for each of the acquired snapshot instances are within a threshold similarity, wherein the snapshot instances are verified according to the entropy values that are within the threshold similarity.
6. The method according to claim 1, further comprising the snapshot orchestrator truncating each snapshot instance to remove changes to the plurality of client keys having corresponding entries in the log prior to a specified entry in the log.
7. To receive notification that the first node has stopped, The stored snapshot instance and its corresponding metadata are retrieved from the storage node. The method according to claim 1, further comprising providing the stored snapshot instance and corresponding metadata to the first node, the first node using the stored snapshot instance and corresponding metadata to recreate the log record specifying the changes to the plurality of client keys.
8. The method according to claim 7, wherein the metadata includes a key for accessing a specified partition of the first node, each partition of the first node independently maintains client logging, and the snapshot orchestrator uses the key to access the partition and provides the stored snapshot instances and corresponding metadata to the partition of the first node.
9. It is a snapshot management system, Processor and A snapshot management system comprising a non-temporary computer-readable medium containing instructions for implementing a snapshot orchestrator, wherein the instructions cause the processor to perform the method described in any one of claims 1 to 8.
10. A program that causes the processor of a snapshot orchestrator to perform the method described in any one of claims 1 to 8.
Citation Information
Patent Citations
Computer, history management method, and program for history management
JP2004246479A
Method and system for data backup
JP2009512077A
Information recovery method and apparatus using a snapshot database
JP2012509520A