Capabilities and Alert Associations

JP2025506460A5Pending Publication Date: 2025-12-24ORACLE INT CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024547127
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-02-03
Filing Date
2023-02-06
Publication Date
2025-12-24

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Techniques are described for monitoring the health of services in a computing environment, such as a data center. More particularly, the present disclosure describes techniques for monitoring the health and availability of capabilities in a computing environment, such as a data center, by enabling association of alerts with the capabilities. A capability represents a set of resources in a data center. By allowing alerts to be associated with a capability, the health or availability of the associated capability may be monitored or confirmed by tracking the state of the alert associated with the capability. For example, if an alert associated with a particular capability is triggered, this may indicate that the particular capability and one or more resources corresponding to the particular capability are not in a healthy state. Thus, the health of the associated capability may be confirmed by monitoring the alerts associated with the capability.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of and priority to the following applications: (1) U.S. Provisional Patent Application No. 63 / 314,972, "Associating Capabilities and Alarms," ​​filed on February 28, 2022 (2) U.S. Provisional Patent Application No. 63 / 308,003, "Techniques for Bootstrapping a Region Build," filed on February 8, 2022 (3) U.S. Provisional Patent Application No. 63 / 312,814, "Techniques for Implementing Virtual Data Centers," filed on February 22, 2022 (4) U.S. Patent Application No. 18 / 164,283, "Associating Capabilities and Alarms," ​​filed on February 3, 2023 (Attorney Docket No. 088325-1325791-306500US). The entire contents of the above applications are incorporated herein by reference.

[0002] Technology Areas The present disclosure relates to improved techniques for monitoring the health of services in a computing environment, such as a data center. More specifically, the present disclosure describes techniques for monitoring the health and availability of capabilities in a computing environment, such as a data center, by enabling association of alerts with capabilities. [Background technology]

[0003] background A cloud infrastructure provider may provide cloud computing infrastructure and related services in many geographic areas around the world. To provide this infrastructure, the cloud infrastructure provider may operate one or more data centers corresponding to a local geographic area. These data centers may be included as part of a "region," which is a logical abstraction of the geographic area and computing resources of the one or more data centers. Building a new region may include provisioning computing resources, configuring infrastructure, and deploying code to those resources. Conventional techniques for building a region involve significant manual operations. Bootstrapping existing services into a new region can be difficult because the services may depend on the functionality of other existing services and / or resources in the region. Relying on manual operations to bootstrap services and / or build a region may not scale well because it incurs significant time costs and introduces risks associated with manual configuration errors. Summary of the Invention [Means for solving the problem]

[0004] overview The present disclosure relates to improved techniques for monitoring the health of services in a computing environment, such as a data center. More particularly, the present disclosure describes techniques for monitoring the health and availability of capabilities in a computing environment, such as a data center, by enabling association of alerts with capabilities. A capability represents a set of resources in a data center. By enabling alerts to be associated with capabilities, the health or availability of the associated capability may be monitored or confirmed by tracking the state of the alert associated with the capability. For example, if an alert associated with a particular capability is triggered, this may indicate that the particular capability and one or more resources corresponding to the particular capability are not in a healthy state. Thus, the health of the associated capability may be confirmed by monitoring the alerts associated with the capability. Various embodiments are described herein, including methods, systems, non-transitory computer-readable storage media storing programs, codes, or instructions executable by one or more processors, and the like.

[0005] In an embodiment, a capability may identify a functional unit associated with a service and represent a set of resources associated with the service in a data center. By allowing alerts to be associated with a capability, the health or availability of the associated capability may be monitored or confirmed by tracking the state of the alert associated with the capability. For example, if an alert associated with a particular capability is triggered, this may indicate that the particular capability and one or more resources corresponding to the particular capability are not in a healthy state. Thus, the health of the associated capability may be confirmed by monitoring the alert associated with the capability. A number of such capabilities and associated alerts may be configured in a computing environment, such as a data center.

[0006] In one embodiment, a technique (e.g., a method, a system, a non-transitory computer-readable storage medium storing executable instructions) is provided for creating an association between an alert and a capability, where the capability corresponds to a set of resources for a service. Once the association is created, the alert can be monitored and, if determined to be in a triggered state, the state of the capability associated with the alert can be changed from healthy to unhealthy. In one embodiment, the association between the alert and the capability is declared in a block configuration associated with the service.

[0007] In one embodiment, the alert-capability association information may be used to determine the health of issued capabilities in a computing environment, such as a data center. A set of issued capabilities in the data center may be determined. A set of alerts associated with the set of capabilities may be identified based on the alert-capability association information. The alerts may be monitored and a first alert in the set of alerts may be determined to be in a triggered state. A subset of the set of issued capabilities is determined that is associated with the first alert in the triggered state. Additionally, a state of each capability in the subset of capabilities is updated from a healthy state to an unhealthy state.

[0008] In one general aspect, the techniques may include identifying an association between an alert and a capability, where the capability corresponds to a function associated with the first service. The techniques may also include monitoring the alert and may further include determining, based on the monitoring, that the alert is in a triggered state. The techniques may also include changing a state of the capability from a healthy state to an unhealthy state in response to the monitoring determination.

[0009] Implementations may include one or more of the following features: In these techniques, the function associated with the first service corresponds to a resource associated with the service. In these techniques, the association between the alert and the capability is declared in a flock configuration of the service that identifies a set of resources associated with the service. In these techniques, the flock configuration identifies one or more parameters associated with the service and a configuration setting of at least one resource associated with the service. These techniques may include determining that a release of the flock for the second service has failed, determining that the flock depends on a capability, and outputting a message indicating an unhealthy state of the capability and the alert as a reason for the failure to release the flock for the second service. These techniques may include delaying the release of the flock based on changing a state of the capability from healthy to unhealthy. These techniques may also include retrying the release of the flock for the second service upon determining that the state of the capability is healthy. The techniques may include, in response to changing a state of a capability to unhealthy, determining one or more dependent capabilities that depend on the capability whose state was changed from healthy to unhealthy, and changing a state of each of the one or more dependent capabilities to unhealthy. The techniques may include an implementation in which the monitoring is performed by a telemetry service. The techniques may include creating an association between the alert and a second capability, where the second capability corresponds to a function associated with a second service, where the second service may be different from the first service. Additionally, the techniques may include identifying a set of one or more capabilities issued at the data center that are marked as healthy and associated with the alert, and changing a state of each capability in the set of one or more capabilities from healthy to unhealthy.

[0010] In one general aspect, the techniques may include determining a set of direct capabilities for a flock configuration of a flock to be scheduled for release. The techniques may also include identifying a set of indirect capabilities that depend on the set of direct capabilities. Some techniques may further include determining a set of alerts associated with at least one of the set of direct capabilities or the set of indirect capabilities. The techniques may also include monitoring a set of alert states for the set of alerts. Some techniques may further include determining, based on the monitoring, that at least one alert of the set of alerts is in a triggered state. The techniques may also include changing a state of a capability associated with the triggered alert from healthy to unhealthy in response to determining that the at least one alert is in a triggered state.

[0011] Implementations may include one or more of the following features: The techniques include updating a set of direct capabilities and a health state of the set of direct capabilities from a healthy state to an unhealthy state. The techniques include delaying the release of a flock based on changing a state of a capability from healthy to unhealthy. Implementations of the described techniques may include hardware, techniques or processes, or computer tangible media.

[0012] The above, together with other features and embodiments, will become more apparent with reference to the following specification, claims, and accompanying drawings.

[0013] To easily identify the discussion of any particular element or operation, the most significant digit or digits of a reference number refer to the figure number in which that element is first introduced. [Brief description of the drawings]

[0014] [Figure 1]FIG. 1 is a block diagram of an environment in which a cloud infrastructure orchestration service (CIOS) may operate to dynamically provide bootstrap services in a region, according to at least one embodiment. [Diagram 2] FIG. 1 is a block diagram illustrating an environment and method for building a Virtual Bootstrap Environment (ViBE) according to at least one embodiment. [Diagram 3] 1 is a block diagram illustrating an environment and method for bootstrapping a service to a target region using ViBE, according to at least one embodiment. [Figure 4] 1 is a simplified flowchart illustrating a process performed to monitor issued capabilities in a data center using alert-capability association information, according to one embodiment. [Diagram 5] 1 is a simplified flowchart illustrating a process performed to monitor issued capabilities in a data center using alert-capability association information, according to one embodiment. [Figure 6] 1 is a simplified flowchart illustrating the processes performed during release and execution of a flock, according to one embodiment. [Figure 7] 1 is a swim-lane flowchart illustrating the processing performed during execution of a flock release and how possible reasons for execution failure may be identified through the use of alert-capability associations according to one embodiment. [Figure 8] 1 is a simplified flowchart illustrating the process performed to check the sanity of capabilities prior to scheduling a flock release, according to one embodiment. [Figure 9] FIG. 1 is a block diagram illustrating a pattern for implementing a service-based cloud infrastructure system in accordance with at least one embodiment. [Figure 10] FIG. 2 is a block diagram illustrating another pattern for implementing a service-based cloud infrastructure system in accordance with at least one embodiment. [Figure 11]FIG. 2 is a block diagram illustrating another pattern for implementing a service-based cloud infrastructure system in accordance with at least one embodiment. [Figure 12] FIG. 2 is a block diagram illustrating another pattern for implementing a service-based cloud infrastructure system in accordance with at least one embodiment. [Figure 13] FIG. 1 is a block diagram illustrating an exemplary computer system according to at least one embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0015] Detailed Description In the following description, for purposes of explanation, specific details are set forth in order to provide a thorough understanding of certain embodiments. It will be apparent, however, that various embodiments may be practiced without these specific details. The drawings and this specification are not intended to be limiting. Use of the word "exemplary" herein means "serving as an example, instance, or illustration." Any embodiment or design described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments or designs.

[0016] Techniques are described for monitoring the health of services in a computing environment, such as a data center. More particularly, the present disclosure describes techniques for monitoring the health and availability of capabilities in a computing environment, such as a data center, by enabling association of alerts with the capabilities. A capability represents a set of resources in a data center. By allowing alerts to be associated with a capability, the health or availability of the associated capability may be monitored or confirmed by tracking the state of the alert associated with the capability. For example, if an alert associated with a particular capability is triggered, this may indicate that the particular capability and one or more resources corresponding to the particular capability are not in a healthy state. Thus, the health of the associated capability may be confirmed by monitoring the alerts associated with the capability.

[0017] Example automated data center construction (region construction) infrastructure In recent years, the adoption of cloud services has been growing rapidly. Currently, various types of cloud services are offered by various different cloud service providers (CSPs). The term cloud service is generally used to describe a service or functionality that is made available on demand (e.g., through a subscription model) to a user or customer by a CSP through the use of systems and infrastructure (cloud infrastructure) that the CSP provides. Typically, the servers and systems that make up the CSP's infrastructure and that are used to provide cloud services to customers are separate from the customer's own on-premise servers and systems. This allows customers to use cloud services offered by CSPs without having to purchase separate hardware and software resources for the services. Cloud services are designed to provide easy, scalable, and on-demand access to applications and computing resources to subscribing customers, without the customer having to invest in procuring the infrastructure used to provide the services or functionality. Cloud services can be offered in various different types or models, such as software as a service (SaaS), platform as a service (PaaS), infrastructure as a service (IaaS), etc. A customer can subscribe to one or more cloud services offered by a CSP. A customer can be any entity, such as an individual, an organization, or a business.

[0018] As mentioned above, a CSP is responsible for providing the infrastructure and resources used to provide cloud services to its subscribing customers. The resources provided by a CSP may include both hardware and software resources. These resources may include, for example, compute resources (e.g., virtual machines, containers, applications, processors), memory resources (e.g., databases, data stores), networking resources (e.g., routers, host machines, load balancers), identities, and other resources. In one embodiment, the resources provided by a CSP to provide a set of cloud services are organized as a datacenter. A datacenter may be configured to provide a particular set of cloud services. A CSP is responsible for populating its datacenter with the infrastructure and resources used to provide the particular set of cloud services. A CSP may build one or more datacenters.

[0019] Datacenters provided by a CSP may be hosted in different regions. A region is a local geographic area and may be identified by a region name. Regions are generally independent from each other and may be separated by large distances, for example across countries or continents. Regions are grouped into realms. Examples of CSP regions include US West, US East, Australia East, Australia Southeast, etc.

[0020] A region may include one or more data centers, which are located within a geographic area corresponding to the region. As an example, the data centers in a region may be located in cities within the region. For example, for a particular CSP, the data center for the US West region may be located in San Jose, California, the data center for the US East region may be located in Ashburn, Virginia, the data center for the Australia East region may be located in Sydney, Australia, the data center for the Australia Southeast region may be located in Melbourne, Australia, etc.

[0021] Data centers within a region may be organized as one or more availability domains that are used for high availability and disaster recovery. An availability domain may contain one or more data centers within a region. Availability domains within a region are isolated from each other, fault tolerant, and designed to make it highly unlikely that data centers in multiple availability domains will fail simultaneously. For example, availability domains within a region may be structured to make it less likely that a failure in one availability domain within a region will affect the availability of data centers in other availability domains in the same region.

[0022] When a customer or subscriber registers or signs up for one or more services offered by a CSP, the CSP creates a tenancy for the customer. A tenancy is like an account created for a customer. In one embodiment, the tenancy for a customer exists in a single realm and has access to all regions that belong to that realm. And the customer's users can access the services for which the customer has registered under this tenancy.

[0023] As mentioned above, a CSP provides cloud services to its customers by building or deploying data centers. As a CSP's customer base expands, the CSP typically builds new data centers in new regions or expands the capacity of existing data centers to accommodate and serve the growing demands of the customers and to better serve the customers. The data centers are preferably built in close geographic proximity to the locations of the customers served by the data center. The geographical proximity between the data center and the customers served by the data center helps in more efficient use of resources and faster and more reliable service to the customers. Thus, a CSP typically builds new data centers in new regions in geographical areas that are geographically close to the customers served by the data center. For example, if the customer base is expanding in Germany, the CSP may build one or more data centers in new regions in Germany.

[0024] Building a datacenter(s) in a region is sometimes referred to as building a region. The term "region build" is used to refer to building one or more datacenters in a region. Building a datacenter in a region involves provisioning or creating a new set of resources that are required or used to provide the set of services that the datacenter is configured to provide. The end result of the region build process is the creation of a datacenter in the region that is capable of providing the set of services intended for the datacenter and includes a set of resources that are used to provide this set of services.

[0025] Building a new datacenter in a region is a very complex task that requires extensive coordination between various bootstrapping activities. At a high level, this involves performing and coordinating various tasks such as identifying the set of services to be provided by the datacenter, identifying the various resources required to provide the set of services, creating, provisioning, and deploying the identified resources, and properly describing the resources to enable their intended use. Each of these tasks has further subtasks that require coordination, which further increases the complexity. Due to this complexity, currently, building a datacenter in a region involves multiple manual initiation or control tasks that require careful manual coordination. As a result, the task of building a new region (building one or more datacenters in a region) requires a lot of time. Building a datacenter can take many months, for example. This process is also highly error-prone and may require multiple iterations before the desired configuration of the datacenter is achieved, further lengthening the time it takes to build the datacenter. These constraints and issues severely limit the ability of CSPs to scale computing resources in a timely manner to meet growing customer needs.

[0026] This disclosure describes techniques for reducing the construction time, thereby reducing the waste of computing resources and reducing the risks associated with the construction of one or more data centers in a region. Whereas previously it took weeks and months to build a data center in a region, by using the techniques described herein, a new data center can be built in a region in a relatively much shorter time with less risk of error than traditional approaches.

[0027] Disclosed herein is a cloud infrastructure orchestration service (CIOS) configured to bootstrap (e.g., provision and deploy) services to a new datacenter based on predefined configuration files that identify resources (e.g., infrastructure components and software to be deployed) for implementing a given change to the datacenter. The CIOS can identify dependencies between resources, execution targets, phases, and flocks by parsing and analyzing the configuration files (e.g., flock configurations). The CIOS may generate specific data structures from the analysis and may use these data structures to drive operations and manage the order in which services are bootstrapped to regions. The CIOS may utilize these data structures to identify when a service can be bootstrapped, when bootstrapping is blocked, and / or when a bootstrapping operation associated with a previously blocked service can be resumed. Advantageously, the CIOS can identify circular dependencies in the data structures and perform operations to eliminate / resolve these circular dependencies prior to executing the task. Using these techniques, the CIOS significantly reduces the risk of executing a task before the resources on which the task depends are available.

[0028] Using the techniques disclosed herein, CIOS may perform datacenter modifications through parallel processing optimizations while preventing tasks from starting until functionality on which they depend is available in the region. In this manner, CIOS allows for more efficient execution of region construction, significantly reducing the time required to construct datacenters and the wasted computing resource usage seen with traditional approaches.

[0029] Certain definitions A "region" is a logical abstraction that corresponds to a geographic location. A region may include any suitable number of one or more execution targets. In some embodiments, an execution target may correspond to a data center.

[0030] An "execution target" represents the smallest unit of change for executing a release. A "release" represents a statement of intent to orchestrate a particular change to a service (e.g., version 8 deployment, "add internal DNS record", etc.). For most services, an execution target represents an "instance" of the service. A single service can be bootstrapped onto one or more execution targets each. An execution target may be associated with a set of devices (e.g., a data center).

[0031] "Bootstrapping" is intended to refer to the collective tasks associated with the provisioning and deployment of any suitable number of resources (e.g., infrastructure components, artifacts, etc.) that correspond to a single service.

[0032] A "service" refers to a functionality provided by a set of resources. The set of resources for a service includes any suitable combination of cloud provider hosted infrastructure, platform, or software (e.g., applications) that can be configured to provide the functionality of the service. A service may be made available to users over the Internet.

[0033] "Artifact" refers to an infrastructure component or code that is deployed to a Kubernetes engine cluster, which may include software (e.g., applications), configuration information for the infrastructure component (e.g., configuration files), etc.

[0034] A "flock config" refers to a configuration file (or a set of configuration files) that describes a set of all resources (e.g., infrastructure components and artifacts) associated with a single service. A flock config may contain declarative statements that specify one or more aspects that correspond to a desired state of the service's resources.

[0035] A "service state" represents a point-in-time snapshot of all resources (e.g., infrastructure resources, artifacts, etc.) associated with a service. A service state indicates the status corresponding to provisioning and / or deployment tasks associated with the service resources.

[0036] IaaS provisioning (or "provisioning") refers to obtaining a computer or virtual host for use, and even installing the necessary libraries or services on it. The expression "provisioning a device" refers to the evolution of a device to a state where it can be used by an end user for a specific purpose. A device that has gone through the provisioning process may also be referred to as a "provisioned device". Preparing a provisioned device (installing libraries and daemons) may be part of provisioning, but this is different from deploying a new application or a new version of an application to a prepared device. In most cases, deployment does not include provisioning, which may need to be performed first. A prepared device may also be referred to as an "infrastructure component".

[0037] IaaS deployment (or "deployment") refers to the process of provisioning and / or installing a new application or a new version of an application onto a provisioned infrastructure component. Once the infrastructure component has been provisioned (e.g., acquired, allocated, prepared, etc.), additional software may be deployed (e.g., provided and installed onto the infrastructure component). After provisioning and deployment are complete, the infrastructure component may be referred to as a "resource." Examples of resources include, but are not limited to, virtual machines, databases, object storage, block storage, load balancers, etc.

[0038] A "capability" identifies a unit of functionality associated with a service, which may be some or all of the functionality provided by the service. As an example, a capability may be published that indicates that a resource is available for authorization / authentication processing (e.g., a subset of the functionality provided by the resource). As another example, a capability may be published that indicates that all functionality of a service is available. Capabilities can be used to identify functionality on which a resource or service depends and / or the functionality of the resource or service that is available.

[0039] A "virtual bootstrap environment" (ViBE) represents a virtual cloud network that is provisioned in an overlay of an existing region (e.g., a "host region"). The provisioned ViBE is connected to the new region using a communication channel (e.g., IPSec Tunnel VPN). Certain essential core services (or "seed" services) can be provisioned in the ViBE, such as a deployment orchestrator, public key infrastructure (PKI) services, etc. These services can provide the capabilities required to bring hardware online, establish a chain of trust to the new region, and deploy other services in the new region. The use of a virtual bootstrap environment can prevent circular dependencies between bootstrap resources by utilizing resources in the host region. Services can be staged and tested in the ViBE before the physical region (e.g., the target region) is available.

[0040] "Cloud Infrastructure Orchestration Service" (CIOS) may refer to a system configured to manage the provisioning and deployment operations of any suitable number of services as part of a region construction.

[0041] A multi-flock orchestrator (MFO) may be a computing component (e.g., a service) that coordinates events between components of the CIOS to provision and deploy services to a target region (e.g., a new region). The MFO tracks events related to each service in the region build and takes action in response to the events.

[0042] A "host region" refers to a region that hosts a Virtual Bootstrap Environment (ViBE). A host region may be used to bootstrap a ViBE.

[0043] "Target region" refers to the region being constructed. "Publishing a capability" refers to "publishing" as used in "publisher-subscriber" computing design, or providing an indication that a particular capability is available (or unavailable). Capabilities provide an indication that a resource / service's functionality is available by "publishing" (e.g., collected by a capability service, provided to a capability service, pushed, pulled, etc.). In some embodiments, capabilities may be published / transmitted by an event, notification, data transmission, function call, API call, etc. An event (or other notification / data transmission, etc.) indicating availability of a particular capability can be broadcast / addressed (e.g., published) to a capability service.

[0044] A "capabilities service" may be a flock that is configured to model dependencies between different flocks. Capabilities services may be provided within a cloud infrastructure orchestration service and may define the capabilities, services, and features that are made available in a region.

[0045] A "Real-time Regional Data Distributor" (RRDD) may be a service or system configured to manage regional data that can be injected into a flock configuration to dynamically create new regional execution targets.

[0046] In some examples, techniques are described herein for implementing a cloud infrastructure orchestration service (CIOS). Such techniques can be configured to manage the bootstrapping (e.g., software provisioning and deployment) of infrastructure components within a cloud environment (e.g., region), as briefly described above. In some cases, the CIOS can include computing components (e.g., CIOS central and CIOS regional (both of which are described in more detail below)) that can be configured to manage the bootstrapping tasks (provisioning and deployment) for a given service as well as a multi-flock orchestrator (also described in more detail below) that is configured to initiate / manage region construction (e.g., bootstrapping operations corresponding to multiple services).

[0047] CIOS enables region builds and global infrastructure provisioning and code deployment with minimal manual effort by service teams (e.g., beyond initial approval and / or physical shipping of hardware, as the case may be). High level responsibilities of CIOS include, but are not limited to, coordinating region builds, providing users with a current view of the resources it manages (e.g., within regions, across regions, globally, etc.), and managing the bootstrap operations to bootstrap resources within regions.

[0048] The CIOS may provide view reconciliation that can reconcile a view of a desired state (e.g., desired configuration) of a resource with the current / actual state (e.g., current configuration) of the resource. In some cases, view reconciliation may include obtaining state data that identifies the actual operating resources and their respective current configurations and / or states. Reconciliation may be performed at various levels of granularity, such as at the service level.

[0049] The CIOS can perform plan generation in which differences between desired and current states of resources are identified. Part of plan generation can include identifying actions that need to be performed to move resources from their current state to the desired state. In some examples, the CIOS can present the generated plan to a user for approval. In these examples, the CIOS can allow the user to accept or reject the plan based on user input from the user. This allows the user to spend less time reasoning about the plan, and since the plan is machine-generated, it is more accurate. Most of the plans are too detailed for human consumption, but the CIOS can provide this data through an advanced user interface (UI).

[0050] In some instances, the CIOS can handle change control execution by executing approved plans. Once an execution plan has been created and approved, engineers may not need to participate in change control unless the CIOS initiates a rollback. The CIOS can handle rollbacks to previous service versions (e.g., if it detects degradation of service health during execution) by generating a plan to revert the service to a previous (e.g., pre-release) state.

[0051] CIOS can measure service health by monitoring alerts and running integration tests. CIOS can help teams quickly prescribe and later execute rollback behavior in case of service degradation. CIOS can generate and display plans as well as track approvals. CIOS can combine provisioning and deployment capabilities in a single system that coordinates these tasks across region builds. CIOS also supports discovery of flocks (e.g., service resources such as flock configurations corresponding to any suitable number of services), artifacts, resources, and dependencies. CIOS can discover dependencies between execution tasks at any level (e.g., resource level, execution target level, phase level, service level, etc.) by static analysis (e.g., including parsing and processing the contents) of one or more configuration files. Using these dependencies, CIOS can generate various data structures from these dependencies that can be used to drive task execution (e.g., tasks related to provisioning infrastructure resources and deployment of artifacts across regions).

[0052] FIG. 1 is a block diagram of an environment 100 in which a cloud infrastructure orchestration service (CIOS) 102 may operate to dynamically provide bootstrap services in a region, according to at least one embodiment. CIOS 102 may include components such as, but not limited to, a real-time regional data distributor (RRDD) 104, a multi-flock orchestrator (MFO) 106, a CIOS central 108, a CIOS regional 110, and a capability service 112. The specific functions of CIOS central 108 and CIOS regional 110 are provided in more detail in U.S. patent application Ser. No. 17 / 016,754, entitled "Techniques for Deploying Infrastructure Resources with a Declarative Provisioning Tool," the entire contents of which are incorporated herein by reference. In some embodiments, any suitable combination of the components of CIOS 102 may be provided as a service. In some embodiments, any portion of CIOS 102 may be deployed to a region (e.g., a data center represented by host region 103). In some embodiments, CIOS 102 may include any suitable number of cloud services (not shown in FIG. 1), as discussed in more detail below with respect to U.S. patent application Ser. No. 17 / 016,754 and FIGS. 2 and 3.

[0053] The real-time regional data distributor (RRDD) 104 may be configured to maintain and provide regional data identifying realms, regions, execution targets, and availability domains. In some cases, the regional data may be in any suitable form (e.g., JSON format, data object / container, XML, etc.). The regional data maintained by the RRDD 104 may include any suitable number of data subsets that may be individually referenced by a corresponding identifier. As an example, an identifier "all_regions" may be associated with a data structure (e.g., a list, structure, object, etc.) that includes metadata for all defined regions. As another example, an identifier such as "realm" may be associated with a data structure that identifies metadata for a number of realms and a set of regions corresponding to each realm. In general, the regional data may maintain any suitable attributes of one or more realms, regions, availability domains (ADs), execution targets (ETs), etc. (identifiers, DNS suffixes, state (e.g., state of the region), etc.). The RRDD 104 may be configured to manage a regional state as part of the regional data. The regional state may include any suitable information indicative of a state of bootstrap within the region. By way of example, some exemplary region states may include "initial," "constructing," "creating," "suspended," or "decommissioned." The "initial" state may indicate a region that has not yet been bootstrapped. The "constructing" state may indicate that bootstrapping of one or more blocks in the region has begun. The "creating" state may indicate that bootstrapping is complete and the region is ready for validation. The "suspended" state may indicate that CIOS Central 108 or CIOS Regional 110 has suspended internal interactions with the regional stack, possibly due to operational issues. The "decommissioned" state may indicate that the region has been decommissioned and may be unavailable and / or unable to be reconnected.

[0054] CIOS central 108 may be configured to provide any suitable number of user interfaces through which a user (e.g., user 109) may interact with CIOS 102. As an example, a user may modify region data through a user interface provided by CIOS central 108. CIOS central 108 may additionally provide various interfaces that allow a user to review changes made to flock configurations and / or artifacts, generate and review plans, approve / reject plans, and review status regarding plan execution (e.g., corresponding to infrastructure provisioning, deployment, tasks involving region construction, and / or a desired state of any suitable number of resources managed by CIOS 102). CIOS central 108 may implement a control plane that is configured to manage any suitable number of CIOS regional 110 instances. CIOS central 108 may provide one or more user interfaces that allow user 109 to review and / or modify region data through presentation of the region data. CIOS central 108 may be configured to invoke RRDD 104 functions through any suitable number of interfaces. In general, CIOS central 108 may be configured to manage region data directly or indirectly (e.g., via RRDD 104). CIOS central 108 may be configured to inject region data into a flock configuration as a variable upon compilation of the flock configuration.

[0055] Each instance of CIOS regional 110 may correspond to a module configured to perform bootstrapping tasks associated with a single service of the region. CIOS regional 110 may receive desired state data from CIOS central 108. In some embodiments, the desired state data may include a flock configuration that declares (e.g., by declarative statements) a desired state of resources associated with the service. CIOS central 108 may maintain current state data that indicates any suitable aspect of the current state of resources associated with the service. In some embodiments, CIOS regional 110 may identify one or more resources as requiring modification by comparing the desired state data and the current state data. For example, CIOS regional 110 may determine that provisioning of one or more infrastructure components, deployment of one or more artifacts, or any suitable modification of the resources of the service is required to bring the state of the resources into line with the desired state. When CIOS regional 110 performs the bootstrapping operation, it may publish data indicating various capabilities of the resources that have been made available. A "capability" identifies a unit of functionality associated with a service, which may be some or all of the functionality provided by the service. As an example, a capability may be published that indicates that a resource is available for authorization / authentication processing (e.g., a subset of the functionality provided by the resource). As another example, a capability may be published that indicates that all functionality of a service is available. Capabilities can be used to identify functionality on which a resource or service depends and / or the functionality of the resource or service that is available.

[0056] The capability service 112 is configured to maintain capability data that indicates: 1) capabilities of various services that are currently available; 2) whether any resources / services are waiting for a particular capability; 3) specific resources and / or services waiting for a given capability; or any suitable combination thereof. The capability service 112 may provide an interface through which the capability data may be requested. The capability service 112 may provide one or more interfaces (e.g., application programming interfaces) that allow for sending the capability data to the MFO 106 and / or CIOS regional 110 (e.g., each instance of the CIOS regional 110). In some embodiments, any suitable component or module of the MFO 106 and / or CIOS regional 110 may be configured to request the capability data from the capability service 112.

[0057] In some embodiments, a multi-flock orchestrator (MFO) 106 may be configured to drive region construction attempts. In some embodiments, the MFO 106 may manage information representing the flock / flock configuration version and / or artifact version utilized to bootstrap a given service in the region (or configure a change unit for the target region). In some embodiments, the MFO 106 may be configured to monitor (or be notified of) changes to the region data managed by the real-time regional data distributor 104. In some embodiments, region construction may be triggered by the MFO 106 upon receiving an indication that the region data has changed. In some embodiments, the MFO 106 may collect various flock configurations and artifacts used in region construction. Some or all of the flock configurations may be configured to be region agnostic; that is, the flock configurations may not explicitly identify the region to which the flock is bootstrapped. In some embodiments, the MFO 106 may initiate a data injection process in which the collected flock configurations are recompiled (e.g., by the CIOS central 108). During recompilation, execution of an action (e.g., by CIOS Central 108) injects the configuration file with region data maintained by the real-time regional data distributor 104. The flock configuration can reference the region data through variables / parameters without requiring hard-coded identification of the region data. This data injection allows the flock configuration to be dynamically modified at run time without hard-coding the region data, making it more difficult to change.

[0058] The multi-flock orchestrator 106 may perform static flock analysis, analyzing the flock configuration to identify dependencies between resources, execution targets, phases, and flocks, in particular, identifying circular dependencies that need to be eliminated. In some embodiments, the MFO 106 may generate any suitable number of data structures based on the identified dependencies. These data structures (e.g., directed acyclic graphs, linked lists, etc.) may be utilized by the cloud infrastructure orchestration service 102 to drive operations to perform region construction. As an example, these data structures may collectively specify the order in which services are bootstrapped within a region. An example of such a data structure is discussed separately below with respect to the construction dependency graph 338 of FIG. 3. If circular dependencies (e.g., service A requires service B and vice versa) exist and are identified by the static flock analysis and / or graph, the MFO may be configured to notify any suitable service team that a corresponding change in the flock configuration is required to correct these circular dependencies. The MFO 106 may be configured to manage the order in which services are bootstrapped into a region by traversing one or more data structures. The MFO 106 can identify capabilities available within a given region at any given time (e.g., using data obtained from the capability service 112). The MFO 106 can use this data to identify when a service can be bootstrapped, when bootstrap is blocked, and / or when a bootstrap operation associated with a previously blocked service can be resumed. Based on this passage, the MFO 106 can perform various releases in which the MFO 106 sends instructions to the CIOS central 108 to perform bootstrap operations corresponding to any suitable number of flock configurations. In some examples, the MFO 106 can be configured to identify that one or more flock configurations may require multiple releases due to the presence of circular dependencies in the graph.As a result, the MFO 106 may send multiple sets of instructions to the CIOS central 108 for a given block configuration to resolve circular dependencies identified in the graph.

[0059] In some embodiments, a user may request the construction of a new region (e.g., target region 114), which may include bootstrapping resources corresponding to various services. In some embodiments, target region 114 may not be reachable (and / or secure) at the time the region construction request is initiated. Rather than deferring bootstrapping until target region 114 is available and configured to perform the bootstrapping operation, CIOS 102 may initiate region construction using a virtual bootstrap environment 116. The virtual bootstrap environment (ViBE) 116 may be an overlay network hosted by host region 103 (an existing region that has previously been configured with a core set of services, is reachable, and is secure). MFO 106 may utilize resources from host region 103 to bootstrap resources into ViBE 116 (commonly referred to as "constructing the ViBE"). As an example, MFO 106, through CIOS Central 108, can provide instructions to an instance of CIOS Regional 110 in a host region (e.g., host region 103) to bootstrap another instance of CIOS Regional in ViBE 116. Once the CIOS Regional in ViBE is available for processing, bootstrapping of services to target region 114 can continue in ViBE 116. Previously bootstrapped services in ViBE 116 can be moved to target region 114 when target region 114 is available to perform the bootstrap operation. By utilizing these techniques, CIOS 102 can greatly increase the speed of region construction by significantly reducing the need for any manual input and / or provision of configuration.

[0060] 2 is a block diagram illustrating an environment 200 and method for constructing a virtual bootstrap environment (ViBE) 202 (an example of ViBE 116 of FIG. 1) according to at least one embodiment. ViBE 202 represents a virtual cloud network that is provisioned in an overlay of an existing region (e.g., hosted region 204, which is an example of hosted region 103 of FIG. 1 and, in one embodiment, a hosted region service enclave). ViBE 202 represents an environment in which services may be staged to a target region (e.g., a region under construction, such as target region 114 of FIG. 1) before the target region is available.

[0061] To bootstrap a new region (e.g., target region 114 of FIG. 1), a core set of services may be bootstrapped. These core set of services are present in the host region 204, but are not present in ViBE (nor in the target region). These essential core services provide the functionality required for provisioning devices, establishing a chain of trust to the new region, and deploying other services (e.g., flocks) to the region. ViBE 202 may be a tenancy deployed in the host region 204. This is considered a virtual region.

[0062] When a target region is available for bootstrapping operations, ViBE202 can connect to the target region so that services in the ViBE can interact with services and / or infrastructure components of the target region. This allows for deployment of generation-level services instead of self-contained seed services as in traditional systems and requires connectivity to the target region over the Internet. Traditionally, seed services are deployed as part of a collection of containers and used to bootstrap dependencies required to build the region. Using the existing regional infrastructure / tools, resources can be bootstrapped (e.g., provisioned and deployed) to ViBE202 and connected to the service enclaves of the region (e.g., host region 204) for hardware provisioning and service deployment until the target region is self-sufficient and can communicate directly. Utilization of ViBE202 allows for establishment of dependencies and services required to enable infrastructure provisioning / preparation and software deployment while utilizing the host region's resources breaks circular dependencies for core services.

[0063] A multi-flock orchestrator (MFO) 206 may be configured to perform operations to build (e.g., configure) ViBE 202. MFO 206 may obtain applicable flock configurations corresponding to various resources to be bootstrapped into a new region (in this case, a ViBE region of ViBE 202). As an example, MFO 206 may obtain a flock configuration (e.g., a "ViBE flock configuration") that identifies aspects of bootstrap capability service 208 and worker 210. As another example, MFO 206 may obtain another flock configuration corresponding to bootstrapping domain name service (DNS) 212 into ViBE 202.

[0064] In step 1, MFO 206 may issue instructions to CIOS central 214 (e.g., an example of CIOS central 108 and CIOS central 214 of FIGS. 1 and 2, respectively). For example, MFO 206 may send a request (e.g., including a ViBE flock configuration) to bootstrap capability services 208 and workers 210 that do not yet exist in ViBE 202 at this point. In some embodiments, CIOS central 214 may have access to all flock configurations. Thus, in some examples, MFO 206 may send an identifier for the ViBE flock configuration rather than the file itself, and CIOS central 214 may independently retrieve it from storage (e.g., DB 308 or flock DB 312 of FIG. 3).

[0065] In step 2, CIOS central 214 may provide the ViBE flock configuration to CIOS regional 216 via a corresponding request. In step 3, CIOS regional 216 may analyze the ViBE flock configuration to identify and perform specific infrastructure provisioning and deployment operations.

[0066] In some embodiments, CIOS regional 216 may utilize additional corresponding services for provisioning and deployment. For example, in step 4, CIOS regional 216 may instruct deployment orchestrator 218 (e.g., an example of a core service or other writing, building, and deploying application software in host region 204) to execute instructions to bootstrap capability service 208 and worker 210 in ViBE 202.

[0067] In step 5, a capability may be sent to capability service 208 (e.g., from CIOS regional 216, deployment orchestrator 218, via worker 210) indicating that resources corresponding to the ViBE block are available. Capability service 208 may persist this data. In some embodiments, capability service 208 adds this information to a list of capabilities available in ViBE. As an example, the capability provided to capability service 208 in step 5 may indicate that capability service 208 and worker 210 are available for processing.

[0068] In step 6, the MFO 206 may, based on receiving or obtaining data (an identifier corresponding to a capability) from the capability service 208, identify that the capability indicates that the capability service 208 and worker 210 are available.

[0069] In step 7, as a result of receiving / obtaining the data in step 6, MFO 206 may instruct CIOS Central 214 to bootstrap a DNS service (e.g., DNS 212) into ViBE 202. These instructions may identify or include a particular block configuration that corresponds to the DNS service.

[0070] In step 8, CIOS central 214 may instruct CIOS regional 216 to deploy DNS 212 to ViBE 202. In some embodiments, CIOS central 214 provides a DNS flocking configuration for DNS 212.

[0071] In step 9, a worker 210 deployed on ViBE 202 may be assigned the task of deploying DNS 212 by CIOS regional 216. The worker may execute a declarative infrastructure provisioner as described above in connection with FIGURE 3 to identify a set of operations that need to be performed for the deployment of DNS 212 (e.g., by comparing the flock configuration (desired state) against the current state of the (non-existent) resources associated with the flock).

[0072] At step 10, deployment orchestrator 218 may instruct worker 210 to deploy DNS 212 according to the actions identified at step 9. As shown, worker 210 proceeds to perform the actions of deploying DNS 212 to ViBE 202 at step 11. At step 12, worker 210 notifies capability service 208 that DNS 212 is available on ViBE 202. MFO 206 may then identify the ViBE flock configuration and resources associated with the DNS flock configuration as available, and proceed to bootstrap any suitable number of additional resources into ViBE.

[0073] After steps 1-12 are completed, the process for building ViBE 202 is complete and ViBE 202 may be considered built.

[0074] FIG. 3 is a block diagram illustrating an environment 300 and method for bootstrapping a service to a target region using ViBE, according to at least one embodiment.

[0075] In step 1, user 302 may modify region data using any suitable user interface provided by CIOS central 304 (an example of CIOS central 108 and CIOS central 214 in FIGS. 1 and 2, respectively). As an example, user 302 may create a new region into which a number of services are bootstrapped.

[0076] In step 2, CIOS central 304 may perform an operation to send the changes to RRDD 306 (an example of RRDD 104 in FIG. 1). In step 3, RRDD 306 may store the received region data in database 308, which is a data store configured to store region data including any suitable identifiers, attributes, states, etc., such as region, AD, realm, ET, etc. In some embodiments, updater 307 may be utilized to store the region data in database 308 or any suitable data store that may make such updates accessible (e.g., by a service team). In some embodiments, updater 307 may be configured to notify updates to database 308 (e.g., by any suitable electronic notification).

[0077] In step 4, the MFO 310 (an example of the MFOs 106 and 206 in FIGS. 1 and 2, respectively) may detect changes in the region data. In some embodiments, the MFO 310 may be configured to poll the RRDD 306 for changes in the region data. In some embodiments, the RRDD 306 may be configured to publish or notify the MFO 310 of the region changes.

[0078] In step 5, detecting a change in region data may trigger MFO 310 to retrieve a version set (e.g., a version set associated with a particular identifier, such as a "golden version set" identifier) ​​that identifies the specific version of each flock (e.g., service) and each artifact corresponding to that flock to be bootstrapped into the new region. The version set may be retrieved from DB 312. As a flock evolves and changes, the versions of each corresponding setting and artifact used to build the region may also change. These changes may be persisted in flock DB 312 so that MFO 310 may identify the version of the flock setting and artifact to use to build the region (e.g., ViBE region, target region / non-ViBE region, etc.). Flock settings (e.g., all versions of a flock setting) and / or artifacts (e.g., all versions of an artifact) may be stored in DB 308, DB 312, or any suitable data store accessible to CIOS Central 304 and / or MFO 310.

[0079] In step 6, the MFO 310 may request that the CIOS Central 304 recompile each of the flock settings associated with the version set with the current region data. In some embodiments, this request may indicate the version of each flock setting and / or the artifacts that correspond to those flock settings.

[0080] In step 7, the CIOS Central 304 may obtain the current regional data from the DB 308 (eg, directly or via the real-time regional data distributor 306) and retrieve any suitable flock settings and artifacts according to the version requested by the MFO 310.

[0081] In step 8, the CIOS central 304 may inject the current region data into the flock configuration by recompiling the flock configuration with the region data obtained in step 7. The CIOS central 304 may return the compiled flock configuration to the MFO 310. In some embodiments, the CIOS central 304 may only indicate that compilation has occurred, and the MFO 310 may access the recompiled flock configuration via the RRDD 306.

[0082] In step 9, the MFO 310 may perform a static analysis of the recompiled flock configuration. As part of the static analysis, the MFO 310 may identify dependencies between flocks by analyzing the flock configuration (e.g., using a library associated with a declarative infrastructure provisioner (e.g., Terraform, etc.)). From this analysis and the identified dependencies, the MFO 310 may generate a build dependency graph 338. The build dependency graph 338 may be a directed acyclic graph that identifies the order in which flocks are bootstrapped into new regions (and / or changes indicated in the flock configuration are applied). Each node in the graph may correspond to the bootstrapping of any suitable portion of a particular flock. The specific bootstrap order may be identified based at least in part on the dependencies. In some embodiments, the dependencies may be expressed as attributes of the nodes and / or specified by the edges of the graph connecting the nodes. The MFO 310 may drive the region build operation by traversing the graph (e.g., starting from a start node).

[0083] In some embodiments, MFO 310 may utilize a cycle detection algorithm to detect whether there is a cycle (e.g., service A depends on service B, and vice versa). MFO 310 may identify orphan capability dependencies. For example, MFO 310 may identify orphan nodes in construction dependency graph 338 that are not connected to other nodes. MFO 310 may identify capabilities that have been issued improperly (e.g., when a capability is issued prematurely and the corresponding functionality is not yet actually available). MFO 310 may detect from the graph that there are one or more instances that issue the same capability. In some embodiments, any suitable number of these errors may be detected, and MFO 310 (or another suitable component, such as CIOS Central 304) may be configured to notify or present this information to a user (e.g., via electronic notification, a user interface, etc.). In some embodiments, MFO 310 may be configured to resolve circular dependencies by force-deleting / recreating resources and may redirect CIOS Central 304 to perform bootstrap operations for those resources and / or corresponding flock configurations.

[0084] The initiating node may correspond to bootstrapping the ViBE flock, and the second node may correspond to bootstrapping the DNS. Steps 10-15 correspond to deployment (by deployment orchestrator 317, which is an example of deployment orchestrator 218 of FIG. 2) of the ViBE flock to ViBE 316 (e.g., an example of ViBE 116 and 202 of FIGS. 1 and 2, respectively). That is, steps 10-15 of FIG. 3 generally correspond to steps 1-6 of FIG. 2. Upon being notified that capabilities exist corresponding to deployment of the ViBE flock (e.g., indicating that capability service 318 and worker 320, which correspond to capability service 208 and worker 210 of FIG. 2, are available), MFO 310 resumes traversing build dependency graph 338 and identifies the next operation to perform.

[0085] As an example, MFO 310 may continue traversing construction dependency graph 338 to identify DNS blocks to be deployed. Steps 16-21 may be performed to deploy DNS 322 (an example of DNS 212 in FIG. 2). These operations may generally correspond to steps 7-12 in FIG. 2.

[0086] At step 21, a capability may be stored indicating that DNS 322 is available. Upon detecting this capability, MFO 310 may resume traversing build dependency graph 338. During this traversal, MFO 310 may identify any suitable portion of an instance of CIOS regional (e.g., an instance of CIOS regional 314) to be deployed to ViBE 316. In some embodiments, steps 16-21 may be substantially repeated with respect to the deployment of CIOS regional (ViBE) 326 (CIOS regional 314, an instance of CIOS regional 110 in FIG. 1) and worker 328 to ViBE 316. Additionally, a capability indicating that CIOS regional (ViBE) 326 is available may be sent to capability service 318.

[0087] Upon detecting that CIOS Regional (ViBE) 326 is available, MFO 310 may resume traversing build dependency graph 338. During this traversal, MFO 310 may identify a deployment orchestrator (e.g., deployment orchestrator 330, which is an example of deployment orchestrator 317) to be deployed to ViBE 316. In some embodiments, steps 16-21 may be substantially repeated for the deployment of deployment orchestrator 330. Additionally, information may be sent to capability service 318 identifying capabilities indicating that deployment orchestrator 330 is available.

[0088] After deployment orchestrator 330 is deployed, ViBE 316 may be considered available for processing subsequent requests. Upon detecting that deployment orchestrator 330 is available, MFO 310 may direct the routing of subsequent bootstrap requests to ViBE components rather than using the host region components (host region 332 components). Thus, MFO 310 may continue traversing build dependency graph 338 at each node directing flock deployment to ViBE 316 via CIOS Central 304. CIOS Central 304 may request CIOS Regional (ViBE) 326 to deploy resources according to the flock configuration.

[0089] At some point in this process, the target region 334 may become available. An indication that the target region is available may be discernible from region data for the target region 334 provided by the user 302 (e.g., as an update to the region data). The availability of the target region 334 may depend on the establishment of a network connection between the target region 334 and an external network (e.g., the Internet). The network connection may be supported over a public network (e.g., the Internet), but may also include the use of software security tools (e.g., IPSec) to provide one or more encrypted tunnels (e.g., IPSec tunnels such as tunnel 336) from ViBE 316 to the target region 334. As used herein, “IPSec” refers to a suite of protocols for authenticating and encrypting network traffic on networks that use the Internet Protocol (IP), and may include one or more available implementations of the suite of protocols (e.g., Openswan, Libreswan, strongSwan, etc.). The network may connect ViBE 316 to a service enclave in the target region 334.

[0090] Prior to the establishment of the IPSec tunnel, the initial network connection to the target region 334 may be sufficient connectivity (e.g., an out-of-band VPN tunnel) to allow bootstrapping of networking services until IPSec gateways are deployed to assets (e.g., bare metal assets) in the target region 334. To bootstrap the network resources of the target region 334, the deployment orchestrator 330 may deploy IPSec gateways at the assets in the target region 334. The deployment orchestrator 330 may then deploy VPN hosts in the target region 334 that are configured to terminate the IPSec tunnels from ViBE 316. Once the services in ViBE 316 (e.g., deployment orchestrator 330, Service A, etc.) can establish IPSec connections with the VPN hosts in the target region 334, the bootstrap operation from ViBE 316 to the target region 334 may begin.

[0091] In some embodiments, the bootstrap operation may begin with services in ViBE 316 that support hosting instances of core services deployed from ViBE 316 by provisioning resources in the target region 334. For example, the host provisioning service may allocate computing resources for VMs by provisioning a hypervisor on infrastructure (e.g., bare metal hosts) in the target region 334. When the host provisioning service completes the allocation of physical resources in the target region 334, it may publish information indicating a capability indicating that the physical resources in the target region 334 have been allocated. This capability may be published (e.g., by worker 328) to capability service 318 via CIOS regional (ViBE) 326.

[0092] Once the hardware allocation for the target region 334 has been established and posted to the capability service 318, CIOS regional (ViBE) 326 can orchestrate the deployment of instances of core services from ViBE 316 to the target region 334. This deployment may be similar to the process described above with respect to building ViBE 316, but using ViBE components (e.g., CIOS regional (ViBE) 326, worker 328, deployment orchestrator 330) instead of the service enclave components of the host region 332. The deployment operations may generally correspond to steps 16-21 described above.

[0093] When a service is deployed from ViBE 316 to a target region 334, a DNS record associated with the service may correspond to an instance of the service in ViBE 316. The DNS record associated with the service may be updated later to complete the deployment of the service to the target region 334. In other words, the instance of the service in ViBE 316 may continue to receive traffic (e.g., requests) to the service until the DNS record is updated. The service may be partially deployed to the target region 334 and may publish information (e.g., to capability service 318) indicating a capability that the service is partially deployed. For example, a service running in ViBE 316 may be deployed to the target region 334 along with corresponding compute instances, load balancers, and associated applications and other software, but may need to wait for database data to move to the target region 334 before completing the deployment. A DNS record (e.g., managed by DNS 322) may still be associated with the service in ViBE 316. Once the data movement for the service is complete, the DNS record may be updated to point to the operational service deployed in target region 334. Thereafter, the deployed service in target region 334 will receive traffic (e.g., requests) for that service, while the instance of the service in ViBE 316 may not receive traffic for that service.

[0094] Capabilities and Alert Associations The present disclosure relates to improved techniques for monitoring the health of services in a computing environment, such as a data center. More specifically, the present disclosure describes techniques for monitoring the health and availability of capabilities in a computing environment, such as a data center, by enabling association of alerts with capabilities.

[0095] In an embodiment, a capability may identify a functional unit associated with a service and represent a set of resources associated with the service. By allowing alerts to be associated with a capability, the health or availability of the associated capability may be monitored or confirmed by tracking the state of the alert associated with the capability. For example, if an alert associated with a particular capability is triggered, this may indicate that the particular capability and one or more resources corresponding to the particular capability are not in a healthy state. Thus, by monitoring the alerts associated with the capability, the health of the associated capability may be confirmed. A number of such capabilities and associated alerts may be configured in a computing environment, such as a data center.

[0096] The ability to create associations between alerts and capabilities can be used for a variety of applications and use cases. In one embodiment (e.g., the embodiment shown in FIG. 1), a telemetry system may be provided in a data center (e.g., in the host region 103 or the target region 114) that includes an infrastructure configured to monitor alerts and determine the status of the alerts in one or more data centers in the region. The telemetry system may recognize alerts that are configured for the data center. In a simple example, an alert may have two states: (1) a normal state and (2) a trigger or abnormal state that indicates a deviation from the normal state and may trigger the alert due to some underlying problem. In one embodiment, the alert may continue to be triggered until the underlying condition that triggered the alert is resolved.

[0097] In one embodiment, techniques are described that allow for the association of various health states with a capability. In one embodiment, a health state associated with a capability can have one of two values: (1) a healthy state, indicating that the capability is issued and available in its correct intended state, or (2) an unhealthy state, indicating an error state of the capability. In one embodiment, when a capability is marked as unhealthy, it is also marked as unissued (i.e., not issued) and unavailable.

[0098] An alert can be triggered for a variety of reasons. For example, an alert can be triggered by a canary (an application that sends synthetic traffic to a service to monitor / measure the health of the service). The canary can generate a baseline for the health of the service. The canary can trigger an alert if it has a problem or if the service is found to have a problem. As another example, an alert can be triggered if a resource is inaccessible or has not responded for a while. An alert can be triggered for a variety of other reasons.

[0099] In some embodiments, canaries can also be associated with capabilities. The discussion herein of alert-capability associations may apply to canary-alert associations as well. In some other embodiments, capabilities can also be associated with other resources in the data center environment.

[0100] The data center may also be provided with a health monitoring system configured to monitor the status of capabilities in the data center. The health monitoring system may communicate with the telemetry system to determine the status of alerts in the data center monitored by the telemetry system. If the health monitoring system determines that a particular alert has been triggered, it may determine one or more capabilities (if any) associated with the alert and mark the capabilities as unhealthy (or unavailable or in error). The health monitoring system may then initiate or perform one or more actions in response to determining that the capabilities are in an unhealthy state. The one or more actions may include, for example, returning an unhealthy capability to a healthy state. The monitoring of alerts and associated health monitoring of capabilities may be used in various scenarios, as described below.

[0101] In some embodiments, the health monitoring system 116 may be part of a telemetry system. In some other embodiments, the health monitoring system 116 may be part of some other system in the data center. For example, the functionality of the health monitoring system 116 may be implemented by the capability service 112 shown in FIG. 1.

[0102] For the purposes of this application, a capability identifies a functional unit associated with a service. A capability may represent a set of one or more resources, which may be software or hardware resources, functions, services, etc. A capability may be identified by a capability label. Examples of capability labels include "Capability_A", "Capability_B", "Capability_C", etc. For example, a capability labeled Capability_A corresponds to resources R1 and R2 and may be denoted as Capability_A={R1,R2}. Similarly, Capability_B={R3}, Capability_C={R4,R5,R6}, etc.

[0103] A capability is considered to be issued (or enabled or available) in a computing environment, such as a datacenter, when all resources corresponding to that capability have been created, deployed, and are available for use for their intended purposes in the datacenter. In the above example, Capability_A may be considered to be issued with respect to a datacenter when resources R1 and R2 corresponding to Capability_A have been created and deployed in the datacenter and are ready for use for their intended purposes.

[0104] A data center may provide various services. For a service, a set of one or more resources used or required to provide the service is referred to as a flock for the service. A flock may include one or more resources. A flock typically represents a set of managed resources responsible for providing a service. Flock-related information for a service is declared as part of the flock configuration for the service. A flock configuration for a service is typically stored in the form of a file, and is therefore also referred to as a flock configuration file for the service. A flock configuration for a service describes the flocks (e.g., a set of resources including infrastructure components and artifacts associated with the service) associated with the service. A flock configuration for a service may include declarative statements that specify one or more aspects corresponding to a desired state of the resources associated with the service. A flock configuration may identify various settings and configurations associated with flock resources and other parameters associated with the service. In one embodiment, a flock configuration includes Terraform code, and a flock configuration file is a Terraform file.

[0105] In a typical datacenter environment configured to provide multiple services, services may depend on other services or resources. For example, service B may require service A to be present before service B can be bootstrapped (potentially because service B uses service A). Service C may require service B to be present before service C can be bootstrapped in the datacenter. Thus, there may be multiple dependencies between services. In one embodiment, these dependencies are explicitly expressed through the use of capabilities. For example, a flocking of a service may identify any capabilities on which the service depends. The capabilities on which a service depends are referred to as the capability dependencies of the service (also referred to as the capability dependencies of the flocking of the service).

[0106] A flock associated with a service (also referred to as a service flock) is described in a flock configuration specified for that service. A flock configuration may be in the form of a flock configuration file. A capability dependency of a service is also referred to as a capability dependency of the flock for that service. A capability dependency of a service (e.g., a capability dependency of a service flock) may be a required capability dependency or an optional capability dependency. A required capability dependency of a service (or a service flock) is one that requires a capability to be published or enabled in a data center before a release can be scheduled for that service or service flock. An optional capability dependency of a service flock is one in which if the optional capability dependency is not published in a data center, a release of that flock can still be scheduled to run as long as the optional required capability dependency for that service flock is satisfied. Thus, an optional capability dependency does not block a release of a service. If an optional capability dependency becomes available or published, another release of the service block is scheduled and executed. In this way, a block for a service can potentially be released and executed multiple times (sometimes referred to as reentrancy of service block releases). In general, a release of a service is scheduled and executed when the required capability dependencies of the service have been published, but one or more optional capability dependencies have not yet been published. A release and subsequent execution of a service block where one or more optional capability dependencies have not been published or are not available will result in the instantiation of a version of the service that may have some, but not all, capabilities. Note that not all service blocks need to have capability dependencies, whether required or optional.

[0107] A service's capability dependencies may be explicitly declared in the service's flock configuration, or may be implicitly determined from the service's flock configuration (by an MFO, as described below). If a capability dependency is explicitly declared in a flock configuration, the flock configuration contains code or metadata that explicitly declares the capability dependency. In other cases, a capability dependency may be inferred from code in the flock configuration (e.g., by the MFO 106, if the MFO 106 reads and parses the flock configuration). A service's flock configuration may identify zero or more required capabilities for the service flocks specified by the flock configuration. Thus, a service is not required to have capability dependencies.

[0108] As a result of the release and execution of a block for a service, one or more capabilities may (but are not necessarily) issued at the datacenter. Information identifying any capabilities issued by the execution of a release of a service block is also declared in the block configuration of the service. In some implementations (such as the embodiment shown in FIG. 1), the MFO 106 schedules the release of a service block. The release is then executed, for example, by CIOS central 108 and CIOS regional 110 of FIG. 1, and zero or more capabilities may be issued at the datacenter upon successful execution of the release. Thus, the release and corresponding execution of a service block results in additional capabilities being issued at the datacenter.

[0109] In some implementations (such as the embodiment shown in FIG. 1 ), as part of the process of configuring a datacenter in a region where the datacenter will provide a particular set of services, the MFO 106 is provided with a flock configuration of a particular set of services. The MFO 106 then reads and parses these flock configuration files to identify, for each flock described by the flock configuration, the capability dependencies of the flock, if any, and the capabilities issued upon execution of one or more releases of the flock, if any. The MFO 106 then constructs an acyclic dependency graph to represent the various capability dependencies between the flock configurations and the corresponding flocks. The dependency graph also identifies a dependency ordering of the capabilities identified by the multiple flock configurations. The MFO 106 may generate any suitable number of data structures based on the identified dependencies. These data structures (e.g., directed graphs, directed acyclic graphs, linked lists, etc.) may be utilized by CIOS 102 (e.g., MFO 106, CISO central 108, CIOS regional 110) to drive operations to build one or more datacenters in a region (also referred to as executing a region build). The process performed by MFO 106 in analyzing the flock configuration and generating the acyclic dependency graph may also be referred to as static flock analysis.

[0110] The capability dependencies between different flocks specified in the flock configurations corresponding to different services can be illustrated by the following example: A flock configuration (FC_A) for service A specifies a flock (F_A) for service A. FC_A does not identify any capability dependencies of F_A, but does identify that Capability_A is issued upon successful release and execution of flock F_A. A flock configuration (FC_B) specifies information related to a flock (F_B) for service B and may indicate that Capability_A is a required capability dependency and that Capability_B is issued upon successful execution of flock F_B for service B. A flock configuration (FC_C) specifies information for a flock (F_C) for service C and may indicate that Capability_B is a required capability dependency and that Capability_C is issued to the data center upon successful release and execution of flock F_C for service C. Given these flock configurations, the dependency graph generated by the MFO 106 identifies the following: FC_A(F_A)→Capability_A→FC_B(F_B)→Capability_B→FC_C(F_C)→Capability_C, If F_A does not have a capability dependency, then in the datacenter, release and execution of F_A will issue Capability_A, which is a required capability dependency of F_B, release and execution of F_B will issue Capability_B, which is a required capability dependency of F_C, and release and execution of F_C will issue Capability_C. For capabilities, the dependency graph is as follows: Capability_B depends on Capability_A, Capability_C depends on Capability_B, and is represented as Capability_A → Capability_B → Capability_C. Thus, the dependency graph generated by the MFO 106 for the datacenter identifies dependencies between capabilities declared in or inferred from the flock configurations corresponding to various services instantiated in the datacenter.

[0111] For the example dependency relationship (Capability_A → Capability_B → Capability_C), Capability_B is referred to as a direct dependency of Capability_C. Similarly, Capability_A is referred to as a direct dependency of Capability_B. Capability_A is also referred to as an indirect dependency of Capability_C. If Capability_A depends on another capability, then this other capability also has an indirect dependency on Capability_C, and so on. In general, in a hierarchical dependency graph, for a particular capability X in the graph, the parent capability in the graph of capability X is referred to as a direct capability dependency of capability X. All ancestor capabilities in the graph of the parent capability are referred to as indirect capability dependencies of capability X. Capabilities in a dependency graph can have multiple levels of indirect capability dependencies.

[0112] The MFO then schedules the release of each service block based on the availability of the published capabilities at the data center. In the above example, assuming that all the capability dependencies are necessary capability dependencies, where Capability_B depends on Capability_A and Capability_C depends on Capability_B, the MFO would first schedule the release of block F_A for service A. If the release is successful, Capability_A would be published at the data center. The MFO 106 would then schedule the release of block F_B for service B after determining that the necessary capability dependency (i.e., Capability_A) of block F_B for service B is satisfied. If the release of service B is successful, Capability_B would be published at the data center. Then, after the MFO 106 determines that the required capability dependency of F_C for service C (i.e., Capability_B) is satisfied, it will schedule a release of F_C for service C, which results in the publication of Capability_C. In this manner, as more releases of flocks are scheduled and executed, more capabilities are published in the data center environment, triggering additional releases in the data center environment.

[0113] In some embodiments, the association between alerts and capabilities is explicitly declared in the service's configuration. An example of such a declaration is shown below: Alarm_XYZ { ... labels = (“abc”,“def”,“Capability_X”,“Capability_Y”). ... } In the above example, an alarm identified by the label "Alarm_XYZ" is associated with a capability labeled Capability_X, and a second association is created between Alarm_XYZ and the capability labeled Capability_Y. In the process of reading and analyzing the flock configuration, the MFO learns these associations that are declared in the flock configuration file. The MFO may then store this alarm-capability association information for use by other systems in the data center. For example, a telemetry system, a health monitoring system, or other systems in the data center may use this association information to monitor alarms and use the alarm information to determine the health of the associated capabilities. In some other embodiments, the MFO 106 may provide the alarm-capability association information to interested systems in the data center.

[0114] In one embodiment, associations learned by the MFO are added to the alert definition information for the data center. An alert can be associated with zero, one, or more capabilities. A capability can be associated with zero, one, or more alerts. As described above, these associations between alerts and capabilities can be declared or specified in one or more flock configuration files.

[0115] The alert-capability association information may be used for a variety of different applications and use cases. In one embodiment, the alert-capability association information may be used to monitor the health of issued capabilities in a computing environment, such as a data center. FIG. 4 is a simplified flowchart 400 illustrating a process performed to monitor issued capabilities in a data center using alert-capability association information, according to one embodiment. The process illustrated in FIG. 4 may be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of the respective systems, may be implemented using hardware, or a combination thereof. The software may be stored in a non-transitory storage medium (e.g., memory device). The method presented in FIG. 4 and described below is exemplary and not intended to be limiting. Although FIG. 4 illustrates various process steps occurring in a particular order or sequence, this is not intended to be limiting. In an alternative embodiment, the process may be performed in some different order, or some steps may be performed in parallel. In some embodiments (such as the embodiment shown in FIG. 1), the process shown in FIG. 4 may be performed by capability service 112. In some implementations, the process shown in FIG. 4 may be performed by a health monitoring system. In yet other implementations, the process shown in FIG. 4 may be performed by a telemetry system located at a data center.

[0116] At 402, a set of capabilities that have been issued in the data center environment and are marked as healthy is determined. As mentioned above, in one embodiment, the capability service 112 shown in Figure 1 is responsible for tracking the issued capabilities. Thus, at 402, the set of issued capabilities may be determined by querying the capability service 112.

[0117] At 404, a set of alerts associated with the capabilities identified at 402 is determined. As described above, in one embodiment, the MFO 106 may identify and learn alert-capability associations when retrieving and analyzing flock configurations during a static flock analysis process. The MFO may then provide this learned information to a system configured to perform the process illustrated in FIG. 4. In one embodiment, the association information may be stored by the MFO and then available to an entity performing the process illustrated in FIG. 4. A set of alerts associated with the issued capabilities identified at 402 may be determined based on this learned alert-capability association information.

[0118] At 406, the alerts identified at 406 are monitored. The monitoring may be performed over a period of time. In one embodiment, the monitoring may be performed by a telemetry system, and the monitoring observations may be provided by the telemetry system to a system performing the process shown in FIG.

[0119] At 408, it is determined that a particular alert from the set of alerts identified at 404 has been triggered (i.e., the alert is in a triggered state) based on the monitoring performed at 406. The alert may have been triggered due to some problem condition. In some embodiments, the alert may remain triggered (e.g., in a triggered state) while the underlying problem persists. In some embodiments, if the alert is in a triggered state, the telemetry system may communicate information regarding the triggered alert to a system that performs the process shown in FIG. 4.

[0120] At 410, all capabilities associated with the particular triggered alert are determined from the set of capabilities identified at 402. The triggered alert may be associated with one or more of the issued capabilities identified at 402.

[0121] At 412, for each capability identified at 410, the state of the capability is changed from healthy to unhealthy (or from available to unavailable).

[0122] In an embodiment, for each capability that is marked as unhealthy, one or more actions may be triggered. These actions may include, for example, notifying an interested party (e.g., a system administrator at the data center) of the capability's status change from healthy to unhealthy, one or more actions to identify the underlying problem that caused a particular alert to be triggered and actions to correct the underlying problem, recording the capability's status change in a log file, notifying the capability service 110 of the capability's status change from healthy to unhealthy, etc. The process illustrated in FIG. 4 and described above may be performed periodically at the data center.

[0123] In some embodiments, a change in the status of a first capability from healthy to unhealthy may also change the health status of capabilities that depend on the first capability. For example, capability A may have a dependent capability B, which has a dependent capability C, which has a dependent capability D, and so on. Capability B has a direct dependency on capability A, while C and D have indirect dependencies on capability A. In some embodiments, a change in the health status of capability A may also change the health status of both capability A's directly and indirectly dependent capabilities (e.g., the health status of capabilities B, C, and D change). In some other embodiments, only the health status of direct capability dependencies changes in response to a change in capability A's health status (e.g., the health status of capability B changes).

[0124] FIG. 5 is a simplified flowchart 500 illustrating a process performed to monitor issued capabilities in a data center using alert-capability association information, according to one embodiment. The process illustrated in FIG. 5 may be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of the respective systems, may be implemented using hardware, or a combination thereof. The software may be stored in a non-transitory storage medium (e.g., memory device). The method presented in FIG. 5 and described below is exemplary and not intended to be limiting. FIG. 5 depicts various process steps occurring in a particular order or sequence, but is not intended to be limiting. In an alternative embodiment, the process may be performed in some different order, or some steps may be performed in parallel. In some embodiments (such as the embodiment illustrated in FIG. 1), the process illustrated in FIG. 5 may be performed by the capability service 112. In some implementations, the process illustrated in FIG. 5 may be performed by a health monitoring system. In yet another embodiment, the process shown in FIG. 4 may be performed by a telemetry system located at a data center.

[0125] At 502, a set of alerts declared for the data center is identified and monitored. In one embodiment, the alerts are declared in the flock configurations of various services provided by the data center. The MFO 106 may identify these alerts when retrieving and analyzing the flock configurations during a static flock analysis process. Information identifying the set of alerts declared for the data center may then be stored.

[0126] At 504, it is determined that a particular alert from the set of monitored alerts has been triggered based on the monitoring performed at 502. For example, if an alert is triggered, the telemetry system may communicate information regarding the triggered alert to a system performing the process shown in FIG.

[0127] At 506, all issued capabilities associated with the particular triggered alert are identified. Alert-capability association information may be used to identify all issued capabilities associated with the particular triggered alert.

[0128] At 508, the health state of each capability identified at 506 is changed from healthy to unhealthy. As described above with respect to Figure 4, for capabilities whose state is changed from healthy to unhealthy, one or more actions may be triggered in response to the change.

[0129] In one embodiment, the alert-capability relationships may be used when performing a release of a flock for a service. As described above, if the capability dependencies of the flock for the service are satisfied, the MFO schedules the release of the flock. The release is then performed by CIOS central 108 in cooperation with CIOS regional 110. In the event that the release fails to execute, the MFO may use the alert-capability relationship information to determine the likely cause of the execution failure. For example, the MFO may determine that the flock release may have failed due to one of the dependent capabilities being in an unhealthy state.

[0130] FIG. 6 is a simplified flowchart 600 illustrating a process performed during flock release and execution, according to one embodiment. The process illustrated in FIG. 6 may be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of the respective systems, may be implemented using hardware, or a combination thereof. The software may be stored in a non-transitory storage medium (e.g., memory device). The method presented in FIG. 6 and described below is exemplary and not intended to be limiting. FIG. 6 shows various process steps occurring in a particular order or sequence, but this is not intended to be limiting. In an alternative embodiment, the process may be performed in some different order, or some steps may be performed in parallel. In some embodiments (such as the embodiment shown in FIG. 1), the process illustrated in FIG. 6 may be performed in coordination by MFO 106, CIOS Central 108, CIOS Regional 110, and Capability Service 112.

[0131] At 602, a multi-flock orchestrator (MFO) 106 schedules the release of a flock for a service. The MFO may schedule the release of a flock after determining that all of the flock's required capability dependencies have been satisfied.

[0132] At 604, the release may be received and executed by CIOS central 108 in cooperation with CIOS regional 110, but the execution of the release fails. In one embodiment, the execution of the flock release may include determining the current state of the target datacenter for the service, determining the desired target state of the datacenter as specified in the flock configuration file for the flock to be released, generating a plan to get from the current state to the target state, and then executing the plan. In one embodiment, the Terraform tool may be used to execute the release. For purposes of the process illustrated in FIG. 6, it is assumed that the execution of the release fails.

[0133] At 606, the MFO 106 may receive a failure code from the CIO central / regional indicating a failure to execute the release.

[0134] At 608, the MFO 106 may cause the CIOS central / regional to retry executing the release until the release is successfully executed and a success code is received or a maximum number of retries is exceeded, in one embodiment, the maximum number of retries is set to some threshold number, such as 3, 5, etc.

[0135] At 610, assume that the MFO 106 determines that the maximum number of retries has been exceeded and the release execution still indicates a failed status.

[0136] At 612, the MFO 106 determines a set of capability dependencies for the flock to be released. In some embodiments, both direct and indirect capability dependencies are determined at 612. For example, a flock configuration may identify capability A as a direct dependency, however, capability A may depend on capability B, which depends on capability C, which depends on capability D, and so on. Capabilities B, C, and D are indirect dependencies of the flock. In some embodiments, as part of the process at 612, both direct and indirect capability dependencies are identified for the flock to be released (e.g., capabilities A, B, C, and D are identified). In some other embodiments, only direct capability dependencies are identified (e.g., capability A is identified).

[0137] At 614, the MFO 106 identifies a set of alerts that are associated with the one or more capabilities identified at 612. In one embodiment, the alert-capability association information may be used to identify the alerts that correspond to the capabilities determined at 612.

[0138] At 616, the MFO 106 determines the status of each alert identified at 614 and identifies any alerts that have been triggered. In one embodiment, the MFO 106 may determine the status of the alerts by querying a telemetry system.

[0139] At 618, for the alert identified at 616 as being triggered, the MFO 106 determines a set of one or more capabilities identified at 612 that are associated with the triggered alert. The alert-capability association information may be used to identify the capabilities associated with the triggered alert.

[0140] At 620, the MFO 106 generates an error message that identifies the capabilities identified at 618 and the associated trigger alerts detected at 616 as possible reasons for the failure of the flock release execution. At 622, the MFO 106 may output the error message generated at 620. At 624, the MFO terminates the flock release execution.

[0141] In the embodiment described above, the process of flowchart 600 (except for 604) is performed by MFO 106. In other embodiments, the process may be performed by MFO 106 in cooperation with other systems. For example, some or all of the process may be performed by CIOS central 108 and / or CIOS regional 110.

[0142] 7 is a swim-lane flowchart 700 illustrating the processing performed during execution of a flock release and how possible reasons for execution failure may be identified through the use of alert-capability relationships, according to one embodiment. FIG 7 illustrates the interactions between MFO 701, CIOS 720 (including CIOS central 108 and CIOS regional 110), and other components.

[0143] In step 1, MFO 710 can create a release that is sent to CIOS 720. In step 2, CIOS 720 can apply the release to downstream services 730. In this case, the release fails and an error is returned from downstream services 730 to CIOS 720 in step 3. In response to the error in step 3, CIOS 720 sends a message to MFO 710 indicating the failure of the release in step 4.

[0144] In step 5, MFO 710 can instruct CIOS 720 to retry the release up to three times, and CIOS 720 can attempt to apply the release to downstream services 730 in step 6. In block 7, the application of the release can fail and an error message can be sent from downstream services 730 to CIOS 720 in 7. In step 8, an error message can be returned from CIOS 720 to MFO 710. In step 8, MFO 710 can instruct CIOS 720 of the failed release, and CIOS 720 can return the failed release to MFO 710 in step 10.

[0145] In step 11, the MFO 710 can request an alert associated with the failed release from a service that manages alerts (e.g., T2 / alarms 720) and determine if the capability is unhealthy. In step 12, T2 / alarms 740 returns the requested alert to the MFO 710, which can determine if a capability is associated with the triggered alert (e.g., whether the capability is unhealthy). In step 13, the MFO 710 can request that the ticket management service 720 create a ticket for the unhealthy capability. In step 14, the ticket management service 750 can return the ticket to the MFO 710.

[0146] In step 15, the MFO may store the release so that application of the release may be attempted at a later time. In step 16, any stored releases may be processed after being reverted to a healthy state. In step 17, the MFO 710 may determine if the capability associated with the release is healthy by fetching alerts, and in step 18, T2 / Alerts 740 may return any alerts associated with the capability. In step 19, if there are no alerts associated with the capability in step 18 and the release is healthy, the MFO 710 may retry the release by sending a command to CIOS 720. In step 20, CIOS 720 may return an acknowledgement and attempt the release.

[0147] In one embodiment, the MFO 106 may check the health of issued capabilities through the use of alert-capability associations prior to scheduling the release of a flock. FIG. 8 is a simplified flowchart 800 illustrating a process performed to check the health of capabilities prior to scheduling the release of a flock, according to one embodiment. The process illustrated in FIG. 8 may be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of the respective systems, may be implemented using hardware, or a combination thereof. The software may be stored in a non-transitory storage medium (e.g., memory device). The method presented in FIG. 8 and described below is exemplary and not intended to be limiting. Although FIG. 8 depicts various process steps occurring in a particular order or sequence, this is not intended to be limiting. In an alternative embodiment, the process may be performed in some different order, or some steps may be performed in parallel. In some embodiments (such as the embodiment illustrated in FIG. 1), the process illustrated in FIG. 8 may be performed by the MFO 106.

[0148] At 802, for a flock for a service to be scheduled for release, the MFO 106 determines the direct capability dependencies (e.g., direct capabilities) of the flock configuration. At 804, based on the direct capability dependencies determined at 804, the MFO 106 determines a set of indirect capabilities that are also required before the flock can be scheduled for release.

[0149] At 806, the MFO 106 determines a set of alerts associated with the capabilities determined at 602 and 604. In one embodiment, the MFO 106 may determine the set of alerts at 806 by using the alert-capability association information.

[0150] At 808, the MFO determines the status of each alert determined at 806. In one embodiment, the MFO 106 may determine the status of the alerts via a telemetry system configured to monitor the alert status of alerts configured for the data center.

[0151] At 810, the MFO 106 identifies any triggered alerts based on the alert conditions determined at 808. If at 810 it is determined that there are no triggered alerts, processing continues at 818 where the MFO 106 continues scheduling the release of flock configurations. If at least one triggered alert is determined to exist at 810, the MFO 106 identifies at 812 one or more capabilities identified at 802 and 804 that are associated with the one or more triggered alerts.

[0152] At 814, the MFO 106 causes the health state associated with the capability identified at 812 to be marked as unhealthy. At 816, scheduling of flock releases is stopped.

[0153] As described above, the MFO may use the alert-capability association information to perform a capability check prior to scheduling the release of a service's block configuration.

[0154] The various systems described above in Figures 1, 2, and 3 and the functions described above in Figures 4, 5, 6, 7, and 8 may be provided as part of an infrastructure provided by a cloud service provider (CSP) for provisioning one or more data centers that may provide one or more cloud services. One or more of these cloud services may be subscribable by one or more customers of the CSP. Figures 9, 10, 11, and 12, and the accompanying description below, describe examples of infrastructure that may be used to implement cloud services, including infrastructure as a service. Figure 13 illustrates an exemplary computer system that may be used to execute and implement the various functions described in this disclosure.

[0155] Exemplary Cloud Service Infrastructure Architecture As mentioned above, infrastructure as a service (IaaS) is a particular type of cloud computing. IaaS can be configured to provide virtualized computing resources over a public network (e.g., the Internet). In the IaaS model, a cloud computing provider can host infrastructure components (e.g., servers, storage devices, network nodes (e.g., hardware), deployment software, platform virtualization (e.g., hypervisor layer), etc.). In some cases, an IaaS provider may also provide various services (e.g., billing, monitoring, logging, load balancing, clustering, etc.) that are associated with these infrastructure components. Thus, since these services can be policy-driven, an IaaS user can maintain application availability and performance by implementing policies that drive load balancing.

[0156] In some cases, IaaS customers can access resources and services over a wide area network (WAN) such as the Internet and use the cloud provider's services to install other elements of the application stack. For example, a user can log into an IaaS platform to create virtual machines (VMs), install operating systems (OS) on each VM, deploy middleware such as databases, create storage buckets for workloads and backups, and even install enterprise software on the VMs. The customer can then use the provider's services to perform a variety of functions, including balancing network traffic, troubleshooting application issues, monitoring performance, managing disaster recovery, etc.

[0157] In most cases, the cloud computing model requires the participation of a cloud provider, which may be, but need not be, a third-party service dedicated to providing IaaS (e.g., providing, renting, selling). An entity may also choose to deploy a private cloud and become its own provider of infrastructure services.

[0158] In some examples, IaaS deployment is the process of putting a new application or a new version of an application onto a prepared application server, etc. It may also include the process of preparing the server (e.g., installing libraries, daemons, etc.). This is often managed below the hypervisor layer (e.g., server, storage, network hardware, and virtualization) by the cloud provider. Thus, the customer may be responsible for handling (e.g., on self-service virtual machines (which can be spun up on demand)), middleware, and / or application deployment, etc.

[0159] In some instances, IaaS provisioning may refer to obtaining a computer or virtual host to use and even installing the necessary libraries or services on them. In most cases, deployment does not include provisioning, which may need to be performed first.

[0160] Sometimes there are two different challenges in IaaS provisioning. First, there is the initial challenge of provisioning an initial set of infrastructure before anything works. Second, there is the challenge of evolving the existing infrastructure after all the provisioning (e.g. adding new services, modifying services, removing services, etc.). Sometimes these two challenges can be addressed by allowing the configuration of the infrastructure to be defined declaratively. In other words, the infrastructure (e.g. the components required and how they interact) can be specified by one or more configuration files. Thus, the overall topology of the infrastructure (e.g. which resources depend on which resources and how they work together) can be described declaratively. Sometimes, once the topology is specified, workflows can be generated to create and / or manage the various components described in the configuration files.

[0161] In some examples, the infrastructure may have many interconnected elements. For example, there may be one or more virtual private clouds (VPCs) (e.g., potential on-demand pools of configurable and / or shared computing resources), also known as a core network. Also, in some examples, there may be one or more inbound / outbound traffic group rules provisioned to define how the network inbound and / or outbound traffic is configured and one or more virtual machines (VMs). Other infrastructure elements such as load balancers, databases, etc. may also be provisioned. If there is a desire and / or addition of more infrastructure elements, the infrastructure may evolve in increments.

[0162] In some cases, the employment of continuous deployment techniques may enable deployment of infrastructure code across various virtual computing environments. The described techniques may also enable infrastructure management within these environments. In some instances, a service team may write code that is desired to be deployed to one or more (but often many) different production environments (e.g., across various geographic locations, possibly even across the globe). In some instances, however, the infrastructure into which the code will be deployed must first be set up. In some cases, provisioning may be performed manually, provisioning tools may be used to provision resources, and / or deployment tools may be used to deploy the code after the infrastructure has been provisioned.

[0163] 9 is a block diagram 900 illustrating an example pattern of an IaaS architecture according to at least one embodiment. A service operator 902 may be communicatively coupled to a secure host tenancy 904, which may include a virtual cloud network (VCN) 906 and a secure host subnet 908. In some examples, the service operator 902 may employ one or more client computing devices, which may be portable handheld devices (e.g., iPhones, mobile phones, iPads, computing tablets, personal digital assistants (PDAs)) or wearable devices (e.g., Google Glass head-mounted displays, etc.) running software such as Microsoft Windows Mobile and / or a variety of mobile operating systems such as iOS, Windows Phone, Android, BlackBerry 8, Palm OS, etc., and capable of using the Internet, email, short message service (SMS), BlackBerry, or other communications protocols. Alternatively, the client computing devices may be general purpose personal computers, examples of which include personal and / or laptop computers running various versions of the Microsoft Windows, Apple Macintosh, and / or Linux operating systems. The client computing devices may be workstation computers running any of a variety of commercially available UNIX or UNIX-like operating systems, including, but not limited to, various GNU / Linux operating systems such as Google Chrome OS.Alternatively or additionally, the client computing device may be any other electronic device, such as a thin-client computer, an Internet-enabled gaming system (e.g., a Microsoft Xbox gaming console with or without a Kinect® gesture input device), and / or a personal messaging device, capable of communicating over a network that can access the VCN 906 and / or the Internet.

[0164] The VCN 906 may include a local peering gateway (LPG) 910, which may be communicatively coupled to a secure shell (SSH) VCN 912 via an LPG 910 included in the SSH VCN 912. The SSH VCN 912 may include an SSH subnet 914, which may also be communicatively coupled to a control plane VCN 916 via an LPG 910 included in the control plane VCN 916. The SSH VCN 912 may also be communicatively coupled to a data plane VCN 918 via the LPG 910. The control plane VCN 916 and the data plane VCN 918 may be included in a service tenancy 919, which may be owned and / or operated by the IaaS provider.

[0165] The control plane VCN 916 may include a control plane demilitarized zone (DMZ) tier 920 that operates as a perimeter network (e.g., a portion of an enterprise network between an enterprise intranet and an external network). DMZ-based servers may help limit liability and limit intrusions. The DMZ tier 920 may also include one or more load balancer (LB) subnets 922, a control plane app tier 924 that may include an app subnet 926, and a control plane data tier 928 that may include a database (DB) subnet 930 (e.g., a front-end DB subnet and / or a back-end DB subnet). The LB subnet 922 included in the control plane DMZ tier 920 may be communicatively coupled to the app subnet 926 included in the control plane app tier 924 and an Internet gateway 934 that may be included in the control plane VCN 916, and the app subnet 926 may be communicatively coupled to the DB subnet 930, a service gateway 936, and a network address translation (NAT) gateway 938 included in the control plane data tier 928. The control plane VCN 916 may include a service gateway 936 and a NAT gateway 938 .

[0166] The control plane VCN 916 can include a data plane mirror app layer 940 that can include an app subnet 926. The app subnet 926 included in the data plane mirror app layer 940 can include a virtual network interface controller (VNIC) 942 that can run a compute instance 944. The compute instance 944 can communicatively couple the app subnet 926 of the data plane mirror app layer 940 to the app subnet 926 that can be included in the data plane app layer 946.

[0167] The data plane VCN 918 may include a data plane app layer 946, a data plane DMZ layer 948, and a data plane data layer 950. The data plane DMZ layer 948 may include a LB subnet 922, which may be communicatively coupled to an app subnet 926 of the data plane app layer 946 and an Internet gateway 934 of the data plane VCN 918. The app subnet 926 may be communicatively coupled to a service gateway 936 of the data plane VCN 918 and a NAT gateway 938 of the data plane VCN 918. Additionally, the data plane data layer 950 may include a DB subnet 930, which may be communicatively coupled to the app subnet 926 of the data plane app layer 946.

[0168] The internet gateways 934 of the control plane VCNs 916 and data plane VCNs 918 may be communicatively coupled to a metadata management service 952, which may be communicatively coupled to the public internet 954. The public internet 954 may be communicatively coupled to a NAT gateway 938 of the control plane VCNs 916 and data plane VCNs 918. The service gateways 936 of the control plane VCNs 916 and data plane VCNs 918 may be communicatively coupled to cloud services 956.

[0169] In some examples, the service gateways 936 of the control plane VCNs 916 and the data plane VCNs 918 can make application programming interface (API) calls to the cloud services 956 without going over the public Internet 954. The API calls from the service gateways 936 to the cloud services 956 can be unidirectional. The service gateways 936 can make API calls to the cloud services 956, and the cloud services 956 can send the requested data to the service gateways 936. However, the cloud services 956 do not have to initiate the API calls to the service gateways 936.

[0170] In some examples, the secure host tenancy 904 may be directly connected to an otherwise isolated service tenancy 919. The secure host subnet 908 may communicate with the SSH subnet 914 through an LPG 910, which may allow bidirectional communication through an otherwise isolated system. By connecting the secure host subnet 908 to the SSH subnet 914, the secure host subnet 908 may be accessible to other entities in the service tenancy 919.

[0171] The control plane VCN 916 may enable configuration or provisioning of desired resources by users of the service tenancy 919. The desired resources provisioned in the control plane VCN 916 may be deployed or used in the data plane VCN 918. In some examples, the control plane VCN 916 may be isolated from the data plane VCN 918, and a data plane mirror app layer 940 of the control plane VCN 916 may communicate with a data plane app layer 946 of the data plane VCN 918 via a VNIC 942 that may be included in the data plane mirror app layer 940 and the data plane app layer 946.

[0172] In some examples, a user or customer of the system may make a request, e.g., a create, read, update, or delete (CRUD) operation, through the public Internet 954, which may send the request to the metadata management service 952. The metadata management service 952 may send the request to the control plane VCN 916 through the Internet Gateway 934. The request may be received by the LB Subnet 922 included in the control plane DMZ layer 920. The LB Subnet 922 may determine that the request is valid, and in response to this determination, the LB Subnet 922 may send the request to the app subnet 926 included in the control plane app layer 924. If the request is validated and a call to the public Internet 954 is required, the call to the public Internet 954 may be sent to the NAT Gateway 938, which may make the call to the public Internet 954. Memory that may be desirable to store due to the request may be stored in the DB Subnet 930.

[0173] In some examples, the data plane mirror app layer 940 may facilitate direct communication between the control plane VCN 916 and the data plane VCN 918. For example, it may be desirable to apply configuration changes, updates, or other suitable modifications to resources included in the data plane VCN 918. The control plane VCN 916 can perform the configuration changes, updates, or other suitable modifications of the resources by communicating directly with the resources included in the data plane VCN 918 via the VNIC 942.

[0174] In some embodiments, the control plane VCN 916 and the data plane VCN 918 may be included in the service tenancy 919. In this case, the user or customer of the system may not own or operate either the control plane VCN 916 or the data plane VCN 918. Alternatively, the IaaS provider may own or operate both the control plane VCN 916 and the data plane VCN 918, and both may be included in the service tenancy 919. This embodiment may allow for network isolation that may prevent users or customers from interacting with the resources of other users or customers. This embodiment may also allow for private storage of databases by users or customers of the system without having to rely on the public internet 954, which may not have the desired level of threat prevention for storage.

[0175] In another embodiment, the LB subnet 922 included in the control plane VCN 916 can be configured to receive signals from the service gateway 936. In this embodiment, the control plane VCN 916 and the data plane VCN 918 can be configured to be called by the IaaS provider's customers without calling the public internet 954. The IaaS provider's customers may desire this embodiment because databases they use can be stored in a service tenancy 919 that is controlled by the IaaS provider and can be isolated from the public internet 954.

[0176] 10 is a block diagram 1000 illustrating another example pattern of an IaaS architecture, according to at least one embodiment. A service operator 1002 (e.g., service operator 902 of FIG. 9 ) may be communicatively coupled to a secure host tenancy 1004 (e.g., secure host tenancy 904 of FIG. 9 ), which may include a virtual cloud network (VCN) 1006 (e.g., VCN 906 of FIG. 9 ) and a secure host subnet 1008 (e.g., secure host subnet 908 of FIG. 9 ). The VCN 1006 may include a local peering gateway (LPG) 1010 (e.g., LPG 910 of FIG. 9 ), which may be communicatively coupled to a secure shell (SSH) VCN 1012 via an LPG 910 included in the SSH VCN 1012 (e.g., SSH VCN 912 of FIG. 9 ). SSH VCN 1012 may include an SSH subnet 1014 (e.g., SSH subnet 914 in FIG. 9 ), and SSH VCN 1012 may be communicatively coupled to a control plane VCN 1016 via an LPG 1010 that is included in a control plane VCN 1016 (e.g., control plane VCN 916 in FIG. 9 ). The control plane VCN 1016 may be included in a service tenancy 1019 (e.g., service tenancy 919 in FIG. 9 ), and the data plane VCN 1018 (e.g., data plane VCN 918 in FIG. 9 ) may be included in a customer tenancy 1021, which may be owned or operated by a user or customer of the system.

[0177] The control plane VCN 1016 may include a control plane DMZ tier 1020 (e.g., control plane DMZ tier 920 of FIG. 9 ) that may include a LB subnet 1022 (e.g., LB subnet 922 of FIG. 9 ), a control plane app tier 1024 (e.g., control plane app tier 924 of FIG. 9 ) that may include an app subnet 1026 (e.g., app subnet 926 of FIG. 9 ), and a control plane data tier 1028 (e.g., control plane data tier 928 of FIG. 9 ) that may include a DB subnet 1030 (e.g., similar to database (DB) subnet 930 of FIG. 9 ). The LB subnet 1022 included in the control plane DMZ layer 1020 may be communicatively coupled to an app subnet 1026 included in the control plane app layer 1024 and an Internet gateway 1034 (e.g., Internet gateway 934 in FIG. 9 ) that may be included in the control plane VCN 1016, and the app subnet 1026 may be communicatively coupled to a DB subnet 1030, a service gateway 1036 (e.g., service gateway 936 in FIG. 9 ), and a network address translation (NAT) gateway 1038 (e.g., NAT gateway 938 in FIG. 9 ) included in the control plane data layer 1028. The control plane VCN 1016 may comprise the service gateway 1036 and the NAT gateway 1038.

[0178] The control plane VCN 1016 may include a data plane mirror app layer 1040 (e.g., data plane mirror app layer 940 of FIG. 9 ), which may include an app subnet 1026. The app subnet 1026 included in the data plane mirror app layer 1040 may include a virtual network interface controller (VNIC) 1042 (e.g., the VNIC of 942) on which a compute instance 1044 (e.g., similar to the compute instance 944 of FIG. 9 ) may run. The compute instance 1044 may facilitate communication between the app subnet 1026 of the data plane mirror app layer 1040 and the app subnet 1026 included in the data plane app layer 1046 via the VNIC 1042 included in the data plane mirror app layer 1040 and the VNIC 1042 included in the data plane app layer 1046 (e.g., data plane app layer 946 of FIG. 9 ).

[0179] The Internet gateway 1034 included in the control plane VCN 1016 may be communicatively coupled to a metadata management service 1052 (e.g., metadata management service 952 of FIG. 9 ), which may be communicatively coupled to a public Internet 1054 (e.g., public Internet 954 of FIG. 9 ). The public Internet 1054 may be communicatively coupled to a NAT gateway 1038 included in the control plane VCN 1016. The service gateway 1036 included in the control plane VCN 1016 may be communicatively coupled to cloud services 1056 (e.g., cloud services 956 of FIG. 9 ).

[0180] In some examples, the data plane VCN 1018 may be included in the customer tenancy 1021. In this case, the IaaS provider may provide a control plane VCN 1016 for each customer, and the IaaS provider may configure a unique compute instance 1044 for each customer, which is included in the service tenancy 1019. Each compute instance 1044 may enable communication between the control plane VCN 1016 in the service tenancy 1019 and the data plane VCN 1018 in the customer tenancy 1021. The compute instance 1044 may enable deployment or use of resources provisioned in the control plane VCN 1016 in the service tenancy 1019 in the data plane VCN 1018 in the customer tenancy 1021.

[0181] In another example, the IaaS provider's customer may have a database that resides in the customer tenancy 1021. In this example, the control plane VCN 1016 may include a data plane mirror app layer 1040 that may include the app subnet 1026. The data plane mirror app layer 1040 may reside in the data plane VCN 1018, but may not reside in the data plane VCN 1018. That is, the data plane mirror app layer 1040 may be accessible to the customer tenancy 1021, but may not reside in the data plane VCN 1018, and may not be owned or operated by the IaaS provider's customer. The data plane mirror app layer 1040 may be configured to make calls to the data plane VCN 1018, but may not be configured to make calls to any entities included in the control plane VCN 1016. A customer may desire deployment or use of resources in the data plane VCN 1018 that have been provisioned in the control plane VCN 1016, and the data plane mirror app layer 1040 may facilitate the deployment or other use of the resources that the customer desires.

[0182] In some embodiments, the IaaS provider's customer can apply filters to the data plane VCN 1018. In this embodiment, the customer can determine what the data plane VCN 1018 can access, and the customer may limit access from the data plane VCN 1018 to the public Internet 1054. The IaaS provider may not be able to apply filters or control the data plane VCN 1018's access to any external networks or databases. The customer's application of filters and controls to the data plane VCN 1018 in the customer tenancy 1021 can help isolate the data plane VCN 1018 from other customers and the public Internet 1054.

[0183] In some embodiments, the cloud services 1056 can access services that may not be in the public internet 1054, the control plane VCN 1016, or the data plane VCN 1018 by making calls through the service gateway 1036. The connection between the cloud services 1056 and the control plane VCN 1016 or the data plane VCN 1018 may not be live or continuous. The cloud services 1056 may be on different networks owned or operated by the IaaS provider. The cloud services 1056 may be configured to receive calls from the service gateway 1036 or may not be configured to receive calls from the public internet 1054. Some cloud services 1056 may be isolated from other cloud services 1056, and the control plane VCN 1016 may be isolated from cloud services 1056 that may not be in the same region as the control plane VCN 1016. For example, control plane VCN 1016 may be located in "Region 1," and cloud service "Deployment 9" may be located in Region 1 and "Region 2." If a call to deployment 9 is made by a service gateway 1036 included in control plane VCN 1016 located in Region 1, the call may be sent to deployment 9 in Region 1. In this example, control plane VCN 1016 or deployment 9 in Region 1 may not be communicatively coupled to or in communication with deployment 9 in Region 2.

[0184] 11 is a block diagram 1100 illustrating another example pattern of an IaaS architecture, according to at least one embodiment. A service operator 1102 (e.g., service operator 902 of FIG. 9 ) may be communicatively coupled to a secure host tenancy 1104 (e.g., secure host tenancy 904 of FIG. 9 ), which may include a virtual cloud network (VCN) 1106 (e.g., VCN 906 of FIG. 9 ) and a secure host subnet 1108 (e.g., secure host subnet 908 of FIG. 9 ). The VCN 1106 may include an LPG 1110, which may be communicatively coupled to an SSH VCN 1112 via an LPG 1110 (e.g., LPG 910 of FIG. 9 ) included in the SSH VCN 1112 (e.g., SSH VCN 912 of FIG. 9 ). SSH VCN 1112 may include an SSH subnet 1114 (e.g., SSH subnet 914 in FIG. 9 ), and SSH VCN 1112 may be communicatively coupled to a control plane VCN 1116 via an LPG 1110 included in the control plane VCN 1116 (e.g., control plane VCN 916 in FIG. 9 ) and to a data plane VCN 1118 via an LPG 1110 included in the data plane VCN 1118 (e.g., data plane VCN 918 in FIG. 9 ). The control plane VCN 1116 and the data plane VCN 1118 may be included in a service tenancy 1119 (e.g., service tenancy 919 in FIG. 9 ).

[0185] The control plane VCN 1116 may include a control plane DMZ tier 1120 (e.g., control plane DMZ tier 920 of FIG. 9 ) that may include a load balancer (LB) subnet 1122 (e.g., LB subnet 922 of FIG. 9 ), a control plane app tier 1124 (e.g., control plane app tier 924 of FIG. 9 ) that may include an app subnet 1126 (e.g., similar to app subnet 926 of FIG. 9 ), and a control plane data tier 1128 (e.g., control plane data tier 928 of FIG. 9 ) that may include a DB subnet 1130. The LB subnet 1122 included in the control plane DMZ layer 1120 may be communicatively coupled to an app subnet 1126 included in the control plane app layer 1124 and an Internet gateway 1134 (e.g., Internet gateway 934 in FIG. 9 ) that may be included in the control plane VCN 1116, and the app subnet 1126 may be communicatively coupled to a DB subnet 1130, a service gateway 1136 (e.g., service gateway in FIG. 9 ), and a network address translation (NAT) gateway 1138 (e.g., NAT gateway 938 in FIG. 9 ) included in the control plane data layer 1128. The control plane VCN 1116 may comprise the service gateway 1136 and the NAT gateway 1138.

[0186] The data plane VCN 1118 may include a data plane app layer 1146 (e.g., data plane app layer 946 of FIG. 9 ), a data plane DMZ layer 1148 (e.g., data plane DMZ layer 948 of FIG. 9 ), and a data plane data layer 1150 (e.g., data plane data layer 950 of FIG. 9 ). The data plane DMZ layer 1148 may include a LB subnetwork 1122 that may be communicatively coupled to a trusted app subnetwork 1160 and a non-trusted app subnetwork 1162 of the data plane app layer 1146 and an Internet gateway 1134 included in the data plane VCN 1118. The trusted app subnetwork 1160 may be communicatively coupled to a service gateway 1136 included in the data plane VCN 1118, a NAT gateway 1138 included in the data plane VCN 1118, and a DB subnetwork 1130 included in the data plane data layer 1150. The untrusted app subnet 1162 may be communicatively coupled to a service gateway 1136 included in the data plane VCN 1118 and a DB subnet 1130 included in the data plane data layer 1150. The data plane data layer 1150 may include a DB subnet 1130 that may be communicatively coupled to a service gateway 1136 included in the data plane VCN 1118.

[0187] The untrusted app subnet 1162 may include one or more primary VNICs 1164(1)-1164(N), which may be communicatively coupled to tenant virtual machines (VMs) 1166(1)-1166(N). Each tenant VM 1166(1)-1166(N) may be communicatively coupled to a respective app subnet 1167(1)-1167(N), which may be included in a respective container egress VCN 1168(1)-1168(N), which may be included in a respective customer tenancy 1170(1)-1170(N). Each secondary VNIC 1172(1)-1172(N) may facilitate communication between the untrusted app subnet 1162 included in the data plane VCN 1118 and the app subnets included in the container egress VCNs 1168(1)-1168(N). Each container egress VCN 1168(1)-1168(N) may include a NAT gateway 1138 that may be communicatively coupled to the public Internet 1154 (e.g., public Internet 954 in FIG. 9).

[0188] An Internet gateway 1134 included in the control plane VCN 1116 and the data plane VCN 1118 may be communicatively coupled to a metadata management service 1152 (e.g., metadata management system 952 of FIG. 9 ), which may be communicatively coupled to the public Internet 1154. The public Internet 1154 may be communicatively coupled to a NAT gateway 1138 included in the control plane VCN 1116 and the data plane VCN 1118. A service gateway 1136 included in the control plane VCN 1116 and the data plane VCN 1118 may be communicatively coupled to cloud services 1156.

[0189] In some embodiments, the data plane VCN 1118 may be integrated with a customer tenancy 1170. This integration may be useful or desirable for an IaaS provider's customer, such as when they may want support for running code. A customer may provide code to run, which may be disruptive, may communicate with other customer resources, or may have undesirable effects. In response, the IaaS provider may determine whether to run the code provided to it by the customer.

[0190] In some examples, a customer of an IaaS provider may grant temporary network access to the IaaS provider and request a capability to be granted to the data plane app layer 1146. The code to execute this capability may be configured to execute in VMs 1166(1)-1166(N) and may not be configured to execute anywhere else on the data plane VCN 1118. Each VM 1166(1)-1166(N) may be connected to one customer tenancy 1170. Each container 1171(1)-1171(N) contained in VMs 1166(1)-1166(N) may be configured to execute code. In this case, there may be a double isolation (the containers 1171(1)-1171(N) executing the code may be contained in VMs 1166(1)-1166(N) contained in at least the untrusted app subnet 1162) which may help prevent erroneous or unwanted code from damaging the IaaS provider's network or a different customer's network. The containers 1171(1)-1171(N) may be communicatively coupled to the customer tenancy 1170 and may be configured to send or receive data to the customer tenancy 1170. The containers 1171(1)-1171(N) may not be configured to send or receive data to any other entity in the data plane VCN 1118. Upon completion of the code execution, the IaaS provider may disable or discard the containers 1171(1)-1171(N).

[0191] In some embodiments, the Trusted App Subnet 1160 may execute code that may be owned or operated by the IaaS provider. In this embodiment, the Trusted App Subnet 1160 may be communicatively coupled to the DB Subnet 1130 and may be configured to perform CRUD operations on the DB Subnet 1130. The Non-Trusted App Subnet 1162 may be communicatively coupled to the DB Subnet 1130, but in this embodiment, may be configured to perform read operations on the DB Subnet 1130. The Containers 1171(1)-1171(N) that may be included in each customer's VMs 1166(1)-1166(N) and that may execute code from the customer may not be communicatively coupled to the DB Subnet 1130.

[0192] In other embodiments, the control plane VCN 1116 and the data plane VCN 1118 may not be directly communicatively coupled. In this embodiment, there may not be direct communication between the control plane VCN 1116 and the data plane VCN 1118. However, communication may occur indirectly in at least one manner. An LPG 1110 may be established by an IaaS provider that may facilitate communication between the control plane VCN 1116 and the data plane VCN 1118. In another example, the control plane VCN 1116 or the data plane VCN 1118 may make a call to a cloud service 1156 via a service gateway 1136. For example, a call from the control plane VCN 1116 to the cloud service 1156 may include a request for a service that may communicate with the data plane VCN 1118.

[0193] 12 is a block diagram 1200 illustrating another example pattern of an IaaS architecture, according to at least one embodiment. A service operator 1202 (e.g., service operator 902 of FIG. 9 ) may be communicatively coupled to a secure host tenancy 1204 (e.g., secure host tenancy 904 of FIG. 9 ), which may include a virtual cloud network (VCN) 1206 (e.g., VCN 906 of FIG. 9 ) and a secure host subnet 1208 (e.g., secure host subnet 908 of FIG. 9 ). VCN 1206 may comprise an LPG 1210, which may be communicatively coupled to an SSH VCN 1212 via an LPG 1210 (e.g., LPG 910 of FIG. 9 ) included in SSH VCN 1212 (e.g., SSH VCN 912 of FIG. 9 ). SSH VCN 1212 may include SSH subnet 1214 (e.g., SSH subnet 914 in FIG. 9 ), and SSH VCN 1212 may be communicatively coupled to control plane VCN 1216 via LPG 1210 included in control plane VCN 1216 (e.g., control plane VCN 916 in FIG. 9 ) and to data plane VCN 1218 via LPG 1210 included in data plane VCN 1218 (e.g., data plane VCN 918 in FIG. 9 ). Control plane VCN 1216 and data plane VCN 1218 may be included in service tenancy 1219 (e.g., service tenancy 919 in FIG. 9 ).

[0194] The control plane VCN 1216 may include a control plane DMZ layer 1220 (e.g., control plane DMZ layer 920 of FIG. 9 ) that may include a LB subnet 1222 (e.g., LB subnet 922 of FIG. 9 ), a control plane app layer 1224 (e.g., control plane app layer 924 of FIG. 9 ) that may include an app subnet 1226 (e.g., app subnet 926 of FIG. 9 ), and a control plane data layer 1228 (e.g., control plane data layer 928 of FIG. 9 ) that may include a DB subnet 1230 (e.g., DB subnet 1130 of FIG. 11 ). The LB subnet 1222 included in the control plane DMZ layer 1220 may be communicatively coupled to an app subnet 1226 included in the control plane app layer 1224 and an Internet gateway 1234 (e.g., Internet gateway 934 in FIG. 9 ) that may be included in the control plane VCN 1216, and the app subnet 1226 may be communicatively coupled to a DB subnet 1230, a service gateway 1236 (e.g., service gateway in FIG. 9 ), and a network address translation (NAT) gateway 1238 (e.g., NAT gateway 938 in FIG. 9 ) included in the control plane data layer 1228. The control plane VCN 1216 may comprise the service gateway 1236 and the NAT gateway 1238.

[0195] The data plane VCN 1218 can include a data plane app layer 1246 (e.g., data plane app layer 946 of FIG. 9 ), a data plane DMZ layer 1248 (e.g., data plane DMZ layer 948 of FIG. 9 ), and a data plane data layer 1250 (e.g., data plane data layer 950 of FIG. 9 ). The data plane DMZ layer 1248 can include a trusted app subnet 1260 (e.g., trusted app subnet 1160 of FIG. 11 ) and a non-trusted app subnet 1262 (e.g., non-trusted app subnet 1162 of FIG. 11 ) of the data plane app layer 1246 and a LB subnet 1222 that can be communicatively coupled to an Internet gateway 1234 included in the data plane VCN 1218. The trusted app subnet 1260 may be communicatively coupled to a service gateway 1236 included in the data plane VCN 1218, a NAT gateway 1238 included in the data plane VCN 1218, and a DB subnet 1230 included in the data plane data layer 1250. The untrusted app subnet 1262 may be communicatively coupled to a service gateway 1236 included in the data plane VCN 1218 and a DB subnet 1230 included in the data plane data layer 1250. The data plane data layer 1250 may comprise a DB subnet 1230 that may be communicatively coupled to a service gateway 1236 included in the data plane VCN 1218.

[0196] The untrusted app subnet 1262 can include primary VNICs 1264(1)-1264(N) that can be communicatively coupled to tenant virtual machines (VMs) 1266(1)-1266(N) that reside within the untrusted app subnet 1262. Each tenant VM 1266(1)-1266(N) can execute code in a respective container 1267(1)-1267(N) and can be communicatively coupled to an app subnet 1226 that can be included in a data plane app layer 1246 that can be included in a container egress VCN 1268. The secondary VNICs 1272(1)-1272(N) can facilitate communication between the untrusted app subnet 1262 included in the data plane VCN 1218 and the app subnet included in the container egress VCN 1268. The container egress VCN may include a NAT gateway 1238 that may be communicatively coupled to the public Internet 1254 (e.g., the public Internet 954 in FIG. 9).

[0197] An Internet gateway 1234 included in the control plane VCN 1216 and the data plane VCN 1218 may be communicatively coupled to a metadata management service 1252 (e.g., metadata management system 952 of FIG. 9 ), which may be communicatively coupled to the public Internet 1254. The public Internet 1254 may be communicatively coupled to a NAT gateway 1238 included in the control plane VCN 1216 and the data plane VCN 1218. A service gateway 1236 included in the control plane VCN 1216 and the data plane VCN 1218 may be communicatively coupled to cloud services 1256.

[0198] In some examples, the pattern illustrated by the architecture of block diagram 1200 of FIG. 12 may be an exception to the pattern illustrated by the architecture of block diagram 1100 of FIG. 11 and may be desirable for customers of an IaaS provider when the IaaS provider cannot communicate directly with the customer (e.g., in a non-connected region). Each of containers 1267(1)-1267(N) contained in VMs 1266(1)-1266(N) for each customer may be accessible in real time by the customer. Containers 1267(1)-1267(N) may be configured to make calls to each of secondary VNICs 1272(1)-1272(N) contained in app subnet 1226 of data plane app tier 1246 that may be contained in container egress VCN 1268. Secondary VNICs 1272(1)-1272(N) may send the call to NAT gateway 1238, which may send the call to public Internet 1254. In this example, containers 1267(1)-1267(N) that are accessible in real time by customers may be isolated from control plane VCN 1216 and may be isolated from other entities included in data plane VCN 1218. Containers 1267(1)-1267(N) may also be isolated from resources of other customers.

[0199] In another example, a customer can use containers 1267(1)-1267(N) to call cloud service 1256. In this example, the customer can execute code in containers 1267(1)-1267(N) that requests a service from cloud service 1256. Containers 1267(1)-1267(N) can send the request to secondary VNICs 1272(1)-1272(N), which can send the request to a NAT gateway, which can send the request to public internet 1254. Public internet 1254 can send the request to LB subnet 1222 included in control plane VCN 1216 via internet gateway 1234. In response to determining that the request is valid, the LB subnet can send the request to the app subnet 1226, which can send the request to the cloud service 1256 via the service gateway 1236.

[0200] It should be understood that the IaaS architectures 900, 1000, 1100, 1200 depicted in the figures may have components other than those depicted. Additionally, the embodiments depicted in the figures are merely some examples of cloud infrastructure systems that may incorporate an embodiment of the present disclosure. In other embodiments, an IaaS system may have more or fewer components than depicted, may combine two or more components, or may have a different configuration or arrangement of components.

[0201] In one embodiment, the IaaS system described herein may include a self-service, subscription-based, elastically scalable, reliable, highly available and secure offering of a suite of application, middleware and database services delivered to customers. Oracle Cloud Infrastructure (OCI), offered by the Assignee, is one example of such an IaaS system.

[0202] 13 illustrates an exemplary computer system 1300 upon which various embodiments may be implemented. The system 1300 may be used to implement any of the computer systems described above. As shown in the figure, the computer system 1300 includes a processing unit 1304 that communicates with a number of peripheral subsystems via a bus subsystem 1302. The peripheral subsystems may include a processing acceleration unit 1306, an I / O subsystem 1308, a storage subsystem 1318, and a communication subsystem 1324. The storage subsystem 1318 includes a tangible computer readable storage medium 1322 and a system memory 1310.

[0203] Bus subsystem 1302 provides a mechanism for allowing the various components and subsystems of computer system 1300 to communicate with each other as desired. Although bus subsystem 1302 is shown diagrammatically as a single bus, alternative embodiments of the bus subsystem may utilize multiple buses. Bus subsystem 1302 may be any of a number of types of bus structures including a memory bus or memory controller, a peripheral bus, and a local bus using any of a variety of bus architectures. For example, such architectures may include an Industry Standard Architecture (ISA) bus, which may be implemented as a mezzanine bus manufactured in accordance with the IEEE P1386.1 standard, a MicroChannel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.

[0204] Processing unit 1304, which may be implemented as one or more integrated circuits (e.g., conventional microprocessors or microcontrollers), controls the operation of computer system 1300. Processing unit 1304 may include one or more processors. These processors may include single-core or multi-core processors. In some embodiments, processing unit 1304 may be implemented as one or more independent processing units 1332 and / or 1334, each including a single-core or multi-core processor. In other embodiments, processing unit 1304 may be implemented as a quad-core processing unit formed by incorporating two dual-core processors on a single chip.

[0205] In various embodiments, the processing unit 1304 may execute various programs in response to program code and may maintain multiple simultaneously executing programs or processes. At any given time, some or all of the program code being executed may reside on the processor 1304 and / or on the storage subsystem 1318. With suitable programming, the processor 1304 may provide the various functions discussed above. The computer system 1300 may also include a processing acceleration unit 1306, which may include a digital signal processor (DSP), a special purpose processor, and / or the like.

[0206] The I / O subsystem 1308 may include user interface input devices and user interface output devices. User interface input devices may include pointing devices such as keyboards, mice or trackballs, touch pads or touch screens integrated into displays, scroll wheels, click wheels, dials, buttons, switches, keypads, audio input devices with voice command recognition systems, microphones, and other types of input devices. User interface input devices may include motion sensing and / or gesture recognition devices such as Microsoft Kinect® motion sensors that enable user control and interaction with input devices such as Microsoft Xbox® 360 game controllers through a natural user interface using gestures and voice commands. User interface input devices may also include eye gesture recognition devices such as the Google Glass® blink detector that detects a user's eye activity (e.g., "blinking" during filming and / or menu selection) and translates eye gestures as input to an input device (e.g., Google Glass®). The user interface input devices may also include a voice recognition sensing device that allows a user to interact with a voice recognition system (eg, the Siri® navigator) via voice commands.

[0207] User interface input devices may also include, but are not limited to, three-dimensional (3D) mice, joysticks or pointing sticks, game pads and graphic tablets, as well as audio / visual devices such as speakers, digital cameras, digital video cameras, portable media players, webcams, image scanners, fingerprint scanners, barcode readers 3D scanners, 3D printers, laser range finders, and eye-tracking devices. User interface input devices may also include medical imaging input devices such as, for example, computed tomography, magnetic resonance imaging, position emission tomography, and medical ultrasound devices. User interface input devices may also include audio input devices such as, for example, MIDI keyboards, digital musical instruments, and the like.

[0208] User interface output devices may include non-visual displays such as a display subsystem, indicator lights, or audio output devices. The display subsystem may be a flat panel device such as using a cathode ray tube (CRT), liquid crystal display (LCD) or plasma display, a projection device, a touch screen, etc. In general, use of the term "output device" is intended to include all possible types of devices and mechanisms for outputting information from computer system 1300 to a user or to another computer. For example, user interface output devices may include, but are not limited to, a variety of display devices that visually convey text, graphics, and audio / video information, such as monitors, printers, speakers, headphones, automobile navigation systems, plotters, audio output devices, and modems.

[0209] Computer system 1300 may also include a storage subsystem 1318 that includes software elements illustrated here as located within system memory 1310. System memory 1310 may store program instructions that can be loaded and executed by processing unit 1304, as well as data generated during the execution of these programs.

[0210] Depending on the configuration and type of computer system 1300, system memory 1310 may be volatile (such as random access memory (RAM)) and / or non-volatile (such as read-only memory (ROM), flash memory, etc.). RAM typically contains data and / or program modules that are immediately accessible to and / or currently being operated on and executed by processing unit 1304. In some embodiments, system memory 1310 may include a number of different types of memory, such as static random access memory (SRAM) or dynamic random access memory (DRAM). In some embodiments, ROM may typically store a basic input / output system (BIOS) containing the basic routines that help transfer information between elements within computer system 1300, such as during start-up. Also, by way of non-limiting example, system memory 1310 illustrates application programs 1312, program data 1314, and operating system 1316, which may include client applications, web browsers, mid-tier applications, relational database management systems (RDBMS), and the like. By way of example, operating systems 1316 may include various versions of Microsoft Windows, Apple Macintosh, and / or Linux operating systems, various commercially available UNIX or UNIX-like operating systems (including, but not limited to, various GNU / Linux operating systems, Google Chrome OS, etc.), and / or mobile operating systems such as iOS, Windows Phone, Android OS, BlackBerry OS, and Palm OS operating systems.

[0211] Storage subsystem 1318 may also provide a tangible computer readable storage medium for storing basic programming and data constructs that provide the functionality of some embodiments. Storage subsystem 1318 may store software (programs, code modules, instructions) that, when executed by a processor, provide the functionality described above. These software modules or instructions may be executed by processing unit 1304. Storage subsystem 1318 may also provide a repository for storing data used in accordance with the present disclosure.

[0212] Storage subsystem 1300 may also include a computer readable storage medium reader 1320 that may be further coupled to a computer readable storage medium 1322. Together, optionally, computer readable storage medium 1322, in combination with system memory 1300, may comprehensively represent remote, local, fixed, and / or removable storage devices, as well as storage media for temporarily and / or permanently containing, storing, transmitting, and retrieving computer readable information.

[0213] Additionally, the computer readable storage medium 1322 containing the code or portions of code may include any suitable medium known or used in the art (including storage media and communication media), including, but not limited to, volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing and / or transmitting information. This may include tangible computer readable storage media, such as RAM, ROM, Electronically Erasable Programmable ROM (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or other tangible computer readable media. This may also include non-tangible computer readable media, such as a data signal, data transmission, or any other medium usable to transmit the desired information and accessible by computer system 1300.

[0214] By way of example, the computer readable storage medium 1322 may include hard disk drives that read from and write to non-removable, non-volatile magnetic media, magnetic disk drives that read from and write to removable, non-volatile magnetic disks, and optical disk drives that read from and write to removable, non-volatile optical disks, such as CD ROMs, DVDs, Blu-Ray disks, or other optical media. The computer readable storage medium 1322 may include, but is not limited to, Zip drives, flash memory cards, Universal Serial Bus (USB) flash drives, Secure Digital (SD) cards, DVD disks, digital video tapes, and the like. The computer readable storage medium 1322 may also include flash memory-based SSDs, enterprise flash drives, solid-state drives (SSDs) based on non-volatile memory, such as solid-state ROMs, solid-state RAMs, dynamic RAMs, static RAMs, DRAM-based SSDs, magnetoresistive RAM (MRAM) SSDs, and hybrid SSDs that use a combination of DRAM and flash memory-based SSDs. The disk drives and their associated computer-readable media may provide non-volatile storage of computer readable instructions, data structures, program modules, and other data for the computer system 1300.

[0215] The communications subsystem 1324 provides an interface to other computer systems and networks. The communications subsystem 1324 serves as an interface for the transmission and reception of data between the computer system 1300 and other systems. For example, the communications subsystem 1324 may enable the computer system 1300 to connect to one or more devices via the Internet. In some embodiments, the communications subsystem 1324 may comprise a wireless voice and / or data network (e.g., using cellular technology, 3G, 4G, or Enhanced Data Rates for Global Evolution (EDGE), WiFi (advanced data network technology such as the IEEE 802.11 family of standards, or other mobile communications technology, or any combination thereof), a global positioning system (GPS) receiver component, and / or a radio frequency (RF) transceiver component for accessing other components. In some embodiments, the communications subsystem 1324 may provide a wired network connection (e.g., Ethernet) in addition to or as an alternative to a wireless interface.

[0216] Also, in some embodiments, the communications subsystem 1324 can receive incoming communications in the form of structured and / or unstructured data feeds 1326, event streams 1328, event updates 1330, etc., on behalf of one or more users who may be using the computer system 1300.

[0217] As an example, the communications subsystem 1324 may be configured to receive data feeds 1326 in real time from users of other communications services, such as social networks and / or web feeds, such as Twitter® feeds, Facebook® updates, RSS (Rich Site Summary) feeds, and / or real-time updates from one or more third party information sources.

[0218] The communications subsystem 1324 may also be configured to receive data in the form of a continuous data stream, which may include an event stream 1328 of real-time events and / or event updates 1330, which may be continuous or effectively infinite with no apparent end. Examples of applications that generate continuous data include, for example, sensor data applications, financial tickers, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, automobile traffic monitoring, and the like.

[0219] The communications subsystem 1324 may also be configured to output structured and / or unstructured data feeds 1326, event streams 1328, event updates 1330, etc. to one or more databases that may be in communication with one or more streaming data source computers coupled to the computer system 1300.

[0220] The computer system 1300 can be one of a variety of types, including a portable handheld device (such as an iPhone® mobile phone, an iPad® computing tablet, a PDA, etc.), a wearable device (such as a Google Glass® head mounted display), a PC, a workstation, a mainframe, a kiosk, a server rack, or any other data processing system.

[0221] Due to the ever-changing nature of computers and networks, the description of the illustrated computer system 1300 is intended as an example only. Many other configurations are possible, whether with more or fewer components than the illustrated system. For example, customized hardware may also be used, and / or particular elements may be implemented in hardware, firmware, software (including applets), or a combination. Additionally, connections to other computing devices, such as network input / output devices, may be employed. Based on the disclosure and teachings provided herein, one of ordinary skill in the art will recognize other ways and / or methods of implementing various embodiments.

[0222] Although specific embodiments have been described above, various modifications, variations, alternative configurations, and equivalents are within the scope of the present disclosure. The embodiments are not limited to operating in a particular data processing environment, but may freely operate in multiple data processing environments. Furthermore, while the embodiments have been described using a specific sequence of transactions and steps, it will be apparent to those skilled in the art that the scope of the present disclosure is not limited to the sequence of transactions and steps described. Various features and aspects of the above-described embodiments may be used individually or together.

[0223] Furthermore, while embodiments have been described using a particular combination of hardware and software, it will be appreciated that other combinations of hardware and software are within the scope of the present disclosure. Embodiments may be implemented solely in hardware, solely in software, or by using a combination thereof. The various processes described herein may be implemented on the same processor or on different processors in any combination. Thus, when a component or module is described as being configured to perform an operation, such configuration may be achieved, for example, by designing an electronic circuit to perform the operation, by programming a programmable electronic circuit (such as a microprocessor) to perform the operation, or any combination thereof. Processes may communicate using a variety of techniques, including, but not limited to, conventional techniques for inter-process communication. Also, different pairs of processes may use different techniques, or the same pair of processes may use different techniques at different times.

[0224] Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive. However, it will be apparent that additions, differences, deletions, and other modifications and alterations may be made without departing from the broad spirit and scope of the appended claims. Thus, although specific embodiments of the present disclosure have been described, they are not intended to be limiting. Various modifications and equivalents are intended to be within the scope of the following claims.

[0225] The embodiments may be realised by using a computer program product comprising a computer program / instructions which, when executed by a processor, cause the processor to perform any of the methods described in the present disclosure.

[0226] Use of the terms "a," "an," and "the," and similar referents in the context of describing the disclosed embodiments (particularly in the context of the claims below) shall be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. The terms "comprising," "having," "including," and "containing" shall be construed as open-ended terms (i.e., meaning "including, but not limited to"), unless otherwise noted. The term "connected" shall be construed as partly or wholly contained in, attached to, or integrally connected to, even if there is something intervening. The recitation of ranges of values ​​herein is merely intended to serve as a shorthand method of individually referring to each separate value falling within the range, unless otherwise indicated herein, and each separate value is incorporated herein as if it were a separate recitation herein. All methods described herein may be performed in any suitable order, unless otherwise indicated herein or clearly contradicted by context. The use of any and all examples or exemplary language (e.g., "such as") herein is intended to facilitate understanding of the embodiments only and does not limit the scope of the disclosure unless otherwise asserted. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the disclosure.

[0227] Unless otherwise noted, disjunctive language, such as the phrase "at least one of X, Y, or Z," is intended to be understood in the context in which it is commonly used to indicate that an item, term, etc. can be either X, Y, Z, or any combination thereof (e.g., X, Y, and / or Z). Thus, such disjunctive language is generally not intended to, and should not, imply that an embodiment requires the presence of at least one of X, at least one of Y, or at least one of Z, respectively.

[0228] Preferred embodiments of the present disclosure are described herein, including the best mode known for carrying out the present disclosure. Modifications of these preferred embodiments may become apparent to those skilled in the art upon reading the above description. Such modifications may be adopted by those skilled in the art as necessary, and the present disclosure may be practiced differently from the specific descriptions herein. Accordingly, the present disclosure includes all modifications and equivalents of the subject matter recited in the claims appended hereto as permitted by applicable law. Moreover, unless otherwise indicated herein, the present disclosure includes any combination of the above-described elements in all possible variations thereof.

[0229] All references cited in this specification, including publications, patent applications, and patents, are herein incorporated by reference to the same extent as if each reference was individually and specifically indicated to be incorporated by reference in its entirety.

[0230] Although aspects of the disclosure are described in the above specification with reference to specific embodiments, those skilled in the art will recognize that the disclosure is not limited thereto. The various features and aspects of the disclosure described above may be used individually or together. Moreover, the embodiments may be utilized in any number of environments and applications beyond those described herein without departing from the broad spirit and scope of the specification. Accordingly, the specification and drawings are to be regarded as illustrative, and not limiting.

Claims

1. 1. A computer-implemented method comprising: including identifying links between warnings and capabilities; the capability corresponds to a function associated with a first service; The method comprises: monitoring said alert; determining that the alert is in a triggered state based on the monitoring; and In response to determining that the alert is in a triggered state, changing a state of the capability from a healthy state to an unhealthy state.

2. The method of claim 1 , wherein the functionality associated with the first service corresponds to a resource associated with the first service.

3. 3. The method of claim 1, wherein the association between the alert and the capability is declared in a flock configuration of the first service, the flock configuration identifying a set of resources associated with the first service.

4. The method of claim 3 , wherein the flock configuration identifies one or more parameters associated with the first service and a configuration setting for at least one resource associated with the first service.

5. determining that the release of the second service flock has failed; determining that the floc depends on the capability; 3. The method of claim 1, further comprising: outputting a message indicating the unhealthy state of the capability and the warning as a reason for the failure of the release of the block for the second service.

6. The method of claim 5 , further comprising delaying the release of the floc based on changing the state of the capability from healthy to unhealthy.

7. The method of claim 5 , further comprising retrying the release of the block for the second service upon determining that the capability state is healthy.

8. In response to changing the state of the capability to an unhealthy state, determining one or more dependent capabilities that depend on the capability whose state has changed from healthy to unhealthy; The method of claim 1 or claim 2, further comprising: changing the state of each of the one or more dependent capabilities to unhealthy.

9. The method of claim 1 or claim 2, wherein the monitoring is performed by a telemetry service.

10. 3. The method of claim 1 or claim 2, further comprising creating an association between the alert and a second capability, the second capability corresponding to a function associated with a second service, the second service being different from the first service.

11. identifying a set of one or more capabilities issued at a data center that are marked as healthy and associated with the alert; 3. The method of claim 1 or claim 2, further comprising: changing the state of each capability in the set of one or more capabilities from healthy to unhealthy.

12. determining a set of direct capabilities for a flock configuration of a flock to be scheduled for release; identifying a set of indirect capabilities that depend on the set of direct capabilities; determining a set of alerts associated with at least one of the set of direct capabilities or the set of indirect capabilities; monitoring a set of alert conditions for said set of alerts; determining, based on said monitoring, that at least one alert of said set of alerts is in a triggered state; In response to determining that the at least one alert is in a triggered state, changing a state of a capability associated with the triggered alert from healthy to unhealthy.

13. The method of claim 12 , further comprising updating a health state of the set of direct capabilities and the set of indirect capabilities from a healthy state to an unhealthy state.

14. 14. The method of claim 12 or claim 13, further comprising delaying release of the floc based on changing the state of the capability from healthy to unhealthy.

15. 1. A system comprising: Includes memory, the memory stores an association between an alert and a first capability, the first capability corresponding to a function corresponding to a set of resources for a service; The system comprises: The method further includes one or more processors configured to identify an association between an alert and a capability, the capability corresponding to a function associated with the first service, and the one or more processors configured to: monitoring said alert; determining that the alert is in a triggered state based on the monitoring; and and, in response to determining that the alert is in a triggered state, changing a state of the capability from a healthy state to an unhealthy state.

16. The system of claim 15 , wherein the functionality associated with the first service corresponds to a resource associated with the first service.

17. 17. The system of claim 15 or claim 16, wherein the association between the alert and the capability is declared in a flock configuration of the first service, the flock configuration identifying a set of resources associated with the first service.

18. 20. The system of claim 17, wherein the flock configuration identifies one or more parameters associated with the first service and a configuration setting for at least one resource associated with the first service.

19. the one or more processors: determining that the release of the second service flock has failed; determining that the floc depends on the capability; 17. The system of claim 15 or claim 16, further configured to: output a message indicating the unhealthy state of the capability and the warning as a reason for the failure to release the block for the second service.

20. 20. The system of claim 19, wherein the one or more processors are further configured to delay release of the floc based on changing the state of the capability from healthy to unhealthy.