Application provisioning using an active and available inventory

The orchestrator system addresses the challenge of automating application deployment in dynamic computing environments by deriving AAI from server and cloud data, optimizing resource allocation and management without requiring agents, enhancing efficiency and automation.

JP2025529853AActive Publication Date: 2025-09-09RAKUTEN SYMPHONY INC
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
JP2025511323
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-12-08
Publication Date
2025-09-09
Estimated Expiration
2042-12-08

AI Technical Summary

Technical Problem

Modern computing environments face challenges in automating the deployment of applications due to dynamic scaling and the lack of agents on servers, especially edge servers, which complicates the determination of active and available inventory (AAI) for computing resources.

Method used

An orchestrator system that collects log data from servers and cloud platforms to derive AAI without requiring agents on the servers, using vector log agents and log processors to enrich data and identify available resources, allowing for automated deployment and management of computing resources.

Benefits of technology

Enables efficient and automated deployment of applications by accurately determining AAI, optimizing resource utilization, and managing computing resources across complex environments, including edge servers, reducing the need for manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025529853000001_ABST
    Figure 2025529853000001_ABST
Patent Text Reader

Abstract

A computer system pulls observability data (metrics, logs, events, alerts, inventory) for multiple components from a remote server, which may be part of a cloud computing platform. The components may be application instances, containers, storage volumes, pods, or other components. The computer system derives utilization metrics for each component and for each of one or more types of computing resources (compute, memory, and storage). The utilization metrics are compared to the available inventory of computing resources to obtain an active and available inventory (AAI). Components may be redeployed and allocated reduced computing resources based on the AAI. Components may be grouped into clusters, and components may be consolidated into a reduced number of clusters based on the AAI. Applications may be provisioned and deployed on clusters in groups of different types (points, triangles, lines, graphs) with different runtime requirements based on location, latency, hardware resources, and / or round-robin allocation.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to provisioning applications using an active and available inventory. [Background technology]

[0002] Whether processing e-commerce transactions, streaming content, providing back-end data management for mobile applications, or other services, modern enterprises require large amounts of computing resources, including processor time, memory, and persistent data storage. The amount of computing resources changes over time. Modern computing environments can dynamically scale up and down to adapt to changes in usage. For example, Kubernetes is a popular orchestrator for adding and removing application instances based on usage. Modern computing environments may be managed by many people, which creates even more opportunities to add or remove application instances and other components of the computing environment.

[0003] It would be an advancement in the art to enable automated deployment of applications in complex computing environments. Summary of the Invention [Means for solving the problem]

[0004] The apparatus comprises a computing device including one or more processing devices and one or more memory devices operably coupled to the one or more processing devices. The one or more memory devices store executable code that, when executed by the one or more processing devices, causes the one or more processing devices to receive a specification of one or more application instances that includes both (a) one or more computing resource requirements and (b) one or more cluster runtime requirements. The executable code causes the one or more processing devices to identify, for each application instance of the one or more application instances, one or more clusters of one or more hosts that satisfy (a) and (b), and deploy the one or more application instances to the one or more hosts.

[0005] In order that the advantages of the present invention may be readily understood, a more particular description of the invention briefly described above will be rendered by reference to specific embodiments which are illustrated in the accompanying drawings, the invention being described and explained with additional specificity and detail through the use of the accompanying drawings, with the understanding that these drawings depict only typical embodiments of the invention and therefore should not be considered as limiting its scope. [Brief explanation of the drawings]

[0006] [Figure 1] FIG. 1 is a schematic block diagram of a network environment in which active and available inventory (AAI) discovery may be performed, according to one embodiment.

[0007] [Figure 2] FIG. 2 is a schematic block diagram illustrating components for collecting and processing log data according to one embodiment.

[0008] [Figure 3] FIG. 3 is a schematic block diagram illustrating sources of provisioning data according to one embodiment.

[0009] [Figure 4] FIG. 4 is a schematic block diagram illustrating components illustrating the processing of log data to obtain AAI according to one embodiment.

[0010] [Figure 5] FIG. 5 is a process flow diagram of a method for collecting provisioning data according to one embodiment.

[0011] [Figure 6] FIG. 6 is a process flow diagram of a method for deriving an AAI according to one embodiment.

[0012] [Figure 7] FIG. 7 is a schematic block diagram illustrating the derivation of relationships between components according to one embodiment.

[0013] [Figure 8] FIG. 8 is a schematic block diagram of a topology of components of a network environment, according to one embodiment.

[0014] [Figure 9] FIG. 9 is a process flow diagram of a method for identifying relationships between components according to a manifest and dynamic provisioning data, according to one embodiment.

[0015] [Figure 10] FIG. 10 is a process flow diagram of a method for identifying session relationships between components according to one embodiment.

[0016] [Figure 11] FIG. 11 is a process flow diagram of a method for identifying access relationships between components according to one embodiment.

[0017] [Figure 12]FIG. 12 is a process flow diagram of a method for identifying network relationships according to one embodiment.

[0018] [Figure 13] FIG. 13 is a process flow diagram of a method for generating a representation of a topology according to one embodiment.

[0019] [Figure 14A] FIG. 14A is an exemplary representation of a topology according to one embodiment.

[0020] [Figure 14B] FIG. 14B is an exemplary diagram of application data according to one embodiment.

[0021] [Figure 14C] FIG. 14C is an exemplary diagram of cluster data according to one embodiment.

[0022] [Figure 14D] FIG. 14D is an exemplary diagram illustrating the importance of storage volumes, according to one embodiment.

[0023] [Figure 15] FIG. 15 is a diagram illustrating data used to redeploy applications and perform cluster integration, according to one embodiment.

[0024] [Figure 16A] FIG. 16A illustrates an exemplary application redeployment and cluster integration according to one embodiment. [Figure 16B] FIG. 16B illustrates an exemplary application redeployment and cluster integration according to one embodiment. [Figure 16C] FIG. 16C illustrates an exemplary application redeployment and cluster integration according to one embodiment.

[0025] [Figure 17A] FIG. 17A is a process flow diagram of an exemplary method for performing redeployment of an application according to one embodiment.

[0026] [Figure 17B] FIG. 17B is a process flow diagram of an exemplary method for performing redeployment of an application according to one embodiment.

[0027] [Figure 18] FIG. 18 is a process flow diagram of a method for merging clusters according to one embodiment of the present invention.

[0028] [Figure 19] FIG. 19 is a process flow diagram of a method for identifying candidate cluster mergers.

[0029] [Figure 20] FIG. 20 is a schematic block diagram illustrating a topology change according to one embodiment.

[0030] [Figure 21] FIG. 21 is a process flow diagram of a method for locking a topology according to one embodiment.

[0031] [Figure 22] FIG. 22 is a process flow diagram of a method for preventing topology changes according to one embodiment.

[0032] [Figure 23] FIG. 23 is a process flow diagram of a method for detecting topology changes according to one embodiment.

[0033] [Figure 24] FIG. 24 is a schematic diagram illustrating the deployment of multiple applications on multiple clusters according to one embodiment.

[0034] [Figure 25] FIG. 25 is a schematic block diagram illustrating a cluster specification according to one embodiment.

[0035] [Figure 26] FIG. 26 is a schematic block diagram illustrating a dot application specification according to one embodiment.

[0036] [Figure 27] FIG. 27 is a schematic block diagram illustrating a triangle application specification according to one embodiment.

[0037] [Figure 28] FIG. 28 is a schematic block diagram illustrating a line application specification according to one embodiment.

[0038] [Figure 29] FIG. 29 is a process flow diagram of a method for provisioning a dot application according to one embodiment.

[0039] [Figure 30] FIG. 30 is a process flow diagram of a method for provisioning a triangular application according to one embodiment.

[0040] [Figure 31] FIG. 31 is a process flow diagram of a method for provisioning a line application according to one embodiment.

[0041] [Figure 32] FIG. 32 is a process flow diagram of a method for provisioning a graph application according to one embodiment.

[0042] [Figure 33] FIG. 33 is a diagram illustrating the division of a graph application into line and triangle applications according to one embodiment.

[0043] [Figure 34] FIG. 34 is a schematic block diagram of an exemplary computing device suitable for implementing methods according to embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0044] 1 illustrates an exemplary network environment 100 in which the systems and methods disclosed herein may be used. The components of network environment 100 may be connected to each other by a network, such as a local area network (LAN), a wide area network (WAN), the Internet, a chassis backplane, or other type of network. The components of network environment 100 may be connected by wired or wireless network connections.

[0045] The network environment 100 includes multiple servers 102. Each of the servers 102 may include one or more computing devices, such as a computing device having some or all of the attributes of computing device 3400 of FIG. 34. Each server 102 lacks an agent to coordinate the execution of management tasks. The systems and methods described herein allow for active and available inventory (AAI) determinations to be performed for servers 102 that lack an agent to assist in determining the AAI.

[0046] As used herein, "active and available inventory" (AAI) refers to computing resources available for allocation to application instances, including some or all of the storage on physical storage devices installed on the server 102, the memory of the server 102, the processing cores of the server 102, and the networking bandwidth of network connections between the server 102 and other servers 102 or other computing devices.

[0047] Computing resources may be allocated within a cloud computing platform 104, such as Amazon Web Services (AWS), GOOGLE CLOUD, AZURE, or other cloud computing platforms. Cloud computing resources may include physical storage, processor time, memory, and / or networking bandwidth purchased in units specified by a provider on the cloud computing platform.

[0048] In some embodiments, some or all of the servers 102 may function as edge servers within a telecommunications network. For example, some or all of the servers 102 may be coupled to a baseband unit (BBU) 102a that provides conversion between radio frequency signals output and received by an antenna 102b and digital data transmitted and received by the servers 102. For example, each BBU 102a may perform this conversion according to a cellular wireless data protocol (e.g., 4G, 5G, etc.). Servers 102 that function as edge servers may have limited computing resources or may be heavily loaded, making it infeasible for the servers 102 to run agents that collect data for obtaining AAI. Similarly, if there are a large number of servers 102, installing agents for data collection may be a time-consuming task.

[0049] The orchestrator 106 provisions computing resources to application instances of one or more different application executables, such as according to a manifest that defines the computing resource requirements for each application instance. The manifest can define dynamic requirements that define scaling up the number of application instances and corresponding computing resources depending on usage. The orchestrator 106 can include or work with utilities such as KUBERNETES to perform dynamic scaling up and down of the number of application instances.

[0050] The orchestrator 106 runs on a computer system separate from the server 102 and is connected to the server 102 using a network that requires the use of a destination address for communication, such as a network including other protocols including the Ethernet protocol, Internet Protocol (IP), Fibre Channel, or any higher-level protocols built upon the aforementioned protocols, such as User Datagram Protocol (UDP), Transport Control Protocol (TCP), etc.

[0051] The orchestrator 106 can cooperate with the servers 102 to initialize and configure the servers 102. For example, each server 102 can cooperate with the orchestrator 106 to obtain a gateway address to use for outbound communications and a source address assigned to the server 102 to use for inbound communications. The servers 102 can cooperate with the orchestrator 106 to install an operating system on the servers 102. For example, the gateway address and source address may be provided, and the operating system may be installed using techniques described in U.S. Patent Application No. 16 / 903,266, entitled "AUTOMATED INITIALIZATION OF SERVERS," filed June 16, 2020, which is incorporated herein by reference in its entirety.

[0052] The orchestrator 106 may be accessible via an orchestrator dashboard 108. The orchestrator dashboard 108 may be implemented as a web server or other server-side application accessible via a browser or client application running on a user computing device 110, such as a desktop computer, laptop computer, mobile phone, tablet computer, or other computing device.

[0053] The orchestrator 106 can cooperate with the servers 102 to provision computing resources for the servers 102 and instantiate components of the distributed computing system on the servers 102 and / or the cloud computing platform 104. For example, the orchestrator 106 can import manifests that define the provisioning of computing resources and instantiation of components, such as clusters 111, pods 112 (e.g., KUBERNETES pods), containers 114 (e.g., DOCKER containers), storage volumes 116, and application instances 118. The orchestrator can then allocate computing resources and instantiate the components according to the manifests.

[0054] The manifest may define requirements such as network latency requirements, affinity requirements (same node, same chassis, same rack, same data center, same cloud region, etc.), anti-affinity requirements (different nodes, different chassis, different racks, different data centers, different cloud regions, etc.), as well as minimum provisioning requirements (number of cores, amount of memory, etc.), performance or quality of service (QoS) requirements, or other constraints. Thus, the orchestrator 106 can provision computing resources to meet or nearly meet the requirements of the manifest.

[0055] Component instantiation and component management may be performed by workflows. A workflow is a set of tasks, executables, configurations, parameters, and other computing capabilities that are predefined and stored in the workflow repository 120. Workflows may be defined to instantiate each type of component (e.g., clusters 111, pods 112, containers 114, storage volumes 116, application instances), monitor the performance of each type of component, repair each type of component, upgrade each type of component, replace each type of component, copy (e.g., snapshot, backup), restore each type of component from the copy, and perform other tasks. Some or all of the tasks performed by the workflows may be implemented using Kubernetes or other utilities to perform some or all of the tasks.

[0056] The orchestrator 106 can instruct the workflow orchestrator 122 to perform a task with respect to a component. In response, the workflow orchestrator 122 retrieves a workflow corresponding to the task (e.g., task type and component type, such as instantiate, monitor, upgrade, replace, copy, restore, etc.) from the workflow repository 120. The workflow orchestrator 122 then selects a worker 124 from a worker pool and instructs the worker 124 to perform the workflow with respect to the server 102 or cloud computing platform 104. The instruction from the orchestrator 106 can specify a particular server 102, cloud region or cloud provider, or other location for executing the workflow. The worker 124, which may be a container, then performs the workflow's functions with respect to the location instructed by the orchestrator 106. In one embodiment, the worker 124 can also perform the task of retrieving the workflow from the workflow repository 120 as instructed by the workflow orchestrator 122.

[0057] In one embodiment, the container implementing the worker 124 is remote from the server 102 on which the worker 124 implements the workflow. The worker 124 can still execute some or all of the workflow without an agent installed on the server 102 or cloud computing platform 104 that is programmed to collaborate with the worker 124 to execute the workflow. For example, the worker 124 can establish a secure command line interface (CLI) connection to the server 102 or cloud computing platform 104. For example, a secure shell (ssh), remote login (rlogin), remote procedure call (RPC), or other interface provided by the operating system of the server 102 or cloud computing platform 104 can be used to send instructions and verify completion of the instructions on the server 102 or cloud computing platform 104.

[0058] One workflow may involve monitoring computing resource usage by each component (hereinafter, the "monitoring workflow"). The monitoring workflow may be invoked periodically by the orchestrator 106 for each component, or the monitoring workflow may be a persistent process that runs periodically with periods of inactivity between them.

[0059] The monitoring workflow may include establishing a secure connection to each component, reading one or more log files for each component, and passing the log files to a vector log agent 126. The vector log agent 126 may perform initial processing on the data in the log files to obtain enriched data. The vector log agent 126's processing may include enriching the data in the log files (e.g., providing contextual information indicating the component, time, identifier of the originating server 102, hosting container 114, cluster 111, pod 112, virtual machine, unit of computing resource of the cloud computing platform 104, etc.), performing map-reduce functions on messages in the log files, combining messages in the log files into aggregate representations of messages, and other functions. The vector log agent 126 may process the log files according to one or more Vector Remap Language (VRL) statements. The vector log agent 126 may run independently of the worker 124, or the monitoring workflow may include running an instance of the vector log agent 126. For example, each monitoring workflow may include a set of VRL statements that correspond to the types of components that the monitoring workflow is configured to monitor, and each monitoring workflow may then include processing log files according to the VRL statements of the monitoring workflow.

[0060] The enriched data output by the vector log agent 126 can be stored in the log store 128. The log processor 130 reads the enriched data from the log store and derives an active and available inventory (AAI), which is a list of computing resources available for allocation to components. How the log processor 130 obtains the AAI is described in more detail below. The log processor 130 passes the AAI to the orchestrator 106. The orchestrator 106 can use the AAI to perform various functions on the components, such as adding, removing, or redeploying them to a different location.

[0061] FIG. 2 illustrates the collection of log files 200 from various components. The log files 200 may be collected using each component's monitoring workflow or other techniques for collecting log files. The log files 200 may include log files generated by an operating system 202 running on the server 102. Alternatively, the cloud computing platform 104 may generate log files 200 that describe the state of units of computing resources and / or executables running on the cloud computing platform 104. Virtual machines on which components run may also generate log files 200. In the following description, reference is made to the log files 200 with the understanding that any observability data represented as log files or other formats may be collected and processed in a similar manner. In particular, metrics, events, alerts, inventory, and other data may be collected instead of or in addition to the log files 200 and processed in a similar manner.

[0062] A cluster 111 is a collection of hosts (servers 102 and / or one or more units of computing resources on a cloud computing platform) managed as a unit. Each host includes a master running on one of the hosts that manages the deployment of pods 112, containers 114, and application instances 118 on the hosts of the cluster. The master manages the scaling up, scaling down, and redeployment of the application instances 118. As used herein, actions performed by and with respect to a cluster 111 may be understood to be performed by or with respect to the master managing the cluster 111. Each cluster 111 may generate one or more log files 200 that describe the operation of the cluster 111.

[0063] Kubelet 204 is a KUBERNETES agent that runs on a node and implements instructions from a server 102 or cluster 111 on a cloud computing platform to instantiate, monitor, and manage pods 112. Each Kubelet 204 can generate one or more log files 200 that describe the operation of the Kubelet 204 and each pod 112 running within it. A pod 112 is a group of one or more containers 114 with shared storage, network resources, and execution context. A pod 112 can generate one or more log files 200 that describe the state of the pod 112 and the execution of the containers 114 in the pod 112. Each container 114 can generate one or more log files 200 that describe the execution of the container and any application instances 118 running within the container 114. Each application instance 118 can also generate one or more log files that describe the operation of the application instance 118. A storage volume 116 may be a unit of virtualized storage, and a storage manager that implements the storage volume 116 may also generate one or more log files 200 that describe the operation of the storage volume 116 .

[0064] The log files 200 are pulled from the server 102 or cloud computing platform 104 where they are stored and processed by the vector log agent 126 to generate enriched data. The enriched data is processed by the log processor 130 to obtain the AAI. The orchestrator 106 receives the AAI and manages the provisioning of unused computing resources identified in the AAI for use by the components.

[0065] 3, the data included in the log file 200 may be related to provisioning data 300 to obtain AAI. The provisioning data 300 includes identifiers of components instantiated by the orchestrator 106 and allocation data indicating the computing resources allocated to each component. For example, the on-premise provisioning data 302 may describe provisioning for one or more servers 102. For example, the on-premise provisioning data 302 may include multiple entries, each including a node identifier (i.e., an identifier of the server 102), a computing allocation (e.g., a number of processor cores), a memory allocation (e.g., a number of megabytes (MB), a number of gigabytes (GB), or other units of memory), a storage allocation (e.g., a number of megabytes (MB), a number of gigabytes (GB), or other units of storage), and a component identifier (e.g., an identifier of the cluster 111, pod 112, container 114, storage volume 116, or application instance 118) to which the allocation belongs. The component identifier may be in the form of a universally unique identifier (UUID) that is centrally assigned, such as by the orchestrator 106 or other central component, to all components that belong to a common namespace. An entry may reference multiple components. For example, provisioning may occur at the level of a cluster 111, such that all pods 112, containers 114, storage volumes 116, and application instances 118 of the cluster 111 are referenced in the entry for the cluster 111.

[0066] The provisioning data 300 may further include cloud provisioning data 304. The cloud provisioning data 304 may describe provisioning for one or more units of computing resources on the cloud computing platform 104. The cloud provisioning data 304 may include multiple entries, each including a unit identifier that identifies a unit of cloud computing resources. The identifier of the unit of computing resources may further identify a cloud computing provider (e.g., AWS, AZURE, GOOGLE CLOUD), a region of the cloud computing platform 104, and / or other data. Each entry may further include data describing an allocation of computing, memory, and storage. Each entry may further include identifiers of one or more components to which the allocation belongs, as described above with respect to the on-premises provisioning data 302.

[0067] Note that the on-premise provisioning data 302 and the cloud provisioning data 304 are dynamic: the orchestrator 106 can scale up and down the number of application instances 118 of any given executable, as well as the number of pods 112, containers 114, and storage volumes 116 used by the application instances.

[0068] In addition to the provisioning data 300, the AAI may also be determined using other data, such as hardware inventory data 306 and cloud inventory data 308. The hardware inventory data 306 may include an entry for each server 102. Each entry may indicate the computing (e.g., total number of processing cores, graphics processing unit (GPU) cores, or other computing components), memory, and storage available on the server 102, as well as the node identifier of the server 102. The cloud inventory data 308 similarly includes entries that include identifiers of units of cloud computing resources and the computing, memory, and storage available for that unit. The hardware inventory data 306 and the cloud inventory data 308 may indicate current availability; i.e., an entry may be removed or flagged as unavailable in response to the server 102 or cloud computing platform 104 referenced by the entry becoming unavailable due to a failure or lack of network connectivity. The availability of the server 102 or cloud computing platform 104 may be determined by performing a health check, sending a ping message, measuring traffic latency, detecting a failed network connection, or any other technique for determining the status and accessibility of a computing device.

[0069] 4 illustrates an approach for calculating the AAI. The log file 200 includes multiple log messages 400. Each message may include a text string including a component identifier and a value, such as a usage value. The entry identifier may also be obtained from the directory location of the log file or the name of the log file. The usage value may include some or all of the following: an indicator of the processor time spent executing the component identified by the entry identifier; an amount of memory occupied by the component identified by the component identifier; and an amount of storage used (e.g., written) by the component identified by the component identifier. For example, there may be separate entries, each indicating separate information regarding the component identifier: one entry indicating processor time and another entry indicating memory used. In one implementation, a log message 400 includes one or more usage values, and another log message 400 includes a process identifier and a component identifier executing the process identified by the process identifier.

[0070] The log message 400 is processed by the vector agent 126 to obtain enriched data 402. For example, items of the enriched data 402 may include a component identifier and usage metrics (processor time, memory, storage) for that component identifier. The vector agent 126 may obtain the enriched data 402 by executing one or more VRL statements on the log message 400. For example, a log message 400 that associates a process identifier with a usage value may be mapped by the vector agent 126 to a log message that associates the process identifier with the component identifier. The vector agent 126 may execute a map-reduce function to aggregate the usage values ​​into aggregated usage metrics for the component identifier.

[0071] The enrichment data 402 may then be processed by the log processor 130 along with the provisioning data 300 to obtain an active and available inventory (AAI) 406. For example, the provisioning data 300 may include provisioning entries 404 that include a node identifier of a server 102 or an identifier of a unit of computing resource within a cloud computing platform. Each provisioning entry 404 may include a component identifier, i.e., an identifier of a cluster 111, a pod 112, a container 114, a storage volume 116, or an application instance 118. Each provisioning entry 404 may include a value indicating an allocation, i.e., the computing, memory, and / or storage allocated to the component identified by the component identifier.

[0072] Thus, the log processor 130 may obtain one or more provisioning entries 404 that include the component identifier, and items of enrichment data 402 that include the same component identifier. For a given computing resource on a host (a unit of computing resource within the server 102 or cloud computing platform 104), let U(t,i) denote the utilization of that computing resource reported at a given time (t) for component i, let P(t,i) denote the current allocation of that computing resource to component i, and let T denote the inventory of that computing resource available on the host. Thus, the AAI of that computing resource on the host is:

number

number

[0073] 5 and 6 illustrate methods 500 and 600, respectively, that may be performed using network environment 100 to obtain AAI. Methods 500 and 600 may be performed by one or more computing devices 3400 (see description of FIG. 34 below), such as one or more computing devices executing orchestrator 106 and / or log processor 130.

[0074] Referring specifically to FIG. 5 , method 500 may include step 502 of obtaining component identifiers for statically defined components, such as those referenced in a manifest ingested by orchestrator 106. Method 500 may include step 504 of obtaining component identifiers for dynamically created components. Dynamically created components may be those instantiated to scale up capacity. Dynamically created components may be created by orchestrator 106 or KUBERNETES. Component identifiers for dynamically created components may be obtained from log files 200 generated by KUBERNETES, i.e., the KUBERNETES master, Kubelet, or other component of the KUBERNETES installation that performs component instantiation. Note that dynamically created components may also be deleted. Thus, the current set of component identifiers obtained in steps 502 and 504 may be updated to remove component identifiers for components that were dynamically deleted due to a scale-down, host failure, or other event.

[0075] Method 500 may include obtaining 506 static provisioning for each component identifier of each statically defined component and obtaining 508 dynamic provisioning for each component identifier of each dynamically created component. The provisioning for each component identifier may include a host identifier (an identifier for a server 102 or a unit of computing resource in a cloud computing platform) and an allocation of one or more computing resources (computing power, memory, and / or storage). Method 500 may further include obtaining a total available inventory. The total available inventory may include an inventory of each host currently available (functional and accessible via a network connection). The inventory of each host may include total processor cores, memory, and / or storage capacity.

[0076] 6, method 600 may include deriving 602 usage data for each component identifier identified in steps 502 and 504. As described above, deriving 602 usage data may include retrieving log file 200, enriching log file 200 to obtain enriched data 402, and aggregating enriched data 402 to obtain usage metrics for each component identifier.

[0077] Method 600 may include deriving 604 usage data for each host. For example, usage metrics for each component running on each host may be aggregated (e.g., summed) to obtain total metrics for each host, i.e., total computing power usage, total memory usage, total storage usage. As used herein, "computing power" may be defined as the amount of processor time used, the number of processor cycles used, and / or the percentage of processor cycles or time used.

[0078] Method 600 may include step 606 of retrieving static and dynamic provisioning data for each component identifier (see description of steps 506 and 508) and an inventory of each host (see description of step 510). An AAI may then be derived (608). As described above, step 608 may include calculating some or all of the AAI(t), O(t,i), and O(t) for each computing resource (computing power, memory, storage) of each host.

[0079] The method 600 may further include a step 610 of using the AAI to modify provisioning in the network environment 100. A non-limiting list of modifications may include: Provisioning additional components (clusters 111, pods 112, containers 114, storage volumes 116, and / or application instances 118) to utilize the computing resources identified in the AAI according to the manifest. Redeploy a component to a different host to more closely meet the performance, quality of service, affinity, anti-affinity, latency, or other requirements expressed in the manifest. Remove underutilized components. Removing underutilized components that are distributed across multiple servers 102 or multiple units of computing resources within the cloud computing platform 104 and redeploying some or all of the underutilized components onto a reduced number of hosts. Redeploying underutilized components (e.g., (O(t,i) / P(t,i))<0.5) to a server 102 or to a cloud computing platform 104 that has higher latency and / or fewer computing resources than the underutilized component's current host. Redeploying overutilized components (e.g., (O(t,i) / P(t,i))<0.9) to servers 102 that have lower latency and / or more computing resources than the overutilized component's current host.

[0080] 7 , the log processor 130, the orchestrator 106, and / or some other component may further process the provisioning data 300 and the log file 200 to identify relationships between component identifiers. For example, the provisioning data 300 may indicate a hosting relationship 700. As used herein, a “hosting relationship” refers to a component running on or within another component, such as a server 102, a cluster 111 or a pod 112 hosted by a unit of computing resources of the cloud computing platform 104, a container 114 running in a pod 112, or an application instance 118 running in a container. A storage volume 116 may be considered to have a hosting relationship 700, i.e., a relationship of being hosted by the container 114 or pod 112 on which the storage volume 116 resides. The hosting relationship 700 may be derived from instructions in a manifest that defines the instantiation of a second component on a first component, thereby defining the hosting relationship 700 between the first and second components. Hosting relationships may be derived from the log file 200 in a similar manner, i.e., a record of instantiating a second component on a first component establishes a hosting relationship between the first and second components.

[0081] The provisioning data 300 may further indicate environment variable relationships 702. The manifest may include instructions to configure one or more environment variables of a first component to reference a second component, e.g., to configure the first component to use the services of or provide services to the second component. The log file 200 may record configuring one or more environment variables of a first component to reference a second component in a similar manner.

[0082] The provisioning data 300 may further indicate a network relationship 704. The manifest may include instructions for configuring a first component to use an IP address or other type of address belonging to a second component, thereby establishing the network relationship 704 between the first and second components. The log file 200 may record that a first component has been configured to reference an address of a second component in a similar manner. Establishing the network relationship 704 may be a multi-step process: 1) determining that the first component is configured to use the first address; and 2) mapping the first address to an identifier of the second component.

[0083] As noted above, provisioning data 300 is dynamic and may change over time. Thus, some or all of hosting relationships 700, environment variable relationships 702, and network relationships 704 may be re-derived on a fixed, recurring basis or in response to detection of records in log file 200 that indicate actions that may affect any of these relationships 702-704.

[0084] The log files 200 may also be evaluated to identify other types of relationships between components. For example, the log files 200 may be evaluated to identify session relationships 706. When a first component establishes a session at the application level to use an application instance 118 that is or is hosted by a second component, one or more log files 200 generated by the second component may record this fact. Thus, to obtain the current session relationship 706 between the pair of components, the log files 200 may be analyzed to identify session creations and terminations.

[0085] The log files 200 may be evaluated to identify access relationships 708. When a first component accesses a session of an application instance 118 that is or is hosted by a second component, one or more log files 200 generated by the second component may record this fact. The access may include generating a request for a service provided by the second component, reading data from the second component, writing data to the second component, or other interaction between the first and second components. Accordingly, the log files 200 may be analyzed to identify access by the first component to the second component. Whether an access indicates a current access relationship may be handled in various ways. That is, in response to identifying a record of an access, an access relationship 708 may be created between the first and second components and accessed by the second component, and this access relationship may (a) be maintained as long as the first and second components exist, or (b) be deleted if no access is recorded in the log files 200 for a threshold period.

[0086] The log files 200 may be evaluated to identify network connection relationships 708. For example, when a first component establishes a network connection to a second component, the log files 200 of one or both of the first and second components may record this fact. Accordingly, the log files 200 may be analyzed to identify the establishment of the network connection between the first and second components and the termination (if any) of the network connection between the first and second components. In this manner, all active network connections between the components may be identified as network connection relationships 710. In response to identifying the creation of a network connection between the first and second components, the network connection relationship 710 may be created between the first and second components, and the network connection relationship 710 may (a) be maintained as long as the first and second components exist, (b) be deleted when the network connection is terminated, or (c) expire if a new network connection is not established within a threshold time after the network connection is terminated.

[0087] Network connection relationship 710 may be distinguished from network relationship 704 in the sense that network connection relationship 710 refers to an actual network connection, while network relationship 704 refers to configuring a first component with a network address of a second component, regardless of whether a network connection is established. In some embodiments, only network connection relationship 710 is used.

[0088] Referring to FIG. 8 , the log processor 130, the orchestrator 106, and / or some other component may further generate a topology representation 800. The topology 800 may be represented as a graph including nodes and edges. Each node may be a component identifier for a component. The components may include a host 802 (e.g., a server 102 or a unit of computing resources in a cloud computing platform), a cluster 111, a pod 112, a container 114, a storage volume 116, an application instance 118, or other components. The edges of the topology connect the nodes and represent relationships between the nodes, such as any of the hosting relationships 700, environment variable relationships 702, network relationships 704, session relationships 706, access relationships 708, and network connection relationships 710. The edges may be unidirectional, indicating that a first node depends on a second node for correct functioning, but that the second node does not depend on the first node. The edges may be bidirectional, indicating that the first node and the second node depend on each other. For example, hosting relationship 700 may be unidirectional, indicating a dependency of a second component on a first component that is a host for the second component. Network relationship 704 or network connection relationship 710 may be bidirectional, because both components must function for a network connection to exist.

[0089] FIG. 9 illustrates a method 900 for processing provisioning data 300. Method 900 may be performed by log processor 130, orchestrator 106, and / or some other component. Provisioning data 300 is retrieved (902). Retrieving 902 may include pulling the provisioning data from a manifest ingested by orchestrator 106 and pulling log files 200 from the component as described above with respect to FIG. 2. Retrieving 902 may include an enrichment step in which data from the manifest and / or log files 200 is processed by vector log agent 126 to add additional information, perform map-reduce operations, or perform other operations. For example, enrichment may include adding an identifier of the source of log file 200, the directory location of log file 200, or other data to facilitate associating the data in log file 200 with a particular component identifier. Retrieving 902 may include processing the manifest and / or log file 200 according to one or more VRL statements.

[0090] The method 900 may include extracting 904 a hosting relationship 700. Extracting the hosting relationship 700 may include parsing sentences of the form "<instantiation instruction>...<host component identifier>...<hosted component identifier>". For example, there may be a set of keywords that indicate instantiation that can be identified, and lines of code or log messages that contain these keywords can be processed to obtain the identifiers of the host component and the hosted component. A hosting relationship 700 can then be created that references the identifiers of the host component and the hosted component.

[0091] Extracting hosting relationships 904 may further include deleting hosting relationships 700 where a hosted or host component has been deleted. Log messages containing instructions to delete a component may be identified, an identifier for the deleted component may be extracted, and any hosting relationships 700 that reference the identifier for the deleted component may be deleted.

[0092] The method 900 may include extracting 906 environment variable relationships 702. Extracting environment variable relationships 702 may include parsing statements of the form "<configuration instruction>...<configured component identifier>...<referenced component identifier>." For example, a set of keywords may be found in a statement or log message regarding the setting of an environment variable. These keywords may be identified, and the lines of code or log messages containing these keywords may be processed to identify the identifiers of the configured component, i.e., the component for which the environment variable was set, and the identifiers of the referenced component, i.e., the components referenced by the configured component's environment variables. An environment variable relationship 702 may then be created that references the identifiers of the configured component and the referenced component, and possibly one or more environment variables of the configured component that are configured to reference the referenced component.

[0093] A statement in the log file 200 that creates an environment variable relationship 702 may modify a previously existing environment variable relationship. For example, the environment variable relationship 702 may record the name of an environment variable for a configured component. A first environment variable relationship 702 for a configured component that includes a variable name may be deleted in response to a subsequently identified environment variable relationship 702 for the configured component that references the same variable name. An exception to this approach may be implemented if an environment variable can store multiple values. For example, an explicit delete command including the variable name, configured component identifier, and referenced component identifier is required before deleting an environment variable relationship 702 that includes the variable name, configured component identifier, and referenced component identifier.

[0094] Method 900 can include a step 908 of extracting network relationship 704. The step of extracting network relationship 704 can include steps of parsing sentences in the form of "<network configuration instruction>...<configured component identifier>...<IP address, domain name, URL, etc.>" and sentences in the form of "<address assignment instruction>…<referenced component identifier>...<IP address, domain name, URL, etc.>", and these sentences may be located at different positions within the manifest or log file 200. For example, a set of keywords may be found in the instruction sentences or log messages related to the assignment of network addresses to the referenced components and the configuration of components configured to communicate with the addresses of the referenced components. These keywords can be identified, and the code lines or log messages containing these keywords can be processed to identify the network addresses and identifiers of the configured components and the referenced components, that is, the referenced component is the component to which the network address is assigned, and the configured component is the component configured to use its network address to send data to and / or receive data from the referenced component. Then, referring to the identifiers of the configured components and the referenced components, a network relationship 704 may be created, optionally including the network address. Additional information can include the protocol used, port numbers, network relationships (e.g., whether the referenced component functions as a network gateway, proxy, etc.).

[0095] A statement in the log file 200 may be configured to change the configuration of a configured component such that the configured component uses a different referenced component's network address. Such a statement may be parsed, and a new network relationship 704 may be created in a manner similar to that described above. A previously created network relationship 704 for the configured component may be deleted or may continue to exist. For example, there may be an explicit instruction to delete the configured component's configuration to use the referenced component's network address referenced by the previously created network relationship 704. In response to recording the execution of such an instruction, the previously created network relationship 704 may be deleted.

[0096] FIG. 10 illustrates a method 1000 for extracting session relationships 706. Method 1000 may be performed by log processor 130, orchestrator 106, and / or some other component. Method 1000 includes retrieving 1002 a log file 200. Retrieving 1002 a log file 200 may include pulling the log file 200 from a component as described above with respect to FIG. 2. Retrieving 1002 may include an enrichment step in which data from log file 200 is processed by vector log agent 126 to add additional information, perform a map-reduce operation, or perform other operations. For example, enrichment may include adding an identifier of the source of log file 200, the directory location of log file 200, or other data to facilitate associating the data in log file 200 with a particular component identifier. Retrieving 1002 may include processing log file 200 according to one or more VRL statements.

[0097] Method 1000 may include a step 1004 of obtaining a session setup message from log file 200 before or after any enrichment step of log file 200. The session setup message may be a message indicating that a session has been successfully initiated and may include identifiers of the server component (i.e., the component providing the service) and the client component (i.e., the component requesting the service).

[0098] Method 1000 may include step 1006 of retrieving a session termination message from log file 200 before or after any enrichment step of log file 200. The session termination message may be a message indicating that a session has terminated in response to either an instruction from a client component, an instruction from a server component, expiration of a timeout period, failure of an intermediate component or network connection between the client and server components, restart or failure of the client or server component, or other cause. The session termination message may also include identifiers of the server component (i.e., the component providing the service) and the client component (i.e., the component requesting the service). If the session termination is due to a failure (network connection, intermediate component, client component, or server component), only the server component or the client component may be referenced by the log message. In such a case, all session relationships referencing the components referenced in the log message may be considered terminated and deleted.

[0099] The method 1000 may include a step 1008 of updating the session relationship 706 by adding a session relationship 706 corresponding to the session identified in the setup message to be created. The session relationship 706 may include identifiers of the server and client components, and may include other information such as a timestamp from the setup message, an identifier for the session itself, the type of session, or other data.

[0100] Updating 1008 the session relationship 706 may include deleting the session relationship 706 corresponding to a session identified as terminated in a session close message (including a message indicating a failure). For example, if the session has a unique session identifier, the session relationship 706 including the session identifier included in the session close message may be deleted. Alternatively, if the session close message references a set of a client identifier and a server component identifier, the session relationship 706 including the same client identifier and server component identifier may be deleted. In some implementations where the session has a known time to live (TTL), the session relationship 706 may be deleted based on the expiration of the TTL, regardless of whether a session close message corresponding to the session relationship has been received.

[0101] FIG. 11 illustrates a method 1100 for extracting access relationships 708. Method 1100 may be performed by log processor 130, orchestrator 106, and / or some other component. Method 1100 includes retrieving 1102 a log file 200. Retrieving 1102 a log file 200 may include pulling the log file 200 from a component as described above with respect to FIG. 2. Retrieving 1102 may include an enrichment step in which data from log file 200 is processed by vector log agent 126 to add additional information, perform a map-reduce operation, or perform other operations. For example, enrichment may include adding an identifier of the source of log file 200, the directory location of log file 200, or other data to facilitate associating the data in log file 200 with a particular component identifier. Retrieving 1102 may include processing log file 200 according to one or more VRL statements.

[0102] Method 1100 may include step 1104 of extracting access relationships 708 from log file 200 either before or after enriching log file 200. Access relationships 708 may be identified in various ways, such as by parsing log messages in log file 200 of a server component (i.e., a component that provides a service) indicating requests from a client component (i.e., a component that requests a service), log messages in log file 200 of a client component indicating requests from a client component to a server component, log messages of another component that stores results of access requests from the client component to the server component, etc. Access relationships 708 may include an identifier of the server component, an identifier of the client component, and one or more timestamps or other metadata for one or both of (a) each request from the client component to the server component and (b) each response from the server component to the client component.

[0103] The method 1100 may include identifying 1106 an expired access relationship 708. An expired access relationship 708 may be defined as one whose most recent timestamp (for a request and / or response) is older than a threshold time, e.g., 1 minute, 5 minutes, 1 hour, 1 day, etc. The threshold time may be unique for each type of component; for example, an instance 118 of one application may have a different threshold than an instance of another application. The threshold time may be automatically derived as a multiple of the average time between requests for each client of the server component.

[0104] Method 1100 may then include step 1108 of updating access relationships 708 to add the access relationships detected in step 1104. Step 1108 of updating access relationships may include deleting expired access relationships. Step 1108 of updating access relationships may include step 708 of merging access relationships. For example, if a pair of access relationships 708 reference the same server and client component identifiers, the access relationships 708 may be combined into a single access relationship 708 that includes the most recent timestamps of the pair of access relationships 708. Access relationships 708 may include records of access requests and / or responses between the client and server component, and the records of the pair of access requests are combined during merging. Alternatively, each access relationship 708 may include statistical characteristics of past requests and / or responses, and the merged access request includes a combination of the statistical characteristics of the pair of access relationships 708. In one embodiment, merging is performed before step 1106 of identifying expired relationships.

[0105] FIG. 12 illustrates a method 1200 for extracting network connection relationships 710. Method 1200 may be performed by log processor 130, orchestrator 106, and / or some other component. Method 1200 includes retrieving 1102 a log file 200. Retrieving 1202 a log file 200 may include pulling the log file 200 from a component as described above with respect to FIG. 2. Retrieving 1202 may include an enrichment step in which data from log file 200 is processed by vector log agent 126 to add additional information, perform a map-reduce operation, or perform other operations. For example, enrichment may include adding an identifier of the source of log file 200, the directory location of log file 200, or other data to facilitate associating the data in log file 200 with a particular component identifier. Retrieving 1202 may include processing log file 200 according to one or more VRL statements.

[0106] Method 1200 may include a step 1204 of obtaining a connection setup message from log file 200 before or after any enrichment step of log file 200. The session setup message may be a record of an exchange of handshake messages or other messages indicating that a network connection has been successfully established between a first component and a second component.

[0107] Method 1200 may include a step 1206 of retrieving a connection termination message from log file 200 before or after any enrichment step of log file 200. The connection termination message may be a message indicating that the network connection has been terminated in response to either an instruction from a client component, an instruction from a server component, expiration of a timeout period, failure of an intermediate component, or network connection between the client and server components, or other cause. In one embodiment, the session termination message may include a message indicating failure of a physical link between a first component and a second component, a restart of the first component or the second component, and a failure or restart of a component hosting the first component or the second component.

[0108] Method 1000 may include identifying 1208 an expired network connection relationship 710. Identifying 1208 an expired network connection relationship 710 may include identifying (a) a pair of components for which a network connection relationship 710 exists, (b) a pair of components for which there is no current network connection as indicated by a connection termination message, and (c) a pair of components for which a predetermined period of time has expired since the last connection termination message was received for the pair of components. With regard to (c), some connections have a predetermined TTL, and a network connection relationship 710 expires when a predetermined period of time greater than the TTL has expired since the last connection setup message for the pair of components.

[0109] As an alternative to the above approach, all network connection relationships 710 expire as soon as the network connection represented by the network connection relationship 710 terminates due to TTL expiration or explicit termination as indicated in a connection termination message.

[0110] Method 1200 may include step 1210 of updating network connection relationships 710 by deleting expired network connection relationships 710 and adding new network connection relationships 710 indicated by the connection setup messages from step 1204. It is possible for a first component and a second component to have multiple network connection relationships, such as connections to different ports by different applications. Thus, a separate network connection relationship 710 may exist for each network connection, or a single network connection relationship 710 may be created to represent all network connections between a pair of components. Network connection relationships 710 may include data describing each connection (such as a setup timestamp, protocol, and port). This data may be updated to remove data describing a connection when that connection is terminated. Similarly, network connection relationships 710 may be updated to add data describing a connection between a pair of components represented by network connection relationship 710 when the connection is set up.

[0111] Referring to Figures 13 and 14A-14D, the illustrated method 1300 can be used to generate a visual representation of a topology that is displayed on a display device, such as a user device 110, via an orchestrator dashboard 108.

[0112] 13 and 14A , method 1300 may be performed by orchestrator 106, and visual representation 1400 may be provided to user computing device 110 via orchestrator dashboard 108. User computing device 110 may then display visual representation 1400, receive user interactions with visual representation 1400, and report the user interactions to orchestrator 106 for processing. A user may request generation of visual representation 1400 via orchestrator dashboard 108. The retrieval and processing of provisioning data 300 and log files 200 to generate the visual representation may be performed in response to a request from a user.

[0113] The method 1300 may include step 1302 of extracting component identifiers from the provisioning data, as described above. Each component identifier is then used as a node in the graph. The method 1300 may then include step 1304 of adding edges between nodes for hosting relationships 700 between the component identifiers represented by the nodes. The method 1300 may include step 1306 of adding edges between nodes for environmental variable relationships 702 between the component identifiers represented by the nodes. The method 1300 may include step 1308 of adding edges between nodes for network relationships 704 between the component identifiers represented by the nodes. The method 1300 may include step 1308 of adding edges between nodes for network relationships 704 between the component identifiers represented by the nodes. The method 1300 may include step 1310 of adding edges between nodes for session relationships 706 between the component identifiers represented by the nodes. The method 1300 may include step 1312 of adding edges between nodes for access relationships 708 between the component identifiers represented by the nodes. Method 1300 may include adding 1314 edges between nodes for network connection relationships 710 between component identifiers represented by the nodes. The relationships between components described herein are exemplary only, and method 1300 may include adding edges for other types of relationships between components.

[0114] A visual representation 1400 of the topology represented by the graph may then be displayed 1316. An example visual representation 1400 is shown in FIG. 14. Graphical elements may be displayed to represent components such as hosts 802, pods 112, containers 114, storage volumes 116, and application instances 118. The graphical elements may include images and / or text, such as the UUID of each component.

[0115] The visual representation 1400 may include lines 1402 between the graphical elements representing the components, where the lines 1402 represent the edges of the graph. The lines 1402 may be color-coded, with each color representing a relationship type 702-710. A pair of components may have multiple relationships, such as some or all of an environment variable relationship 702, a network relationship 704, a session relationship 706, an access relationship 708, and a network connection relationship. A separate line 1402 may be displayed to represent each type of relationship, or a single line may represent all of the relationships between the components represented by the pair of graphical elements.

[0116] The graphical elements or lines 1402 may be augmented with additional visual data describing the components or relationships represented by the graphical elements or lines 1402. For example, the additional visual data may be displayed upon clicking on the graphical elements or lines 1402, hovering over the graphical elements or lines 1402, or upon other interaction. The additional data may be collected from the log 200 and may include component usage and / or AAI data, as described above.

[0117] For example, for a graphical element representing a host 802, the additional data may include AAI data for the host, such as status 1404 (up, critical, down, unreachable, etc.), and available and / or in-use computing power 1406 (processor cores, processor time, processor cycles, etc.), available and / or in-use memory 1408, and available and / or in-use storage 1410. For a graphical element representing a storage volume 116, the additional data may include status 1412, available storage 1414 and / or storage usage, and IOP (input / output operations) usage 1416 and / or availability. For a graphical element representing a cluster 111, pod 112, container 114, or application instance 118, the additional data may include status 1418, computing power usage 1420, memory usage 1422, and storage usage 1424. For a cluster 111 and / or a pod 112, the computing power usage 1420, memory usage 1422, and storage usage 1424 may be an aggregate of the computing resources used by all containers 114, application instances 118, and storage volumes 116 managed by the cluster 111 and / or pod 112, as well as the cluster 111 and / or pod 112 itself.

[0118] For line 1402, the additional data may include data describing one or more relationships represented by line 1402, such as a list of each type of relationship 700-710 represented by the line, a status 1426 of each relationship, and a usage amount 1428 of each relationship. Relationship usage may include, for example, the amount of data sent over a network connection, the number or frequency of requests for a session or access relationship, the latency of the network connection, the latency of responses to requests for a session or access relationship, or other data.

[0119] The graphical elements or lines 1402 may also be expanded with an action menu 1430, such as in response to user interaction with the graphical elements or lines 1402. The action menu 1430 may include graphical elements that, when selected by a user, invoke one or both of: (a) an action that modifies information shown in the visual representation 1400; and (b) an action that performs an action with respect to the component represented by the graphical element or line 1402. For example, the action menu 1430 may include elements for deleting the component, restarting the component, creating a relationship 702-710 between the component and another component, creating a snapshot or backup copy of the component, duplicating the component, duplicating the component, or invoking other actions. Thus, the method 1300 may include a step 1318 of receiving an interaction with the visual representation 1400 of the topology and, in response, performing an action, such as a step 1320 of modifying the information displayed in the visual representation and / or modifying the component represented by the visual representation of the topology. Actions invoked with respect to a component may be performed with respect to other components, such as those hosted by the component. For example, an action invoked with respect to cluster 111 may be performed on all pods 112, containers 114, storage volumes 116, and application instances 118 hosted by cluster 111.

[0120] 14B shows an application browsing interface 1432 that may be displayed to a user, such as using data obtained according to method 1300 or some other technique. The application browsing interface 1432 may include one or more cluster elements 1434 representing the cluster 111. The user may select one of the cluster elements 1434 to invoke a display of additional information about the cluster 111. For example, a display of one or more namespace elements 1436, such as a list of names in the namespace of the cluster 111, each name representing a component 112, 114, 116, 118 of the cluster 111 or another variable, service, or other entity accessible to the component of the cluster 111. The interface 1432 may display a selector element 1438 that allows the user to enter criteria for filtering or selecting names from the namespace of the cluster 111. For example, the user can select based on version (e.g., which HELM release of Kubernetes the component belongs to or was deployed by), type of application (database, web server, etc.), executable image, instantiation data, or any other criteria.

[0121] For each application instance 118 that meets the criteria entered by the user into selector element 1438, application viewing interface 1432 can display various items of information for the application instance 118. Exemplary information items may include daemon set 1440a, deployment data 1440b, stateful set 1440c, replica set 1440d, config map 1440e, one or more secrets 1440f, or other data 1440g. Some or all of the items may be selected by the user to invoke the display of additional data. For example, the user may invoke the display of pod data 1442 for the pod 112 that hosts the application instance 118, container data 1444 describing the container 114 that hosts the application instance 118, persistent volume claim (PVC) data 1446 for the storage volume 116 accessed by the application instance 118, and volume data 1448 describing the storage volume 116 accessed by the application instance 118.

[0122] For each element selectable in the application viewing interface 1432, selecting the element can invoke a display of elements associated with that element, and can also invoke a display of real-time data for each element, such as any of the observability data for each element (e.g., log data 200) that may be collected, processed (aggregated, formatted, etc.), and displayed as observability data is generated for each element.

[0123] 14C shows yet another interface 1450 that may be used to visually represent the topology and receive user input to invoke a display of additional information about a cluster 111, a host 1452 running one or more components of the cluster 111, and a storage device 1454 of one of the hosts 1452. The interface 1450 may include a cluster element 1456 that represents the cluster 111, a namespace element 1458 that represents the namespace of the cluster 111, a composite application element 1460 that represents two or more application instances 118 that together define a bundled application, and a single application element 1462 that represents the single application instance 118.

[0124] Selecting a given element 1456, 1458, 1460, 1462 can invoke a display of additional information; selecting cluster element 1456 can invoke a display of namespace elements 1458, selecting a name from namespace elements 1458 can invoke a display of composite application elements 1460, and selecting a name from composite application elements 1460 can invoke a display of single application elements 1462.

[0125] Selecting a single application element 1462 can invoke a display of data describing the application instance 118 represented by the single application instance 118. For example, the data can include other data such as element 1464 indicating config map data, element 1466 indicating various sets (replica set, deployment set, stateful set, daemon set, etc.), element 1468 indicating secrets, or any observability data for the application instance 118.

[0126] Selecting elements 1462, 1464, 1466 can invoke the display of additional data, such as a pod element 1470 containing data describing the pod 112, a PVC element 1480 describing the PVC, and a volume element 1482 describing the storage volume 116 (such as data describing the amount of data used by the storage volume 116 and the storage device that stores the data for the storage volume 116).

[0127] Interface 1450 can be used to assess the criticality of components of cluster 111. For example, selecting namespace element 1458 can invoke a display of aggregated data 1484, such as aggregated logs (e.g., log files combined by chronologically ordering the messages in the log files), aggregated metrics (aggregated processor usage, memory usage, storage usage), aggregated alerts and / or events (e.g., events and / or alerts combined and ordered by time of occurrence), aggregated access logs (e.g., allowing for tracking of user behavior with respect to cluster 111 or components of cluster 111), etc. The aggregated data 1484 can be used in combination with topology data, as described in U.S. Patent Application No. 16 / 561,994, filed September 5, 2019, and entitled "PERFORMING ROOT CAUSE ANALYSIS IN A MULTI-ROLE APPLICATION," which is incorporated herein by reference in its entirety, to perform root cause analysis (RCA).

[0128] Selecting a single application element 1462 can invoke a display of importance 1486 of the application instance 118 represented by the single application element 1462. The importance 1486 can be a metric that is a function of the number of other application instances 118 that depend on the application instance 118, for example, having relationships 700-710 with the application instance 118. The importance 1486 can include a "blast radius" of the application instance 118 (see FIG. 14D and corresponding discussion).

[0129] Selecting a pod element 1470 can invoke a display of the pod density 1488 (e.g., number of pods) of the host running the pod 112 represented by the pod element 1470. The pod density 1488 can be used to determine the importance of the host and whether the host may be overloaded.

[0130] Selecting the PVC element 1480 can invoke a display of the volume density 1490 (e.g., number of storage volumes 116, total size of storage volumes 116) stored on the storage device or individual storage devices of the host. The volume density 1490 may be used to determine the criticality of the host and whether the host's storage devices may be overloaded.

[0131] FIG. 14D illustrates yet another interface 1492 that can be used to visually represent the topology. Interface 1492 can include visual representations of the illustrated components. Storage devices 1494 (e.g., hard disk drives, solid-state drives) store data for storage volumes 116 and are used by application instances 118, which can have one or more relationships, e.g., relationships 700-710, with other application instances 118, which themselves have relationships 700-710 with other application instances. In particular, one or more application instances 118 that are not running on the same host as the storage volume may be represented in interface 1492. Interface 1492 can also be a “blast radius” representation that indicates the impact that a failure of storage device 1494 would have on other application instances 118 or other components of cluster 111 containing storage volume 116 or one or more other clusters 111.

[0132] 15-19, using AAI, computing resources allocated to components in network environment 100 may be reduced based on the usage of computing resources by application instance 118. Cloud computing platform 104 may charge for computing resources purchased regardless of actual usage. Thus, AAI may be used to identify changes to the deployment of application instances to reduce computing resources purchased.

[0133] 15, the orchestrator 106 or another component may calculate a cluster host inventory 1502a-1502c for each of the multiple clusters 111a-111c. The cluster host inventory 1502a-1502c is the number of processing cores, amount of memory, and amount of storage on the servers 102 allocated to the particular cluster 111a-111c. In the case of a cloud computing platform 104, the cluster host inventory 1502a-1502c may include the cloud computing platform's computing power, memory, and amount of storage allocated to the cluster 111a-111c.

[0134] The orchestrator 106 or another component may further calculate cluster provisioning 1504a-1504c for each cluster 111a-111c. The cluster provisioning 1504a-1504c is the computing resources (computing power, memory, and / or storage) allocated to components (e.g., pods 112a-112c, containers 114, storage volumes 116, or application instances 118a-118l) within the clusters 111a-111c. In some cases, the cluster provisioning 1504a-1504c is identical to the cluster host inventory 1502a-1502c and is omitted. In other examples, the cluster provisioning 1504a-1504c includes the computing resources allocated to individual components (pods 112a-112c, storage volumes 116, application instances 118a-118l) of the clusters 111a-111c.

[0135] The orchestrator 106 or another component may further calculate cluster usage 1506a-1506c for each cluster 111a-111c. The cluster usage 1506a-1506c for a cluster 111a-111c may include, for each computing resource (computing power, memory, storage), the total usage of that computing resource by all components within the cluster 111a-111c, including the cluster itself. The cluster usage 1506a-1506c may be obtained from the log file 200, as described above. The cluster usage 1506a-1506c for a cluster 111a-111c may include a list of the amount of each computing resource used by the individual components of the cluster 111a-111c and the cluster 111a-111c itself.

[0136] The orchestrator 106 or another component can further calculate a cluster AAI 1508a-1508c for each cluster 111a-111c. The cluster AAI 1508a-1508c can include the AAI(t), O(t,i), and O(t) calculated as above, except that the hardware inventory is limited to the cluster host inventory 1502a-1502c and only the usage of components within the cluster 111a-111c and the cluster itself is used in the calculation.

[0137] FIG. 16A is a simplified diagram of available computing resources and their usage. Each bar in FIG. 16A represents either the amount of computing resources (hardware inventory 1502a-c, cluster AAI 1508a-c) or the usage of computing resources (application instance 118a-c). The illustrated representation is simplified in that only one computing resource is shown, omitting other usages (pods 112a-c, storage volume 116, clusters 111a-c themselves), although these usages and computing resources may actually be included. As can be seen, each cluster has a cluster AAI 1508a-c of computing resources that represents the difference between the cluster host inventory 1502a-c and the usage by the various components of each cluster 111a-c.

[0138] Continuing with reference to FIG. 16A, and with reference to FIG. 16B, one or more components may be redeployed from one cluster 111a-111c to another cluster. For example, application 118d on cluster 111a is consuming significantly more computing resources than the other applications 118a-118c on cluster 111a. In contrast, cluster 111b has cluster AAI 1508b with sufficient computing resources to host application 118d. Thus, application 118d may be redeployed on cluster 111b.

[0139] In a cloud computing environment 104 in which computing resources are virtualized, the amount of cluster host inventory 1502a-1502c for some or all of the clusters 111a-111c can be reduced, thereby reducing the amount charged for the cluster host inventory 1502a-1502c. In particular, because the usage of the cluster host inventory 1502a is significantly reduced by removing the usage of the application instance 118d, significant cost savings can be achieved by reducing the cluster host inventory 1502a.

[0140] Redeployment of application instance 118d to another cluster 111b may be contingent on satisfying one or more constraints. Failure to satisfy a constraint may prevent redeployment. For example, there may be a requirement that the receiving cluster 111b have a sufficient amount of computing resources (computing power, memory, and storage) to receive application instance 118d. There may be a requirement that moving application instance 118d to cluster 111b does not violate any affinity requirements with respect to application instances 118a-118c remaining on the original cluster 111a. There may be a constraint that moving application instance 118d to cluster 111b does not violate any anti-affinity requirements with respect to application instances 118e-118h running on the receiving cluster 111b. Redeployment of application instance 118d to receiving cluster 111b may include adding application instance 118d to pods 112c, 112d of receiving cluster 111b or creating a new pod on receiving cluster 111b.

[0141] Redeployment of an application instance 118, such as application instance 118d in the illustrated example, may include redeploying the application instance 118 from server 102 to cloud computing platform 104, or vice versa. For example, application instance 118d may be hosted on cloud computing platform 104. Application instance 118d may be moved to server 102 because it is using an amount of computing resources above a threshold and would have better performance if hosted locally on server 102 and would be less costly if billing from cloud computing platform 104 for application instance 118d were eliminated. Similarly, an application instance 118 with usage below a minimum threshold may be moved from server 102 to the cloud to provide local computing resources on server 102 to an application instance on cloud computing platform 104 with usage above a maximum threshold.

[0142] 16C , in another example, cluster consolidation may be performed by moving all application instances 118a-118d, which may be deployed to one or more other clusters 111b, 111c, subject to any affinity and anti-affinity constraints and provided the other clusters 111b, 111c have sufficient cluster AAIs 1508b, 1508c. In that case, the entire cluster host inventory 1502a may be deleted, along with the corresponding costs of the cluster host inventory 1502a.

[0143] 17A illustrates an example method 1700a that may be performed by the orchestrator 106 or another component to redeploy an application instance 118 to a different cluster 111. To facilitate understanding of the method, reference is made to the components illustrated in FIG. 15 as a non-limiting example. In particular, any number of clusters 111 hosting any number of components may be processed according to method 1700a.

[0144] The method 1700a may include determining 1702 usage and cluster AAIs for each cluster 111, such as component usage 1506a-1506c for the plurality of clusters 111a-111c and cluster AAIs 1508a-1508c for the plurality of clusters 111a-111c. The method 1700a may include identifying 1704 candidate redeployment. The step 1704 identifying candidate redeployment may be limited to evaluating usage of the application instance 118 with respect to the cluster AAI of the cluster 111 to determine whether redeployment is possible. The candidate redeployment may include transferring a particular application instance 118 (e.g., application instance 118d) to a receiving cluster 111 (e.g., cluster 111b) that has a sufficient cluster AAI to receive the application instance 118. A candidate redeployment may include replacing a first application instance 118 on a first cluster 111 with a second application instance 118 on a second cluster 111, where the second cluster has a larger host AAI than the first cluster and the first application instance 118 has a larger usage than the second application instance 118. A candidate redeployment may include removing the first application instance 118 on the first cluster 111, where the second application instance 118 on the second cluster 111 is in a load balancing relationship with the second application instance 118 and the second cluster 111 has a cluster AAI sufficient to receive the usage of the first application instance 118, and possibly a larger cluster AAI than the first cluster 111. When identifying candidate redeployment instances 118 (1704), multiple application instances 118 of a cluster 111 that have affinity constraints with each other may be treated as a unit, i.e., the receiving cluster 111 must have sufficient cluster AAIs to receive all of the multiple application instances 118.

[0145] Method 1700a may include step 1706 of filtering candidate redeployment candidates based on constraints, such as anti-affinity requirements, latency requirements, or other requirements. For example, if redeploying application instance 118d to cluster 111b would violate the anti-affinity constraint of application instance 118d with respect to application instance 118e, then such redeployment of application instance 118d is filtered out in step 1706. Similarly, if redeploying application instance 118d to cluster 111b would exceed the minimum latency required for application instance 118d with respect to application instances 118i-118l in cluster 111c, then such redeployment is filtered out in step 1706. The anti-affinity and latency requirements are exemplary only, and other constraints may be imposed in step 1706.

[0146] The method 1700a may include calculating 1708 a billing reduction achievable by the candidate redeployment, i.e., the amount by which the candidate redeployment would reduce the cluster host inventory 1502a-1502c of the modified cluster if the candidate redeployment were performed. If the billing reduction is found to be greater than a minimum threshold, the candidate redeployment is implemented by performing a transfer, replacement, or deletion of the candidate redeployment (1712). A redeployment involving moving an application instance 118 from a first cluster 111 to a second cluster 111 may include installing the new application instance 118 on the second cluster (creating a container and installing the application instance 118 in the container), stopping the original application instance 118 on the first cluster 111, and starting the new application instance 118 running on the second cluster 111. Other configuration changes may be required to configure other components to access the new application instance 118 on the second cluster 111.

[0147] Method 1700a may further include step 1714 of reducing an amount of cloud computing resources used by one or more clusters 111. For example, in the example of FIG. 16B , the computing resources allocated to cluster 111a may be reduced after application instance 118d is redeployed to cluster 111b. The reduction amount may be such that the cluster AAI of each cluster 111 is lowered to a zero or non-zero threshold (e.g., a percentage of usage of the components deployed to each cluster) for one or more computing resources (computing power, memory, storage), assuming that the usage of the cluster's components after the redeployment remains the same as the usage values ​​used to calculate the cluster AAI of cluster 111.

[0148] 17B illustrates an alternative method 1700b for redeploying an application instance 118. The method 1700b may be performed by the orchestrator 106 or other component to redeploy the application instance 118 to a different cluster 111.

[0149] The method 1700a may include determining 1702 usage amounts and cluster AAIs for each cluster 111, such as component usage amounts 1506a-1506c for the plurality of clusters 111a-111c and cluster AAIs 1508a-1508c for the plurality of clusters 111a-111c.

[0150] Method 1700a may include step 1704 of re-planning the placement of components using the components' computing resource usage instead of provisioning requirements. When initially instantiating components 111, 112, 114, 116, 118 in network environment 100, orchestrator 106 may perform a planning process to place the components based on required computing resources, affinity requirements, anti-affinity requirements, latency requirements, or other requirements. Orchestrator 106 further attempts to improve the performance of components working together by reducing latency and using computing resources as efficiently as possible.

[0151] As an example, the orchestrator 106 may use a planning algorithm such as that disclosed in U.S. Patent No. 10,817,380 B2, filed October 27, 2020, entitled "IMPLEMENTING AFFINITY AND ANTI-AFFINITY CONSTRAINTS IN A BUNDLED APPLICATION," which is incorporated herein by reference in its entirety. In contrast to the initial plan, the provisioning requirements in step 1716 for each component may be set to be the compute resource usage measured for each component as described above using log data pulled from the component's host. Alternatively, the provisioning requirements may be set to an intermediate value between the component's provisioning defined by the manifest and the measured usage for that component, such as usage scaled by a number greater than 1, such as a number between 1.1 and 2.

[0152] The result of step 1716 may be one or more plans that define where each component is to be placed (e.g., on which server 102 or which unit of computing resources of the cloud computing platform, which pod 112, which cluster 111, etc.). The billing savings achieved by each plan may be calculated (1708) and evaluated (1710) to determine whether the plan provides at least a threshold reduction in computing resource allocation over the current configuration of the components based on the usage of each component measured in step 1702. As discussed above, reducing the allocation of computing resources results in a reduction in the costs of the cloud computing platform 104.

[0153] If so, one of the plans, such as the plan that provides the greatest cost savings, may be implemented 1712. Implementing the plan 1712 may include moving the components one at a time to the locations defined in the plan to avoid disruptions, or pausing all components, redeploying the components as defined in the plan, and restarting all components. Redeploying each component may be performed as described above with respect to step 1712 of method 1700a.

[0154] Following or during the redeployment performance 1712, method 1700b may include reducing allocated cloud computing resources 1714 from the cloud computing platform. The reduction may be such that the cluster AAI of each cluster 111 is lowered to a zero or non-zero threshold (e.g., a percentage of usage of each cluster after the redeployment) for one or more computing resources (computing power, memory, storage), assuming that usage of the cluster's components after the redeployment remains the same as the usage values ​​used to calculate the cluster AAI of the cluster 111.

[0155] 18 illustrates an alternative method 1800 for redeploying application instances 118 to consolidate the number of clusters 111 of an original configuration, such as the illustrated reduction of clusters shown in FIGS. 16A and 16C. Method 1700b may be performed by orchestrator 106 or another component.

[0156] The method 1800 may include determining 1802 usage and cluster AAIs for each cluster 111 in the original configuration, such as component usages 1506a-1506c for the plurality of clusters 111a-111c and cluster AAIs 1508a-1508c for the plurality of clusters 111a-111c.

[0157] Method 1800 may include attempting to identify consolidations 1804. Consolidation is the placing of components of multiple clusters into a subset of the multiple clusters, where one or more clusters of the multiple clusters and one or more hosts of the multiple clusters are excluded. A method for attempting to identify consolidations is described below with respect to FIG. 19.

[0158] If a consolidation is found (1806), the consolidation can be implemented (1808). If multiple consolidations are found, the consolidation that achieves the highest cost savings can be implemented (1808). The consolidation can include a plan that defines the location of each component on the surviving cluster 111. Thus, the components may be re-instantiated, configured, and started on the surviving cluster. In one embodiment, only components that are in a different location in the plan relative to their original configuration are redeployed to a different location. The original components may be shut down while the consolidation is implemented. Alternatively, the components may continue to operate and be migrated one at a time until the plan is implemented (1808).

[0159] Computing resources allocated to clusters that are removed as part of consolidation implementation 1808 may be reduced 1810. In the case of on-premise equipment, the servers 102 may be taken offline or allocated for other uses. Payment for units of cloud computing resources on the cloud computing platform 104 for use of one or more units of cloud computing resources allocated to the removed cluster may be terminated, or other action may be taken to terminate acquisition of one or more units of cloud computing resources.

[0160] FIG. 19 illustrates a method 1900 that may be used to identify potential cluster consolidations. Method 1900 may be performed by orchestrator 106 or another component. Method 1900 may include treating each cluster 111 as a “target cluster” (1902) and replanning 1904 excluding the target cluster 111, i.e., without the cluster host inventory currently assigned to the target cluster 111. The replanning may be performed with respect to the cluster host inventory of clusters 111 other than the target cluster 111 (“surviving clusters”), as described above with respect to step 1716 of method 1700b. As described above, the replanning may include identifying a location for each component on hosts of the surviving clusters using a planning algorithm such as that disclosed in U.S. Pat. No. 10,817,380 B2, where each component is allocated computing resources at least as large as its usage, and where the location of each component satisfies any affinity, anti-affinity, latency, or other requirements with respect to the locations of other components.

[0161] If it is found that no plan exists to eliminate the target cluster 111 (1906), the method 1900 ends with respect to the target cluster 111. If one or more plans are found to exist, each plan is added to the set of candidate mergers (1908).

[0162] After processing each cluster 111 as a target cluster, if one or more plans are found to eliminate the target cluster, method 1900 may be recursively repeated using the set of clusters 111 excluding the target cluster. For example, assume there are clusters 111a-111f, and a plan is found to eliminate the cluster host inventory of cluster 111a. Method 1900 may be repeated to determine whether the cluster host inventory of any of clusters 111b-111f can be eliminated. This process may be repeated until method 1900 identifies no more possible consolidations.

[0163] After treating each cluster 111 as a target cluster and performing any recursive iterations, the result is either no possible candidate consolidations or a set of one or more candidate consolidations. If there are multiple candidate consolidations, then in step 1808 the candidate consolidation that provides the greatest billing reduction may be selected for implementation.

[0164] Referring to FIG. 20, as noted throughout the above description, the topology is dynamic. The components of topology 2000 (clusters 111, pods 112, containers 114, storage volumes 116, and application instances 118) can change at any time. Sources of change include automatic scaling up or down of components based on usage by the orchestrator 106, such as using a tool like KUBERNETES. In particular, for each cluster 111, KUBERNETES, alone or in cooperation with the orchestrator 106, manages the scaling up or down of the number of pods 112 and corresponding containers 114, storage volumes 116, and application instances. An administrator can also manually add or remove components and relationships between components.

[0165] For example, as indicated by the dotted line representations, pod 112, container 114, storage volume 116, and application instance 118 can be added. Similarly, components and relationships marked with an "X" (represented by line 2002) represent components and relationships between components that can be removed from topology 2000.

[0166] In production facilities where stability is important, changes to topology 2000 may be prohibited or subject to one or more constraints to reduce the risk of changes that could cause crashes, overloads, or other types of instability.

[0167] 21 , for example, the illustrated method 2100 may be performed by the orchestrator 106 in cooperation with the orchestrator dashboard 108 or some other component. The method 2100 may include receiving 2102 a topology lock definition, such as from a user device 110 via the orchestrator dashboard 108. The topology lock definition may define the scope of the topology lock, for example, the entire topology, a particular cluster 111 or set of clusters 111, a particular host or set of hosts (a server 102 or unit of computing resource on the cloud computing platform 104), hosts located in a particular geographic area or facility, a particular area of ​​the cloud computing platform 104, or other definition.

[0168] A topology lock definition can further include restrictions on a particular type of component (cluster 111, pod 112, container 114, storage volume 116, application instance 118) or a particular type of relationship. With respect to application instances 118, the restrictions may refer to instances of a particular executable or class of executables. Restrictions may specify that for a particular type of component, instance of a particular executable, or relationship of a particular type, (a) the number cannot change, (b) the number cannot increase, (c) the number cannot decrease, or (d) the number cannot increase faster than a predetermined rate or the number cannot decrease faster than a predetermined rate.

[0169] Method 2100 may include receiving 2104 a topology policy for each topology lock definition. The topology policy defines actions to be taken to either or both (a) prevent violations of the topology lock definition or (b) handle violations of the topology lock definition.

[0170] The method 2100 may include configuring 2106 some or all of the orchestrator 106, a workflow in the workflow repository 120, or other components to implement each topology lock definition and its corresponding topology policy.

[0171] For example, the use of a workflow to instantiate or de-instantiate (i.e., delete) a component of a certain type may be modified to reference a topology policy corresponding to a topology lock that references components of that type, such that instantiation or de-instantiation of components of that type, if required according to the corresponding policy, is not permitted to complete in violation of the topology lock. In another example, a workflow that violates a topology lock generates an alert.

[0172] In another example, a container 114 may be configured to reference a container network interface (CNI), a container runtime interface (CRI), or a container storage interface (CSI) that is invoked by the container 114 during instantiation and / or startup. Any of the CNI, CRI, and CSI may be an agent of the orchestrator and may be modified to respond to the instantiation of a container 114 hosting an application instance 118 modifying the topology lock, whereby (a) preventing the instantiation if required by the corresponding topology policy or (b) generating an alert.

[0173] The above examples are merely examples of how a topology lock can be enforced, and any other aspect of the instantiation or de-instantiation of a component may be modified to include evaluating whether the instantiation or de-instantiation violates the topology lock and implementing any action required by the corresponding topology policy.

[0174] 22 illustrates a method 2200 for preventing violations of a topology lock with a corresponding topology policy. Method 2200 may be performed by the orchestrator 106, a CRI, a CNI, a CSI, or other components. Method 2200 includes receiving 2202 a request to create a component. Note that a request to delete a component may be processed similarly.

[0175] The request may be evaluated with respect to the topology lock and corresponding policies (2204). For example, step 2204 may include evaluating whether the request is to create a component in a portion of the topology referenced by the topology lock (e.g., in a particular cluster 111, a particular set of servers 102, a particular region or data center, a particular region of a cloud computing platform, etc.) and whether the component is of a type of component referenced by the topology lock. Step 2204 may include evaluating whether the request to create or delete a component is a prohibited action of the topology lock. For example, if changes are not allowed, the request to create or delete a component is prohibited. If only decreases are prohibited, the request to create a component may be permitted. If rate-limited increases are allowed, step 2204 may include evaluating whether creating a component would exceed a rate limit. If the request is a request to delete a component and only increases are prohibited, the request to delete the component may be permitted. If rate-limited decreases are allowed, step 2204 may include evaluating whether deleting the component would exceed a rate limit.

[0176] If the creation or deletion request is found to be permitted 2206, the request is implemented 2210. If not permitted, the method 2200 may include blocking the implementation of the request. Blocking may include one or more of the following: · Finishing the workflow required to implement the request. Preventing a CNI, CRI, or CSI from completing the setup of a component that is being created or a container that hosts a component that is being created.

[0177] Note that creating or deleting relationships between components may be handled similarly. A request to create or delete a relationship may be evaluated (2204) with respect to one or more topology locks and, if not permitted according to the topology locks, may either be implemented (2210) or blocked (2208). Blocking may be implemented using a modified workflow, CNI, CRI, or CSI. Blocking may also be performed in other ways, such as blocking network traffic to set up a session relationship 706, an access relationship 708, or a network connection relationship 710.

[0178] 23 illustrates a method 2300 for handling topology locks and corresponding policies. Method 2300 may be performed by orchestrator 106 or other components. Method 2300 may be performed in addition to or as an alternative to method 2200. For example, a policy corresponding to a topology lock may specify that changes that violate the topology lock should be blocked so that method 2200 is performed. A policy corresponding to a topology lock may detect a violation of the topology lock after it has occurred and specify that an alert should be issued or the violation should be reverted so that method 2300 is performed.

[0179] Method 2300 may include step 2302 of generating a current topology of the installation, such as according to method 1300 of FIG. 13 or some other approach. Method 2300 may include step 2304 of comparing the current topology, either at the time of initial instantiation of the installation or at a time after the initial installation, with a previous topology of the installation at a previous time. For example, the previous topology may be the topology that existed at or before a first time when the topology lock was created, and the current topology is obtained from provisioning data 300 and / or log files 200 generated at a second time after the first time.

[0180] A topology lock may have a scope that is smaller than all of the entire topology (see the description of step 2102 of method 2100). Thus, portions of the current and previous topologies that correspond only to that scope may be compared in step 2304. A topology lock may be limited to components of a particular type, such that only components of the current topology that have a particular type are compared in step 2304. If the topology lock references a type of relationship, relationships of that type in the current and previous topologies may be compared.

[0181] Method 2300 may include step 2306 of evaluating whether the current topology violates one or more topology locks relative to a previous topology. For example, whether new components of a particular type have been added to part of an installation (e.g., a cluster 111, a server 102, a data center, a cloud computing area, etc.). For example, component identifiers for each component of each type referenced by the topology lock may be compiled for the current and previous topologies. Component identifiers for the current topology that are not included in the component identifiers for the previous topology may be identified. Similarly, component identifiers for the previous topology that are not present in the current topology may be identified if the topology lock prevents their removal.

[0182] If the topology lock references a relationship type, each relationship in the current topology is attempted to match a relationship in the previous topology, i.e., to see if it has the same component identifier and type as the relationship in the previous topology. Relationships that do not have a corresponding patch in the previous topology may be considered new. Similarly, relationships from the previous topology that lack a match in the current topology may be considered deleted. In step 2306, it may be determined whether the new or deleted relationship violates policy.

[0183] For a topology lock violated in step 2306, method 2300 may include step 2308 of evaluating a topology policy corresponding to the topology lock. Actions indicated in the topology policy may then be implemented. For example, if a policy is found to require a change that violates the topology lock to be reverted (2310), method 2300 may include invoking a workflow to revert the change (2314). The workflow may be a workflow for removing a component or relationship that violates the topology lock. Such a workflow may be the same as the workflow used to remove components or relationships of that type when scaling down due to lack of usage. The workflow may be a series of steps for removing a component or relationship in an orderly and non-disruptive manner, i.e., processing pending transactions and migrating workload to another component. If a component or relationship is deleted in violation of the topology lock, the workflow may reinstantiate the component or relationship. The workflow for reinstantiating a component or relationship may be the same as the workflow used to create an initial instance of a component or relationship of that type or to scale up the number of components or relationships of that type.

[0184] If indicated by the topology policy corresponding to the topology lock, method 2300 may include generating 2312 an alert. The alert may be directed to the user device 110 or the administrator's user account, the individual who invoked the change to the policy that violated the topology lock, or another user. The alert may convey information such as the topology lock that was violated, the number of components or relationships that violated the policy, a graphical representation of the change to the policy (see, e.g., the graphical representation in FIG. 20), or other data.

[0185] 24, application instances 118 can have various relationships to one another. As described herein, application instances 118 are categorized as either dot application instances 2400, triangle application instances 2402, line application instances 2404, and graph application instances.

[0186] A dot application instance 2400 is an application instance 118 that does not have relationships (e.g., relationships 700-710) with other application instances 118. For example, an application instance 2400 may be an instance of an application that provides a standalone service. A dot application instance 2400 may be an application instance that does not have a particular type of relationship to other application instances 118. For example, a dot application instance 2400 may lack a hosting relationship 700, an environment variable relationship 702, or a network relationship 704 with another application instance 118. In an embodiment, one or more of a session relationship 706, an access relationship 708, and a network connection relationship 710 may still exist with respect to the dot application instance 2400 and another application instance 118.

[0187] Triangle application instance 2402 includes at least three application instances 118 that all have relationships with one another, such as any of relationships 700-710. Although "triangle application instance" is used throughout, this term should be understood to include any number of application instances 118, where each application instance 118 is dependent on all the other application instances 118.

[0188] In one example of a triangle application instance 2402, the application instances 118 may be replicas of each other, with one of the application instances 118 being a primary replica that processes production requests, and two or more other application instances 118 being backup replicas that mirror the state of the primary replica. Thus, each change to the state of the primary replica must be propagated to and confirmed by each backup replica. Health checks may be performed by the backup replicas with respect to each other and with respect to the primary replica to determine whether a backup replica should become the primary replica. Thus, the above-described relationship between the primary replica and the backup replica becomes a triangle application instance 2402. In the illustrated example, each application instance 118 in the set of triangle application instances 2402 runs on a different cluster 111.

[0189] The line application instance 2404 includes multiple application instances 118 arranged in a pipeline such that an input to a first application instance results in a corresponding output that is received as an input to a second application instance, and similarly for any number of application instances. As an example, the application instances 118 of the line application instance 2404 may include a web server, a backend server, and a database server. A web request received by the web server may be translated by the web server into one or more requests to the backend server. The backend server may process the one or more requests, which may require one or more queries to the database server. Responses from the database server are processed by the backend server to obtain responses that are sent to the web server. The web server may then generate a web page that includes the response and send the web page as a response to the web request. In the illustrated example, each application instance 118 of the line application instance 2404 runs on a different cluster 111.

[0190] The graph application instance 2406 includes multiple application instances 118, including a line application instance 2404 and / or a triangle application instance 2402, connected by one or more relationships, such as one or more relationships 700-710. For example, the application instance 118 for a first line application instance 2404 can receive output from the application instance 118 for a second line application instance 2404, thereby creating a branch. Similarly, the application instances 118 for a first triangle application instance 2402 can generate output that is received by the application instance for the line application instance 2404 or the application instances 118 for another set of triangle application instances 2402. The application instances 118 for the first triangle application instance 2402 can receive output from the application instance for the line application instance 2404 or the application instances 118 for another set of triangle application instances 2402.

[0191] 25, a cluster 111 may have a corresponding cluster specification 2500. The cluster specification 2500 may be created before or after the creation of the cluster 111 and includes information useful for provisioning components (pods 112, containers 114, storage volumes 116, and / or application instances 118) on the cluster 111.

[0192] For example, cluster specification 2500 for cluster 111 may include an identifier 2502 for cluster 111 and a location identifier 2504. Location identifier 2504 may include one or both of a name assigned to the geographic area in which one or more hosts on which cluster 111 executes are located and data describing the geographic area in which one or more hosts are located, such as in the form of the name of a city, state, country, zip code, or some other political or geographic entity. Location identifier 2504 may include coordinates (latitude and longitude or Global Positioning System) describing the location of one or more hosts. If there are multiple geographically dispersed hosts, the location (political or geographic name and / or coordinates) of each host may be included in location identifier 2504.

[0193] The cluster specification 2500 may include a list of computing resources 2506 for one or more hosts. The computing resources may include the number of processing cores, the amount of memory, and the amount of storage available on the one or more hosts. For example, the computing resources may include a cluster host inventory for the cluster 111, as described above. If the cluster 111 already hosts one or more components, the computing resources 2506 may additionally or alternatively include a cluster AAI for the one or more hosts, as defined above.

[0194] 26, a dot application specification 2600 may include an identifier 2602 for an application instance 118 that is created in accordance with the dot application specification 2600. The dot application specification 2600 may include one or more runtime requirements 2604. For example, the runtime requirements 2604 may include a location requirement 2606. For example, the location requirement 2606 may include the name of a political or geographic entity within which a host executing the application instance 118 must be located. The location requirement 2606 may be specified in terms of coordinates and a radius around the coordinates within which a host executing the application instance 118 must be located.

[0195] Runtime requirements 2604 may further include availability requirements 2608. Availability requirements 2608 may be values ​​from a set of possible values ​​that indicate the required availability of application instance 118 of dot application specification 2600. For example, such values ​​may include “high availability,” “intermittent availability,” and “low availability.” Orchestrator 106 may then interpret availability requirements 2608 when selecting a host for application instance 118 and configuring application instance 118 on the selected host.

[0196] The runtime requirements 2604 may further include cost requirements 2610. The cost requirements 2610 may indicate the permitted cost for running the application instance 118 of the dot application specification 2600. For example, a cloud computing provider may charge for some or all of the computing power (e.g., processor cores), memory, and storage used by the application instance 118. Thus, the cost requirements 2610 may specify a maximum amount that can be spent running the application instance 118, such as the amount that can be spent per day, month, or other period.

[0197] The dot application specification 2600 may further include computing resource requirements 2612 that specify the amount of processing power, memory, and / or storage required to execute the application instance 118 of the dot application specification 2600. The computing resource requirements 2612 may be a static definition or may be dynamic, such as an annotation indicating that provisioning may be dynamically changed based on initial provisioning requirements and usage (e.g., as described above with reference to FIGS. 15-19).

[0198] The dot application specification 2600 may further include a tolerance 2614 that specifies whether exceptions to any of the above-mentioned requirements 2604, 2612 are allowed. For example, the tolerance 2614 may indicate that the application instance 118 of the dot application specification 2600 should not be deployed unless all of the requirements 2604, 2612 are met. The tolerance 2614 may indicate that if a cluster 111 that meets the requirements 2604, 2612 is not found, the application instance 118 may be deployed to the closest alternative (“best fit”). The tolerance may indicate an allowed deviation from any of the requirements 2604, 2612 if a cluster 111 that meets the requirements 2604, 2612 is not found.

[0199] The dot application specification 2600 defines the provisioning of the dot application specification's application instance 118. Other parameters that define the instantiation and configuration of the application instance 118 on the selected host may be included in the manifest ingested by the orchestrator 106 in addition to the dot application specification 2600. Alternatively, the dot application specification 2600 may be part of the manifest.

[0200] 27 , a triangle application specification 2700 may include an identifier 2702 for a set of application instances 118 to be created in accordance with the triangle application specification 2700. The triangle application specification 2700 may include one or more runtime requirements 2704. For example, the runtime requirements 2704 may include a location requirement 2706. For example, the location requirement 2706 may include the name of a political or geographic entity in which a host executing the set of application instances 118 must be located. The location requirement 2706 may be specified in terms of coordinates and a radius around the coordinates in which one or more hosts executing one or more application instances 118 of the hierarchy must be located. The location requirement 2706 may include an individual location for each application instance 118 in the set of application instances 118.

[0201] Runtime requirements 2704 may further include availability requirements 2708. Availability requirements 2708 may be values ​​from a set of possible values ​​that indicate the required availability for the set of application instances 118 of triangle application specification 2700. For example, such values ​​may include “high availability,” “intermittent availability,” and “low availability.” Orchestrator 106 may then interpret availability requirements 2708 when selecting hosts for the set of application instances 118 and configuring the set of application instances 118 on the selected hosts. Availability requirements 2708 may include individual availability requirements for each application instance 118 in the set of application instances 118.

[0202] The runtime requirements 2704 may further include cost requirements 2710. The cost requirements 2710 may indicate the permitted cost for executing the set of application instances 118 of the triangle application specification 2700. For example, a cloud computing provider may charge for some or all of the computing power (e.g., processor cores), memory, and storage used by each application instance 118 in the set of application instances 118. Thus, the cost requirements 2710 may specify a maximum amount that may be spent on executing the set of application instances 118, such as an amount that may be spent per day, month, or other period. The cost requirements 2710 may include individual cost requirements for each application instance 118 in the set of application instances 118.

[0203] Runtime requirements 2704 may further include latency requirements 2712. Because each application instance 118 in a set of application instances 118 depends on every other application instance in the set, proper functioning may require the latency to be below a specified maximum latency in terms of time, such as 10 ms, 20 ms, or some other time value. Latency requirements 2712 may be specified for each pair of application instances 118 in the set, i.e., the maximum latency allowed between the application instances 118 of each possible pair of application instances 118.

[0204] The triangle application specification 2700 may further include computing resource requirements 2714 that specify the amount of processing power, memory, and / or storage required to execute each application instance 118 in the set of application instances 118 of the triangle application specification 2700. The computing resource requirements 2714 may be a static definition or may be dynamic, for example, an annotation indicating that provisioning may be dynamically changed based on initial provisioning requirements and usage (e.g., as described above with reference to Figures 15-19).

[0205] The triangle application specification 2700 may further include a replication requirement 2716 that specifies the number of application instances 118 to be included in the set of application instances, e.g., a value greater than or equal to 3. Thus, if an application instance 118 fails, the orchestrator 106 creates a new application instance 118 to satisfy the replication requirement 2716.

[0206] The triangle application specification 2700 may further include tolerances 2718 that specify whether exceptions to any of the above-mentioned requirements 2704, 2714, 2716 are allowed. For example, the tolerances 2718 may indicate that the application instance 118 of the triangle application specification 2700 should not be deployed unless all of the requirements 2704, 2714, 2716 are met. The tolerances 2718 may indicate that if a cluster 111 that satisfies the requirements 2704, 2714, 2716 is not found, the application instance 118 may be deployed to the closest alternative (“best fit”). The tolerances may indicate allowed deviations from any of the requirements 2704, 2714, 2716 if a cluster 111 that satisfies the requirements 2704, 2714, 2716 is not found.

[0207] The triangle application specification 2700 defines the provisioning of a set of application instances 118. The instantiation and configuration of each application instance 118 on a selected host, as well as the creation of any relationships 700-710 between the application instances 118, may be performed according to a manifest captured by the orchestrator 106 in addition to the triangle application specification 2700. Alternatively, the triangle application specification 2700 may be part of the manifest.

[0208] 28, a line application specification 2800 can include multiple tier specifications 2802. Each tier specification 2802 corresponds to a different tier in the pipeline defined by the line application specification 2800. Each tier specification 2802 can include a specification of the type of application instance 118 to be instantiated for that tier. Each tier can include multiple application instances 118 of the same or different types.

[0209] Each tier specification 2802 may include identifiers 2804 of one or more application instances 118 to be created in accordance with the tier specification 2802. The tier specification 2802 may include one or more runtime requirements 2806. For example, the runtime requirements 2806 may include a location requirement 2808. For example, the location requirement 2808 may include the name of a political or geographic entity in which one or more hosts executing one or more application instances 118 of the tier must be located. The location requirement 2808 may be specified in terms of coordinates and a radius around the coordinates in which all hosts executing one or more application instances 118 of the tier must be located. The location requirement 2808 may include individual locations for one or more application instances 118 of the tier.

[0210] Runtime requirements 2806 may further include availability requirements 2810. Availability requirements 2810 may be a value from a set of possible values ​​that indicates the required availability for one or more application instances 118 of the tier. For example, such values ​​may include “high availability,” “intermittent availability,” and “low availability.” Orchestrator 106 may then interpret availability requirements 2810 when selecting one or more hosts for one or more application instances 118 of the tier and configuring the one or more application instances 118 on the selected hosts. Availability requirements 2810 may include individual availability requirements for each application instance 118 of the one or more application instances 118.

[0211] Runtime requirements 2806 may further include cost requirements 2812. Cost requirements 2812 may indicate the permitted cost for running one or more application instances 118 of a tier. For example, a cloud computing provider may charge for some or all of the computing power (e.g., processor cores), memory, and storage used by each application instance 118 of one or more application instances 118. Thus, cost requirements 2812 may specify a maximum amount that may be spent running one or more application instances 118 of a tier, such as an amount that may be spent per day, month, or other period. Cost requirements 2812 may include individual cost requirements for each application instance 118 of the one or more application instances 118 of a tier.

[0212] Runtime requirements 2806 may further include latency requirements 2814. Latency requirements 2814 may one or both of (a) defining a maximum tolerable latency between application instances of the same tier and (b) defining a maximum latency for application instances 118 of preceding and / or succeeding tiers.

[0213] The tier specification 2802 may further include computing resource requirements 2816 that specify the amount of processing power, memory, and / or storage required to run each of the one or more application instances 118 of the tier. The computing resource requirements 2816 may be a static definition or may be dynamic, such as an annotation indicating that provisioning may be dynamically changed based on initial provisioning requirements and usage (e.g., as described above with reference to FIGS. 15-19).

[0214] The tier specification 2802 may further include tolerances 2818 that specify whether exceptions to any of the above-mentioned requirements 2806, 2816 are allowed. For example, the tolerances 2818 may indicate that one or more application instances 118 of a tier should not be deployed unless all of the requirements 2806, 2816 are met. The tolerances 2818 may indicate that if a cluster 111 that meets the requirements 2806, 2816 is not found, then one or more application instances 118 may be deployed to the closest alternative (“best fit”). The tolerances may indicate allowed deviations from any of the requirements 2806, 2816 if a cluster 111 that meets the requirements 2806, 2816 is not found.

[0215] As described above, a graph application instance 2406 includes multiple application instances 118, including multiple line application instances 2404 and / or triangle application instances 2402. Thus, a specification of a graph application instance may include a collection of specifications 2700, 2800 of the line application instances 2404 and / or triangle application instances 2402 that make up the graph application instance.

[0216] 29 shows a method 2900 for deploying a dot application instance 2400. The method 2900 may be performed by the orchestrator 106. For example, the orchestrator 106 may invoke execution of a workflow from the workflow repository 120 by a worker 124 to perform some or part of the method 2900. The method 2900 may be performed in response to the orchestrator 106 receiving the dot application specification 2600 from a user or as part of a manifest.

[0217] The method 2900 may include determining 2902 computing resource requirements 2612 for the dot application instance 2400 and determining 2904 one or more runtime requirements 2604 for the dot application instance 2400. The method 2900 may then include evaluating 2906 the cluster specifications 2500 of the available clusters 111 to determine whether any of the clusters 111 have sufficient computing resources 2506 to meet the computing resource requirements 2612 and meet the runtime requirements 2604. As mentioned above, the available computing resources evaluated may be either the cluster host inventory of the cluster 111 or the cluster AAI of the cluster 111 on which one or more components are already running.

[0218] If one or more matching clusters are found in step 2906, method 2900 may include deploying the application instance 118 corresponding to the dot application instance 2400 to one of the one or more clusters. If multiple clusters are found in step 2906, one cluster 111 may be selected based on one or more criteria, such as geographic proximity, performance, available cluster inventory or cluster AAI, or other criteria.

[0219] If no matching clusters 111 are found in step 2906, method 2900 may include step 2910 of evaluating whether the dot application specification 2600 defines tolerances 2718. Step 2910 may further include evaluating whether any of the available clusters 111 are within defined tolerances for the computing resource requirements 2612 and / or runtime requirements 2604 of the dot application specification 2600. If the dot application specification 2600 does not provide tolerances or the clusters 111 are not within the defined tolerances, the operation fails 2914 and an error message may be returned to the user, the orchestrator 106, the log file 200, or other destination.

[0220] If the dot application specification 2600 provides tolerances and / or there are one or more clusters 111 that are within any defined tolerances, a compromise cluster 111 may be selected (2912). The compromise cluster 111 may be the cluster 111 that most closely matches one or both of the computing resource requirements 2612 and the runtime requirements 2604. For example, from among the clusters 111 having cluster host inventories and / or cluster AAIs that satisfy the computing resource requirements 2612, the cluster 111 that most closely satisfies the runtime requirements 2604 may be selected. For example, the runtime requirements 2604 may be ranked such that the cluster 111 that satisfies the highest-ranked runtime requirements 2604 is selected (2912). Once a compromise cluster is selected, the application instances 118 of the dot application instance 2400 are deployed on the compromise cluster (2908).

[0221] 30 shows a method 3000 for deploying a triangle application instance 2402. Method 3000 may be performed by orchestrator 106. For example, orchestrator 106 may invoke execution of a workflow from workflow repository 120 by worker 124 to perform some or part of method 3000. Method 3000 may be performed in response to orchestrator 106 receiving triangle application specification 2700 from a user or as part of a manifest.

[0222] The method 3000 may include step 3002 of determining computing resource requirements 2714 for the triangle application instance 2402 and step 3004 of determining one or more runtime requirements 2704 for the triangle application instance 2402. The method 3000 may then include step 3006 of evaluating the cluster specifications 2500 of the available clusters 111 to determine whether any of the clusters 111 have sufficient computing resources 2506 to meet the computing resource requirements 2714 and meet the runtime requirements 2704. The evaluation of step 3006 may be performed for each application instance 118 of the triangle application instance 2402 to identify, for each application instance 118, any clusters 111 that have sufficient computing resources 2506 and meet the runtime requirements 2704 for that application instance 118.

[0223] Any matching clusters 111 identified in step 3006 may then be further evaluated to determine 3008 the inter-cluster latency of the matching clusters 111. The inter-cluster latency may have been previously calculated and obtained, or may be tested as part of step 3008.

[0224] The method 3000 may then include evaluating 3010 whether any cluster group can be found among the matching clusters that meets the latency requirements 2712 of the triangle application instance 2402. For example, given that the application instances 118 of the triangle application instance 2402 are designated as A, B, and C, the matching cluster group may be cluster C that meets the computing resource requirements 2714 and runtime requirements 2704 of application instance A. A , Cluster C, which matches the compute resource requirements 2714 and runtime requirements 2704 of application instance B. B, and cluster C that matches the computing resource requirements 2714 and runtime requirements 2704 of application instance C. C and between each of these clusters (C A and C B Between C B and C C Between and C A and C C The latency between (and) meets latency requirement 2712.

[0225] If one or more matching cluster groups are found in step 3010, method 3000 may include step 3012 of deploying application instance 118 of triangle application instance 2402 on one cluster 111 of the one or more matching cluster groups. If multiple cluster groups are found in step 3010, one cluster group may be selected based on one or more criteria such as average inter-cluster latency, geographic proximity, performance, available cluster inventory or cluster AAI, or other criteria.

[0226] If step 3006 does not find a matching cluster or the number of matching clusters is less than the number needed to implement the triangle application instance 2402, the method 3000 may include step 3014 of evaluating whether the triangle application specification 2700 defines a tolerance range 2718. Step 3014 may further include evaluating whether any of the available clusters 111 are within a defined tolerance range for the computing resource requirements 2714 and / or runtime requirements 2704 of the triangle application specification 2700. If the triangle application specification 2700 does not provide a tolerance range or the clusters 111 are not within the defined tolerance range, the operation may fail (3018) and an error message may be returned to the user, the orchestrator 106, the log file 200, or other destination.

[0227] If the triangle application specification 2700 provides a tolerance range and / or there are clusters 111 that fall within any defined tolerance range, one or more compromise clusters 111 may be selected 3016. The compromise clusters 111 may be the clusters 111 that most closely match one or both of the computing resource requirements 2714 and the runtime requirements 2704. For example, from among the clusters 111 having cluster host inventories and / or cluster AAIs that satisfy the computing resource requirements 2714, the clusters 111 that most closely satisfy the runtime requirements 2704 may be selected. For example, the runtime requirements 2704 may be ranked such that the clusters 111 that satisfy the highest-ranked runtime requirements 2704 are selected 3016. Any compromise clusters selected in step 3016 may then be processed in step 3008, which may include processing the compromise clusters along with any matching clusters identified in step 3006.

[0228] If no matching cluster groups are found in step 3010, method 3000 may include step 3020 of evaluating whether the triangle application specification 2700 defines a tolerance 2718 for the latency requirement 2712. Step 3020 may further include evaluating whether any of the inter-cluster latencies of any of the non-matching cluster groups are within a defined tolerance for the latency requirement 2712. If the triangle application specification 2700 does not provide a tolerance or the cluster groups are not within a defined tolerance, the operation may fail 3018 and an error message may be returned to the user, the orchestrator 106, the log file 200, or other destination.

[0229] If the triangle application specification 2700 provides a tolerance range and / or there is at least one cluster group that is within any defined tolerance range, a compromise cluster group may be selected (3022) and the application instances 118 of the triangle application instance 2402 may be deployed on the clusters 111 of the selected compromise cluster group. The compromise cluster group may be the cluster group that most closely matches the latency requirement 2712. If one or more cluster groups include the compromise cluster selected in step 3016, selecting a compromise cluster group 3022 may also include evaluating the combination of inter-cluster latency for each cluster group and how closely each cluster in each cluster group meets the computing resource requirement 2714 and the runtime requirement 2704.

[0230] 31 shows a method 3100 for deploying a line application instance 2404. The method 3100 may be performed by the orchestrator 106. For example, the orchestrator 106 may invoke execution of a workflow from the workflow repository 120 by a worker 124 to perform some or part of the method 3100. The method 3100 may be performed in response to the orchestrator 106 receiving a line application specification 2800 from a user or as part of a manifest.

[0231] The method 3100 may include step 3102 of determining computing resource requirements 2816 for the line application instance 2404 and step 3104 of determining one or more runtime requirements 2806 for the line application instance 2404. The method 3100 may then include step 3006 of evaluating the cluster specifications 2500 of the available clusters 111 to determine whether any of the clusters 111 have sufficient computing resources 2506 to meet the computing resource requirements 2816 and meet the runtime requirements 2806. The evaluation of step 3106 may be performed for each application instance 118 of the line application instance 2404 to identify, for each application instance 118, any clusters 111 that have sufficient computing resources 2506 and meet the runtime requirements 2806 for that application instance 118.

[0232] The groups of matching clusters 111 identified in step 3106 may then be evaluated (3108) to determine a cost function for each group of matching clusters. The cost function for a group of clusters may include an evaluation of a monetary cost, such as the total monetary cost of deploying the application instances 118 of the line application instance 2404 on the clusters 111 of the group, or the monetary cost of deploying the most resource-intensive of the application instances 118 of the line application instance 2404. For example, the application instance 118 that hosts a database is the most resource-intensive in most applications; as a result, the cost function may be limited to evaluating the monetary cost of deploying the application instance 118 that hosts a database on a cluster 111 of a given cluster group that meets the computing resource requirements 2816 and one or more runtime requirements 2806 of the application instance 118 that hosts the database.

[0233] The method 3100 may then include evaluating 3110 whether there are any cluster groups that match the selection criteria. For example, the selection criteria may be a cost function of any cluster groups that is below a predetermined threshold.

[0234] If one or more matching cluster groups are found in step 3110, method 3100 may include step 3112 of deploying application instance 118 of line application instance 2404 on one cluster 111 of the one or more matching cluster groups. If multiple cluster groups are found in step 3110, one cluster group may be selected based on one or more criteria such as a cost function, average inter-cluster latency, geographic proximity, performance, available cluster inventory or cluster AAI, or other criteria.

[0235] If step 3106 does not find a matching cluster or the number of matching clusters is less than the number required to implement the line application instance 2404, the method 3100 may include step 3114 of evaluating whether the line application specification 2800 defines a tolerance range 2818. Step 3114 may further include evaluating whether any of the available clusters 111 are within a defined tolerance range for the computing resource requirements 2816 and / or runtime requirements 2806 of the line application specification 2800. If the line application specification 2800 does not provide a tolerance range or the cluster 111 is not within the defined tolerance range, the operation fails (3118) and an error message may be returned to the user, the orchestrator 106, the log file 200, or other destination.

[0236] If the line application specification 2800 provides tolerances and / or there are clusters 111 that are within any defined tolerances, one or more compromise clusters 111 may be selected (3116). The compromise clusters 111 may be the clusters 111 that most closely match one or both of the computing resource requirements 2816 and the runtime requirements 2806. For example, from among the clusters 111 having cluster host inventories and / or cluster AAIs that satisfy the computing resource requirements 2816, the clusters 111 that most closely satisfy the runtime requirements 2806 may be selected. For example, the runtime requirements 2806 may be ranked such that the clusters 111 that satisfy the highest-ranked runtime requirements 2704 are selected (3116). Any compromise clusters selected in step 3116 may then be processed in step 3108, which may include processing the compromise clusters along with any matching clusters identified in step 3106.

[0237] If no matching cluster groups are found in step 3110, method 3100 may include step 3120 of evaluating whether line application specification 2800 defines tolerances 2818 with respect to cost requirements 2812. Step 3120 may further include evaluating whether any cost functions of the non-matching cluster groups are within the defined tolerances for cost requirements 2812. If line application specification 2800 does not provide tolerances or the cluster groups are not within the defined tolerances, the operation may fail 3118 and an error message may be returned to the user, orchestrator 106, log file 200, or other destination.

[0238] If the line application specification 2800 provides a tolerance range and / or there is at least one cluster group that is within any defined tolerance range, a compromise cluster group may be selected (3122), and the application instances 118 of the line application instance 2404 may be deployed on the clusters 111 of the selected compromise cluster group. The compromise cluster group may be the cluster group that most closely matches the latency requirements 2712. If one or more cluster groups include the compromise cluster selected in step 3116, selecting the compromise cluster group may also evaluate the combination of inter-cluster latency for each cluster group and how closely each cluster in each cluster group meets the computing resource requirements 2816 and the runtime requirements 2806.

[0239] 32 and 33 illustrate a method 3200 for deploying a graph application instance 2406. The method 3200 may include a step 3202 of splitting the graph application instance 2406 into one or more triangle application instances 2402 and line application instances 2404, as shown in FIG. 33. The splitting 3202 may be performed taking into account specifications of the graph application instance 2406, including explicitly defined triangle application specifications 2700 and / or line application specifications 2800. The splitting may also include analyzing a graph representing the application instance 118 of the graph application instance 2406 to identify the triangle application instances 2402 and the line application instances 2404.

[0240] The method 3200 may include provisioning and deploying the triangle application instance 2402, such as according to the method 3000. The method 3200 may include provisioning and deploying the line application instance 2404, such as according to the method 3100.

[0241] Methods 3000, 3100 may be modified in one or more respects when deploying a graph application instance 2406. Method 3000 includes step 3010 of evaluating whether there are matching cluster groups, and method 3100 includes step 3110 of evaluating whether there are matching cluster groups. For method 3200, a "matching cluster group" may be defined as a matching cluster group that includes clusters of each application instance 118 of all triangle application instances 2402 and line application instances 2404 of a graph application instance 2406. Thus, in some embodiments, a matching cluster group must simultaneously satisfy the requirements of all triangle application instances 2402 and line application instances 2404 of a graph application instance 2406. In an alternative approach, the triangle application instances 2402 and line application instances 2404 of the graph application instance 2406 are processed one at a time, such as from largest to smallest (by number of application instances 118) of the triangle application instances 2402 and line application instances 2404 of the graph application instance 2406, or some other order. In either approach, if any set of triangle application instances 2402 or line application instances 2404 cannot be provisioned, i.e., if the operation fails (3018, 3118), the method 3200 fails for the graph application instance. Alternatively, partial failures may be allowed, such that a first portion of the triangle application instances and / or line application instances 2404 is deployed even if the second portion cannot be deployed.

[0242] 34 is a block diagram illustrating an example computing device 3400. The computing device 3400 can be used to perform various procedures as described herein. The server 102, the orchestrator 106, the workflow orchestrator 122, the vector log agent 126, the log processor 130, and the cloud computing platform 104 may each be implemented using one or more computing devices 3400. The orchestrator 106, the workflow orchestrator 122, the vector log agent 126, and the log processor 130 may be implemented on different computing devices 3400, or a single computing device 3400 may host two or more of the orchestrator 106, the workflow orchestrator 122, the vector log agent 126, and the log processor 130.

[0243] Computing device 3400 includes one or more processors 3402, one or more memory devices 3404, one or more interfaces 3406, one or more mass storage devices 3408, one or more input / output (I / O) devices 3410, and a display device 3430, all coupled to a bus 3412. Processor 3402 includes one or more processors or controllers that execute instructions stored in memory device(s) 3404 and / or mass storage 3408. Processor 3402 may also include various types of computer-readable media, such as cache memory.

[0244] The memory device 3404 includes a variety of computer-readable media, such as volatile memory (e.g., random access memory (RAM) 3414) and / or non-volatile memory (e.g., read-only memory (ROM) 3416). The memory device 3404 may also include re-writable ROM, such as flash memory.

[0245] Mass storage 3408 includes various computer-readable media such as magnetic tape, magnetic disks, optical disks, solid-state memory (e.g., flash memory), etc. As shown in Figure 34, a particular mass storage is a hard disk drive 3424. Various drives may also be included in mass storage 3408 to allow reading from and / or writing to various computer-readable media. Mass storage 3408 includes removable media 3426 and / or non-removable media.

[0246] The I / O devices 3410 include various devices that allow data and / or other information to be input to or retrieved from the computing device 3400. Example I / O devices 3410 include cursor control devices, keyboards, keypads, microphones, monitors or other display devices, speakers, printers, network interface cards, modems, lenses, CCD or other image capture devices, etc.

[0247] Display device 3430 includes any type of device capable of displaying information to one or more users of computing device 3400. Display device 3430 may be, for example, a monitor, a display terminal, a video projection device, etc.

[0248] The interface 3406 includes various interfaces that allow the computing device 3400 to interact with other systems, devices, or computing environments. The exemplary interface 3406 includes any number of different network interfaces 3420, such as interfaces to a local area network (LAN), a wide area network (WAN), a wireless network, and the Internet. Other interfaces include a user interface 3418 and a peripheral interface 3422. The interface 3406 may also include one or more peripheral interfaces, such as interfaces for a printer, a pointing device (mouse, trackpad, etc.), a keyboard, etc.

[0249] The bus 3412 allows the processor 3402, memory device 3404, interface 3406, mass storage device 3408, I / O devices 3410, and display device 3430 to communicate with each other and with other devices or components coupled to the bus 3412. The bus 3412 represents one or more of several types of bus structures, such as a system bus, a PCI bus, an IEEE 1394 bus, a USB bus, etc.

[0250] For purposes of illustration, programs and other executable program components are illustrated herein as separate blocks, with the understanding that such programs and components may reside at various times in different storage components of computing device 3400 and are executed by processor 3402. Alternatively, the systems and procedures described herein may be implemented in hardware or a combination of hardware, software, and / or firmware. For example, one or more application-specific integrated circuits (ASICs) can be programmed to execute one or more of the systems and procedures described herein.

[0251] In the foregoing disclosure, reference has been made to the accompanying drawings, which form a part hereof, and which show, by way of illustration, specific embodiments in which the present disclosure may be practiced. It is understood that other embodiments may be utilized and structural changes may be made without departing from the scope of the present disclosure. References herein to "one embodiment," "an embodiment," "an example embodiment," or the like indicate that the described embodiment may include a particular feature, structure, or characteristic, but not all embodiments necessarily include the particular feature, structure, or characteristic. Moreover, such phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is believed to be within the knowledge of one skilled in the art to affect such feature, structure, or characteristic in connection with other embodiments, whether or not explicitly stated.

[0252] Implementations of the systems, devices, and methods disclosed herein may include or utilize special-purpose or general-purpose computers, including computer hardware such as one or more processors and system memory, as described herein. Implementations within the scope of the present disclosure may also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. Such computer-readable media may be any available media accessible by a general-purpose or special-purpose computer system. Computer-readable media that store computer-executable instructions are computer storage media (devices). Computer-readable media that carry computer-executable instructions are transmission media. Thus, by way of example and not limitation, implementations of the present disclosure may include at least two distinctly different types of computer-readable media: computer storage media (devices) and transmission media.

[0253] Computer storage media (devices) include RAM, ROM, EEPROM, CD-ROM, solid-state drives ("SSD") (e.g., RAM-based), flash memory, phase-change memory ("PCM"), other types of memory, other optical disk storage, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store desired program code means in the form of computer-executable instructions or data structures and that can be accessed by a general-purpose or special-purpose computer.

[0254] Embodiments of the devices, systems, and methods disclosed herein can communicate over a computer network. A "network" is defined as one or more data links that enable the transfer of electronic data between computer systems and / or modules and / or other electronic devices. When information is transferred or provided to a computer over a network or another communications connection (either hardwired, wireless, or a combination of hardwired or wireless), the computer properly views the connection as a transmission medium. Transmission media can be used to carry desired program code means in the form of computer-executable instructions or data structures and can include networks and / or data links that can be accessed by a general-purpose or special-purpose computer. Combinations of the above should also be included within the scope of computer-readable media.

[0255] Computer-executable instructions include, for example, instructions and data that, when executed by a processor, cause a general-purpose computer, special-purpose computer, or special-purpose processing device to perform a certain function or group of functions. Computer-executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, or source code. While the present subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the present subject matter defined in the appended claims is not necessarily limited to the features or acts described above. Rather, the described features and acts are disclosed as example forms of implementing the claims.

[0256] Those skilled in the art will appreciate that the present disclosure may be implemented in networked computing environments having many types of computer system configurations, including in-dash vehicle computers, personal computers, desktop computers, laptop computers, message processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, cell phones, PDAs, tablets, pagers, routers, switches, various storage, etc. The present disclosure may also be implemented in distributed system environments where tasks are performed by both local and remote computer systems that are linked via a network (either by hardwired data links, wireless data links, or a combination of hardwired and wireless data links). In a distributed system environment, program modules may be located in both local and remote memory storage.

[0257] Additionally, where appropriate, the functions described herein may be performed by one or more of hardware, software, firmware, digital components, or analog components. For example, one or more application-specific integrated circuits (ASICs) may be programmed to execute one or more of the systems and procedures described herein. Certain terms are used throughout this specification and claims to refer to particular system components. As will be understood by those skilled in the art, components may be referred to by different names. This specification does not intend to distinguish between components that differ in name but not function.

[0258] It should be noted that the sensor embodiments described above may include computer hardware, software, firmware, or any combination thereof to perform at least a portion of their functionality. For example, the sensor may include computer code configured to run on one or more processors and may include hardware logic / electrical circuitry controlled by the computer code. These exemplary devices are provided herein for illustrative purposes and are not intended to be limiting. Embodiments of the present disclosure may be implemented in additional types of devices, as known to those skilled in the art.

[0259] At least some embodiments of the present disclosure are directed to computer program products including such logic (e.g., in the form of software) stored on any computer-usable medium that, when executed on one or more data processing devices, causes the devices to operate as described herein.

[0260] While various embodiments of the present disclosure have been described above, it should be understood that they have been presented by way of example only, and not limitation. It will be apparent to those skilled in the art that various changes in form and detail can be made without departing from the spirit and scope of the present disclosure. Thus, the breadth and scope of the present disclosure should not be limited by any of the above-described exemplary embodiments, but should be defined only in accordance with the following claims and their equivalents. The foregoing description has been presented for purposes of illustration and description. It is not intended to be exhaustive or to limit the disclosure to the precise form disclosed. Many modifications and variations are possible in light of the above teachings. Furthermore, it should be noted that any or all of the foregoing alternative embodiments may be used in any combination desired to form additional hybrid embodiments of the present disclosure.

Claims

1. 1. An apparatus comprising a computing device including one or more processing devices and one or more memory devices operably coupled to the one or more processing devices, The one or more memory devices, when executed by the one or more processing devices, cause the one or more processing devices to: receiving a specification of one or more application instances, the specification including both (a) one or more computing resource requirements and (b) one or more cluster runtime requirements; for each application instance of the one or more application instances, identifying one or more clusters of one or more hosts that satisfy (a) and (b); and deploying the one or more application instances on the one or more hosts; a computing device that stores executable code; Device.

2. the one or more application instances include three or more application instances; the one or more clusters include three or more clusters; and the one or more cluster runtime requirements include a latency requirement for latency between each possible pair of clusters of the three or more clusters; 10. The apparatus of claim 1.

3. the one or more application instances include two or more application instances arranged in a hierarchical manner; the one or more clusters include two or more clusters; and the one or more cluster runtime requirements include a cost function for at least one of the two or more clusters; 10. The apparatus of claim 1.

4. the cost function is the cost of hosting the most resource-intensive of the two or more application instances; 4. The apparatus of claim 3.

5. the most resource-intensive of the two or more application instances is a database application instance; 5. The apparatus of claim 4.

6. the one or more computing resource requirements include a processing power requirement; 10. The apparatus of claim 1.

7. the one or more computing resource requirements include a memory requirement; 10. The apparatus of claim 1.

8. the one or more computing resource requirements include a storage requirement; 10. The apparatus of claim 1.

9. The executable code further, when executed by the one or more processing devices, causes the one or more processing devices to: determining that none of the available clusters meets the one or more computing resource requirements and the one or more cluster runtime requirements; (c) determining that the specification defines a tolerance range; and (c).

10. The apparatus of claim 1.

10. the one or more cluster runtime requirements include a location requirement; 10. The apparatus of claim 1.

11. the one or more cluster runtime requirements include an availability requirement; 10. The apparatus of claim 1.

12. The specification is: dot application instance, a triangle application instance, and Line Application Instance define one or more of:

10. The apparatus of claim 1.

13. the one or more hosts are servers; 10. The apparatus of claim 1.

14. the one or more hosts are part of a cloud computing platform; 10. The apparatus of claim 1.

15. receiving, by a computer system, specifications of a plurality of application instances, the specifications including both (a) one or more computing resource requirements and (b) one or more cluster runtime requirements; Dividing the plurality of application instances into two or more groups of two or more different types; and For each group of the two or more groups: identifying, by the computer system, for each application instance in each group, one or more clusters of one or more hosts that satisfy (a) and (b); and deploying, by the computer system, one or more application instances of each group onto the one or more hosts; method.

16. the two or more different types include a triangle application instance; 16. The method of claim 15.

17. the two or more different types include a line application instance; 16. The method of claim 15.

18. the two or more different types include a triangle application instance and a line application instance; 16. The method of claim 15.

19. the one or more cluster runtime requirements include one or more first runtime requirements for a first type of the two or more different types and one or more second runtime requirements for a second type of the two or more different types; the one or more second runtime requirements are different from the one or more first runtime requirements; 16. The method of claim 15.

20. the one or more first runtime requirements include an inter-cluster latency requirement; the one or more second runtime requirements include a cost requirement; 20. The method of claim 19.

Citation Information

Patent Citations

  • Gateway control method

    JP2007316937A

  • Program managing system, program managing method, client, and program

    JP2011170638A

  • Information processing system, information processor and program

    JP2014170442A

  • Image formation apparatus and program

    JP2018056648A

  • Information processing device and program

    JP2019057037A