Redeploying Applications Using Active and Available Inventory

An agentless method determines AAI to efficiently manage and redeploy computing resources, addressing the challenge of high variability in usage and improving resource allocation in modern computing environments.

JP2025536893APending Publication Date: 2025-11-12RAKUTEN SYMPHONY INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025520138
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-12-16
Publication Date
2025-11-12

AI Technical Summary

Technical Problem

Managing on-premise and cloud computing resources is challenging due to high variability in usage, making it difficult to efficiently allocate and redeploy applications in modern computing environments.

Method used

An agentless approach is employed to determine active and available inventory (AAI) of computing resources by processing log data from multiple hosts, enabling redeployment of components based on utilization, using an orchestrator to manage resources and allocate them efficiently.

Benefits of technology

This method allows for better management of computing resources by redeploying components to optimize utilization, reduce underutilization, and improve performance by reallocating resources to meet performance and latency requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025536893000001_ABST
    Figure 2025536893000001_ABST
Patent Text Reader

Abstract

A computer system pulls observability data (metrics, logs, events, alerts, inventory) for multiple components from a remote server, which may be part of a cloud computing platform. The components may be application instances, containers, storage volumes, pods, or other components. The computer system derives utilization metrics for each component and for each of one or more types of computing resources (compute, memory, and storage). The utilization metrics are compared to an available inventory of computing resources to obtain an active and available inventory (AAI). Components can be redeployed and allocated reduced computing resources based on the AAI. Components can be grouped into clusters, and components can be consolidated into a reduced number of clusters based on the AAI.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to agentless application redeployment based on an active and available inventory in a distributed computing system. [Background technology]

[0002] Whether processing e-commerce transactions, streaming content, providing back-end data management for mobile applications, or other services, modern enterprises require large amounts of computing resources, including processor time, memory, and persistent data storage. The amount of computing resources changes over time. Modern computing environments can dynamically sell and scale down to adapt to changes in usage. For example, Kubernetes is a popular tool for adding and removing application instances based on usage. High variability in computing resource usage makes it difficult to manage on-premise computing hardware and purchased cloud computing resources.

[0003] It would be an advancement in the art to enable better management of on-premise computing hardware and purchased cloud computing resources. Summary of the Invention

[0004] The apparatus includes a computing device including one or more processing devices and one or more memory devices operably coupled to the one or more processing devices. The one or more memory devices store executable code that, when executed by the one or more processing devices, causes the one or more processing devices to receive log data from multiple hosts over a network. The one or more processing devices process the log data to obtain utilization of computing resources of one or more remote servers by multiple components executing on the multiple hosts. An active and available inventory of computing resources of the one or more remote servers is determined according to the utilization. Based on the active and available inventory, one component of the multiple components on a first host of the multiple hosts may be redeployed to a second host of the multiple hosts.

[0005] In order that the advantages of the present invention may be readily understood, the invention, briefly described above, will now be more particularly described by reference to specific embodiments thereof which are illustrated in the accompanying drawings, the invention being described and explained with additional specificity and detail through the use of the accompanying drawings, with the understanding that these drawings depict only typical embodiments of the invention and therefore should not be considered as limiting its scope. [Brief explanation of the drawings]

[0006] [Figure 1] 1 is a schematic block diagram of a network environment in which active and available inventory (AAI) discovery may be performed, according to one embodiment. [Figure 2] FIG. 1 is a schematic block diagram illustrating components for collecting and processing log data, according to one embodiment. [Figure 3] FIG. 2 is a schematic block diagram illustrating sources of provisioning data, according to one embodiment. [Figure 4]FIG. 2 is a schematic block diagram illustrating components illustrating processing of log data to obtain AAI, according to one embodiment. [Figure 5] FIG. 1 is a process flow diagram of a method for collecting provisioning data, according to one embodiment. [Figure 6] FIG. 1 is a process flow diagram of a method for deriving an AAI, according to one embodiment. [Figure 7] FIG. 1 is a schematic block diagram illustrating the derivation of relationships between components according to one embodiment. [Figure 8] 1 is a schematic block diagram of a topology of components of a network environment, according to one embodiment. [Figure 9] FIG. 1 is a process flow diagram of a method for identifying relationships between components according to a manifest and dynamic provisioning data, according to one embodiment. [Figure 10] FIG. 2 is a process flow diagram of a method for identifying session relationships between components, according to one embodiment. [Figure 11] FIG. 2 is a process flow diagram of a method for identifying access relationships between components, according to one embodiment. [Figure 12] FIG. 1 is a process flow diagram of a method for identifying network relationships, according to one embodiment. [Figure 13] FIG. 1 is a process flow diagram of a method for generating a representation of a topology, according to one embodiment. [Figure 14A] 1 is an exemplary representation of a topology, according to one embodiment. [Figure 14B] FIG. 2 is an exemplary diagram of application data, according to one embodiment. [Figure 14C] FIG. 2 is an exemplary diagram of cluster data, according to one embodiment. [Figure 14D] FIG. 2 is an exemplary diagram illustrating the importance of storage volumes, according to one embodiment. [Figure 15] FIG. 1 illustrates data used to redeploy applications and perform cluster integration, according to one embodiment. [Figure 16A] FIG. 1 illustrates an exemplary application redeployment and cluster integration, according to one embodiment. [Figure 16B] FIG. 1 illustrates an exemplary application redeployment and cluster integration, according to one embodiment. [Figure 16C] FIG. 1 illustrates an exemplary application redeployment and cluster integration, according to one embodiment. [Figure 17A] FIG. 2 is a process flow diagram of an exemplary method for performing application redeployment, according to one embodiment. [Figure 17B] FIG. 2 is a process flow diagram of an exemplary method for performing application redeployment, according to one embodiment. [Figure 18] FIG. 2 is a process flow diagram of a method for merging clusters, according to one embodiment of the present invention. [Figure 19] FIG. 1 is a process flow diagram of a method for identifying candidate cluster mergers. [Figure 20] FIG. 1 is a schematic block diagram illustrating topology modification according to one embodiment. [Figure 21] FIG. 1 is a process flow diagram of a method for locking a topology, according to one embodiment. [Figure 22] FIG. 1 is a process flow diagram of a method for preventing topology modification, according to one embodiment. [Figure 23] FIG. 1 is a process flow diagram of a method for detecting topology changes, according to one embodiment. [Figure 24] FIG. 1 is a schematic diagram illustrating the deployment of multiple applications on multiple clusters, according to one embodiment. [Figure 25] FIG. 1 is a schematic block diagram illustrating a cluster specification, according to one embodiment. [Figure 26] FIG. 2 is a schematic block diagram illustrating a dot application specification, according to one embodiment. [Figure 27] FIG. 1 is a schematic block diagram illustrating a triangle application specification according to one embodiment. [Figure 28]FIG. 1 is a schematic block diagram illustrating a line application specification, according to one embodiment. [Figure 29] FIG. 1 is a process flow diagram of a method for provisioning a dot application, according to one embodiment. [Figure 30] FIG. 1 is a process flow diagram of a method for provisioning a triangular application, according to one embodiment. [Figure 31] FIG. 2 is a process flow diagram of a method for provisioning a line application, according to one embodiment. [Figure 32] FIG. 1 is a process flow diagram of a method for provisioning a graph application, according to one embodiment. [Figure 33] FIG. 1 illustrates the division of a graph application into a line application and a triangle application, according to one embodiment. [Figure 34] FIG. 1 is a schematic block diagram of an exemplary computing device suitable for implementing methods according to embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0007] 1 illustrates an exemplary network environment 100 in which the systems and methods disclosed herein may be used. The components of network environment 100 may be connected to each other by a network, such as a local area network (LAN), a wide area network (WAN), the Internet, a chassis backplane, or other type of network. The components of network environment 100 may be connected by wired or wireless network connections.

[0008] The network environment 100 includes multiple servers 102. Each of the servers 102 may include one or more computing devices, such as a computing device having some or all of the attributes of computing device 3400 of FIG. 34. Each server 102 lacks an agent to coordinate the execution of management tasks. The systems and methods described herein enable active and available inventory (AAI) determination to be performed for servers 102 that lack an agent supporting AAI determination.

[0009] As used herein, "active and available inventory" (AAI) refers to computing resources available for allocation to application instances, including some or all of the storage on physical storage devices mounted on a server 102, the memory of a server 102, the processing cores of a server 102, and the networking bandwidth of a network connection between a server 102 and another server 102 or other computing device.

[0010] Computing resources may also be allocated within a cloud computing platform 104, such as Amazon Web Services (AWS), GOOGLE CLOUD, AZURE, or other cloud computing platforms. Cloud computing resources may include physical storage, processor time, memory, and / or networking bandwidth purchased by the cloud computing platform in units specified by the provider.

[0011] In some embodiments, some or all of the servers 102 may function as edge servers within a telecommunications network. For example, some or all of the servers 102 may be coupled to baseband units (BBUs) 102a that provide conversion between radio frequency signals output and received by antennas 102b and digital data transmitted and received by the servers 102. For example, each BBU 102a may perform this conversion according to a cellular wireless data protocol (e.g., 4G, 5G, etc.). Servers 102 that function as edge servers may have limited computing resources or be heavily loaded, making it infeasible for the servers 102 to run agents that collect data for obtaining AAI. Similarly, if there are a large number of servers 102, installing agents for data collection may be a time-consuming task.

[0012] The orchestrator 106 provisions computing resources to application instances of one or more different application executables, such as according to a manifest that defines the computing resource requirements for each application instance. The manifest can define dynamic requirements that define the scaling up of several application instances and corresponding computing resources depending on usage. The orchestrator 106 can include or work with utilities such as KUBERNETES to perform dynamic scaling up and down of the number of application instances.

[0013] The orchestrator 106 runs on a computer system separate from the server 102 and is connected to the server 102 using a network that requires the use of a destination address for communication, such as a network including other protocols including the Ethernet protocol, the internet protocol (IP), Fibre Channel, or any higher-level protocols built upon the aforementioned protocols, such as the user datagram protocol (UDP), the transport control protocol (TCP), etc.

[0014] The orchestrator 106 can cooperate with the servers 102 to initialize and configure the servers 102. For example, each server 102 can cooperate with the orchestrator 106 to obtain a gateway address to use for outbound communications and a source address to be assigned to the server 102 to use for inbound communications. The servers 102 can cooperate with the orchestrator 106 to install an operating system on the servers 102. For example, the gateway address and source address may be provided, and the operating system may be installed using techniques described in U.S. patent application Ser. No. 16 / 903,266, filed June 16, 2020, entitled "AUTOMATED INITIALIZATION OF SERVERS," the contents of which are incorporated herein by reference in their entirety.

[0015] The orchestrator 106 may be accessible via an orchestrator dashboard 108. The orchestrator dashboard 108 may be implemented as a web server or other server-side application accessible via a browser or client application running on a user computing device 110, such as a desktop computer, laptop computer, mobile phone, tablet computer, or other computing device.

[0016] The orchestrator 106 can cooperate with the servers 102 to provision computing resources for the servers 102 and instantiate components of the distributed computing system on the servers 102 and / or the cloud computing platform 104. For example, the orchestrator 106 can import manifests that define the provisioning of computing resources and instantiation of components, such as clusters 111, pods 112 (e.g., KUBERNETES pods), containers 114 (e.g., DOCKER containers), storage volumes 116, and application instances 118. The orchestrator can then allocate computing resources and instantiate the components according to the manifests.

[0017] The manifest may define requirements such as network latency requirements, affinity requirements (same node, same chassis, same rack, same data center, same cloud region, etc.), anti-affinity requirements (different nodes, different chassis, different racks, different data centers, different cloud regions, etc.), as well as minimum provisioning requirements (number of cores, amount of memory, etc.), performance or quality of service (QoS) requirements, or other constraints. Thus, the orchestrator 106 can provision computing resources to meet or nearly meet the requirements of the manifest.

[0018] Component instantiation and component management may be performed by workflows. A workflow is a set of tasks, executables, configurations, parameters, and other computing functions that are predefined and stored in the workflow repository 120. Workflows may be defined to instantiate each type of component (e.g., clusters 111, pods 112, containers 114, storage volumes 116, application instances), monitor the performance of each type of component, repair each type of component, upgrade each type of component, replace each type of component, copy (e.g., snapshot, backup, etc.) and restore each type of component from the copy, and other tasks. Some or all of the tasks performed by the workflows may be implemented using Kubernetes or other utilities to perform some or all of the tasks.

[0019] The orchestrator 106 can instruct the workflow orchestrator 122 to perform a task with respect to the component. In response, the workflow orchestrator 122 retrieves a workflow corresponding to the task (e.g., a task of the type instantiate, monitor, upgrade, replace, copy, restore, etc.) from the workflow repository 120. The workflow orchestrator 122 then selects a worker 124 from a worker pool and instructs the worker 124 to perform the workflow with respect to the server 102 or cloud computing platform 104. The instruction from the orchestrator 106 can specify a particular server 102, cloud region or cloud provider, or other location for executing the workflow. The worker 124, which may be a container, then performs the workflow's functionality with respect to the location instructed by the orchestrator 106. In some implementations, the worker 124 can also perform the task of retrieving the workflow from the workflow repository 120 as instructed by the workflow orchestrator 122.

[0020] In some implementations, the container implementing the worker 124 is remote from the server 102 on which the worker 124 implements the workflow. The worker 124 can still implement some or all of the workflow even if the server 102 or cloud computing platform 104 does not have an agent installed that is programmed to collaborate with the worker 124 to implement the workflow. For example, the worker 124 can establish a secure command line interface (CLI) connection to the server 102 or cloud computing platform 104. For example, a secure shell (ssh), remote login (rlogin), or remote procedure calls (RPC), or other interface provided by the operating system of the server 102 or cloud computing platform 104, can be used to send instructions and verify completion of the instructions on the server 102 or cloud computing platform 104.

[0021] One workflow may involve monitoring computing resource usage by each component (hereinafter, the "monitoring workflow"). The monitoring workflow may be invoked periodically by the orchestrator 106 for each component, or the monitoring workflow may be a persistent process that runs periodically with periods of inactivity between them.

[0022] The monitoring workflow may include establishing a secure connection to each component, reading one or more log files for each component, and passing the log files to a vector log agent 126. The vector log agent 126 may perform initial processing on the data in the log files to obtain enriched data. The vector log agent 126's processing may include enriching the data in the log files (e.g., providing contextual information indicating the component, time, identifiers of the source server 102, hosting container 114, cluster 111, pod 112, virtual machine, or unit of computing resource in the cloud computing platform 104, etc.), performing map / reduce functions on messages in the log files, combining messages in the log files into aggregate representations of messages, and other functions. The vector log agent 126 may process the log files according to one or more vector remap language (VRL) statements. The vector log agent 126 may run independently of the worker 124, or the monitoring workflow may include running an instance of the vector log agent 126. For example, a set of VRL statements may be included in each monitoring workflow that corresponds to the type of component that the monitoring workflow is configured to monitor, and each monitoring workflow may then include processing log files according to the VRL statements of the monitoring workflow.

[0023] The enriched data output by the vector log agent 126 can be stored in the log store 128. The log processor 130 reads the enriched data from the log store and derives an active and available inventory (AAI), which is a list of computing resources available for allocation to components. How the log processor 130 obtains the AAI is described in more detail below. The log processor 130 passes the AAI to the orchestrator 106. The orchestrator 106 can use the AAI to perform various functions on the components, such as adding, removing, or redeploying them to a different location.

[0024] FIG. 2 illustrates the collection of log files 200 from various components. The log files 200 may be collected using a monitoring workflow for each component or other techniques for collecting log files. The log files 200 may include log files generated by an operating system 202 running on the server 102. Alternatively, the cloud computing platform 104 may generate log files 200 that describe the state of units of computing resources and / or executables running on the cloud computing platform 104. Virtual machines on which components execute may also generate log files 200. In the following description, reference is made to log files 200 with the understanding that any observability data represented as log files or other formats may be collected and processed in a similar manner. In particular, metrics, events, alerts, inventory, and other data may be collected instead of or in addition to the log files 200 and processed in a similar manner to the log files 200.

[0025] A cluster 111 is a collection of hosts (servers 102 and / or one or more units of computing resources on a cloud computing platform) managed as a unit. Each host includes a master running on one of the hosts that manages the deployment of pods 112, containers 114, and application instances 118 on the hosts of the cluster. The master manages the scaling up, scaling down, and redeployment of the application instances 118. As used herein, actions performed by and with respect to a cluster 111 may be understood to be performed by or with respect to the master managing the cluster 111. Each cluster 111 may generate one or more log files 200 that describe the operation of the cluster 111.

[0026] Kubelet 204 is a KUBERNETES agent that runs on a node and implements instructions from a server 102 or cluster 111 on a cloud computing platform to instantiate, monitor, and manage pods 112. Each Kubelet 204 can generate one or more log files 200 that describe the operation of the Kubelet 204 and each pod 112 running within it. A pod 112 is a group of one or more containers 114 with shared storage, network resources, and execution context. A pod 112 can generate one or more log files 200 that describe the state of the pod 112 and the execution of the containers 114 of the pod 112. Each container 114 can generate one or more log files 200 that describe the execution of the container and any application instances 118 running within the container 114. Each application instance 118 can also generate one or more log files that describe the operation of the application instance 118. A storage volume 116 may be a unit of virtualized storage, and the storage manager that implements the storage volume 116 may also generate one or more log files 200 that describe the operation of the storage volume 116 .

[0027] Log files 200 are pulled from the server 102 or cloud computing platform 104 where they are stored and processed by the vector log agent 126 to generate enriched data. The enriched data is processed by the log processor 130 to obtain the AAI. The orchestrator 106 receives the AAI and manages the provisioning of unused computing resources identified in the AAI for use by the components.

[0028] 3 , data included in log file 200 may be associated with provisioning data 300 for obtaining AAI. Provisioning data 300 includes identifiers of components instantiated by orchestrator 106 and allocation data indicating computing resources allocated to each component. For example, on-premise provisioning data 302 may describe provisioning for one or more servers 102. For example, on-premise provisioning data 302 may include multiple entries, each including a node identifier (i.e., an identifier of a server 102), a computing allocation (e.g., a number of processor cores), a memory allocation (e.g., a number of megabytes (MB), a number of gigabytes (GB), or other units of memory), a storage allocation (e.g., a number of megabytes (MB), a number of gigabytes (GB), or other units of storage), and an identifier of the component to which the allocation belongs (e.g., an identifier of a cluster 111, a pod 112, a container 114, a storage volume 116, or an application instance 118). The component identifier may be in the form of a universally unique identifier (UUID) that is centrally assigned, such as by the orchestrator 106 or other central component, to all components that belong to a common namespace. An entry may reference multiple components. For example, provisioning may occur at the level of a cluster 111, such that all pods 112, containers 114, storage volumes 116, and application instances 118 of the cluster 111 are referenced in the entry for the cluster 111.

[0029] The provisioning data 300 may further include cloud provisioning data 304. The cloud provisioning data 304 may describe provisioning for one or more units of computing resources on the cloud computing platform 104. The cloud provisioning data 304 may include multiple entries, each including a unit identifier that identifies a unit of cloud computing resources. The identifier of the unit of computing resources may further identify a cloud computing provider (e.g., AWS, AZURE, GOOGLE CLOUD), a region of the cloud computing platform 104, and / or other data. Each entry may further include data describing an allocation of computing, memory, and storage. Each entry may further include identifiers of one or more components to which the allocation belongs, as described above with respect to the on-premises provisioning data 302.

[0030] Note that the on-premise provisioning data 302 and the cloud provisioning data 304 are dynamic: the orchestrator 106 can scale up and down the number of application instances 118 of any given executable, as well as the number of pods 112, containers 114, and storage volumes 116 used by the application instances.

[0031] In addition to the provisioning data 300, the AAI may also be determined using other data, such as hardware inventory data 306 and cloud inventory data 308. The hardware inventory data 306 may include an entry for each server 102. Each entry may indicate the computing (e.g., total number of processing cores, graphics processing unit (GPU) cores, or other computing components), memory, and storage available on the server 102, as well as the node identifier for the server 102. The cloud inventory data 308 similarly includes entries that include identifiers for units of cloud computing resources and the computing, memory, and storage available for that unit. The hardware inventory data 306 and cloud inventory data 308 may indicate current availability; i.e., an entry may be removed or flagged as unavailable in response to the server 102 or cloud computing platform 104 referenced by the entry becoming unavailable due to failure or lack of network connectivity. The availability of the server 102 or cloud computing platform 104 may be determined by performing a health check, sending a ping message, measuring traffic latency, detecting a failed network connection, or any other technique for determining the status and accessibility of a computing device.

[0032] 4 illustrates an approach for calculating the AAI. The log file 200 includes multiple log messages 400. Each message may include a text string including a component identifier and a value, such as a usage value. The entry identifier may also be obtained from the directory location of the log file or the name of the log file. The usage value may include some or all of the following: an indicator of processor time spent executing the component identified by the entry identifier; an amount of memory occupied by the component identified by the component identifier; and an amount of storage used (e.g., written) by the component identified by the component identifier. For example, there may be separate entries, each indicating separate information regarding the component identifier, i.e., one entry indicating processor time and another entry indicating memory used. In some implementations, a log message 400 includes one or more usage values, and another log message 400 includes a process identifier and a component identifier executing the process identified by the process identifier.

[0033] The log message 400 is processed by the vector agent 126 to obtain enriched data 402. For example, an item of the enriched data 402 may include a component identifier and usage metrics (processor time, memory, storage) for that component identifier. The vector agent 126 may obtain the enriched data 402 by executing one or more VRL statements on the log message 400. For example, a log message 400 that associates a process identifier with a usage value may be mapped by the vector agent 126 to a log message that associates the process identifier with a component identifier. The vector agent 126 may perform a map / reduce function to aggregate the usage values ​​into aggregated usage metrics for the component identifier.

[0034] The enriched data 402 may then be processed by the log processor 130 along with the provisioning data 300 to obtain an active and available inventory (AAI) 406. For example, the provisioning data 300 may include provisioning entries 404 that include a node identifier of a server 102 or an identifier of a unit of computing resource within a cloud computing platform. Each provisioning entry 404 may include a component identifier, i.e., an identifier of a cluster 111, a pod 112, a container 114, a storage volume 116, or an application instance 118. Each provisioning entry 404 may include a value indicating an allocation, i.e., the computing, memory, and / or storage allocated to the component identified by the component identifier.

[0035] Thus, the log processor 130 may obtain one or more provisioning entries 404 that include the component identifier, and items of enrichment data 402 that include the same component identifier. For a given computing resource on a host (a unit of computing resource in a server 102 or cloud computing platform 104), let U(t,i) denote the utilization of that computing resource reported at a given time (t) for component i, let P(t,i) denote the current allocation of that computing resource to component i, and let T denote the inventory of that computing resource available on the host. Thus, the AAI of that computing resource on the host is:

number

number

[0036] 5 and 6 illustrate methods 500 and 600, respectively, that may be performed using network environment 100 to obtain AAI. Methods 500 and 600 may be performed by one or more computing devices 3400, such as one or more computing devices executing orchestrator 106 and / or log processor 130 (see description of FIG. 34 below).

[0037] 5 , method 500 may include obtaining 502 component identifiers for statically defined components, such as those referenced in a manifest ingested by orchestrator 106. Method 500 may include obtaining 504 component identifiers for dynamically created components. Dynamically created components may be instantiated to scale up capacity. Dynamically created components may be created by orchestrator 106 or KUBERNETES. Component identifiers for dynamically created components may be obtained from log files 200 generated by KUBERNETES, i.e., the KUBERNETES master, Kubelet, or other component of a KUBERNETES installation that performs component instantiation. Note that dynamically created components may also be deleted. Thus, the current set of component identifiers obtained in steps 502 and 504 may be updated to remove component identifiers for components that were dynamically deleted due to a scale-down, host failure, or other event.

[0038] Method 500 may include obtaining 506 static provisioning for each component identifier of each statically defined component and obtaining 508 dynamic provisioning for each component identifier of each dynamically created component. The provisioning for each component identifier may include a host identifier (an identifier for a unit of computing resources of the server 102 or cloud computing platform) and an allocation of one or more computing resources (computing power, memory, and / or storage). Method 500 may further include obtaining a total available inventory. The total available inventory may include an inventory for each host currently available (functional and accessible via a network connection). The inventory for each host may include a total of processor cores, memory, and / or storage capacity.

[0039] 6, method 600 may include deriving 602 usage data for each component identifier identified in steps 502 and 504. As described above, deriving 602 usage data may include retrieving log file 200, enriching log file 200 to obtain enriched data 402, and aggregating enriched data 402 to obtain usage metrics for each component identifier.

[0040] Method 600 may include deriving 604 usage data for each host. For example, usage metrics for each component running on each host may be aggregated (e.g., added together) to obtain total metrics for each host: total computing power usage, total memory usage, total storage usage. As used herein, "computing power" may be defined as the amount of processor time used, the number of processor cycles used, and / or the percentage of processor cycles or time used.

[0041] Method 600 may include retrieving 606 static and dynamic provisioning data for each component identifier (see description of steps 506 and 508) and an inventory of each host (see description of step 510). An AAI may then be derived 608. As described above, step 608 may include calculating some or all of the AAI(t), O(t,i), and O(t) for each computing resource (computing power, memory, storage) of each host.

[0042] The method 600 may further include modifying 610 the provisioning in the network environment 100 using the AAI. A non-limiting list of modifications may include: Additional components (clusters 111, pods 112, containers 114, storage volumes 116, and / or application instances 118) are provisioned to utilize the computing resources identified in the AAI according to the manifest. Redeploying a component to a different host to more closely meet the performance, quality of service, affinity, anti-affinity, latency, or other requirements expressed in the manifest. Remove underutilized components. Remove underutilized components that are distributed across multiple servers 102 or units of computing resources within the cloud computing platform 104 and redeploy some or all of the underutilized components onto a reduced number of hosts. Redeploying underutilized components (e.g., (O(t,i) / P(t,i))<0.5) to a server 102 or cloud computing platform 104 that has higher latency and / or fewer computing resources than the underutilized component's current host. Redeploying overutilized components (e.g., (O(t,i) / P(t,i))<0.9) to servers 102 that have lower latency and / or more computing resources than the overutilized component's current host.

[0043] 7 , the log processor 130, the orchestrator 106, and / or some other component may further process the provisioning data 300 and the log file 200 to identify relationships between component identifiers. For example, the provisioning data 300 may indicate a hosting relationship 700. As used herein, a “hosting relationship” refers to a component running on or within another component, such as a cluster 111 or a pod 112 hosted by a server 102 or unit of computing resources of the cloud computing platform 104, a container 114 running in a pod 112, or an application instance 118 running in a container. A storage volume 116 may be considered to have a hosting relationship 700, i.e., a relationship of being hosted by the container 114 or pod 112 to which the storage volume 116 is mounted. The hosting relationship 700 may be derived from instructions in a manifest that defines the instantiation of a second component on a first component, thereby defining the hosting relationship 700 between the first and second components. Hosting relationships may be derived from the log file 200 in a similar manner, i.e., a record of instantiating a second component on a first component establishes a hosting relationship between the first and second components.

[0044] The provisioning data 300 may further indicate environment variable relationships 702. The manifest may include instructions for configuring one or more environment variables of a first component to reference a second component, e.g., for configuring the first component to use the services of or provide services to the second component. The log file 200 may record configuring one or more environment variables of a first component to reference a second component in a similar manner.

[0045] The provisioning data 300 may further indicate a network relationship 704. The manifest may include instructions for configuring a first component to use an IP address or other type of address belonging to a second component, thereby establishing a network relationship 704 between the first and second components. The log file 200 may record configuring the first component to reference the address of the second component in a similar manner. Establishing the network relationship 704 may be a multi-step process of 1) determining that the first component is configured to use a first address, and 2) mapping the first address to an identifier for the second component.

[0046] As noted above, provisioning data 300 is dynamic and may change over time. Thus, some or all of hosting relationships 700, environment variable relationships 702, and network relationships 704 may be re-derived on a fixed, recurring basis or in response to detection of records in log file 200 that indicate actions that may affect any of these relationships 702-704.

[0047] The log files 200 may also be evaluated to identify other types of relationships between components. For example, the log files 200 may be evaluated to identify session relationships 706. When a first component establishes a session at the application level to use an application instance 118 that is or is hosted by a second component, one or more log files 200 generated by the second component may record this fact. Thus, to obtain the current session relationships 706 between pairs of components, the log files 200 may be analyzed to identify session creations and terminations.

[0048] The log files 200 may be evaluated to identify access relationships 708. When a first component accesses a session of an application instance 118 that is or is hosted by a second component, one or more log files 200 generated by the second component may record this fact. The access may include generating a request for a service provided by the second component, reading data from the second component, writing data to the second component, or other interaction between the first and second components. Thus, the log files 200 may be analyzed to identify accesses by the first component of the second component. Whether an access indicates a current access relationship may be handled in various ways. In response to identifying a record of an access, an access relationship 708 may be created between the first component and the second component accessed by the second component, and this access relationship may (a) be maintained as long as the first and second components exist, or (b) be deleted if no access is recorded in the log files 200 for a threshold period.

[0049] Log file 200 may be evaluated to identify network connection relationships 708. For example, when a first component establishes a network connection to a second component, log files 200 of one or both of the first and second components may record this fact. Accordingly, log file 200 may be analyzed to identify the establishment of the network connection between the first and second components and the termination (if any) of the network connection between the first and second components. In this manner, all active network connections between components may be identified as network connection relationships 710. Network connection relationship 710 may be created between the first and second components in response to identifying the creation of a network connection between the first and second components, and network connection relationship 710 may (a) be maintained as long as the first and second components exist, (b) be deleted when the network connection is terminated, or (c) expire if a new network connection is not established within a threshold time after the network connection is terminated.

[0050] Network connection relationship 710 can be distinguished from network relationship 704 in the sense that network connection relationship 710 refers to an actual network connection, whereas network relationship 704 refers to configuring a first component with a network address of a second component, regardless of whether a network connection is established. In some implementations, only network connection relationship 710 is used.

[0051] Referring to FIG. 8 , the log processor 130, the orchestrator 106, and / or some other component may further generate a topology representation 800. The topology 800 may be represented as a graph including nodes and edges. Each node may be a component identifier for a component. The components may include a host 802 (e.g., a server 102 or a unit of computing resource in a cloud computing platform), a cluster 111, a pod 112, a container 114, a storage volume 116, an application instance 118, or other components. The edges of the topology connect the nodes and represent relationships between the nodes, such as any of the hosting relationships 700, environment variable relationships 702, network relationships 704, session relationships 706, access relationships 708, and network connection relationships 710. The edges may be unidirectional, indicating that a first node depends on a second node that does not depend on the first node for correct functioning. The edges may be bidirectional, indicating that the first node and the second node depend on each other. For example, hosting relationship 700 may be unidirectional, indicating a dependency of a second component on a first component that is a host for the second component. Network relationship 704 or network connection relationship 710 may be bidirectional, because both components must function for a network connection to exist.

[0052] FIG. 9 illustrates a method 900 for processing provisioning data 300. Method 900 may be performed by log processor 130, orchestrator 106, and / or some other component. Provisioning data 300 is retrieved 902. Retrieving 902 may include pulling provisioning data from a manifest ingested by orchestrator 106 and pulling log files 200 from components as described above with respect to FIG. 2. Retrieving 902 may include an enrichment step in which data from the manifest and / or log files 200 is processed by vector log agent 126 to add additional information, perform map / reduce operations, or perform other operations. For example, enrichment may include adding an identifier of the source of log file 200, a directory location of log file 200, or other data to facilitate associating data in log file 200 with a particular component identifier. Retrieving 902 may include processing the manifest and / or log file 200 according to one or more VRL statements.

[0053] The method 900 may include extracting 904 the hosting relationships 700. Extracting the hosting relationships 700 may include extracting 904 the hosting relationships 700.<instantiation instruction> …<host component identifier> …<hosted component identifier> ". For example, there may be a set of keywords that indicate instantiations that can be identified, and lines of code or log messages that contain these keywords can be processed to obtain identifiers for the host component and the hosted component. Hosting relationships 700 can then be created that reference the identifiers for the host component and the hosted component.

[0054] Extracting 904 the hosting relationships may further include removing hosting relationships 700 where a hosted component or host component has been removed. Log messages containing instructions to remove a component may be identified, an identifier for the removed component may be extracted, and any hosting relationships 700 that reference the identifier for the removed component may be removed.

[0055] The method 900 may include extracting 906 the environmental variable relationships 702. Extracting 906 the environmental variable relationships 702 may include extracting 906 the environmental variable relationships 702.<configuration instruction> …<configured component identifier> …<referenced component identifier> ". For example, a set of keywords may be found within a statement or log message regarding the setting of an environment variable. These keywords can be identified, and the lines of code or log messages containing these keywords can be processed to identify configured components, i.e., the components whose environment variable(s) are set, and referenced components, i.e., the identifiers of components referenced by the configured component's environment variables. Environment variable relationships 702 can then be created that reference the identifiers of the configured component and the referenced component, as well as one or more environment variables of the configured component that are possibly configured to reference the referenced component.

[0056] A statement in log file 200 that creates an environment variable relationship 702 may modify a previously existing environment variable relationship. For example, environment variable relationship 702 may record the name of an environment variable for a configured component. A first environment variable relationship 702 for a configured component that includes a variable name may be deleted in response to a subsequently identified environment variable relationship 702 for the configured component that references the same variable name. An exception to this approach may be implemented if an environment variable can store multiple values. For example, an explicit delete command including the variable name, configured component identifier, and referenced component identifier is required before deleting an environment variable relationship 702 that includes the variable name, configured component identifier, and referenced component identifier.

[0057] The method 900 may include extracting 908 the network relations 704. Extracting 908 the network relations 704 may include extracting 908 the network relations 704.<network configuration instruction> …<configured component identifier> …<IP address,domain name,URL,etc.> " and sentences of the form " <address assignment instruction>…<referenced component identifier> …<IP address,domain name,URL,etc.> " and statements of the form "configure_address_number"," which may be located in different locations within the manifest or log file 200. For example, a set of keywords may be found within a statement or log message related to assigning a network address to a referenced component and configuring a component configured to communicate with the referenced component's address. These keywords can be identified, and lines of code or log messages containing these keywords can be processed to identify the network addresses and identifiers of the configured and referenced components; i.e., the referenced component is the component to which the network address is assigned, and the configured component is the component configured to send data to and / or receive data from the referenced component using the network address. Network relations 704 can then be created that reference the identifiers of the configured and referenced components, and possibly include network addresses. Additional information can include the protocol used, port numbers, and network relations (e.g., whether the referenced component acts as a network gateway, proxy, etc.).

[0058] A statement in log file 200 may change the configuration of a configured component such that the configured component is configured to use a different referenced component's network address. Such a statement may be parsed, and a new network relationship 704 may be created in a manner similar to that described above. A previously created network relationship 704 for a configured component may be deleted or may continue to exist. For example, there may be an explicit instruction to delete the configured component's configuration to use the network address of the referenced component referenced by the previously created network relationship 704. In response to recording the execution of such an instruction, the previously created network relationship 704 may be deleted.

[0059] FIG. 10 illustrates a method 1000 for extracting session relationships 706. Method 1000 may be performed by log processor 130, orchestrator 106, and / or some other component. Method 1000 includes retrieving 1002 log files 200. Retrieving 1002 log files 200 may include pulling log files 200 from a component as described above with respect to FIG. 2. Retrieving 1002 may include an enrichment step in which data from log files 200 is processed by vector log agent 126 to add additional information, perform map / reduce operations, or perform other operations. For example, enrichment may include adding an identifier of the source of log file 200, a directory location of log file 200, or other data to facilitate associating data in log file 200 with a particular component identifier. Retrieving 1002 may include processing log files 200 according to one or more VRL statements.

[0060] The method 1000 may include obtaining 1004 a session setup message from the log file 200 before or after any enrichment of the log file 200. The session setup message may be a message indicating that the session has been successfully initiated and may include identifiers of the server component (i.e., the component providing the service) and the client component (i.e., the component requesting the service).

[0061] Method 1000 may include retrieving 1006 a session termination message from log file 200 before or after any enrichment of log file 200. The session termination message may be a message indicating that a session has terminated in response to either an instruction from a client component, an instruction from a server component, expiration of a timeout period, failure of an intermediate component or network connection between the client and server components, restart or failure of the client or server component, or other cause. The session termination message may also include identifiers of the server component (i.e., the component providing the service) and the client component (i.e., the component requesting the service). If the session termination is due to a failure (network connection, intermediate component, client component, or server component), only the server component or the client component may be referenced by the log message. In such a case, all session relationships referencing the components referenced in the log message may be considered terminated and may be deleted.

[0062] The method 1000 may include updating 1008 the session relation 706 by adding a session relation 706 corresponding to the session identified to be created in the setup message. The session relation 706 may include identifiers of the server and client components, and may include other information such as a timestamp from the setup message, an identifier for the session itself, a type of session, or other data.

[0063] Updating 1008 the session relationship 706 may include deleting the session relationship 706 corresponding to a session identified as terminated in a session closure message (including a message indicating a failure). For example, if a session has a unique session identifier, the session relationship 706 including the session identifier included in the session closure message may be deleted. Alternatively, if the session closure message references a set of a client identifier and a server component identifier, the session relationship 706 including the same client identifier and server component identifier may be deleted. In some implementations where sessions have a known time to live (TTL), the session relationship 706 may be deleted based on the expiration of the TTL, regardless of whether a session closure message corresponding to the session relationship has been received.

[0064] FIG. 11 illustrates a method 1100 for extracting access relationships 708. Method 1100 may be performed by log processor 130, orchestrator 106, and / or some other component. Method 1100 includes retrieving 1102 log file 200. Retrieving 1102 log file 200 may include pulling log file 200 from a component as described above with respect to FIG. 2. Retrieving 1102 may include an enrichment step in which data from log file 200 is processed by vector log agent 126 to add additional information, perform map / reduce operations, or perform other operations. For example, enrichment may include adding an identifier of the source of log file 200, a directory location of log file 200, or other data to facilitate associating data in log file 200 with a particular component identifier. Retrieving 1102 may include processing log file 200 according to one or more VRL statements.

[0065] The method 1100 may include extracting 1104 access relationships 708 from the log file 200 either before or after enriching the log file 200. The access relationships 708 may be identified in various ways, such as by parsing log messages in the log file 200 of a server component (i.e., a component that provides a service) indicating requests from a client component (i.e., a component that requests a service), log messages in the log file 200 of the client component indicating requests from a client component to a server component, or log messages of another component that stores the results of access requests from the client component to the server component. The access relationships 708 may include an identifier of the server component, an identifier of the client component, and one or more timestamps or other metadata for one or both of (a) each request from the client component to the server component and (b) each response from the server component to the client component.

[0066] The method 1100 may include identifying 1106 an outdated access relationship 708. An outdated access relationship 708 may be defined as one whose most recent timestamp (for a request and / or response) is older than a threshold time, e.g., 1 minute, 5 minutes, 1 hour, 1 day, etc. The threshold time may be unique for each type of component; for example, an instance 118 of one application may have a different threshold than an instance of another application. The threshold time may be automatically derived as a multiple of the average time between requests for each client of the server component.

[0067] Method 1100 may then include updating 1108 the access relations 708 to add the access relations detected in step 1104. Updating 1108 the access relations may include deleting expired access relations. Updating 1108 the access relations may include merging 708 the access relations. For example, if a pair of access relations 708 reference the same server and client component identifiers, the access relations 708 may be combined into a single access relation 708 that includes the most recent timestamps of the pair of access relations 708. The access relations 708 may include records of access requests and / or responses between the client and server component such that upon merging, the records of the pair of access requests are combined. Alternatively, each access relation 708 includes statistical characteristics of past requests and / or responses, and the merged access request includes a combination of the statistical characteristics of the pair of access relations 708. In some embodiments, merging is performed before identifying 1106 expired relations.

[0068] FIG. 12 illustrates a method 1200 for extracting network connection relationships 710. Method 1200 may be performed by log processor 130, orchestrator 106, and / or some other component. Method 1200 includes retrieving 1102 log file 200. Retrieving 1202 log file 200 may include pulling log file 200 from a component as described above with respect to FIG. 2. Retrieving 1202 may include an enrichment step in which data from log file 200 is processed by vector log agent 126 to add additional information, perform map / reduce operations, or perform other operations. For example, enrichment may include adding an identifier of the source of log file 200, a directory location of log file 200, or other data to facilitate associating data in log file 200 with a particular component identifier. Retrieving 1202 may include processing log file 200 according to one or more VRL statements.

[0069] The method 1200 may include obtaining 1204 a connection setup message from the log file 200 before or after any enrichment of the log file 200. The session setup message may be a record of an exchange of handshake messages or other messages indicating that a network connection has been successfully established between a first component and a second component.

[0070] Method 1200 may include obtaining 1206 a connection termination message from log file 200 before or after any enrichment of log file 200. The connection termination message may be a message indicating that the network connection has been terminated in response to either an instruction from a client component, an instruction from a server component, expiration of a timeout period, failure of an intermediate component, or failure of the network connection between the client component and the server component, or other cause. In some implementations, the session termination message may include a message indicating failure of a physical link between a first component and a second component, a restart of the first component or the second component, and a failure or restart of a component hosting the first component or the second component.

[0071] The method 1000 may include identifying 1208 an expired network connection relationship 710. Identifying 1208 an expired network connection relationship 710 may include identifying (a) a component pair for which a network connection relationship 710 exists, (b) a component pair for which there is no current network connection as indicated by a connection termination message, and (c) a component pair for which a predetermined period of time has expired since the last connection termination message was received for the component pair. With regard to (c), some connections have a predetermined TTL such that the network connection relationship 710 expires when a predetermined period of time greater than the TTL has expired since the last connection setup message for the component pair.

[0072] As an alternative to the above approach, all network connection relationships 710 expire as soon as the network connection represented by the network connection relationship 710 terminates due to TTL expiration or explicit termination as indicated in the connection termination message.

[0073] Method 1200 may include updating 1210 the network connection relationship 710 by deleting any expired network connection relationship 710 and adding a new network connection relationship 710 indicated by the connection setup message from step 1204. It is possible for a first component and a second component to have multiple network connection relationships, such as connections to different ports by different applications. Thus, a separate network connection relationship 710 may exist for each network connection, or a single network connection relationship 710 may be created to represent all network connections between the pair of components. The network connection relationship 710 may include data describing each connection (such as a setup timestamp, protocol, and port). This data may be updated to remove data describing a connection when that connection is terminated. Similarly, the network connection relationship 710 may be updated to add data describing a connection between the pair of components represented by the network connection relationship 710 when the connection is set up.

[0074] Referring to Figures 13 and 14A-14D, the illustrated method 1300 can be used to generate a visual representation of a topology that is displayed on a display device, such as a user device 110, via an orchestrator dashboard 108.

[0075] 13 and 14A , method 1300 may be performed by orchestrator 106, and visual representation 1400 may be provided to user computing device 110 via orchestrator dashboard 108. User computing device 110 may then display visual representation 1400, receive user interactions with visual representation 1400, and report the user interactions to orchestrator 106 for processing. A user may request generation of visual representation 1400 via orchestrator dashboard 108. Retrieval and processing of provisioning data 300 and log files 200 to generate the visual representation may be performed in response to a request from a user.

[0076] The method 1300 may include extracting 1302 component identifiers from the provisioning data, as described above. Each component identifier is then used as a node in the graph. The method 1300 may then include adding 1304 edges between nodes for hosting relationships 700 between the component identifiers represented by the nodes. The method 1300 may include adding 1306 edges between nodes for environmental variable relationships 702 between the component identifiers represented by the nodes. The method 1300 may include adding 1308 edges between nodes for network relationships 704 between the component identifiers represented by the nodes. The method 1300 may include adding 1308 edges between nodes for network relationships 704 between the component identifiers represented by the nodes. The method 1300 may include adding 1310 edges between nodes for session relationships 706 between the component identifiers represented by the nodes. The method 1300 may include adding 1312 edges between nodes for access relationships 708 between the component identifiers represented by the nodes. The method 1300 may include adding 1314 edges between the nodes for the network connection relationships 710 between the component identifiers represented by the nodes. The relationships between components described herein are merely exemplary, and the method 1300 may include adding edges for other types of relationships between components.

[0077] A visual representation 1400 of the topology represented by the graph may then be displayed 1316. An example visual representation 1400 is shown in Figure 14. Graphical elements may be displayed to represent components such as hosts 802, pods 112, containers 114, storage volumes 116, and application instances 118. The graphical elements may include images and / or text, such as the UUID of each component.

[0078] The visual representation 1400 may include lines 1402 between the graphical elements representing the components, where the lines 1402 represent the edges of the graph. The lines 1402 may be color-coded, with each color representing a type of relationship 702-710. A pair of components may have multiple relationships, such as some or all of an environment variable relationship 702, a network relationship 704, a session relationship 706, an access relationship 708, and a network connection relationship. A separate line 1402 may be displayed to represent each type of relationship, or a single line may represent all of the relationships between the components represented by the pair of graphical elements.

[0079] The graphical elements or lines 1402 may be augmented with additional visual data describing the components or relationships represented by the graphical elements or lines 1402. For example, the additional visual data may be displayed upon clicking on the graphical elements or lines 1402, upon hovering over the graphical elements or lines 1402, or upon other interaction. The additional data may be collected from the log 200 and may include component usage and / or AAI data, as described above.

[0080] For example, for a graphical element representing a host 802, the additional data may include AAI data for the host, such as status 1404 (up, critical, down, unreachable, etc.), available and / or in-use computing power 1406 (processor cores, processor time, processor cycles, etc.), available and / or in-use memory 1408, and available and / or in-use storage 1410. For a graphical element representing a storage volume 116, the additional data may include status 1412, available storage 1414 and / or storage usage, and IOP (input / output operations) usage 1416 and / or availability. For a graphical element representing a cluster 111, pod 112, container 114, or application instance 118, the additional data may include status 1418, computing power usage 1420, memory usage 1422, and storage usage 1424. For a cluster 111 and / or a pod 112, computing power usage 1420, memory usage 1422, and storage usage 1424 may be an aggregate of the computing resources used by all containers 114, application instances 118, and storage volumes 116 managed by the cluster 111 and / or pod 112, as well as the cluster 111 and / or pod 112 itself.

[0081] For line 1402, the additional data may include data describing one or more relationships represented by line 1402, such as a list of each type of relationship 700-710 represented by the line, the status 1426 of each relationship, and the usage of each relationship 1428. The usage of a relationship may include, for example, the amount of data sent over a network connection, the number or frequency of requests for a session or access relationship, the latency of the network connection, the latency of responses to requests for a session or access relationship, or other data.

[0082] The graphical elements or lines 1402 may also be augmented with, for example, an action menu 1430 responsive to user interaction with the graphical elements or lines 1402. The action menu 1430 may include graphical elements that, when selected by a user, invoke one or both of: (a) an action to modify information shown in the visual representation 1400; and (b) an action to perform an action with respect to the component represented by the graphical element or line 1402. For example, the action menu 1430 may include elements for deleting the component, restarting the component, creating a relationship 702-710 between the component and another component, creating a snapshot or backup copy of the component, duplicating the component, duplicating the component, or invoking other actions. Thus, the method 1300 may include receiving 1318 an interaction with the visual representation 1400 of the topology and, in response, performing an action, such as modifying 1320 the information displayed in the visual representation and / or modifying the component represented by the visual representation of the topology. Actions invoked with respect to a component may be performed with respect to other components, such as those hosted by the component. For example, an action invoked with respect to a cluster 111 may be performed with respect to all pods 112 , containers 114 , storage volumes 116 , and application instances 118 hosted by the cluster 111 .

[0083] 14B shows an application browsing interface 1432 that may be displayed to a user, such as using data obtained according to method 1300 or some other technique. The application browsing interface 1432 may include one or more cluster elements 1434 that represent the cluster 111. A user may select one of the cluster elements 1434 to invoke a display of additional information about the cluster 111. One or more namespace elements 1436 may be displayed, such as a list of names in the namespace of the cluster 111, where each name represents a component 112, 114, 116, 118 of the cluster 111 or another variable, service, or other entity accessible to a component of the cluster 111. The interface 1432 may display a selector element 1438 that allows a user to enter criteria for filtering or selecting names from the namespace of the cluster 111. For example, the user can select based on version (e.g., which HELM release of Kubernetes the component belongs to or was deployed by), type of application (database, web server, etc.), executable image, instantiation data, or any other criteria.

[0084] For each application instance 118 that meets the criteria entered by the user into selector element 1438, application viewing interface 1432 can display various items of information for the application instance 118. Exemplary information items can include daemon set 1440a, deployment data 1440b, stateful set 1440c, replica set 1440d, config map 1440e, one or more secrets 1440f, or other data 1440g. Some or all of the items can be selected by the user to invoke the display of additional data. For example, the user can invoke the display of pod data 1442 for the pod 112 that hosts the application instance 118, container data 1444 describing the container 114 that hosts the application instance 118, persistent volume claim (PVC) data 1446 for the storage volume 116 accessed by the application instance 118, and volume data 1448 describing the storage volume 116 accessed by the application instance 118.

[0085] For each element selectable in the application viewing interface 1432, selecting the element can invoke a display of elements associated with that element, and can also invoke a display of real-time data for each element, such as observability data for each element (e.g., log data 200) that may be collected, processed (aggregated, formatted, etc.), and displayed as observability data is generated for each element.

[0086] 14C illustrates yet another interface 1450 that may be used to visually represent the topology and receive user input to invoke the display of additional information about a cluster 111, a host 1452 executing one or more components of the cluster 111, and a storage device 1454 of one of the hosts 1452. The interface 1450 may include a cluster element 1456 representing the cluster 111, a namespace element 1458 representing the namespace of the cluster 111, a composite application element 1460 representing two or more application instances 118 that together define a bundled application, and a single application element 1462 representing the single application instance 118.

[0087] Selecting a given element 1456, 1458, 1460, 1462 can invoke a display of additional information; selecting cluster element 1456 can invoke a display of namespace elements 1458; selecting a name from namespace elements 1458 can invoke a display of composite application elements 1460; selecting a name from composite application elements 1460 can invoke a display of single application elements 1462.

[0088] Selecting a single application element 1462 can invoke a display of data describing the application instance 118 represented by the single application instance 118. For example, the data can include other data such as element 1464 indicating config map data, element 1466 indicating various sets (replica set, deployment set, stateful set, daemon set, etc.), element 1468 indicating secrets, or any observability data for the application instance 118.

[0089] Selecting elements 1462, 1464, 1466 can invoke the display of additional data, such as a pod element 1470 containing data describing the pod 112, a PVC element 1480 describing the PVC, and a volume element 1482 describing the storage volume 116 (such as data describing the amount of data used by the storage volume 116 and the storage device that stores the data for the storage volume 116).

[0090] Interface 1450 can be used to assess the criticality of components of cluster 111. For example, selecting namespace element 1458 can invoke a display of aggregated data 1484, such as aggregated logs (e.g., log files combined by chronologically ordering the messages in the log files), aggregated metrics (aggregated processor usage, memory utilization, storage utilization), aggregated alerts and / or events (e.g., events and / or alerts combined and ordered by time of occurrence), aggregated access logs (e.g., allowing for tracking of user behavior with respect to cluster 111 or components of cluster 111), etc. Aggregated data 1484 can be used in combination with topology data, such as described in U.S. Patent Application No. 16 / 561,994, filed September 5, 2019, and entitled "PERFORMING ROOT CAUSE ANALYSIS IN A MULTI-ROLE APPLICATION," the contents of which are incorporated herein by reference in their entirety, to perform root cause analysis (RCA).

[0091] Selecting a single application element 1462 can invoke a display of importance 1486 for the application instance 118 represented by the single application element 1462. The importance 1486 can be a metric that is a function of several other application instances 118 that depend on the application instance 118, for example, having relationships 700-710 with the application instance 118. The importance 1486 can include a "blast radius" for the application instance 118 (see FIG. 14D and corresponding discussion).

[0092] Selecting a pod element 1470 can invoke a display of the pod density 1488 (e.g., number of pods) of the host running the pod 112 represented by the pod element 1470. The pod density 1488 can be used to determine the importance of the host and whether the host may be overloaded.

[0093] Selecting the PVC element 1480 can invoke a display of the volume density 1490 (e.g., number of storage volumes 116, total size of storage volumes 116) stored on the storage device or individual storage devices of the host. The volume density 1490 can be used to determine the criticality of the host and whether the host's storage devices may be overloaded.

[0094] FIG. 14D illustrates yet another interface 1492 that can be used to visually represent the topology. Interface 1492 can include visual representations of the illustrated components. Storage device 1494 (e.g., hard disk drive, solid state drive) stores data for storage volume 116. Storage volume 116 is used by application instance 118, which can have one or more relationships, e.g., relationships 700-710, with other application instances 118. The other application instances 118 themselves have relationships 700-710 with other application instances. In particular, one or more application instances 118 that are not running on the same host as the storage volume may be represented in interface 1492. Interface 1492 can also include a "blast radius" representation that indicates the impact that a failure of storage device 1494 would have on other application instances 118 or other components of cluster 111 that includes storage volume 116 or one or more other clusters 111.

[0095] 15-19, using AAI, computing resources allocated to components in network environment 100 may be reduced based on the usage of computing resources by application instance 118. Cloud computing platform 104 may charge for computing resources purchased regardless of actual usage. Thus, AAI may be used to identify modifications to the deployment of application instances to reduce purchased computing resources.

[0096] 15, the orchestrator 106 or another component may calculate a cluster host inventory 1502a-1502c for each of the multiple clusters 111a-111c. The cluster host inventory 1502a-1502c is the number of processing cores, amount of memory, and amount of storage on the servers 102 allocated to the particular cluster 111a-111c. For the cloud computing platform 104, the cluster host inventory 1502a-1502c may include the amount of cloud computing platform computing power, memory, and storage allocated to the cluster 111a-111c.

[0097] The orchestrator 106 or another component may further calculate cluster provisioning 1504a-1504c for each cluster 111a-111c. The cluster provisioning 1504a-1504c is the computing resources (computing power, memory, and / or storage) allocated to components (e.g., pods 112a-112c, containers 114, storage volumes 116, or application instances 118a-118l) within the clusters 111a-111c. In some cases, the cluster provisioning 1504a-1504c is identical to the cluster host inventory 1502a-1502c and is omitted. In other examples, the cluster provisioning 1504a-1504c includes the computing resources allocated to individual components (pods 112a-112c, storage volumes 116, application instances 118a-118l) of the clusters 111a-111c.

[0098] The orchestrator 106 or another component may further calculate cluster usage 1506a-1506c for each cluster 111a-111c. The cluster usage 1506a-1506c for a cluster 111a-111c may include, for each computing resource (computing power, memory, storage), the sum of the usage of that computing resource by all components within the cluster 111a-111c, including the cluster itself. The cluster usage 1506a-1506c may be obtained from the log file 200, as described above. The cluster usage 1506a-1506c for a cluster 111a-111c may include a list of the amount of each computing resource used by the individual components of the cluster 111a-111c and the cluster 111a-111c itself.

[0099] The orchestrator 106 or another component can further calculate a cluster AAI 1508a-1508c for each cluster 111a-111c. The cluster AAI 1508a-1508c can include the AAI(t), O(t,i), and O(t) calculated as above, except that the hardware inventory is limited to the cluster host inventory 1502a-1502c and only the usage of components within the cluster 111a-111c and the cluster itself is used in the calculation.

[0100] FIG. 16A is a simplified diagram of available computing resources and their usage. Each bar in FIG. 16A represents either the amount of computing resources (hardware inventory 1502a-c, cluster AAI 1508a-c) or the usage of computing resources (application instance 118a-c). The depicted representation is simplified in that only one computing resource is shown, omitting other usages (pods 112a-c, storage volume 116, clusters 111a-c themselves), although these usages and computing resources may actually be included. As can be seen, each cluster has a cluster AAI 1508a-c of computing resources that represents the difference between the cluster host inventory 1502a-c and the usage by the various components of each cluster 111a-c.

[0101] Continuing with reference to FIG. 16A, and with reference to FIG. 16B, one or more components may be redeployed from one cluster 111a-111c to another cluster. For example, application 118d on cluster 111a consumes significantly more computing resources than the other applications 118a-118c on cluster 111a. In contrast, cluster 111b has cluster AAI 1508b with sufficient computing resources to host application 118d. Therefore, application 118d may be redeployed on cluster 111b.

[0102] In a cloud computing environment 104 in which computing resources are virtualized, the amount of cluster host inventory 1502a-1502c for some or all of the clusters 111a-111c can be reduced, thereby reducing the amount charged for the cluster host inventory 1502a-1502c. In particular, because the usage of the cluster host inventory 1502a is significantly reduced by removing the usage of the application instance 118d, significant cost savings can be achieved by reducing the cluster host inventory 1502a.

[0103] Redeployment of application instance 118d to another cluster 111b may be contingent on satisfying one or more constraints. Failure to satisfy a constraint may prevent redeployment. For example, there may be a requirement that the receiving cluster 111b have a sufficient amount of computing resources (computing power, memory, and storage) to receive application instance 118d. There may be a requirement that moving application instance 118d to cluster 111b does not violate any affinity requirements with respect to application instances 118a-118c remaining on the original cluster 111a. There may be a constraint that moving application instance 118d to cluster 111b does not violate any anti-affinity requirements with respect to application instances 118e-118h running on the receiving cluster 111b. Redeployment of application instance 118d to receiving cluster 111b may also include adding application instance 118d to pods 112c, 112d of receiving cluster 111b or creating a new pod on receiving cluster 111b.

[0104] Redeploying an application instance 118, such as application instance 118d in the illustrated example, may include redeploying the application instance 118 from server 102 to cloud computing platform 104, or vice versa. For example, application instance 118d may be hosted on cloud computing platform 104 and may be moved to server 102 because application instance 118d is using an amount of computing resources above a threshold and would have better performance if hosted locally on server 102 and lower costs if charges from cloud computing platform 104 for application instance 118d are eliminated. Similarly, an application instance 118 with usage below a minimum threshold may be moved from server 102 to the cloud to provide local computing resources on server 102 to an application instance on cloud computing platform 104 with usage above a maximum threshold.

[0105] 16C, in another example, cluster consolidation may be performed by moving all application instances 118a-118d, which may be deployed to one or more other clusters 111b, 111c, subject to any affinity and anti-affinity constraints and provided the other clusters 111b, 111c have sufficient cluster AAIs 1508b, 1508c. In that case, the entire cluster host inventory 1502a may be deleted, along with the corresponding costs of the cluster host inventory 1502a.

[0106] Figure 17A illustrates an example method 1700a that may be performed by the orchestrator 106 or another component to redeploy an application instance 118 to a different cluster 111. To facilitate understanding of the method, reference is made to the components illustrated in Figure 15 as a non-limiting example. In particular, any number of clusters 111 hosting any number of components may be processed according to method 1700a.

[0107] The method 1700a may include determining 1702 usage and cluster AAIs for each cluster 111, such as component usage 1506a-1506c for the plurality of clusters 111a-111c and cluster AAIs 1508a-1508c for the plurality of clusters 111a-111c. The method 1700a may include identifying 1704 candidate redeployment. The identifying candidate redeployment 1704 may be limited to evaluating usage of the application instance 118 with respect to the cluster AAI of the cluster 111 to determine whether redeployment is possible. The candidate redeployment may include transferring a particular application instance 118 (e.g., application instance 118d) to a receiving cluster 111 (e.g., cluster 111b) that has a sufficient cluster AAI to receive the application instance 118. A candidate redeployment may include replacing a first application instance 118 on a first cluster 111 with a second application instance 118 on a second cluster 111, where the second cluster has a larger host AAI than the first cluster and the first application instance 118 has a larger usage than the second application instance 118. A candidate redeployment may include removing the first application instance 118 on the first cluster 111, where the second application instance 118 on the second cluster 111 is in a load balancing relationship with the second application instance 118, where the second cluster 111 has a cluster AAI sufficient to receive the usage of the first application instance 118, and possibly a larger cluster AAI than the first cluster 111. When identifying 1704 candidate redeployment, multiple application instances 118 of a cluster 111 that have affinity constraints with respect to each other may be treated as a unit, i.e., the receiving cluster 111 must have sufficient cluster AAIs to receive all of the multiple application instances 118.

[0108] Method 1700a may include filtering 1706 candidate redeployments based on constraints, such as anti-affinity requirements, latency requirements, or other requirements. For example, if redeployment of application instance 118d to cluster 111b would violate the anti-affinity constraint of application instance 118d with respect to application instance 118e, then such redeployment of application instance 118d is filtered out at step 1706. Similarly, if redeployment of application instance 118d to cluster 111b would exceed the minimum latency required for application instance 118d with respect to application instances 118i-118l in cluster 111c, then such redeployment is filtered out at step 1706. The anti-affinity and latency requirements are exemplary only, and other constraints may be imposed at step 1706.

[0109] The method 1700a may include calculating 1708 a billing reduction achievable by the candidate redeployment, i.e., how much the candidate redeployment would reduce the cluster host inventory 1502a-1502c of the modified cluster if the candidate redeployment were performed. If the billing reduction is found to be greater than a minimum threshold, the candidate redeployment is implemented 1712 by performing a transfer, replacement, or deletion of the candidate redeployment. A redeployment involving moving an application instance 118 from a first cluster 111 to a second cluster 111 may include installing the new application instance 118 on the second cluster (creating a container and installing the application instance 118 in the container), stopping the original application instance 118 on the first cluster 111, and starting the new application instance 118 running on the second cluster 111. Other configuration changes may be required to configure other components to access the new application instance 118 on the second cluster 111.

[0110] Method 1700a may further include reducing 1714 the amount of cloud computing resources used by one or more clusters 111. For example, in the example of FIG. 16B , the computing resources allocated to cluster 111a may be reduced after application instance 118d is redeployed to cluster 111b. The reduction amount may be such that the cluster AAI of each cluster 111 is reduced to a zero or non-zero threshold (e.g., a percentage of the usage of the components deployed to each cluster) for one or more computing resources (computing power, memory, storage), assuming that the usage of the cluster's components after the redeployment remains the same as the usage values ​​used to calculate the cluster AAI of cluster 111.

[0111] 17B illustrates an alternative method 1700b for redeploying an application instance 118. The method 1700b may be performed by the orchestrator 106 or other component to redeploy the application instance 118 to a different cluster 111.

[0112] The method 1700a may include determining 1702 a usage amount and a cluster AAI for each cluster 111, such as component usage amounts 1506a-1506c for the plurality of clusters 111a-111c and cluster AAIs 1508a-1508c for the plurality of clusters 111a-111c.

[0113] The method 1700a may include re-planning 1704 the placement of components using the components' computing resource usage instead of provisioning requirements. When initially instantiating the components 111, 112, 114, 116, 118 in the network environment 100, the orchestrator 106 may perform a planning process to place the components based on required computing resources, affinity requirements, anti-affinity requirements, latency requirements, or other requirements. The orchestrator 106 further attempts to improve the performance of components working together by reducing latency and using computing resources as efficiently as possible.

[0114] As an example, the orchestrator 106 may use a planning algorithm such as that disclosed in U.S. Patent No. 10,817,380 B2, filed October 27, 2020, entitled "IMPLEMENTING AFFINITY AND ANTI-AFFINITY CONSTRAINTS IN A BUNDLED APPLICATION," the contents of which are incorporated herein by reference in their entirety. In contrast to the initial planning, the provisioning requirements in step 1716 for each component may be set to be the compute resource usage measured for each component as described above using log data pulled from the component's host. Alternatively, the provisioning requirements may be set to an intermediate value between the component's provisioning as defined by the manifest and the measured usage for that component, such as usage scaled by a number greater than 1, such as a number between 1.1 and 2.

[0115] The result of step 1716 may be one or more plans that define where each component is to be placed (e.g., on which server 102 or which unit of computing resources of the cloud computing platform, which pod 112, which cluster 111, etc.). The billing savings achieved by each plan may be calculated 1708 and evaluated 1710 to determine whether the plan provides at least a threshold reduction in computing resource allocation over the current configuration of components based on the usage of each component measured in step 1702. As discussed above, reducing the computing resource allocation results in a reduction in the costs of the cloud computing platform 104.

[0116] If so, one of the plans, such as the plan that provides the greatest cost savings, may be implemented 1712. Implementing the plans 1712 may include moving the components one at a time to locations defined in the plans to avoid disruptions, or pausing all components, redeploying the components as defined in the plans, and restarting all components. Redeploying each component may be performed as described above with respect to step 1712 of method 1700a.

[0117] Following or during performing 1712 the redeployment, method 1700b may include reducing cloud computing resources 1714 allocated from the cloud computing platform. The reduction may be such that the cluster AAI of each cluster 111 is reduced to a zero or non-zero threshold (e.g., a percentage of usage of each cluster after redeployment) for one or more computing resources (computing power, memory, storage), assuming that usage of the cluster's components after redeployment remains the same as the usage values ​​used to calculate the cluster AAI of the cluster 111.

[0118] Figure 18 illustrates an alternative method 1800 for redeploying application instances 118 to consolidate the number of clusters 111 of the original configuration, such as the illustrated reduction of clusters shown in Figures 16A and 16C. Method 1700b may be performed by orchestrator 106 or another component.

[0119] The method 1800 may include determining 1802 usage amounts and cluster AAIs for each cluster 111 of the original configuration, such as component usage amounts 1506a-1506c for the plurality of clusters 111a-111c and cluster AAIs 1508a-1508c for the plurality of clusters 111a-111c.

[0120] Method 1800 may include attempting to identify consolidations 1804. Consolidation is the placing of components of multiple clusters into a subset of the multiple clusters, where one or more clusters of the multiple clusters and one or more hosts of the multiple clusters are excluded. Methods for attempting to identify consolidations are described below with respect to FIG. 19.

[0121] If a consolidation is found 1806, the consolidation can be implemented 1808. If multiple consolidations are found, the consolidation that achieves the highest cost savings can be implemented 1808. The consolidation can include a plan that defines the location of each component on the remaining cluster 111. Thus, the components may be re-instantiated, configured, and started on the remaining cluster. In some embodiments, only components that are in a different location in the plan relative to their original configuration are redeployed to a different location. The original components may be shut down while the consolidation is implemented. Alternatively, components may continue to operate and be migrated one at a time until the plan is implemented 1808.

[0122] The computing resources allocated to the cluster that is removed as part of performing consolidation 1808 can be reduced 1810. For on-premise equipment, the server 102 can be taken offline or allocated to other uses. Payment for the use of one or more units of cloud computing resources allocated to the removed cluster can be terminated for units of cloud computing resources on the cloud computing platform 104, or other action can be taken to terminate acquisition of one or more units of cloud computing resources.

[0123] 19 illustrates a method 1900 that may be used to identify potential cluster consolidations. Method 1900 may be performed by orchestrator 106 or another component. Method 1900 may include treating 1902 each cluster 111 as a “target cluster” by replanning 1904 without the target cluster 111, i.e., without the cluster host inventory currently assigned to the target cluster 111. Replanning may be performed with respect to the cluster host inventories of clusters 111 other than the target cluster 111 (“remaining clusters”), as described above with respect to step 1716 of method 1700b. As described above, re-planning may include identifying locations for each component on the hosts of the remaining cluster using a planning algorithm such as that disclosed in U.S. Pat. No. 10,817,380 B2 such that each component is allocated computing resources at least as great as its usage, and such that the location of each component satisfies any affinity, anti-affinity, latency, or other requirements with respect to the locations of other components.

[0124] If it is found 1906 that no plan exists to eliminate the target cluster 111, then the method 1900 ends for the target cluster 111. If it is found that one or more plans exist, then each plan is added 1908 to a set of candidate mergers.

[0125] After processing each cluster 111 as a target cluster, if one or more plans are found to eliminate the target cluster, method 1900 may be recursively repeated using the set of clusters 111 excluding the target cluster. For example, assume there are clusters 111a-111f, and a plan is found to eliminate the cluster host inventory of cluster 111a. Method 1900 may be repeated to determine whether the cluster host inventory of any of clusters 111b-111f can be eliminated. This process may be repeated until method 1900 identifies no more possible consolidations.

[0126] After treating each cluster 111 as a target cluster and performing any recursive iterations, the result is either no possible candidate mergers or a set of one or more candidate mergers. If there are multiple candidate mergers, then in step 1808 the candidate merger that provides the greatest billing reduction may be selected for implementation.

[0127] Referring to FIG. 20, as noted throughout the above description, the topology is dynamic. The components of topology 2000 (clusters 111, pods 112, containers 114, storage volumes 116, and application instances 118) can change at any time. Sources of change can include automatic scaling up or down of components based on usage by orchestrator 106, such as using tools like KUBERNETES. In particular, for each cluster 111, KUBERNETES, alone or in cooperation with orchestrator 106, manages the scaling up or down of the number of pods 112 and corresponding containers 114, storage volumes 116, and application instances. Administrators can also manually add or remove components and relationships between components.

[0128] For example, as indicated by the dotted line representations, pod 112, container 114, storage volume 116, and application instance 118 can be added. Similarly, components and relationships marked with an "X" (represented by line 2002) represent components and relationships between components that can be removed from topology 2000.

[0129] In a production environment where stability is important, modifications to topology 2000 may be prohibited or subject to one or more restrictions to reduce the risk of changes that could cause crashes, overloads, or other types of instability.

[0130] 21 , for example, the illustrated method 2100 may be performed by the orchestrator 106 in cooperation with the orchestrator dashboard 108 or some other component. The method 2100 may include receiving 2102 a topology lock definition, such as from a user device 110 via the orchestrator dashboard 108. The topology lock definition may define the scope of the topology lock, for example, the entire topology, a particular cluster 111 or set of clusters 111, a particular host or set of hosts (servers 102 or units of computing resources on the cloud computing platform 104), hosts located in a particular geographic region or facility, a particular region of the cloud computing platform 104, or other definition.

[0131] A topology lock definition can further include restrictions on specific types of components (clusters 111, pods 112, containers 114, storage volumes 116, application instances 118) or specific types of relationships. With respect to application instances 118, the restrictions can refer to instances of specific executables or classes of executables. Restrictions can specify that for a specific type of component, specific executable instance, or specific type of relationship, (a) the number cannot change, (b) the number cannot increase, (c) the number cannot decrease, or (d) the number cannot increase faster than a specified rate or the number cannot decrease faster than a specified rate.

[0132] Method 2100 may include receiving 2104 a topology policy for each topology lock definition. The topology policy defines actions to be taken to either or both (a) prevent violations of the topology lock definition, or (b) handle violations of the topology lock definition.

[0133] The method 2100 may include configuring 2106 some or all of the orchestrator 106, a workflow in the workflow repository 120, or other component to implement each topology lock definition and its corresponding topology policy.

[0134] For example, a workflow usage to instantiate or de-instantiate (i.e., remove) a component of a certain type may be modified to reference a topology lock and corresponding topology policy that references a component of that type, such that if required according to a corresponding policy, the instantiation or de-instantiation of that type of component is not allowed to complete in violation of the topology lock. In another example, a workflow that would violate a topology lock generates an alert.

[0135] In another example, a container 114 may be configured to reference a container network interface (CNI), a container runtime interface (CRI), or a container storage interface (CSI) that is invoked by the container 114 during instantiation and / or start. Any of the CNI, CRI, and CSI may be an agent of an orchestrator and may be modified to respond to instantiation of a container 114 hosting an application instance 118 modifying a topology lock to either (a) prevent instantiation if required by a corresponding topology policy, or (b) generate an alert.

[0136] The above examples are merely examples of how a topology lock may be enforced, and any other aspect of the instantiation or de-instantiation of a component may be modified to include evaluating whether the instantiation or de-instantiation violates the topology lock and implementing any action required by the corresponding topology policy.

[0137] 22 illustrates a method 2200 for preventing violations of a topology lock with a corresponding topology policy. Method 2200 may be performed by the orchestrator 106, a CRI, a CNI, a CSI, or other component. Method 2200 includes receiving 2202 a request to create a component. Note that a request to delete a component may be processed similarly.

[0138] The request may be evaluated 2204 with respect to the topology lock and corresponding policy. For example, step 2204 may include evaluating whether the request is to create a component in a portion of the topology referenced by the topology lock (e.g., in a particular cluster 111, a particular set of servers 102, a particular region or data center, a particular region of a cloud computing platform, etc.) and whether the component is of a type of component referenced by the topology lock. Step 2204 may include evaluating whether the request to create or delete a component is a prohibited action of the topology lock. For example, if changes are not allowed, the request to create or delete a component is prohibited. If only decreases are prohibited, the request to create a component may be permitted. If increases with a rate limit are allowed, step 2204 may include evaluating whether creating a component would exceed a rate limit. If the request is a request to delete a component and only increases are prohibited, the request to delete the component may be permitted. If decreases with a rate limit are allowed, step 2204 may include evaluating whether deleting the component would exceed a rate limit.

[0139] If the create or delete request is found to be allowed 2206, the request is implemented 2210. If not, the method 2200 may include blocking implementation of the request. Blocking may include one or more of the following: Finishing the workflow required to implement the request. Cause a CNI, CRI, or CSI to prevent the completion of the setup of a component being created or a container that hosts a component being created.

[0140] Note that creating or deleting relationships between components may be handled similarly. A request to create or delete a relationship is evaluated 2204 with respect to one or more topology locks and enforced 2210, or blocked 2208 if not permitted according to the topology locks. Blocking may be implemented using modified workflows, CNIs, CRIs, or CSIs. Blocking may also be performed in other ways, such as blocking network traffic to set up a session relationship 706, an access relationship 708, or a network connection relationship 710.

[0141] 23 illustrates a method 2300 for processing topology locks and corresponding policies. Method 2300 may be performed by orchestrator 106 or other components. Method 2300 may be performed in addition to or instead of method 2200. For example, a policy corresponding to a topology lock may be specified to block changes that violate the topology lock, such that method 2200 is performed. A policy corresponding to a topology lock may be specified to detect a violation of the topology lock after it occurs and to issue an alert or revert the violation, such that method 2300 is performed.

[0142] Method 2300 may include generating 2302 a current topology of the environment, such as according to method 1300 of FIG. 13 or some other approach. Method 2300 may include comparing 2304 the current topology with a previous topology of the environment at a previous time, either at the time of initial instantiation of the environment or at a time after the initial instantiation. For example, the previous topology may be a topology that existed at or before a first time when the topology lock was created, and the current topology is obtained from provisioning data 300 and / or log files 200 generated at a second time after the first time.

[0143] A topology lock may have a scope that is smaller than all of the entire topology (see the description of step 2102 of method 2100). Thus, in step 2304, portions of the current and previous topologies that correspond only to that scope may be compared. The topology lock may be limited to components of a particular type, such that only components of the current topology that have that type are compared in step 2304. If the topology lock references a type of relationship, relationships of that type in the current and previous topologies may be compared.

[0144] Method 2300 may include evaluating 2306 whether the current topology violates one or more topology locks with respect to a previous topology. For example, evaluating whether a new component of a particular type has been added to a portion of the environment (e.g., a cluster 111, a server 102, a data center, a cloud computing area, etc.) may include evaluating whether a new component of a particular type has been added to a portion of the environment (e.g., a cluster 111, a server 102, a data center, a cloud computing area, etc.). For example, component identifiers for each component of each type referenced by the topology lock may be compiled for the current and previous topologies. Component identifiers for the current topology that are not included in the component identifiers for the previous topology may be identified. Similarly, component identifiers for the previous topology that are not in the current topology may be identified if the topology lock prevents deletion.

[0145] If the topology lock references a relationship type, each relationship in the current topology is checked to see if it matches a relationship in the previous topology, i.e., if it has the same component identifier and type as the relationship in the previous topology. A relationship without a corresponding patch in the previous topology may be considered new. Similarly, a relationship in the previous topology that does not have a match in the current topology may be considered deleted. In step 2306, it may be determined whether the new or deleted relationship violates a policy.

[0146] For a topology lock violated in step 2306, method 2300 may include evaluating 2308 a topology policy corresponding to the topology lock. Actions indicated in the topology policy may then be implemented. For example, if a policy is found 2310 to require a change that violates the topology lock to be reverted, method 2300 may include invoking 2314 a workflow to revert the change. The workflow may be a workflow to remove a component or relationship that violates the topology lock. Such a workflow may be the same as the workflow used to remove components or relationships of that type when scaling down due to lack of usage. The workflow may be a series of steps for removing a component or relationship in an orderly and non-disruptive manner, i.e., processing pending transactions and migrating workload to another component. If a component or relationship is deleted in violation of a topology lock, a workflow may reinstantiate the component or relationship. The workflow for reinstantiating a component or relationship may be the same as the workflow used to create an initial instance of a component or relationship of that type or to scale up the number of components or relationships of that type.

[0147] If indicated by the topology policy corresponding to the topology lock, method 2300 may include generating 2312 an alert. The alert may be directed to the user device 110 or the administrator's user account, the individual who invoked the change to the policy that violated the topology lock, or another user. The alert may convey information such as the topology lock that was violated, the number of components or relationships that violated the policy, a graphical representation of the change to the policy (e.g., see the graphical representation in FIG. 20), or other data.

[0148] 24, application instances 118 can have various relationships with respect to one another. As described herein, application instances 118 are categorized as either dot application instances 2400, triangle application instances 2402, line application instances 2404, and graph application instances.

[0149] A dot application instance 2400 is an application instance 118 that does not have relationships (e.g., relationships 700-710) with other application instances 118. For example, an application instance 2400 may be an instance of an application that provides a standalone service. A dot application instance 2400 may be an application instance that does not have a particular type of relationship with respect to other application instances 118. For example, a dot application instance 2400 may lack a hosting relationship 700, an environment variable relationship 702, or a network relationship 704 with another application instance 118. In some embodiments, one or more of a session relationship 706, an access relationship 708, and a network connection relationship 710 may still exist between the dot application instance 2400 and another application instance 118.

[0150] Triangle application instance 2402 includes at least three application instances 118 that all have relationships with one another, such as any of relationships 700-710. Although "triangle application instance" is used throughout, this term should be understood to include any number of application instances 118, where each application instance 118 is dependent on all the other application instances 118.

[0151] In one example of a triangle application instance 2402, the application instances 118 may be replicas of one another, with one of the application instances 118 being a primary replica that processes production requests, and two or more other application instances 118 being backup replicas that mirror the state of the primary replica. Thus, each change to the state of the primary replica must be propagated to and confirmed by each backup replica. Health checks may be performed by the backup replicas with respect to each other and with respect to the primary replica to determine whether a backup replica should become the primary replica. Thus, the above-described relationship between the primary and backup replicas results in a triangle application instance 2402. In the illustrated example, each application instance 118 in the set of triangle application instances 2402 runs on a different cluster 111.

[0152] The line application instance 2404 includes multiple application instances 118 arranged in a pipeline such that an input to a first application instance becomes a corresponding output received as an input to a second application instance, and so on for any number of application instances. As an example, the application instances 118 of the line application instance 2404 may include a web server, a backend server, and a database server. A web request received by the web server may be translated by the web server into one or more requests to the backend server. The backend server may process the one or more requests to request one or more queries to the database server. Responses from the database server are processed by the backend server to obtain responses that are sent to the web server. The web server may then generate a web page including the responses and send the web page as a response to the web request. In the illustrated example, each application instance 118 of the set of line application instance 2404 runs on a different cluster 111.

[0153] The graph application instance 2406 includes multiple application instances 118, including a line application instance 2404 and / or a triangle application instance 2402, connected by one or more relationships, such as one or more relationships 700-710. For example, the application instance 118 for a first line application instance 2404 can receive output from the application instance 118 for a second line application instance 2404, thereby creating a branch. Similarly, the application instances 118 for a first triangle application instance 2402 can generate output that is received by the application instance for the line application instance 2404 or the application instances 118 for another set of triangle application instances 2402. The application instances 118 for the first triangle application instance 2402 can receive output from the application instance for the line application instance 2404 or the application instances 118 for another set of triangle application instances 2402.

[0154] 25, a cluster 111 may have a corresponding cluster specification 2500. The cluster specification 2500 may be created before or after the creation of the cluster 111 and includes information useful for provisioning components (pods 112, containers 114, storage volumes 116, and / or application instances 118) on the cluster 111.

[0155] For example, cluster specification 2500 for cluster 111 may include an identifier 2502 for cluster 111 and a location identifier 2504. Location identifier 2504 may include one or both of a name assigned to the geographic area in which one or more hosts on which cluster 111 executes are located and data describing the geographic area in which one or more hosts are located, such as the name of a city, state, country, zip code, or some other political or geographic entity. Location identifier 2504 may include coordinates (latitude and longitude or Global Positioning System) describing the location of one or more hosts. In the case of multiple geographically dispersed hosts, the location identifier 2504 may include the location (political or geographic name and / or coordinates) of each host.

[0156] Cluster specification 2500 may include a list of computing resources 2506 for one or more hosts. Computing resources may include the number of processing cores, amount of memory, and amount of storage available on one or more hosts. For example, computing resources may include a cluster host inventory for cluster 111, as described above. If cluster 111 already hosts one or more components, computing resources 2506 may additionally or alternatively include a cluster AAI for one or more hosts, as defined above.

[0157] 26, dot application specification 2600 may include an identifier 2602 of an application instance 118 that is created in accordance with dot application specification 2600. Dot application specification 2600 may include one or more runtime requirements 2604. For example, runtime requirement 2604 may include location requirement 2606. For example, location requirement 2606 may include the name of a political or geographic entity within which a host executing application instance 118 must be located. Location requirement 2606 may be specified in terms of coordinates and a radius around the coordinates within which a host executing application instance 118 must be located.

[0158] Runtime requirements 2604 may further include availability requirements 2608. Availability requirements 2608 may be values ​​from a set of possible values ​​that indicate the required availability of application instance 118 of dot application specification 2600. For example, such values ​​may include "high availability," "intermittent availability," and "low availability." Orchestrator 106 may then interpret availability requirements 2608 when selecting a host for application instance 118 and configuring application instance 118 on the selected host.

[0159] Runtime requirements 2604 may further include cost requirements 2610. Cost requirements 2610 may indicate the permitted cost for running application instance 118 of dot application specification 2600. For example, a cloud computing provider may charge for some or all of the computing power (e.g., processor cores), memory, and storage used by application instance 118. Thus, cost requirements 2610 may specify a maximum amount that can be spent running application instance 118, such as the amount that can be spent per day, month, or other period.

[0160] The dot application specification 2600 may further include computing resource requirements 2612 that specify the amount of processing power, memory, and / or storage required to execute the application instance 118 of the dot application specification 2600. The computing resource requirements 2612 may be a static definition or may be dynamic, such as annotations where provisioning may be dynamically modified based on initial provisioning requirements and usage (e.g., as described above with reference to FIGS. 15-19).

[0161] The dot application specification 2600 may further include tolerances 2614 that specify whether exceptions to any of the above-mentioned requirements 2604, 2612 are allowed. For example, the tolerances 2614 may indicate that the application instance 118 of the dot application specification 2600 should not be deployed unless all of the requirements 2604, 2612 are met. The tolerances 2614 may indicate that if a cluster 111 that satisfies the requirements 2604, 2612 is not found, the application instance 118 may be deployed to the closest alternative (“best fit”). The tolerances may indicate an allowed deviation from any of the requirements 2604, 2612 if a cluster 111 that satisfies the requirements 2604, 2612 is not found.

[0162] The dot application specification 2600 defines the provisioning of the dot application specification's application instance 118. Other parameters that define the instantiation and configuration of the application instance 118 on the selected host may be included in the manifest ingested by the orchestrator 106 in addition to the dot application specification 2600. Alternatively, the dot application specification 2600 may be part of the manifest.

[0163] 27 , a triangle application specification 2700 may include an identifier 2702 for a set of application instances 118 to be created in accordance with the triangle application specification 2700. The triangle application specification 2700 may include one or more runtime requirements 2704. For example, the runtime requirements 2704 may include a location requirement 2706. For example, the location requirement 2706 may include the name of a political or geographic entity in which a host executing the set of application instances 118 must be located. The location requirement 2706 may be specified in terms of coordinates and a radius around the coordinates in which one or more hosts executing one or more application instances 118 of the hierarchy must be located. The location requirement 2706 may include an individual location for each application instance 118 in the set of application instances 118.

[0164] Runtime requirements 2704 may further include availability requirements 2708. Availability requirements 2708 may be values ​​from a set of possible values ​​that indicate the required availability for the set of application instances 118 of triangle application specification 2700. For example, such values ​​may include “high availability,” “intermittent availability,” and “low availability.” Orchestrator 106 may then interpret availability requirements 2708 when selecting hosts for the set of application instances 118 and configuring the set of application instances 118 on the selected hosts. Availability requirements 2708 may include individual availability requirements for each application instance 118 in the set of application instances 118.

[0165] Runtime requirements 2704 may further include cost requirements 2710. Cost requirements 2710 may indicate the permitted cost for executing the set of application instances 118 of triangle application specification 2700. For example, a cloud computing provider may charge for some or all of the computing power (e.g., processor cores), memory, and storage used by each application instance 118 in the set of application instances 118. Thus, cost requirements 2710 may specify a maximum amount that can be spent on executing the set of application instances 118, such as an amount that can be spent per day, month, or other period of time. Cost requirements 2710 may include individual cost requirements for each application instance 118 in the set of application instances 118.

[0166] Runtime requirements 2704 may further include latency requirements 2712. Because each application instance 118 in a set of application instances 118 depends on every other application instance in the set, proper functioning may require the latency to be below a specified maximum latency in terms of time, such as 10 ms, 20 ms, or some other time value. Latency requirements 2712, i.e., the maximum latency allowed between each possible pair of application instances 118, may be specified for each pair of application instances 118 in the set.

[0167] The triangle application specification 2700 may further include computing resource requirements 2714 that specify the amount of processing power, memory, and / or storage required to execute each application instance 118 in the set of application instances 118 of the triangle application specification 2700. The computing resource requirements 2714 may be a static definition or may be dynamic, such as annotations where provisioning may be dynamically modified based on initial provisioning requirements and usage (e.g., as described above with reference to Figures 15-19).

[0168] The triangle application specification 2700 may further include replication requirements 2716 that specify the number of application instances 118 to be included in the set of application instances, for example, a value greater than or equal to 3. Thus, if an application instance 118 fails, the orchestrator 106 creates a new application instance 118 to satisfy the replication requirements 2716.

[0169] The triangle application specification 2700 may further include tolerances 2718 that specify whether exceptions to any of the above-mentioned requirements 2704, 2714, 2716 are allowed. For example, the tolerances 2718 may indicate that the application instance 118 of the triangle application specification 2700 should not be deployed unless all of the requirements 2704, 2714, 2716 are met. The tolerances 2718 may indicate that if a cluster 111 that satisfies the requirements 2704, 2714, 2716 is not found, the application instance 118 may be deployed to the closest alternative (“best fit”). The tolerances may indicate an allowed deviation from any of the requirements 2704, 2714, 2716 if a cluster 111 that satisfies the requirements 2704, 2714, 2716 is not found.

[0170] The triangle application specification 2700 defines the provisioning of a set of application instances 118. The instantiation and configuration of each application instance 118 on a selected host, as well as the creation of any relationships 700-710 between the application instances 118, may be performed according to a manifest that is captured by the orchestrator 106 in addition to the triangle application specification 2700. Alternatively, the triangle application specification 2700 may be part of the manifest.

[0171] 28, a line application specification 2800 can include multiple tier specifications 2802. Each tier specification 2802 corresponds to a different tier in the pipeline defined by the line application specification 2800. Each tier specification 2802 can include a specification of the type of application instance 118 to be instantiated for that tier. Each tier can include multiple application instances 118 of the same or different types.

[0172] Each tier specification 2802 may include identifiers 2804 of one or more application instances 118 created in accordance with the tier specification 2802. The tier specification 2802 may include one or more runtime requirements 2806. For example, the runtime requirements 2806 may include location requirements 2808. For example, the location requirements 2808 may include the name of a political or geographic entity in which one or more hosts executing one or more application instances 118 of the tier must be located. The location requirements 2808 may be specified in terms of coordinates and a radius around the coordinates in which all hosts executing one or more application instances 118 of the tier must be located. The location requirements 2808 may include individual locations for one or more application instances 118 of the tier.

[0173] Runtime requirements 2806 may further include availability requirements 2810. Availability requirements 2810 may be a value from a set of possible values ​​that indicates the required availability for one or more application instances 118 of the tier. For example, such values ​​may include “high availability,” “intermittent availability,” and “low availability.” Orchestrator 106 may then interpret availability requirements 2810 when selecting one or more hosts for one or more application instances 118 of the tier and configuring the one or more application instances 118 on the selected hosts. Availability requirements 2810 may include individual availability requirements for each application instance 118 of the one or more application instances 118.

[0174] Runtime requirements 2806 may further include cost requirements 2812. Cost requirements 2812 may indicate the allowed cost for running one or more application instances 118 of a tier. For example, a cloud computing provider may charge for some or all of the computing power (e.g., processor cores), memory, and storage used by each application instance 118 of one or more application instances 118. Thus, cost requirements 2812 may specify a maximum amount that can be spent running one or more application instances 118 of a tier, such as an amount that can be spent per day, month, or other period of time. Cost requirements 2812 may include individual cost requirements for each application instance 118 of the one or more application instances 118 of a tier.

[0175] Runtime requirements 2806 may further include latency requirements 2814. Latency requirements 2814 may one or both of (a) defining a maximum tolerable latency between application instances of the same tier and (b) defining a maximum latency for application instances 118 of preceding and / or succeeding tiers.

[0176] The tier specification 2802 may further include computing resource requirements 2816 that specify the amount of processing power, memory, and / or storage required to run each of the one or more application instances 118 of the tier. The computing resource requirements 2816 may be static definitions or may be dynamic, such as annotations where provisioning may be dynamically modified based on initial provisioning requirements and usage (e.g., as described above with reference to FIGS. 15-19).

[0177] The tier specification 2802 may further include tolerances 2818 that specify whether exceptions to any of the above-mentioned requirements 2806, 2816 are allowed. For example, the tolerances 2818 may indicate that one or more application instances 118 of a tier should not be deployed unless all of the requirements 2806, 2816 are met. The tolerances 2818 may indicate that if no cluster 111 is found that satisfies the requirements 2806, 2816, then one or more application instances 118 may be deployed to the closest alternative (“best fit”). The tolerances may indicate an allowed deviation from any of the requirements 2806, 2816 if no cluster 111 is found that satisfies the requirements 2806, 2816.

[0178] As described above, a graph application instance 2406 includes multiple application instances 118, including multiple line application instances 2404 and / or triangle application instances 2402. Thus, a specification of a graph application instance may include a collection of specifications 2700, 2800 of the line application instances 2404 and / or triangle application instances 2402 that make up the graph application instance.

[0179] 29 shows a method 2900 for deploying a dot application instance 2400. Method 2900 may be performed by an orchestrator 106. For example, the orchestrator 106 may invoke execution of a workflow from the workflow repository 120 by a worker 124 to perform some or part of method 2900. Method 2900 may be performed in response to the orchestrator 106 receiving a dot application specification 2600 from a user or as part of a manifest.

[0180] The method 2900 may include determining 2902 computing resource requirements 2612 for the dot application instance 2400 and determining 2904 one or more runtime requirements 2604 for the dot application instance 2400. The method 2900 may then include evaluating 2906 the cluster specifications 2500 of the clusters 111 to determine whether any of the available clusters 111 have sufficient computing resources 2506 to meet the computing resource requirements 2612 and to meet the runtime requirements 2604. As described above, the available computing resources evaluated may be either the cluster host inventory of the cluster 111 or the cluster AAI of the cluster 111 on which one or more components are already running.

[0181] If one or more matching clusters are found in step 2906, method 2900 may include deploying the application instance 118 corresponding to the dot application instance 2400 to one of the one or more clusters. If multiple clusters are found in step 2906, one cluster 111 may be selected based on one or more criteria, such as geographic proximity, performance, available cluster inventory or cluster AAI, or other criteria.

[0182] If no matching clusters 111 are found in step 2906, method 2900 may include evaluating 2910 whether the dot application specification 2600 defines tolerances 2718. Step 2910 may further include evaluating whether any of the available clusters 111 are within defined tolerances for the computing resource requirements 2612 and / or runtime requirements 2604 of the dot application specification 2600. If the dot application specification 2600 does not provide tolerances or the clusters 111 are not within the defined tolerances, the operation may fail 2914 and an error message may be returned to the user, the orchestrator 106, the log file 200, or other destination.

[0183] If the dot application specification 2600 provides tolerances, and / or there are one or more clusters 111 that are within any defined tolerances, a compromise cluster 111 may be selected 2912. The compromise cluster 111 may be the cluster 111 that most closely meets one or both of the computing resource requirements 2612 and the runtime requirements 2604. For example, from among the clusters 111 having cluster host inventories and / or cluster AAIs that meet the computing resource requirements 2612, the cluster 111 that most closely meets the runtime requirements 2604 may be selected. For example, the runtime requirements 2604 may be ranked such that the cluster 111 that meets the highest-ranked runtime requirement(s) 2604 is selected 2912. Once a compromise cluster is selected, the application instance 118 of the dot application instance 2400 is deployed 2908 on the compromise cluster.

[0184] 30 shows a method 3000 for deploying a triangle application instance 2402. Method 3000 may be performed by orchestrator 106. For example, orchestrator 106 may invoke execution of a workflow from workflow repository 120 by worker 124 to perform some or part of method 3000. Method 3000 may be performed in response to orchestrator 106 receiving triangle application specification 2700 from a user or as part of a manifest.

[0185] The method 3000 may include determining 3002 computing resource requirements 2714 for the triangle application instance 2402 and determining 3004 one or more runtime requirements 2704 for the triangle application instance 2402. The method 3000 may then include evaluating 3006 the cluster specifications 2500 of the clusters 111 to determine whether any of the available clusters 111 meet the computing resource requirements 2714 and have sufficient computing resources 2506 to meet the runtime requirements 2704. The evaluation of step 3006 may be performed for each application instance 118 of the triangle application instance 2402, where the evaluation identifies, for each application instance 118, all clusters 111 that have sufficient computing resources 2506 and meet the runtime requirements 2704 for that application instance 118.

[0186] The matching clusters 111 identified in step 3006 may then be further evaluated to determine 3008 the inter-cluster latency of the matching clusters 111. The inter-cluster latency may have been previously calculated and obtained, or may be tested as part of step 3008.

[0187] The method 3000 may then include evaluating 3010 whether any cluster group can be found among the matching clusters that meets the latency requirements 2712 of the triangle application instance 2402. For example, given that the application instances 118 of the triangle application instance 2402 are designated as A, B, and C, a matching cluster group may be cluster C that meets the computing resource requirements 2714 and runtime requirements 2704 of application instance A. A Cluster C that meets the computing resource requirements 2714 and runtime requirements 2704 of application instance B B , and cluster C that meets the computing resource requirements 2714 and runtime requirements 2704 of application instance C. C and between each of these clusters (C A and C B Between C B and C C Between and C A and C C The latency between (and) meets latency requirement 2712.

[0188] If one or more matching cluster groups are found in step 3010, the method 3000 may include deploying 3012 the application instance 118 of the triangle application instance 2402 on one cluster 111 of the one or more matching cluster groups. If multiple cluster groups are found in step 3010, one cluster group may be selected based on one or more criteria such as average inter-cluster latency, geographic proximity, performance, available cluster inventory or cluster AAI, or other criteria.

[0189] If no matching clusters are found in step 3006, or the number of matching clusters is less than the number needed to implement the triangle application instance 2402, the method 3000 may include evaluating 3014 whether the triangle application specification 2700 defines a tolerance range 2718. Step 3014 may further include evaluating whether any of the available clusters 111 are within a defined tolerance range for the computing resource requirements 2714 and / or runtime requirements 2704 of the triangle application specification 2700. If the triangle application specification 2700 does not provide a tolerance range or the clusters 111 are not within the defined tolerance range, the operation may fail 3018 and an error message may be returned to the user, the orchestrator 106, the log file 200, or other destination.

[0190] If the triangle application specification 2700 provides tolerances, and / or there are clusters 111 that fall within any defined tolerances, one or more compromise clusters 111 may be selected 3016. The compromise clusters 111 may be the clusters 111 that most closely match one or both of the computing resource requirements 2714 and the runtime requirements 2704. For example, from among the clusters 111 having cluster host inventories and / or cluster AAIs that satisfy the computing resource requirements 2714, the clusters 111 that most closely satisfy the runtime requirements 2704 may be selected. For example, the runtime requirements 2704 may be ranked such that the clusters 111 that satisfy the highest-ranked runtime requirement(s) 2704 are selected 3016. Any compromise clusters selected in step 3016 may then be processed in step 3008, which may include processing the compromise clusters along with the matching clusters identified in step 3006.

[0191] If no matching cluster groups are found in step 3010, the method 3000 may include evaluating 3020 whether the triangle application specification 2700 defines a tolerance 2718 for the latency requirement 2712. Step 3020 may further include evaluating whether any of the inter-cluster latencies of any of the non-matching cluster groups are within a defined tolerance for the latency requirement 2712. If the triangle application specification 2700 does not provide a tolerance or the cluster groups are not within a defined tolerance, the operation may fail 3018 and an error message may be returned to the user, the orchestrator 106, the log file 200, or other destination.

[0192] If the triangle application specification 2700 provides tolerances and / or there is at least one cluster group that is within any defined tolerances, a compromise cluster group may be selected 3022 and the application instances 118 of the triangle application instance 2402 may be deployed on the clusters 111 of the selected compromise cluster group. The compromise cluster group may be the cluster group that most closely matches the latency requirements 2712. If one or more cluster groups include the compromise cluster selected in step 3016, selecting a compromise cluster group 3022 may also include evaluating the combination of inter-cluster latency for each cluster group and how closely each cluster in each cluster group meets the computing resource requirements 2714 and the runtime requirements 2704.

[0193] 31 shows a method 3100 for deploying a line application instance 2404. The method 3100 may be performed by the orchestrator 106. For example, the orchestrator 106 may invoke execution of a workflow from the workflow repository 120 by a worker 124 to perform some or part of the method 3100. The method 3100 may be performed in response to the orchestrator 106 receiving a line application specification 2800 from a user or as part of a manifest.

[0194] The method 3100 may include determining 3102 computing resource requirements 2816 for the line application instance 2404 and determining 3104 one or more runtime requirements 2806 for the line application instance 2404. The method 3100 may then include evaluating 3006 the cluster specifications 2500 of the clusters 111 to determine whether any of the available clusters 111 meet the computing resource requirements 2816 and have sufficient computing resources 2506 to meet the runtime requirements 2806. The evaluation of step 3106 may be performed for each application instance 118 of the line application instance 2404, where the evaluation identifies, for each application instance 118, all clusters 111 that have sufficient computing resources 2506 and meet the runtime requirements 2806 for that application instance 118.

[0195] The groups of matching clusters 111 identified in step 3106 may then be evaluated 3108 to determine a cost function for each group of matching clusters. The cost function for a group of clusters may include an evaluation of a monetary cost, such as the total monetary cost of deploying the application instances 118 of the line application instance 2404 on the clusters 111 of the group, or the monetary cost of deploying the most resource-intensive of the application instances 118 of the line application instance 2404. For example, the application instance 118 that hosts a database is the most resource-intensive in most applications; as a result, the cost function may be limited to evaluating the monetary cost of deploying the database-hosting application instance 118 on a cluster 111 of a given cluster group that meets the computing resource requirements 2816 and one or more runtime requirements 2806 of the database-hosting application instance 118.

[0196] The method 3100 may then include evaluating whether there are any cluster groups that meet the selection criteria 3110. For example, the selection criteria may be a cost function of any cluster groups that is below a predetermined threshold.

[0197] If one or more matching cluster groups are found in step 3110, the method 3100 may include deploying 3112 the application instance 118 of the line application instance 2404 on one cluster 111 of the one or more matching cluster groups. If multiple cluster groups are found in step 3110, one cluster group may be selected based on one or more criteria such as a cost function, average inter-cluster latency, geographic proximity, performance, available cluster inventory or cluster AAI, or other criteria.

[0198] If no matching clusters are found in step 3106, or the number of matching clusters is less than the number needed to implement the line application instance 2404, the method 3100 may include evaluating 3114 whether the line application specification 2800 defines tolerances 2818. Step 3114 may further include evaluating whether any of the available clusters 111 are within a defined tolerance range for the computing resource requirements 2816 and / or runtime requirements 2806 of the line application specification 2800. If the line application specification 2800 does not provide a tolerance range or the clusters 111 are not within the defined tolerance range, the operation may fail 3118 and an error message may be returned to the user, the orchestrator 106, the log file 200, or other destination.

[0199] If the line application specification 2800 provides tolerances, and / or there are clusters 111 that fall within any defined tolerances, one or more compromise clusters 111 may be selected 3116. The compromise clusters 111 may be the clusters 111 that most closely match one or both of the computing resource requirements 2816 and the runtime requirements 2806. For example, from among the clusters 111 having cluster host inventories and / or cluster AAIs that satisfy the computing resource requirements 2816, the clusters 111 that most closely satisfy the runtime requirements 2806 may be selected. For example, the runtime requirements 2806 may be ranked such that the clusters 111 that satisfy the highest-ranked runtime requirement(s) 2704 are selected 3116. Any compromise clusters selected in step 3116 may then be processed in step 3108, which may include processing the compromise clusters along with the matching clusters identified in step 3106.

[0200] If no matching cluster groups are found in step 3110, method 3100 may include evaluating 3120 whether the line application specification 2800 defines a tolerance 2818 for the cost requirement 2812. Step 3120 may further include evaluating whether any cost functions of the non-matching cluster groups are within the defined tolerance for the cost requirement 2812. If the line application specification 2800 does not provide a tolerance or the cluster groups are not within the defined tolerance, the operation may fail 3118 and an error message may be returned to the user, the orchestrator 106, the log file 200, or other destination.

[0201] If the line application specification 2800 provides tolerances and / or there is at least one cluster group that is within any defined tolerances, a compromise cluster group can be selected 3122 and the application instances 118 of the line application instance 2404 can be deployed on the clusters 111 of the selected compromise cluster group. The compromise cluster group can be the cluster group that most closely matches the latency requirements 2712. If one or more cluster groups include the compromise cluster selected in step 3116, selecting the compromise cluster group can also evaluate the combination of inter-cluster latency for each cluster group and how closely each cluster in each cluster group meets the computing resource requirements 2816 and the runtime requirements 2806.

[0202] 32 and 33 illustrate a method 3200 for unfolding a graph application instance 2406. The method 3200 may include splitting 3202 the graph application instance 2406 into one or more triangle application instances 2402 and line application instances 2404, as shown in FIG. 33. The splitting 3202 may be performed taking into account specifications of the graph application instance 2406, including explicitly defined triangle application specifications 2700 and / or line application specifications 2800. The splitting may also include analyzing a graph representing the application instance 118 of the graph application instance 2406 to identify the triangle application instances 2402 and the line application instances 2404.

[0203] The method 3200 may include provisioning and deploying the triangle application instance 2402, such as according to the method 3000. The method 3200 may include provisioning and deploying the line application instance 2404, such as according to the method 3100.

[0204] Methods 3000, 3100 may be modified in one or more respects when deploying a graph application instance 2406. Method 3000 includes evaluating 3010 whether there is a matching cluster group, and method 3100 includes evaluating 3110 whether there is a matching cluster group. For method 3200, a "matching cluster group" may be defined as a matching cluster group that includes clusters of each application instance 118 of all triangle application instances 2402 and line application instances 2404 of the graph application instance 2406. Thus, in some embodiments, a matching cluster group must simultaneously satisfy the requirements of all triangle application instances 2402 and line application instances 2404 of the graph application instance 2406. In an alternative approach, the triangle application instances 2402 and line application instances 2404 of the graph application instance 2406 are processed one at a time, such as from largest to smallest (by number of application instances 118) of the triangle application instances 2402 and line application instances 2404 of the graph application instance 2406, or in some other order. In either approach, if any set of triangle application instances 2402 or line application instances 2404 cannot be provisioned, i.e., if the operation fails 3018, 3118, the method 3200 will fail for the graph application instance. Alternatively, partial failures may be allowed, such that a first portion of a triangle application instance and / or a line application instance 2404 is deployed even if a second portion cannot be deployed.

[0205] 34 is a block diagram illustrating an example computing device 3400. The computing device 3400 may be used to perform various procedures as described herein. The server 102, the orchestrator 106, the workflow orchestrator 122, the vector log agent 126, the log processor 130, and the cloud computing platform 104 may each be implemented using one or more computing devices 3400. The orchestrator 106, the workflow orchestrator 122, the vector log agent 126, and the log processor 130 may be implemented on different computing devices 3400, or a single computing device 3400 may host two or more of the orchestrator 106, the workflow orchestrator 122, the vector log agent 126, and the log processor 130.

[0206] Computing device 3400 includes one or more processor(s) 3402, one or more memory device(s) 3404, one or more interface(s) 3406, one or more mass storage device(s) 3408, one or more input / output (I / O) device(s) 3410, and a display device 3430, all coupled to a bus 3412. Processor(s) 3402 include one or more processors or controllers that execute instructions stored in memory device(s) 3404 and / or mass storage device(s) 3408. Processor(s) 3402 may also include various types of computer-readable media, such as cache memory.

[0207] The memory device(s) 3404 include a variety of computer-readable media, such as volatile memory (e.g., random access memory (RAM) 3414) and / or non-volatile memory (e.g., read-only memory (ROM) 3416). The memory device(s) 3404 may also include re-writable ROM, such as flash memory.

[0208] The mass storage device(s) 3408 include a variety of computer-readable media, such as magnetic tape, magnetic disks, optical disks, solid-state memory (e.g., flash memory), etc. As shown in Figure 34, a particular mass storage device is a hard disk drive 3424. Various drives may also be included in the mass storage device(s) 3408 to allow reading from and / or writing to a variety of computer-readable media. The mass storage device(s) 3408 include removable media 3426 and / or non-removable media.

[0209] The I / O device(s) 3410 include various devices that allow data and / or other information to be input to or retrieved from the computing device 3400. Exemplary I / O device(s) 3410 include cursor control devices, keyboards, keypads, microphones, monitors or other display devices, speakers, printers, network interface cards, modems, lenses, CCD or other image capture devices, etc.

[0210] Display device 3430 includes any type of device capable of displaying information to one or more users of computing device 3400. Display device 3430 includes, for example, a monitor, a display terminal, a video projection device, etc.

[0211] The interface(s) 3406 include various interfaces that allow the computing device 3400 to interact with other systems, devices, or computing environments. Exemplary interface(s) 3406 include any number of different network interfaces 3420, such as interfaces to local area networks (LANs), wide area networks (WANs), wireless networks, and the Internet. Other interface(s) include a user interface 3418 and a peripheral interface 3422. The interface(s) 3406 may also include one or more peripheral interfaces, such as interfaces for printers, pointing devices (such as a mouse, trackpad, etc.), keyboards, etc.

[0212] The bus 3412 allows the processor(s) 3402, memory device(s) 3404, interface(s) 3406, mass storage device(s) 3408, I / O device(s) 3410, and display device(s) 3430 to communicate with each other and with other devices or components coupled to the bus 3412. The bus 3412 represents one or more of several types of bus structures, such as a system bus, a PCI bus, an IEEE 1394 bus, a USB bus, etc.

[0213] For purposes of illustration, programs and other executable program components are illustrated herein as separate blocks, with the understanding that such programs and components may reside at various times in different storage components of computing device 3400 and are executed by processor(s) 3402. Alternatively, the systems and procedures described herein may be implemented in hardware or a combination of hardware, software, and / or firmware. For example, one or more application specific integrated circuits (ASICs) can be programmed to execute one or more of the systems and procedures described herein.

[0214] In the foregoing disclosure, reference has been made to the accompanying drawings, which form a part hereof, and which show, by way of illustration, specific implementations in which the present disclosure may be practiced. It is understood that other implementations may be utilized and structural changes may be made without departing from the scope of the present disclosure. References herein to "one embodiment," "an embodiment," "an example embodiment," or the like indicate that the described embodiment may include a particular feature, structure, or characteristic, but do not necessarily mean that all embodiments include the particular feature, structure, or characteristic. Moreover, such phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is believed to be within the knowledge of one skilled in the art to affect such feature, structure, or characteristic in connection with other embodiments, whether or not explicitly stated.

[0215] Implementations of the systems, devices, and methods disclosed herein may include or utilize special-purpose or general-purpose computers, including computer hardware such as one or more processors and system memory, as described herein. Implementations within the scope of the present disclosure may also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. Such computer-readable media may be any available media accessible by a general-purpose or special-purpose computer system. Computer-readable media that store computer-executable instructions are computer storage media (devices). Computer-readable media that carry computer-executable instructions are transmission media. Thus, by way of example, and not limitation, implementations of the present disclosure may include at least two distinctly different types of computer-readable media: computer storage media (devices) and transmission media.

[0216] Computer storage media (devices) include RAM, ROM, EEPROM, CD-ROM, solid state drives (“SSDs”) (e.g., RAM-based), flash memory, phase-change memory (“PCM”), other types of memory, other optical disk storage, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store desired program code means in the form of computer-executable instructions or data structures and that can be accessed by a general-purpose or special-purpose computer.

[0217] Implementations of the devices, systems, and methods disclosed herein can communicate over a computer network. A "network" is defined as one or more data links that enable the transfer of electronic data between computer systems and / or modules and / or other electronic devices. When information is transferred or provided to a computer over a network or another communications connection (either hardwired, wireless, or a combination of hardwired or wireless), the computer properly views the connection as a transmission medium. Transmission media can be used to carry desired program code means in the form of computer-executable instructions or data structures and can include networks and / or data links that can be accessed by a general-purpose or special-purpose computer. Combinations of the above should also be included within the scope of computer-readable media.

[0218] Computer-executable instructions include, for example, instructions and data that, when executed by a processor, cause a general-purpose computer, special-purpose computer, or special-purpose processing device to perform a certain function or group of functions. Computer-executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, or source code. While the present subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the present subject matter defined in the appended claims is not necessarily limited to such features or acts. Rather, the described features and acts are disclosed as exemplary forms of implementing the claims.

[0219] Those skilled in the art will appreciate that the present disclosure may be implemented in networked computing environments having many types of computer system configurations, including in-dash vehicle computers, personal computers, desktop computers, laptop computers, message processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, tablets, pagers, routers, switches, various storage devices, etc. The present disclosure may also be implemented in distributed system environments where tasks are performed by both local and remote computer systems that are linked through a network (either by hardwired data links, wireless data links, or a combination of hardwired and wireless data links). In a distributed system environment, program modules may be located in both local and remote memory storage devices.

[0220] Additionally, where appropriate, the functions described herein may be implemented by one or more of hardware, software, firmware, digital components, or analog components. For example, one or more application specific integrated circuits (ASICs) may be programmed to execute one or more of the systems and procedures described herein. Certain terms are used throughout this specification and claims to refer to particular system components. As one skilled in the art will understand, components may be referred to by different names. This specification does not intend to distinguish between components that differ in name but not function.

[0221] It should be noted that the sensor embodiments described above may include computer hardware, software, firmware, or any combination thereof to perform at least a portion of their functionality. For example, the sensor may include computer code configured to run on one or more processors and may include hardware logic / electrical circuitry controlled by the computer code. These exemplary devices are provided herein for illustrative purposes and are not intended to be limiting. Embodiments of the present disclosure may be implemented in additional types of devices, as known to those skilled in the art.

[0222] At least some embodiments of the present disclosure are directed to computer program products including such logic (e.g., in the form of software) stored on any computer-usable medium, such software, when executed on one or more data processing devices, causing the devices to operate as described herein.

[0223] While various embodiments of the present disclosure have been described above, it should be understood that they have been presented by way of example only, and not limitation. It will be apparent to those skilled in the art that various changes in form and detail can be made without departing from the spirit and scope of the present disclosure. Thus, the breadth and scope of the present disclosure should not be limited by any of the above-described exemplary embodiments, but should be defined only in accordance with the following claims and their equivalents. The foregoing description has been presented for purposes of illustration and description. It is not intended to be exhaustive or to limit the disclosure to the precise form disclosed. Many modifications and variations are possible in light of the above teachings. Furthermore, it should be noted that any or all of the above-described alternative implementations may be used in any combination desired to form additional hybrid implementations of the present disclosure.< / address>

Claims

1. 1. A computing device including one or more processing devices and one or more memory devices operably coupled to the one or more processing devices, wherein the one or more memory devices, when executed by the one or more processing devices, cause the one or more processing devices to: Receive observability data from multiple hosts over a network, processing the observability data to obtain utilization of computing resources of a plurality of hosts by a plurality of components executing on the plurality of hosts; determining an active and available inventory of the computing resources of the plurality of hosts according to the utilization rates; a computing device storing executable code that causes a redeployment of one of the plurality of components on a first host of the plurality of hosts to a second host of the plurality of hosts based on the active and available inventory; Device.

2. The apparatus of claim 1 , wherein the plurality of hosts comprises one or more server systems.

3. The apparatus of claim 1 , wherein the plurality of hosts comprises units of computing resources on a cloud computing platform.

4. 10. The apparatus of claim 1, wherein the executable code, when executed by the one or more processing devices, causes the one or more processing devices to receive the observability data from the multiple hosts by pulling the observability data from the multiple hosts without using an agent running on the multiple hosts.

5. The apparatus of claim 1 , wherein the computing resource comprises processor time.

6. The apparatus of claim 1 , wherein the computing resource comprises a memory.

7. The apparatus of claim 1 , wherein the computing resources include storage.

8. The apparatus of claim 1 , wherein each component of the plurality of components is one of an application instance, a container, and a storage volume.

9. The executable code, when executed by the one or more processing devices, causes the one or more processing devices to:

10. The apparatus of claim 1, wherein in response to determining that redeployment of the component from a first host to a second host does not violate an affinity requirement, an anti-affinity requirement, or a latency requirement, the apparatus redeploys the component of the plurality of components on the first host of the plurality of hosts to the second host of the plurality of hosts.

10. The executable code, when executed by the one or more processing devices, causes the one or more processing devices to:

2. The apparatus of claim 1, wherein, in response to determining that redeployment of the component from a first host to a second host results in a billing reduction, the apparatus redeploys the component of the plurality of components on the first host of the plurality of hosts to the second host of the plurality of hosts.

11. receiving, by a computer system, observability data from a plurality of hosts over a network; processing, by the computer system, the observability data to obtain utilization of computing resources of a plurality of hosts by a plurality of components executing on the plurality of hosts; determining, by the computer system, an active and available inventory of the computing resources of the plurality of hosts according to the utilization rates; redeploying, by the computer system, one component of the plurality of components on a first host of the plurality of hosts to a second host of the plurality of hosts based on the active and available inventory; method.

12. The method of claim 11 , wherein the plurality of hosts comprises one or more server systems.

13. The method of claim 11 , wherein the plurality of hosts comprises units of computing resources on a cloud computing platform.

14. 12. The method of claim 11 , further comprising receiving the observability data from the plurality of hosts by pulling the observability data from the plurality of hosts without using an agent running on the plurality of hosts.

15. The method of claim 11 , wherein the computing resource comprises processor time.

16. The method of claim 11 , wherein the computing resource comprises a memory.

17. The method of claim 11 , wherein the computing resource includes storage.

18. The method of claim 11 , wherein each component of the plurality of components is one of an application instance, a container, and a storage volume.

19. 12. The method of claim 11, further comprising: in response to determining that redeployment of the component from the first host to the second host does not violate an affinity requirement, an anti-affinity requirement, or a latency requirement, redeploying the component of the plurality of components on the first host to the second host of the plurality of hosts.

20. 12. The method of claim 11, further comprising: in response to determining that redeploying the component from the first host to the second host results in a billing reduction, redeploying the component of the plurality of components on the first host to the second host of the plurality of hosts.

Citation Information

Patent Citations

  • Program, method and apparatus for managing information processor

    JP2012032877A

  • Information processing system and data arrangement in information processing system

    JP2022100620A