LIFE CYCLE MANAGEMENT FOR WORKLOADS IN HETEROGENEOUS INFRASTRUCTURES
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- HEWLETT PACKARD ENTERPRISE DEV LP
- Filing Date
- 2022-04-08
- Publication Date
- 2026-07-30
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUND Application and infrastructure lifecycle management can encompass a wide range of activities, including discovery, deployment, updates, patching, change management, configuration management, and vulnerability management. Some lifecycle management activities, such as performing upgrades and patch management, have traditionally involved manual, repetitive tasks prone to configuration and implementation errors. This is now addressed by various cloud automation tools and lifecycle management software that enable automated lifecycle management. Automated lifecycle management can be used to efficiently manage applications in a cloud computing infrastructure. Automated lifecycle management, combined with cloud-native application design (e.g., applications designed to be independent of changes in the computer infrastructure lifecycle), is the standard for some lifecycle management concepts. US 2019 / 0 342 375 A1 describes a system and procedure for flexibly and automatically managing the entire lifecycle of user-defined resources within a cloud computing environment. US 2015 / 0 180 949 A1 describes a cloud middleware system designed to automate the deployment and management of applications in complex hybrid cloud environments.WO 2020 / 263374 A1 describes a two-pronged approach to automated resource management in distributed "fog" or "edge" computing environments, combining both real-time workload balancing and capacity planning. WO 2021 / 069683 A1 describes a multi-state desired-state orchestration system that makes infrastructure management both declarative and reactive to real-time events. WO 2021 / 178598 A1 describes a system for managing the lifecycle of workloads in an environment using an orchestration platform from a completely different technological environment. US 2016 / 0078342 A1 describes a comprehensive system for the autonomous management of virtual networks. An object of the invention is to propose a method, a system, and a non-transitory, machine-readable medium for lifecycle management of workloads in heterogeneous infrastructures.This problem is solved by a method according to claim 1, a computer system according to claim 9 and a non-transitory, machine-readable medium according to claim 13. Brief description of the drawings The aspects described herein are illustrated by way of example and are not limited to the figures in the accompanying drawings, where the same reference numbers refer to similar elements. Fig. 1 shows a block diagram of a cloud computing environment according to one example. Fig. 2 shows a block diagram of a NEM (Node Environment Manager) according to one example. Fig. 3 shows a diagram illustrating an object model according to one example. Fig. 4 shows a diagram illustrating a state model according to one example. Fig. 5 shows a flowchart illustrating the processing performed by an NEM according to one example. Fig. 6 shows a flowchart illustrating the processing by an NEM according to another example. Fig. 7 shows a block diagram of a computer system according to one example. DETAILED DESCRIPTION The examples described here relate to the provision of pool-based, automated infrastructure lifecycle management via abstracted resources. While infrastructure lifecycle management typically occurs independently of an application's lifecycle, there are cases where an application must be integrated into the lifecycle of the underlying data processing infrastructure, which may or may not be cloud-like, and which cannot be managed independently without regard to the application's lifecycle. This is the case, for example, when the application is a service (e.g., a service-like application such as Software as a Service (SaaS)) for building and managing complex computing system architectures, such as clusters. In this case, the data processing infrastructure lifecycle must work in conjunction with the application's configuration orchestration.Another example is a stateful application such as a storage service, where the underlying data processing infrastructure must be integrated into the application to ensure the restoration of persistent state information for the storage service in the data processing infrastructure when needed. When an organization seeks to deliver a cloud experience for on-premises workloads, application orchestrations face various complexities in the underlying computing infrastructure of the cloud computing environment, including deployment mechanisms and resource management tools (such as KubeVirt, an open-source project, VMware vCenter, and similar solutions). Additional challenges arise for application orchestration in the context of a limited computing environment and / or a computing environment with heterogeneous infrastructure. For example, a site, installation, or data center (referred to here generally as a computing environment) might contain a limited number (e.g., 20 to 80) of servers, each providing computing resources (e.g., compute, storage, and / or network resources), housed in a limited number (e.g., 1 or 2) of racks.The servers available in a given computing environment may also be manufactured by different vendors, represent different models and / or have different attributes, also referred to as qualities, which may include, for example, capabilities, capacities and / or configurations. Furthermore, service-type and / or stateful applications can include monitoring the status of the underlying infrastructure used for a particular service or solution (e.g., Machine Learning Operations as a Service (MLOpsaaS), Container as a Service (CaaS), Virtual Machine as a Service (VMaaS), Storage as a Service (STaaS), etc.), as well as associated images for handling various changes (e.g., the availability of an operating system update, a planned or unplanned outage). Generally, it would be desirable to isolate applications from such complexities and infrastructure-specific details so that, among other things, the applications can focus on the core functions and features of the respective solution. The examples described here propose simplifying and generalizing the operation of computer infrastructure from the application perspective by introducing a software framework as an intermediary layer. This framework logically sits between the applications and their respective orchestration tools on the one hand, and the computer infrastructure (including its associated deployment mechanisms and resource management tools) on the other. According to various examples, a software framework provides pool-based resource lifecycle management for applications. For instance, elements of a computing infrastructure can be managed as groups of similar units (e.g., compute, storage, and network units that may have similar attributes or qualities) with consistent, automated lifecycle management operations and states. As described below, this framework (which can be referred to here as the NEM (Node Environment Manager)) provides, in one example, a system for resource objects and state models, as well as information for resource configuration and inventory; an infrastructure plugin framework for integrating custom infrastructure tools and management products or resource management functions for cloud services; and an application programming interface (API) for resource consumption and lifecycle events. Through the consumption API, application orchestrations can, for example, indirectly manage and / or manipulate the underlying infrastructure via interactions with logical resource objects (such as nodes and node pools). As part of the NEM framework, a northbound API is available for normalized resource procurement in a cloud computing environment by using a reservation system for procuring abstracted compute resources (e.g., pools, nodes, etc.), a notification mechanism for resource lifecycle changes to trigger integrations with application orchestration tools, and a southbound API for infrastructure plugins to enable the integration of computing infrastructure tools and services (e.g., via various developed plugins) that support the southbound API calls.In this way, application orchestration tools can be equipped with a simple API for acquiring compute resources, eliminating the need to understand the complex details of the underlying computing infrastructure of the cloud computing environment, its deployment mechanisms, and its management tools. Furthermore, complex lifecycle changes, such as an operating system, firmware, or software update (either on bare-metal machines / servers or virtual machines), a planned outage or disruption, and an untrusted startup or activity, can be made almost transparent to the application. For example, in response to notifications from NEM, the application can organize the reconfiguration of application instances without regard to the complexity of the data processing infrastructure or the details of the management tool or service. One or more application orchestration tools can access compute resources in the cloud computing environment through logical pools of nodes, represented in a consumption model and acquired via reservations. The application orchestration tools can configure these acquired resources to create instances of applications. Nodes, which are logical in nature, can represent either Business Gateway machines (BMs) or Virtual Machines (VMs). They can be defined to provide specific functionality to an application instance based on the node's attributes (e.g., capabilities, capacity, and / or configuration, such as compute power, storage, and / or networking). A Network Engineering Manager (NEM) can organize nodes into a pool that is automatically and collectively managed throughout their lifecycle using integrated tools and management products. The following description sets out numerous specific details to enable a comprehensive understanding of the subject matter disclosed herein. However, it will be clear to a person skilled in the art that the implementation of the aspects described here can also be carried out without some of these specific details. The terms “connected” or “coupled” and related terms are used in an operational sense and are not necessarily limited to a direct connection or coupling. For example, two devices may be coupled directly or via one or more intermediary media or devices. Another example is that devices may be coupled in such a way that information can be exchanged between them without them having a physical connection to each other. Based on the disclosure contained herein, a person skilled in the art will recognize a multitude of ways in which a connection or coupling exists within the meaning of the above definition. terminology If the specification states that a component or feature "may" or "could" be included, then that component or feature does not have to be included or exhibit the feature. As used in the present description and in the following claims, the meaning of "a", "an", and "the" includes the plural unless the context clearly indicates otherwise. As used in the present description, the meaning of "in" also includes "in" and "at" unless the context clearly indicates otherwise. The expressions “in an example,” “according to an example,” and the like generally mean that the specific feature, structure, or property following the expression is contained in at least one example of the present disclosure and may be contained in more than one example of the present disclosure. It is important to note that such expressions do not necessarily refer to the same example. As used here, the term "infrastructure" generally refers to physical computing systems (e.g., rack servers, blade servers, blades, storage arrays, storage servers, hyperconverged platforms, etc.) and / or virtual computing systems (e.g., VMs running on physical computing systems) available in a computing environment. Each individual component or infrastructure element can provide differentiated compute resources (e.g., compute resources (for performing calculations), storage resources (for storing data), and / or network resources (for communicating data)) for use by the workloads running in the computing environment. As mentioned earlier, an on-premises computing environment may be heterogeneous (e.g., BM machines and / or VMs with varying compute resources) and / or may include a limited number of servers. In some examples, a NEM may maintain an inventory of the BM / VM infrastructure on the premises. A "node," as used here, represents an abstracted element of the infrastructure or a logical set of one or more compute resources connected to the infrastructure. Depending on the specific implementation, nodes can represent BM machines or VMs and be defined to provide specific functions to an application instance based on node attributes (e.g., capabilities, capacity, and / or configuration, such as compute power, storage, and / or networking). In the examples described here, a NEM can organize nodes into groups whose lifecycles are managed collectively through integrated tools and / or management products. In some examples, the NEM maintains an infrastructure model to track the backup servers (e.g., which rack and which server) in the inventory represented by the nodes in the consumption model. "Orchestration" generally refers to the automated configuration, management and / or coordination of computer systems, applications, computing resources and / or services. Application orchestration or service orchestration refers to the process of automating the deployment, management, scaling, networking, integration, and availability of applications and / or services for processing or synchronizing data in real time. In some cases, the applications consist of a multitude of software containers. A container can contain one or more workloads. An application orchestration tool is a program (usually run by a system administrator in a cloud computing environment) that implements application orchestration. Cloud computing environment Figure 1 shows a block diagram of a cloud computing environment 100 according to an example. In the context of this example, the cloud computing environment 100 comprises a cloud computing system 102, a SaaS portal 114 through which users can interact with and configure various aspects of the cloud computing system 102, and external providers 116 (e.g., BM providers, trust providers, VM providers, and / or storage providers). The cloud computing system 102 can be located on-premises in a data center or a colocation. In some implementations, the infrastructure 112 may be owned or held by a user or user organization and managed by an infrastructure provider (of which the user or user organization is a customer). In some implementations, the infrastructure 112 may be provided to the user or the user's organization (e.g., in a data center or colocation belonging to the user or the user's organization) by the infrastructure provider for use as a service under a pay-per-use financial model. In some implementations, a node environment manager 110, described below, may be provided by the infrastructure provider as part of the management for use as a service. The cloud computing system 102 can run one or more workloads 104. The workloads 104 can consist of data and applications, with the infrastructure 112 providing the workloads 104 with computing resources (e.g., compute, storage, and / or network resources) to perform tasks. The workloads 104 can interface with the runtime 106 (e.g., a runtime environment with an operating system (OS), a virtual machine manager (VMM), a hypervisor, or other system software) to utilize the computing resources within the infrastructure 112. Depending on the location, the infrastructure 112 can comprise a heterogeneous resource pool with differentiated servers and / or VM instances, as well as other electronic and / or mechanical components (e.g., server racks (not shown), power supplies (not shown), etc.). In one example, the Node Environment Manager (NEM) 110 represents a software intermediary layer logically positioned between the application orchestration tools (e.g., the application orchestration tool 108) running in the cloud computing system 102 and the infrastructure 112. As described below, NEM 110 can consist of a number of conceptual components, including inventory management, node pool management, node reservation management, automated node provisioning management, and / or automated pool management. In addition to a consumption API, NEM 110 can, for example, also provide an inventory API, through which an administrative user of the cloud computing system 102, who accesses NEM 110 via the SaaS portal 114, for example, can manage an inventory of racks and differentiated servers of the infrastructure 112 to offer a number of node pools for higher-level orchestrators (e.g., the application orchestration tool 108). In one example, the application orchestration tool 108 can be used by the administrative user to configure the automated lifecycle management of infrastructure 112 for workloads 104. NEM 110 can provide a simplified API for resource consumption that enables the application orchestration tool 108 to easily reserve, acquire, and manage compute resources for workloads 104 supported by the infrastructure 112 of the cloud computing system 102, without requiring knowledge of the complex details of managing compute resources (especially BM resources). As described below, the abstraction layer provided by the consumption API allows workloads to acquire nodes without knowing whether the underlying servers (e.g., for providing the compute resources) are BM machines or VMs, and by describing the node's role in defining a pool in a data-driven manner. Furthermore, NEM 110 can notify workloads 104 of changes in the state of compute resources when nodes or pools are involved in lifecycle changes (e.g., maintenance, maintenance, or replacement).(as a result of the availability of operating system or software (SW) or firmware updates, planned or unplanned outages, untrusted activities, etc.). The application orchestration tool 108 can then resolve the state changes caused by NEM 110 by reserving and acquiring new nodes and removing old nodes from their pools (e.g., through immutable node handling), thereby avoiding a deviation in the workload configuration and thus the need to reconfigure workload instances. In one example, NEM 110 can also automate the background provisioning of nodes. For instance, nodes with a defined role (e.g., as part of a predefined workload configuration) can be created before they are captured (e.g., before or in response to a reservation) to facilitate immediate availability for use by workloads 104 as soon as they are captured by the application orchestration tool 108. This can improve node capture performance, as building BM nodes is typically time-consuming. This feature can also be beneficial for building VM nodes, which are generally considered fast to provision but can also be slow to build depending on the complexity of the operating system / software in their respective runtime environments. Node Environment Manager (NEM) Figure 2 shows a block diagram of a Node Environment Manager (NEM) 210 according to an example. The NEM 210 is an unlimited example of the NEM 110 from Figure 1. In the context of this example, NEM 210 includes north-bound APIs 220 (e.g., a consumption API, an inventory API, and a solution configuration API described below), a resource object and state model system 230, and an infrastructure plugin framework with a south-bound API 240. The Infrastructure Plugin Framework 240 can be used to generalize and harmonize various tools, management products, and / or cloud APIs in state and lifecycle definitions associated with NEM 210. For example, the Infrastructure Plugin Framework 240 can integrate tools for managing physical resources and infrastructure (e.g., Resource Manager 211 via BM plugins (e.g., bare-metal deployment tools), trust plugins (e.g., Keylime—an open-source project—or other Trusted Platform Module (TPM) solutions for remote boot attestation and runtime health measurement), VM plugins (e.g., KubeVirt, Kubernetes KVM virtualization management, or other virtualization tools), and / or storage plugins, and the like.While such tools and solutions can alternatively be embedded within NEM 210, using a plugin approach in some implementations has the advantage of allowing NEM 210 to focus on object / state management, making applications easier to build and maintain. Furthermore, with the plugin approach, NEM 210 can be more easily extended with additional lifecycle management capabilities over time, and when integrated into the automated lifecycle of application orchestrations, it enables the creation of flexible and durable application and infrastructure architectures. The resource object and state model system 230 comprises a model and state manager 235, an object model 236, a state model 238, a notifier 232, a provisioner 233, a reservation system 234, and a pool manager 237. In one implementation, the object model 236 and the state model 238 can be implemented in Go (a statically typed, compiled programming language), deployed as stateless Kubernetes pods for scalability, accessed via REST (Representational State Transfer) API calls, and persisted in a database. As described below with reference to Fig. 3, the object model 236 can have several levels, including: (i) an inventory model in which the inventory of the BM / VM infrastructure that is within the cloud computing system (e.g., Cloud Computing System 102) for use by workloads (e.g.,, (ii) a solution model that includes a solution configuration mixed with user-defined settings associated with tools / products integrated into NEM 210, e.g., via the Infrastructure Plugin Framework 240, for a given solution; and (iii) a consumption model that includes a set of logical objects (e.g., nodes and node pools) that can be used to indirectly manipulate the actual computing infrastructure at runtime through application orchestration tools (e.g., the Application Orchestration Tool 108). Based on the object model 236, the persisted logical resource objects can automatically transition through the states defined in the state model 238 either through REST API calls (e.g., to APIs 220) or NEM background processes, which can lead to notifications (e.g., generated by the Notifier 232) to application orchestration tools via APIs 220 and / or calls to the infrastructure plugin framework 240 to invoke suitable tailored management tools or cloud infrastructure APIs. As described below with reference to Fig. 4, the model and state manager 235 can be responsible for orchestrating state changes associated with logical resource objects (e.g., nodes and node pools) as a result of actions or operations performed by application orchestration tools, administrative users, and / or the resource managers 211 connected to infrastructure plugins within the infrastructure plugin framework 240. Under the direction of the model and state manager 235, the state model 238 can maintain the states associated with the logical resource objects. The reservation system 234 can be responsible for handling reservation requests, such as those made via a consumption API 220. Node reservation management can be used in conjunction with node pools to ensure the availability and acquisition (possibly with availability time estimates) of nodes for performing lifecycle implementations (e.g., provisioning nodes for use by a workload and removing nodes currently in use by a workload, i.e., immutable node management). In some examples, a reservation can be made for a desired number of nodes in a pool to ensure that a task (e.g., creating a workload control plane or scaling a workload) can be performed without running out of nodes before the task is complete.Further details on various example types of reservations and various example types of actions that can be carried out via the reservation system 234 are described below with reference to Fig. 3. Provisioner 233 can be responsible for automatically managing node provisioning. For example, automatic node provisioning management includes ensuring that nodes not belonging to any node pool (available nodes) and corresponding to a defined node role are kept in a provisioned state with the latest operating system image / firmware / BIOS settings for that node type. During initial installation in the data processing environment, all nodes can be provisioned with an image to ensure they have the image specified for the solution and model type. Subsequently, when a new image is made available for a model type, the inventory can be updated with the current image.All available nodes that have not been purchased and belong to an existing pool can be reused to ensure that a node requested by a workload always has the latest image for that node's model type. As described below, workloads can decide when they are ready to update the image on their purchased nodes that belong to their node pools and can decommission a node over time and replace it with an available node from the pool. Pool Manager 237 can be responsible for automated pool management. In one example, automated pool management involves the automated management of node pools, which are integrated into workload lifecycle management processes to ensure that nodes and / or node pools remain updated either on command or according to a schedule, or can be removed due to failure. As described below, a pool status can indicate whether nodes belonging to the pool need to be updated or whether maintenance is due. In this way, workload lifecycle automation can iterate through the pool nodes to either perform an update process or remove the node from the pool for maintenance. Once all nodes in the pool are in a "normal" state, the pool status can revert to its "normal" state.Further details on various example types of node / pool states and transitions between states are described below with reference to Fig. 4. The various functional units (e.g., the APIs 220, the notifier 232, the reservation system 234, the provisioner 233, the model and state manager 235, and the pool manager 237) of the NEM 210, described above with reference to Fig. 2 and below with reference to the flowcharts in Figs. 5-6, can be implemented in the form of executable instructions stored on a non-transient, machine-readable medium (e.g., random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, disk drive, or the like) and executed by a hardware-based processing resource (e.g.,The processing can be implemented using a microcontroller, a microprocessor, central processing core(s), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or similar, and / or in the form of other types of electronic circuits. The processing can be performed, for example, by one or more computer systems of various forms (e.g., servers, blades, desktop computers, laptop computers), such as the computer system described in Fig. 7. Object model Figure 3 shows a diagram illustrating an object model 300 according to an example. Object model 300 is a non-restrictive example of object model 236 from Figure 2. Object model 300 can serve as an abstract and rationalized persistent model used by a NEM (e.g., NEM 210) to describe infrastructure elements (e.g., infrastructure 112). Object model 300 can also be used by the NEM for the acquisition or release of infrastructure elements. Based on the definition of pools of logical resource objects (e.g., nodes (e.g., node 310) and pools (e.g., pool 308) of nodes), the object model allows the specification of qualities or attributes of a node via a node or instance typing system, or it can allow distinguishing aspects that are applied to nodes (e.g., network connections, attached storage volumes, etc.) when they are acquired and assigned to a pool.Computing resources can be used in groups, where one or more aspects of the properties of computing resources (e.g., CPU type, memory size, server model, device type, etc.) can be core aspects, while others (e.g., storage resources, block storage, volumes, network resources, networks, virtual network infrastructure components) can be context-dependent, based on the use of the core resource. In the context of the present example, the object model 300 comprises three levels: a consumption model 304, a solution model 312, and an inventory model 322. The consumption model includes reservations (e.g., Reservation 306) and a set of logical resource objects (e.g., nodes (e.g., Node 310) and pools (e.g., Pool 308) of nodes) that can be used to indirectly manipulate the actual data processing infrastructure at runtime through application orchestration tools (e.g., the application orchestration tool 108). In the consumption model 304, a reservation 306 can be made for a group of one or more nodes, resulting in the formation of a pool 308 of one or more nodes 310. A consumption API (e.g., one of the APIs 220) can support operations to create, read, update, and delete (CRUD) nodes, pools, and reservations.In addition, other administrative processes (e.g., listing) and lifecycle implementations (deployment, scaling, etc.) can be provided. In some examples, reservations (e.g., reservation 306) can be one of several reservation types used by lifecycle management implementations to balance lifecycle operational requirements and performance against the impact on node availability in the computing environment. Reservations can be specified as immediate, delayed with high impact, delayed with low impact, and delayed recycled. An immediate reservation is suitable for high-priority tasks (e.g., creating the management layer for a Kubernetes cluster, where it is best if all nodes are immediately available). In one example, creating an immediate reservation will fail if there are currently fewer provisioned servers than required for the reservation.Creating a deferred high-impact or low-impact reservation will fail if the total number of servers (provisioned or unprovisioned) with the same node role as the pool is less than the requested number. Such an error likely means that more nodes need to be added to the NEM 210 pool to meet the reservation request. A high-impact deferred reservation can initially reserve provisioned servers and then procure additional servers, provisioned as needed to meet the demand for unprovisioned servers. This type of reservation is suitable for medium-priority tasks (such as scaling a workload). A low-impact deferred reservation can initially reserve unprovisioned servers and retain all remaining servers to meet the demand for provisioned servers.This type of reservation is suitable for tasks where time is not a critical factor, such as updating the image of nodes in a running workload. Reservations can have a specific duration. If a reservation is not fulfilled within this duration, it can time out, be deleted, and the remaining unclaimed servers can be made available for other reservations. Nodes acquired before the reservation timed out can remain in the pool. The Solution Model 312 (NEM) can be used to define and maintain node criteria (e.g., in the form of attributes) and the management of similar nodes for lifecycle implementations of solution workloads. These attributes can be defined for the solution and made available through the solution management API to orchestrators and IT staff who create and manage pools. Higher-level orchestrators can then use nodes from the pool, knowing that these nodes possess the desired properties / characteristics to fulfill the defined node roles, without having to worry about the node's backup server, location, image type, or image itself. This allows application developers to focus on responding to lifecycle states communicated through the NEM and reconfiguring their solution workloads accordingly. Solution Model 312 can also be used to specify the solution configuration and / or custom settings associated with tools / products integrated into NEM. Solution Model 312 can be tailored to a specific solution to accommodate the different processing styles of the solution workloads (e.g., Workloads 104). In the context of this example, the solution model 312 is represented by class objects (e.g., a server class 314 and a VM class 320), a node role 316, and a startup role 318. The class objects can be used to describe a resource classification (e.g., a server model and configuration number, or a VM resource configuration). Each node can be associated with a specific role (e.g., node role 316). The number and types of roles are specific to the type of solution and can be defined by an administrative user via a SaaS portal (e.g., SaaS Portal 114), which calls a solution management API (e.g., one of the APIs 220), which in turn can perform appropriate CRUD operations on objects within the solution model 312.To illustrate, examples of node roles in the context of a storage solution implementation could be those of a gateway node, a storage protocol translator node, and a data store node. Other solutions may define node roles more generally (e.g., Kubernetes master node vs. Kubernetes worker node). Node role 316 may include a startup role that describes the startup configuration (e.g., the server boot configuration) for the associated resource. In some implementations, criteria for minimum capabilities, capacities, and / or configurations (e.g., regarding compute resources, storage resources, and / or memory resources) may be specified for a node of a particular role.In conjunction with an MLOps solution, an administrator user can, for example, specify the requirement for a graphics processing unit (GPU) and / or a minimum number of CPU cores for a conclusion node to be used for performing the conclusion processing. The inventory model 322 can be used to define and maintain infrastructure details for selecting and creating the infrastructure to deliver solution workloads. An IT administrator or a higher-level orchestrator can use the inventory API (e.g., one of the APIs 220), directly or indirectly via the SaaS portal, to initialize an inventory with information about the racks within the infrastructure 112 and the servers housed in those racks. Changes to the infrastructure associated with the data processing environment (e.g., Data Processing Environment 100) should be logged in the Inventory API to keep the online inventory up to date, for example, with the inventory of the BM / VM infrastructure available in the data processing environment. Inventory model 322 can also be used to track the current image for each node role. The current image to be used for the node role in solution model 312 can be specified when the solution is released. As new images are made available over time to address issues (e.g., security vulnerabilities) and / or provide enhancements, the Inventory API can be used to track the latest image for each node role that should be used for the solution. In the context of this example, inventory model 322 is shown with objects representing a server 324, a VM 328, and a rack 326. Server 324 can have a specific server class (e.g., server class 314) and a specific node role (e.g., node role 316) and can be associated with a specific rack (e.g., rack 326) in which the corresponding physical server is located. Similarly, VM 328 can have a specific VM class (VM class 320) and be associated with a specific node role (e.g., node role 316). This three-tiered model allows a specific node (e.g., node 310) to be assigned to the backup server (BM or VM) and vice versa. State model Figure 4 shows a diagram illustrating a State Model 400 according to an example. State Model 400 is a non-restrictive example of State Model 238 from Figure 2. State Model 400 can be used by a NEM (e.g., NEM 210) to represent the current state of logical resource objects (e.g., nodes and node pools). In an example, state transitions within State Model 400 are determined by the model and the state manager, propagated through the State Model, and then used to trigger: (i) infrastructure plugin functionality, notifications to workloads (e.g., Workloads 104), for example, via application orchestration tools, and (iii) potentially additional orchestrations. The NEM can use an immutable node approach, meaning that state changes such as updates, restarts, trust violations, failures, or disturbances are handled by workload-based orchestration by releasing (deleting) the node from its pool. Deleting a node from a pool can trigger further state changes and additional predefined orchestrations within the NEM, such as the delete handling described in Table 1 (below). Using the State Model 400 simplifies workload-based orchestration relationships to state changes of logical resource objects. As shown in Table 1 (below), for example, workload-based orchestration can be structured to respond to state changes amplified at the pool level (e.g.,(This can mean that the highest priority state among all node states within a pool is used for that pool, thus simplifying the identification of pools and nodes to be processed.) The system then identifies the relevant node within the pool, performs appropriate preprocessing (e.g., pool expansion, workload evacuation, etc.), postprocessing (e.g., pool cleanup), and releases and captures resources according to the state change. For example, a predefined set of preprocessing and / or postprocessing operations to be performed on a pool or node can be specified by an administrator for each node state. In one example, a model and state manager (e.g., Model and State Manager 235), implemented in NEM, can be responsible for determining and orchestrating state changes associated with logical resource objects (e.g., nodes and node pools) that result from actions or operations performed by application orchestration tools (e.g., Application Orchestration 108), administrative users, and / or infrastructure plugins. In the context of this example, the node state can be set to an available state 404, a normal state 406, a replace state 408, a remove state 410, an error state 412, or a failure state 414. In another example, an immediate mode flag can also be part of the node state to indicate the degree of immediacy of the actions to be taken for the replace and remove states. When a new node is created within NEM, either through inventory changes or customer reservations, the underlying server is provisioned accordingly (e.g., a BM server in the case of a BM server), and the node starts in the 404 available state. In the 404 available state, the node is provisioned and ready for use by a workload, but not yet assigned to a pool. Once the node is assigned to a pool, it transitions to the 406 normal state. From the 406 normal state, the node can transition to the 408 replace state, the 410 remove state, or the 412 failure state. The node can enter the Replace state 408, for example, when it is determined that the node should be replaced either immediately (e.g., Immediate mode is true) or at the workload's discretion (Immediate mode is false). Immediate mode can be set to true if the replacement should occur immediately, such as after a hotfix release (e.g., when an urgent patch needs to be applied to a server in the workload) and the server needs to be updated and restarted immediately. Immediate mode can be set to false if the replacement can be performed at the workload's discretion, for example, if the Replace state coincides with a regularly scheduled maintenance window. After the node is removed from its pool, for example, by the application orchestration tool 108 via the Consumption API, and after the deletion process is successfully completed (e.g.,After reimaging, restarting and BIOS configuration, the node can be returned to the available 404 state. The node can, for example, enter the removal state 410 if it is determined that the node should be removed either immediately (e.g., immediate mode is true) or at the workload's discretion (immediate mode is false). Immediate mode can be set to true if removal should occur immediately, for example, due to an unplanned outage or a breach of trust. Immediate mode can be set to false if removal can be performed at the workload's discretion, for example, if the removal state aligns with a regularly scheduled maintenance window. In response to the node being removed from its pool, for example, by the application orchestration tool 108 via the consumption API, and after successful completion of the deletion process (e.g., reimaging), the node can be placed in the failure state 414 until an IT administrator completes the review. The node might enter error state 412, for example, if it detects that the backup server has failed or the trust server has been compromised. In response to a node entering error state 412, NEM can automatically delete the node and automatically place both the node and the backup server into failure state 414. As mentioned earlier, a node can transition from the "Remove" (410) or "Error" (412) state to the "Failure" (414) state. In the context of this example, nodes in the failure state (414) are not automatically restored and returned to the availability state (404). Instead, the IT administrator is expected to review the backup server and release it for reuse before the node can be returned to the availability state (404). From the failure state (414), the IT administrator can either delete the node (e.g., if the error cannot be resolved) or return it to the availability state (404) (e.g., after successfully resolving the error). In the context of this example, the states of the logical NEM resource objects have a priority hierarchy (e.g., Table 1 (below) lists the states in an example order of increasing priority, from lowest priority at the top to highest priority at the bottom), and the pool state is amplified in priority order. In an example of pool state amplification, the state of a given pool (pool state) can be determined based on the most severe state of the respective node states (node state) within that given pool. This allows workload-based orchestrations to efficiently query or listen for pool states (e.g., from a relatively small number of pools, typically fewer than five) to trigger state-change orchestrations, rather than querying or listening at the node level (e.g., with a typically much larger number of nodes).Table 1 - Example of node states and deletion processing in response to the deletion of the node from its pool. NormalK.AKA Normal operating state of a node. Node is connected to a pool. Replace False 1. Image 2. Restart 3. Firmware 1. The node must be updated within a regularly scheduled maintenance window. 1. Image update (update / upgrade) 2. Breach of trust. 1. The workload can be replaced. 2. The workload can be evacuated. 3. The pool can be expanded to replace nodes, or nodes can be recycled. ReplaceTrue1. Image2. Restart3. Firmware1. The node must be replaced immediately.2. The workload can be evacuated.3. The node should be recycled.1. Image update / restart (hotfix)2. Trust breach. Remove False 1. Image 2. Examine 1. The node must be removed within a regularly scheduled maintenance window. 2. The workload can be evacuated. 1. Failure (planned). 3. The pool must be expanded to remove knots; knots cannot be recycled. RemoveTrue1. Image2. Investigate1. The node must be removed immediately.2. The workload can be evacuated.3. The pool must be expanded to remove the node; the node cannot be recycled.1. Failure (immediate)2. Breach of trust. ForcedFaultTrue 1. The backup server failed and the node was immediately deleted by NEM. 2. The workload was lost and must be rebuilt. 1. Monitored server failure. 2. Trusted server compromised. 3. The pool must be cleaned up and expanded after the failure of a node to ensure the scalability of the pool. NEM processing Figure 5 shows a flowchart illustrating the processing performed by a NEM according to an example. The processing can be implemented, for example, without limitation, using the NEM 210 from Figure 2 to facilitate lifecycle management for workloads (e.g., Workloads 104) on heterogeneous infrastructure (e.g., Infrastructure 112). Block 510 maintains a consumption model in which the heterogeneous infrastructure is represented in a generalized form as logical resource objects. In one example, the logical resource objects include nodes and pools of nodes. Each node can have a node role (e.g., node role 316), which is defined, for example, in a solution model (e.g., solution model 312) of an object model (e.g., object model 300) of the NEM. Based on the nodes' respective attributes, the respective node roles can specify certain functionality that the nodes can provide for a workload (e.g., one of the workloads 104). The nodes' attributes can be expressed in terms of their respective compute configuration, storage, and / or network resources. Block 520 manages a state model through which the logical resource objects transition between multiple states. In response to these transitions, notifications can be sent to an application orchestration tool associated with the workload (e.g., application orchestration tool 108). The multiple states and transitions can be described as above with reference to Figure 4. Block 530 abstracts the interactions of the application orchestration tool with elements of the heterogeneous infrastructure used by the workload by providing a consumption API (e.g., one of APIs 220) through which requirements for managing the lifecycle of the heterogeneous infrastructure with respect to the logical resource objects are expressed. As described in Fig. 3, the consumption API can, for example, support various CRUD operations, management operations, and lifecycle implementations on nodes and node pools. Figure 6 shows a flowchart illustrating the processing by a NEM according to another example. The processing can be implemented, for example, and without limitation, using the NEM 210 from Figure 2 to facilitate lifecycle management for workloads (e.g., Workloads 104) on heterogeneous computing resources within an infrastructure (e.g., Infrastructure 112). Block 610 provides a consumption API to receive a request from an application orchestration tool (e.g., application orchestration tool 108) to manage the lifecycle of heterogeneous compute resources for a workload. In one example, the application orchestration tool can use the consumption API to indirectly manage and / or manipulate the underlying infrastructure through interactions with logical resource objects (e.g., nodes and pools of nodes). In block 620, the request is executed by invoking an integrated infrastructure plugin to perform an operation associated with the request. For example, creating a node set or reserving a node set for allocation to a specific pool might cause NEM to immediately or in the background instigate a BM provisioning tool or a virtualization tool to provision the corresponding BM or VM infrastructure. In block 630, in response to the completion of the call, an object model and / or a state model are updated, representing the lifecycle of the heterogeneous compute resources for the workload. For example, in response to the completion of the provisioning of the BM or VM infrastructure, the reserved nodes may transition from an available state (e.g., available state 404) to a normal state (e.g., normal state 406), depending on the node's state. Block 640 notifies the application orchestration tool of state changes to the heterogeneous computing resources within the state model. As described above with reference to Fig. 4, the application orchestration tool can be notified, for example, via automatic notifications from a notifier (e.g., Notifier 232) or REST-based polling, of node state changes that indicate the need to replace, remove, or restore a failed node (e.g., as a result of performing a hotfix, the availability of an image update, regularly scheduled maintenance, or the detection of a server failure). While the examples described with reference to Figures 5-6 contain a number of enumerated blocks, other examples may contain additional blocks before, after, and / or between the enumerated blocks. Likewise, in some examples, one or more of the enumerated blocks may be omitted or executed in a different order. Computer system Figure 7 shows a block diagram of a computer system according to the first example. In the example shown in Figure 7, the computer system 700 comprises a processing resource 710 coupled to a non-transient, machine-readable medium 720 encoded with instructions for performing one or more of the processes described herein. The computer system 700 may be a server, a server cluster, a computer appliance, a workstation, a converged system, a hyperconverged system, or the like. The computer system 700 may be part of the infrastructure to be managed (e.g., infrastructure 112) within a particular computing environment (e.g., cloud computing environment 100). In other examples, the computer system 700 may reside in a cloud (e.g., a public cloud) and be connected to the infrastructure to be managed (e.g.,communicate with infrastructure 112) or be an administration server in the same data center as the infrastructure being managed. The processing resource 710 may comprise a microcontroller, a microprocessor, CPU core(s), GPU core(s), an ASIC, an FPGA, and / or other hardware device suitable for retrieving and / or executing instructions from the machine-readable medium 720 to perform the functions relating to various examples described herein. Additionally or alternatively, the processing resource 710 may include an electronic circuit for executing the functionality of the instructions described herein. The machine-readable medium 720 can be any medium suitable for storing executable instructions. Non-restrictive examples of machine-readable media 720 include RAM, ROM, EEPROM, flash memory, a hard disk drive, an optical disk, or the like. The machine-readable medium 720 can be a non-transitory medium, the term "non-transitory" excluding any transitive transmission signals. The machine-readable medium 720 can be located within the computer system 700, as shown in Fig. 7, in which case the executable instructions can be considered to be "installed" or "embedded" on the computer system 700. Alternatively, the machine-readable medium 720 can be a portable (e.g., external) storage medium and part of an "installation package."The instructions stored on the machine-readable medium 720 may be useful to implement at least part of one or more of the methods described here. In the context of the present example, the machine-readable medium 720 is encoded with a series of executable instructions 730-750. It is understood that some or all of the executable instructions and / or electronic circuits contained in one block may be contained in alternative implementations in another block shown in the figures or in another block not shown. Instructions 730, when executed, can cause the processing resource 710 to maintain a consumption model in which the heterogeneous infrastructure is represented in a generalized form as logical resource objects, including nodes and multiple pools of nodes. As an example, instructions 730 can be useful for executing block 510 of Fig. 5. Instructions 740, when executed, can cause the processing resource 710 to maintain a state model, allowing the logical resource objects to switch between multiple states and triggering notifications to an application orchestration tool associated with the workload. For example, instructions 740 can be useful for executing block 520 of Figure 5. When executed, instructions 750 can cause processing resource 710 to abstract the interactions of the application orchestration tool with the heterogeneous infrastructure used by the workload by providing an API through which requirements for managing the lifecycle of the heterogeneous infrastructure can be expressed with reference to the logical resource objects. As an example, instructions 750 can be useful for executing block 530 of Figure 5. The foregoing description includes numerous details to facilitate understanding of the subject matter disclosed herein. However, the implementation can also be carried out without some or all of these details. Other embodiments may include modifications and deviations from the details described above. It is intended that the following claims cover such modifications and variations.
Claims
A method comprising: maintaining a consumption model (304) in which a heterogeneous infrastructure (112) is represented in a generalized form as logical resource objects, including nodes (310) and a plurality of pools (308) of the nodes (310), wherein the nodes (310) have respective node roles (316) that indicate specific functionality, how the nodes (310) can be operated to provide functionality for a workload (104) based on the respective attributes of the nodes (310); maintaining a state model (238, 400) by which the logical resource objects transition between a plurality of states, and whereupon notifications are provided to an application orchestration tool (108) associated with the workload (104);and abstracting interactions by the application orchestration tool (108) with a subset of the heterogeneous infrastructure (112) used by the workload (104) by providing an application programming interface, API (220), through which requests for managing a lifecycle of the subset of the heterogeneous infrastructure (112) are expressed with reference to the logical resource objects. The method of claim 1, further comprising the execution of the requests by calling integrated infrastructure plugins to perform operations associated with the requests. Method according to claim 2, wherein the integrated infrastructure plugins comprise a bare-metal deployment tool and a virtualization tool. The method of claim 1, further comprising notifying the application orchestration tool (108) of state model changes. Method according to claim 1, wherein the respective attributes are expressed in the form of respective capabilities, capacities or configurations of the nodes (310). The method according to claim 1, which further comprises providing a reservation system (234) through which the nodes (310) are assigned to the plurality of pools (308). Method according to claim 6, wherein the assignment of the nodes (310) to the multiple pools (308) facilitates the management of the nodes (310) as groups with consistent automated lifecycle management operations and states. The method of claim 1, further comprising, as part of a background process and prior to the allocation to the plurality of pools (308), providing a subset of the nodes (310) with a specific node role (316) to facilitate immediate availability for use by the workload (104). A system comprising: a processing resource (710); and a non-transitory, machine-readable medium (720) coupled to the processing resource (710) and containing instructions which, when executed by the processing resource (710), cause the processing resource (710) to: provide a consumption application programming interface, API (220), to receive from an application orchestration tool (108) a request to manage a lifecycle of heterogeneous compute resources for a workload (104); execute the request by invoking an integrated infrastructure plugin to perform an operation associated with the request; update an object model (236, 300) or state model (238, 400) representing a lifecycle of heterogeneous compute resources for the workload (104) in response to the completion of the call;and to notify the application orchestration tool (108) of state changes of the heterogeneous computing resources.; System according to claim 9, wherein the instructions further cause the processing resource (710) to provide availability of the heterogeneous computing resources for workloads (104) managed by the application orchestration tool (108) via a reservation system (234). System according to claim 10, wherein the reservation system (234) allocates the heterogeneous computing resources to a pool (308) to facilitate the management of the heterogeneous computing resources as a group with consistent automated lifecycle management operations and states. System according to claim 9, wherein the integrated infrastructure plugin comprises a bare-metal deployment tool or a virtualization tool. A non-transitory, machine-readable medium (720) that stores instructions which, when executed by a processing resource (710) of a computer system, cause the processing resource (710) to: maintain a consumption model (304) in which a heterogeneous infrastructure (112) is represented in a generalized form as logical resource objects, including nodes (310) and a plurality of pools (308) of the nodes (310), wherein the nodes (310) have respective node roles (316) that indicate a specific functionality, how the nodes (310) can be operated to provide functionality for a workload (104) based on the respective attributes / qualities of the nodes (310);to maintain a state model (238, 400) through which the logical resource objects transition between a variety of states, and whereupon notifications are provided to an application orchestration tool (108) associated with the workload (104); and to abstract interactions by the application orchestration tool (108) with a subset of the heterogeneous infrastructure (112) used by the workload (104) by providing an application programming interface, API (220), through which requests for managing a lifecycle of the subset of the heterogeneous infrastructure (112) with reference to the logical resource objects are expressed. Non-transitory, machine-readable medium (720) according to claim 13, wherein the instructions further cause the processing resource (710) to execute the requests by making calls to integrated infrastructure plugins to perform operations associated with the requests. Non-transitory, machine-readable medium (720) according to claim 14, wherein the integrated infrastructure plugins comprise a bare-metal deployment tool and a virtualization tool. Non-transitory, machine-readable medium (720) according to claim 14, wherein the instructions further cause the processing resource (710) to inform the application orchestration tool (108) about state model changes. Non-transitory, machine-readable medium (720) according to claim 14, wherein the respective attributes / qualities are expressed in the form of respective capabilities, capacities and / or configurations of the nodes (310). Non-transitory, machine-readable medium (720) according to claim 14, wherein the instructions further cause the processing resource (710) to use a reservation system (234) to allocate the nodes (310) to the plurality of pools (308). Non-transitory, machine-readable medium (720) according to claim 18, wherein the assignment of the nodes (310) to the plurality of pools (308) facilitates the management of the nodes (310) as groups with consistent automated lifecycle management operations and states. Non-transitory, machine-readable medium (720) according to claim 14, wherein the instructions further cause the processing resource (710) to provide, as part of a background process and prior to allocation to the plurality of pools (308), a subset of the nodes (310) with a specific node role (316) to facilitate immediate availability for use by the workload (104).