Dynamic POD resource restriction adjustment based on data analysis
By monitoring and predicting the resource usage of pods in a container orchestration platform and dynamically adjusting pod resource limits, the problem of unreasonable resource allocation in existing technologies is solved, achieving efficient utilization of computing resources and cost reduction.
Patent Information
- Application Number
- CN202480029031.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-05-05
- Filing Date
- 2024-05-02
- Publication Date
- 2025-11-28
AI Technical Summary
Existing technologies make it difficult to dynamically adjust pod resource limits in container orchestration platforms, resulting in unreasonable resource allocation, waste, high costs, and difficulty in effectively utilizing computing resources.
By monitoring the runtime resource usage of pods on the container orchestration platform, using trained machine learning models to predict upcoming resource demands, and dynamically adjusting pod resource limits, combined with autoscaling mode, resources can be allocated and utilized on demand.
It enables cost-effective use of computing resources, reduces energy consumption and carbon footprint, improves resource utilization efficiency, and avoids resource waste.
Smart Images

Figure CN121039636A_ABST
Abstract
Description
Background Technology
[0001] One or more aspects generally involve facilitating processing within a computing environment, and specifically involve dynamically adjusting the pod resource limits of a container orchestration platform pod at runtime within the computing environment.
[0002] Container orchestration platforms and / or tools provide a framework for managing containers and microservice architectures on a scale. Container orchestration automates the development, management, scaling, and networking of containers and can be used by enterprises that need to deploy and manage hundreds or even thousands of containers.
[0003] A container provides application deployment units and a self-contained execution environment for microservice-based applications. Containers enable multiple parts of an application to run independently as microservices on the same hardware, providing greater control over individual application components and their lifecycles.
[0004] Container orchestration platforms typically support clusters, a control plane, agents, and one or more pods. A cluster consists of the control plane and one or more computing machines or nodes. The control plane is a collection of processes that control the nodes of the container orchestration platform and is where task assignment begins. An agent is a service that runs on a node, reads the container manifest, and ensures that the defined containers are started and running. A pod is a group of one or more containers deployed to a single node. Containers within a pod share IP addresses, IPCs, hostnames, and other resources. Summary of the Invention
[0005] This paper overcomes certain shortcomings of the prior art by providing a computer-implemented method to facilitate processing within a computing environment, and offers additional advantages. The computer-implemented method includes deploying a container orchestration platform container with one or more container resources in the computing environment. Each pod resource has one or more associated pod resource limits. Furthermore, the computer-implemented method includes monitoring the runtime resource usage of the container orchestration platform pods and predicting upcoming resource usage of the container orchestration platform pods by a trained machine learning model, the prediction using at least part of the monitored runtime resource usage. Additionally, the computer-implemented method includes dynamically adjusting the pod resource limits of the one or more pod resource limits of the container orchestration platform pods in the computing environment. This dynamic adjustment is based on the monitored runtime resource usage and the predicted upcoming resource usage. Advantageously, processing is facilitated in container-based computing environments by providing enhanced use of resources. By dynamically adjusting pod resource limits, cost-effective use of computing resources, as well as efficient and dynamic utilization of resources, is provided, supporting on-demand resource allocation without unnecessary waste. Furthermore, this process saves costs, reduces carbon footprint, and decreases energy consumption by using computing resources more effectively in container-based computing environments.
[0006] In one implementation, the computer-implemented method further includes obtaining a trained machine learning model by training a machine learning model on historical resource usage data of the container orchestration platform pod. Furthermore, in one embodiment, the computer-implemented method further includes using predicted upcoming resource usage of the container orchestration platform pod while continuing to train the trained machine learning model. In one example, the machine learning model includes a linear regression model. Advantageously, using a trained machine learning model to predict upcoming resource usage of the container orchestration platform pod facilitates dynamic adjustment of pod resource constraints to provide, for example, cost-effective use of computing resources and efficient and dynamic utilization of resources, and supports on-demand resource allocation without resource waste.
[0007] In one implementation, the deployment includes initializing and deploying the container orchestration platform pod. Initialization and deployment include, for example, obtaining a dynamic resource definition that will be used to dynamically allocate pod resources for the container orchestration platform pod being deployed. The dynamic resource definition includes a resource usage formula. Furthermore, initialization and deployment include creating a deployment object that will be used when deploying the container orchestration platform pod. In one embodiment, the creation includes resolving the resource usage formula of the dynamic resource definition and, based on the resolved resource usage formula, generating one or more initial pod resource limits for initializing the container orchestration platform pod. Additionally, initialization and deployment include initializing and deploying a container orchestration platform pod having one or more container resources in the computing environment, and applying the generated one or more initial container resource limits to the deployment of the container orchestration platform pod having one or more container resources. Advantageously, generating one or more initial pod resource limits for initializing the container orchestration platform pod based on the resource usage formula of the dynamic resource definition further enhances resource utilization, for example, by setting the initial pod resource limits through the resource usage formula, to efficiently initialize the container orchestration platform pod with appropriate pod resource limits.
[0008] In one implementation, the monitored runtime resource usage includes the runtime CPU usage and runtime memory usage of the container orchestration platform pod. Advantageously, dynamically adjusting pod resource limits based on the runtime CPU and runtime memory usage of the container orchestration platform pod allows for more efficient use of computing resources in the container-based computing environment, thereby saving processing costs, reducing carbon footprint, and decreasing energy consumption.
[0009] In one embodiment, dynamically adjusting includes dynamically increasing the pod resource limit of a container orchestration platform pod during runtime in the computing environment. In another implementation, dynamically adjusting includes dynamically decreasing the pod resource limit of a container orchestration platform pod during runtime in the computing environment. Advantageously, dynamically increasing and / or decreasing the pod resource limit of a container orchestration platform pod during runtime in the computing environment provides cost-effective use of computing resources, as well as efficient and dynamic utilization of resources, and supports on-demand resource allocation without allocating unnecessary resources.
[0010] In one implementation, the dynamic adjustment is also based on the autoscaling mode of the container orchestration platform pods. In one example, when the autoscaling mode is a post-intervention mode, the dynamic adjustment is also based on the number of instances of the container orchestration platform pods remaining at a specified pod instance limit for a set period of time. In another example, when the autoscaling mode is a pre-initiated mode, the dynamic adjustment occurs before the autoscaling of the number of instances of the container orchestration platform pods. Advantageously, dynamically adjusting pod resource limits based on the autoscaling mode provides the ability to dynamically adjust pod resource limits, post-level pod autoscaling, or pre-level pod autoscaling. This further enhances the efficient and dynamic utilization of resources, for example, to better balance load processing and avoid inefficient system operations, such as slow operation of computing systems.
[0011] In one implementation, the computer-implemented method further includes dynamically adjusting pod resource limits based on container orchestration platform pods, dynamically adjusting one or more pod resources from one or more other container orchestration platform pods within the computing environment. Advantageously, dynamically adjusting one or more pod resources or one or more other container orchestration platform pods within the computing environment based on pod resource limits scales resources up and / or down across multiple container orchestration platform pods according to the needs within the container-based computing environment, promoting resource scalability within the computing environment. This facilitates processing within the container-based computing environment and enhances resource utilization. Furthermore, the dynamic utilization of resources across multiple container orchestration platform pods better balances workload processing and enhances the efficient operation of the container-based computing environment.
[0012] This document also describes and claims protection for computer systems and computer program products related to one or more aspects. Furthermore, this document describes and claims protection for services related to one or more aspects.
[0013] Additional features and advantages are achieved through the techniques described herein. Other embodiments and aspects are described in detail herein and are considered part of the claimed aspects. Attached Figure Description
[0014] One or more aspects are specifically pointed out and clearly claimed by way of example in the claims at the end of the specification. The foregoing and objectives, features, and advantages of one or more aspects will become apparent from the following detailed description taken in conjunction with the accompanying drawings:
[0015] Figure 1 An example of a computing environment that includes and / or uses one or more aspects of the present invention is shown;
[0016] Figure 2An embodiment of a computer program product having a pod resource limitation adjustment module according to one or more aspects of the present invention is described;
[0017] Figure 3 An embodiment of a pod resource limit adjustment process according to one or more aspects of the present invention is described;
[0018] Figure 4 Another example of a container-based computing environment that includes and / or uses one or more aspects of the present invention is described;
[0019] Figure 5 An example of a service mesh for a container-based computing environment according to one or more aspects of the present invention, including pod resource constraint adjustment, is described.
[0020] Figure 6 Depicting one or more aspects of the invention Figure 4 An example of the workflow of the control components of a control system;
[0021] Figure 7A Depicting one or more aspects of the invention Figure 4 An example of the workflow of the predictive component of a control system;
[0022] Figure 7B Another example of a computing environment that includes and / or uses one or more aspects of the present invention is described;
[0023] Figure 7C Depicting one or more aspects of the invention Figure 4 Further examples of the prediction workflow for the prediction component;
[0024] Figure 8 Depicting one or more aspects of the invention Figure 4 An example of the workflow of the message service component of a control system;
[0025] Figure 9 Description of one or more aspects of the present invention Figure 4 An example of the workflow of the Vertical Pod Autoscaler (VPA) engine component of the control system; and
[0026] Figure 10 According to one or more aspects of the present invention Figure 4 Another example workflow of the vertical pod autoscaler engine component of the control system. Detailed Implementation
[0027] The accompanying drawings, incorporated in and forming a part of this specification, further illustrate the invention and, together with the detailed descriptions of the invention, serve to explain various aspects of the invention. Note that, in this respect, descriptions of well-known systems, devices, processing techniques, etc., have been omitted to avoid unnecessarily obscuring the details of the invention. However, it should be understood that the detailed descriptions and specific examples, while indicating aspects of the invention, are given by way of illustration only and not by way of limitation. Various substitutions, modifications, additions, and / or other arrangements within the spirit or scope of the basic concept of the invention will be apparent to those skilled in the art from this disclosure. It should also be noted that numerous aspects or features of the invention are disclosed herein, and unless inconsistent, each disclosed aspect or feature may be combined with any other disclosed aspect or feature as required by the specific application of the disclosed concept.
[0028] It should also be noted that specific code, designs, architectures, protocols, layouts, diagrams, or tools are used to describe illustrative embodiments below, and these are merely examples and not intended to limit the scope. Furthermore, for clarity of description, specific software, hardware, tools, or data processing environments are used in certain instances to describe illustrative embodiments, and these are merely examples. Illustrative embodiments can be used in conjunction with other equivalent or similar purpose structures, systems, applications, or architectures. One or more aspects of the illustrative embodiments can be implemented in software, hardware, or a combination thereof.
[0029] As will be understood by those skilled in the art, the program code referenced in this application may include software and / or hardware. For example, the program code in some embodiments of the invention may utilize a software-based implementation of the described functionality, while other embodiments may include fixed-function hardware. Some embodiments may combine both types of program code. Figure 1 Examples of program code, also known as one or more programs, are depicted, including an operating system 122 and a pod resource limit adjustment module 200 stored in a persistent storage device 113.
[0030] One or more aspects of this invention are incorporated into, executed by, and / or used by a computing environment. As examples, the computing environment can be of various architectures and types, including but not limited to: personal computing, client-server, distributed, virtual, simulation, partitioned, non-partitioned, cloud-based, quantum, grid, time-sharing, clustered, peer-to-peer, mobile, having one or more nodes, having one or more processors, and / or any other type of environment and / or configuration capable of performing, for example, a process (or processes) such as the pod resource constraint adjustment processing disclosed herein. Aspects of this invention are not limited to a particular architecture or environment.
[0031] Before further describing detailed embodiments of the present invention, reference is made below. Figure 1 The discussion includes examples of computing environments that use one or more aspects of the present invention.
[0032] Various aspects of this disclosure are described by narrative text, flowcharts, block diagrams of computer systems, and / or block diagrams of machine logic included in embodiments of a computer program product (CPP). Regarding any flowchart, depending on the technology involved, operations may be performed in a different order than that shown in a given flowchart. For example, again according to the technology involved, two operations shown in consecutive flowchart blocks may be performed in reverse order, as a single integrated step, simultaneously, or in a manner that at least partially overlaps in time.
[0033] Computer Program Product Embodiment (“CPP Embodiment” or “CPP”) is a term used in this disclosure to describe any collection of one or more storage media (also referred to as “media”) collectively included in a collection of one or more storage devices, the collection of one or more storage devices collectively including machine-readable code corresponding to instructions and / or data for performing computer operations specified in a given CPP claim. A “storage device” is any tangible device capable of holding and storing instructions used by a computer processor. Without limitation, a computer-readable storage medium can be an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices that include these media include: magnetic disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory sticks, floppy disks, mechanical encoding devices (such as punch cards or pits / platforms formed in the main surface of the disk), or any suitable combination of the foregoing. Computer-readable storage media, as used in this disclosure, should not be construed as storing transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides, optical pulses through fiber optic cables, electrical signals transmitted through wires, and / or other transmission media. As those skilled in the art will understand, data is typically moved at certain incidental points in time during the normal operation of the storage device, such as during access, defragmentation, or garbage collection; however, this does not make the storage device transient, because the data is not transient when it is stored.
[0034] Computing environment 100 includes examples of environments for executing at least some of the computer code involved in performing the methods of the present invention, such as block 200 of a pod resource limit adjustment module. In addition to block 200, computing environment 100 includes, for example, a computer 101, a wide area network (WAN) 102, an end-user device (EUD) 103, a remote server 104, a public cloud 105, and a private cloud 106. In this embodiment, computer 101 includes a processor set 110 (including processing circuitry 120 and a cache 121), communication infrastructure 111, volatile memory 112, persistent storage device 113 (including an operating system 122 and block 180, as described above), a peripheral device set 114 (including a user interface (UI) device set 123, storage device 124, and an Internet of Things (IoT) sensor set 125), and a network module 115. Remote server 104 includes a remote database 130. Public cloud 105 includes gateway 140, cloud coordination module 141, host physical machine set 142, virtual machine set 143, and container set 144.
[0035] Computer 101 can take the form of a desktop computer, laptop computer, tablet computer, smartphone, smartwatch or other wearable computer, mainframe computer, quantum computer, or any other form of computer or mobile device now known or to be developed in the future capable of running programs, accessing networks, or querying databases such as remote database 130. As is well known in the field of computer technology, and depending on the technology, the performance of a computer-implemented method can be distributed across multiple computers and / or multiple locations. On the other hand, in this presentation of computing environment 100, the detailed discussion focuses on a single computer, specifically computer 101, to keep the presentation as simple as possible. Computer 101 can reside in the cloud, even... Figure 1 It is not shown in the cloud, and on the other hand, computer 101 does not need to be in the cloud unless it can be indicated with certainty to any extent.
[0036] Processor assembly 110 includes one or more computer processors of any type now known or to be developed in the future. Processing circuitry 120 may be distributed across multiple packages, such as multiple cooperating integrated circuit chips. Processing circuitry 120 may implement multiple processor threads and / or multiple processor cores. Cache 121 is memory located within the processor chip package and is typically used for data or code that should be readily accessible by the threads or cores running on processor assembly 110. Cache memory is typically organized into multiple levels based on its relative proximity to the processing circuitry. Alternatively, some or all of the cache in the processor assembly may be located “off-chip.” In some computing environments, processor assembly 110 may be designed to work with qubits and perform quantum computing.
[0037] Computer-readable program instructions are typically loaded onto computer 101 to cause the processor set 110 of computer 101 to perform a series of operational steps to implement a computer-implemented method, such that the instructions thus executed instantiate the method specified in the flowcharts and / or descriptive descriptions of the computer-implemented method included in this document (collectively, the “method of the invention”). These computer-readable program instructions are stored in various types of computer-readable storage media, such as cache 121 and other storage media discussed below. The program instructions and associated data are accessed by the processor set 110 to control and direct the execution of the method of the invention. In computing environment 100, at least some of the instructions for performing the method of the invention may be stored in permanent storage device 113, within block 180.
[0038] Communication structure 111 is a signal transmission path that allows the various components of computer 101 to communicate with each other. Typically, this structure consists of switches and conductive paths, such as switches and conductive paths that form buses, bridges, physical input / output ports, etc. Other types of signal communication paths can be used, such as fiber optic communication paths and / or wireless communication paths.
[0039] Volatile memory 112 is any type of volatile memory now known or to be developed in the future. Examples include dynamic random access memory (RAM) or static RAM. Typically, volatile memory 112 is characterized by random access, but this is not necessary unless explicitly stated otherwise. In computer 101, volatile memory 112 is located in a single package and is internal to computer 101; however, alternatively or additionally, volatile memory may be distributed across multiple packages and / or located externally relative to computer 101.
[0040] The persistent storage device 113 is any form of non-volatile memory known now or developed in the future for use with a computer. The non-volatility of this memory means that the stored data is retained regardless of whether power is supplied to the computer 101 and / or directly to the persistent storage device 113. The persistent storage device 113 may be a read-only memory (ROM), but typically at least a portion of the persistent memory allows data to be written, deleted, and rewritten. Some common forms of persistent storage include hard disks and solid-state storage devices. The operating system 122 may take several forms, such as various known proprietary operating systems or operating systems employing an open-source portable operating system interface type with a kernel. The code included in box 126 typically includes at least some of the computer code involved in performing the methods of the present invention.
[0041] Peripheral device set 114 includes a set of peripheral devices for computer 101. Data communication connections between peripheral devices and other components of computer 101 can be implemented in various ways, such as Bluetooth connectivity, near field communication (NFC) connectivity, connections made by cables (such as Universal Serial Bus (USB) type cables), plug-in connections (e.g., secure digital (SD) cards), connections made through local area communication networks, and even connections made through wide area networks such as the Internet. In various embodiments, UI device set 123 may include components such as displays, speakers, microphones, wearable devices (such as goggles and smartwatches), keyboards, mice, printers, touchpads, game controllers, and haptic devices. Storage device 124 is an external storage device, such as an external hard drive, or a pluggable storage device, such as an SD card. Storage device 124 can be permanent and / or volatile. In some embodiments, storage device 124 may take the form of a quantum computing storage device for storing data in the form of qubits. In embodiments where computer 101 requires substantial storage (e.g., where computer 101 locally stores and manages a large database), this storage can be provided by peripheral storage devices designed for storing very large amounts of data, such as a Storage Area Network (SAN) shared by multiple geographically distributed computers. The IoT sensor set 125 comprises sensors that can be used in IoT applications. For example, one sensor could be a thermometer, while another could be a motion detector.
[0042] Network module 115 is a collection of computer software, hardware, and firmware that allows computer 101 to communicate with other computers via WAN 102. Network module 115 may include hardware such as a modem or Wi-Fi transceiver, software for packetizing and / or depacketizing data transmitted over the communication network, and / or web browser software for transmitting data over the Internet. In some embodiments, the network control and network forwarding functions of network module 115 are performed on the same physical hardware device. In other embodiments (e.g., embodiments utilizing Software-Defined Networking (SDN)), the control and forwarding functions of network module 115 are performed on physically separate devices, such that the control function manages several different network hardware devices. Computer-readable program instructions for performing the methods of the present invention can typically be downloaded to computer 101 from an external computer or external storage device via a network adapter card or network interface included in network module 115.
[0043] WAN 102 is any wide area network (e.g., the Internet) capable of transmitting computer data over non-local distances using any technology known now or developed in the future for transmitting computer data. In some embodiments, WAN 102 may be replaced by and / or supplemented by a local area network (LAN) designed to transmit data between devices located in a local area (e.g., a Wi-Fi network). WANs and / or LANs typically include computer hardware such as copper transmission cables, fiber optic transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and edge servers.
[0044] End User Equipment (EUD) 103 is any computer system used and controlled by an end user (e.g., a customer of the enterprise operating computer 101) and can take any of the forms discussed above in conjunction with computer 101. EUD 103 typically receives useful and available data from the operation of computer 101. For example, assuming computer 101 is designed to provide recommendations to the end user, these recommendations are typically transmitted from network module 115 of computer 101 to EUD 103 via WAN 102. In this way, EUD 103 can display or otherwise present the recommendations to the end user. In some embodiments, EUD 103 can be client equipment, such as a thin client, heavy client, mainframe, desktop computer, etc.
[0045] Remote server 104 is any computer system that provides at least some data and / or functionality to computer 101. Remote server 104 can be controlled and used by the same entity operating computer 101. Remote server 104 represents a machine that collects and stores useful and available data used by other computers, such as computer 101. For example, if computer 101 is designed and programmed to provide recommendations based on historical data, that historical data can be provided to computer 101 from a remote database 130 of remote server 104.
[0046] Public cloud 105 is any computer system that can be used by multiple entities, providing on-demand availability of computer system resources and / or other computing capabilities (particularly data storage (cloud storage) and computing power) without the need for direct, active management by users. Cloud computing typically leverages resource sharing to achieve scalability consistency and economy. Direct and active management of the computing resources of public cloud 105 is performed by the computer hardware and / or software of cloud coordination module 141. The computing resources provided by public cloud 105 are typically implemented by virtual computing environments running on various computers constituting the host physical machine set 142, which is the entirety of physical computers in and / or available to the public cloud 105. Virtual computing environments (VCEs) typically take the form of virtual machines from virtual machine set 143 and / or containers from container set 144. It should be understood that these VCEs can be stored as images and can be transferred between various physical machine hosts as images or after the VCEs are instantiated. Cloud coordination module 141 manages the transfer and storage of images, deploys new instantiations of VCEs, and manages the active instantiation of VCE deployments. Gateway 140 is a collection of computer software, hardware, and firmware that allow public cloud 105 to communicate via WAN 102.
[0047] Now, we will provide some further explanation of Virtualized Computing Environments (VCEs). A VCE can be stored as an "image." A new active instance of a VCE can be instantiated from this image. Two common types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to an operating system feature where the kernel allows multiple isolated user-space instances, called containers, to exist. From the perspective of the programs running within them, these isolated user-space instances typically appear as actual computers. Computer programs running on a regular operating system can utilize all the resources of that computer, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, programs running within a container can only use the contents of the container and the devices allocated to the container; this is a characteristic known as containerization.
[0048] Private cloud 106 is similar to public cloud 105, except that computing resources are available only to a single enterprise. While private cloud 106 is depicted as communicating with WAN 102, in other embodiments, private cloud may be completely disconnected from the Internet and accessible only via a local / private network. A hybrid cloud is a combination of multiple clouds of different types (e.g., private, community, or public cloud types) typically implemented by different vendors. Each of the multiple clouds remains a separate and discrete entity, but the larger hybrid cloud architecture is bound together by standardization or proprietary technology that enables coordination, management, and / or data / application portability across the multiple component clouds. In this embodiment, public cloud 105 and private cloud 106 are both part of a larger hybrid cloud.
[0049] The computing environment described above is merely one example of a computing environment for incorporating, executing, and / or using one or more aspects of the present invention. Other examples are possible. Furthermore, in one or more embodiments, Figure 1 One or more components / modules need not be included in the computing environment and / or used in one or more aspects of the invention. Furthermore, in one or more embodiments, additional and / or other components / modules may be used. Other variations are possible.
[0050] As shown, in one example, computing environment 100 supports containers. Containers can be provided in the cloud, such as public cloud (e.g., public cloud 105), private cloud (e.g., private cloud 106), hybrid cloud, and / or on-premises (e.g., computer 101). In one example, containers are managed by one or more container orchestration platforms. An example of such a platform is Kubernetes. ® Kubernetes is an open-source, scalable, and portable container management platform. It is a registered trademark of the Linux Foundation, San Francisco, CA. Other platforms can also be used. For example, in Kubernetes... ® In a container, there is its own shared central processing unit, file system, process space, memory, etc. Furthermore, due to their loose isolation properties, containers can share the operating system (OS) between applications; containers are decoupled from the underlying infrastructure; containers can be distributed and moved across operating systems and to the cloud; and each container is repeatable. Containers are designed to be stateless, and the code running within a container is immutable – it is not modified; instead, new container images are built to include any code changes.
[0051] As described, container orchestration platforms typically support one or more containers, where a pod is a group of one or more containers deployed to a single node. Containers within a pod typically share IP addresses, IPCs, hostnames, and other resources. In one example, the compute environment could employ Kubernetes. ® And / or another platform to manage containers. Kubernetes ® A Kubernetes cluster is a platform used to run and manage containers from multiple container runtimes. The computing environment may include one or more nodes, an operating system shared by those nodes, and the underlying hardware used by those nodes, such as processing units. Nodes can be virtual or physical machines, and they can be in a farm (e.g., in computer 101 and / or other computing devices) and / or in a cloud environment (e.g., public cloud 105, private cloud 106, hybrid cloud environment, and / or other cloud environment). In one example, a node includes a container runtime; one or more pods; a proxy; and a client agent. An example of a proxy is the Kube proxy, a network proxy that runs on each node in the cluster, implementing Kubernetes. ® As part of the service concept, the Kube agent maintains network rules on the nodes, and these rules allow network communication from network sessions inside or outside the cluster to pods. An example of a client agent is a Kubelet running on each node. It can register the node with an Application Programming Interface (API) server using one or more of the hostname, flags, or other indicators, which then validates and configures data for objects (e.g., pods). In other examples, the platform is not Kubernetes. ® On this platform, intermediate proxies and client proxies can be used. Many examples are possible.
[0052] In one example, a container runtime interface is provided, which is a plugin interface that enables client agents (e.g., Kubelet) to use various container runtimes without having to recompile cluster components. Furthermore, in one example, a pod includes one or more containers, and each container includes a container image, for example, containing one or more applications with one or more libraries and / or one or more binary code and / or text resources. The container images can be deployed on nodes.
[0053] Resource control for pods and containers allows developers to specify how many resources (e.g., memory, CPU, storage devices, etc.) a pod needs at runtime. For example, there are typically two different static settings for pod resource control. These include a request or lower bound, which defines how many resources are needed as a minimum requirement (i.e., the lower bound), and a limit, which defines an upper limit (i.e., the upper limit) of resource usage. For example, an exemplary pod specification might include: apiVersion: v1 kind: Pod metadata: name: frontend spec: containers: - name: app image: images.my-company.example / app:v4 resources: lower limits: memory: "64Mi" CPU: "250m" upper limits: memory: "128Mi" CPU: "500m" Note that in the above example pod specification, Mi (megabytes) and m (millibytes) are Kubernetes values. ® Resource units. Container orchestration platform schedulers schedule pod and / or container deployments based on pod resource constraints and available resources on existing nodes. In practice, specific resource constraints for different resource types (such as memory and CPU) are often different.
[0054] Based on current specifications, deployers can optionally provide specific values for lower and upper bounds for different resource types. Essentially, these lower and upper bounds are static values based on estimates of how many resources are needed for pod deployment and runtime processing. This approach presents several challenges. For example, deployers must understand the details of the application and existing infrastructure to perform acceptable estimations, such as how many resources are available in a typical cluster of nodes and how many resources remain on existing nodes. Furthermore, after deployment, incoming jobs or requests for processing can be unpredictable, making it difficult to define static values for pod resource lower and upper bounds. Traditionally, the number of pods can be scaled, but there is no solution to the challenges associated with pod resource lower and upper bounds.
[0055] To address these challenges, this paper discloses one or more embodiments of a pod resource limit adjustment module and processing, and originally referenced... Figure 2-3 Describe it. Figure 2 An embodiment of a pod resource limit adjustment module 200 according to one or more aspects of the present invention is described, the module including code or instructions for performing pod resource limit adjustment processing, and Figure 3 An embodiment of pod resource limit adjustment processing according to one or more aspects of the present invention is described.
[0056] refer to Figure 1 and 2 According to one or more aspects of the present invention, in one example, the pod resource limit adjustment module 200 includes various sub-modules for performing processing. The sub-modules are, for example, computer-readable program code (e.g., instructions) and computer-readable media (e.g., persistent storage devices (e.g., persistent storage device 113, such as a disk) and / or caches (e.g., cache 121, as an example). The computer-readable media may be part of a computer program product and may be executed by and / or using one or more computers (such as (one or more) computers 101), processors (e.g., processors of processor set 110), and / or processing circuitry (e.g., processing circuitry of processor set 110, etc.).
[0057] exist Figure 2In an embodiment, an example submodule of the pod resource limit adjustment module 200 includes, for example: a dynamic resource definition acquisition submodule 202, for acquiring a dynamic resource definition used to dynamically allocate pod resources for a container orchestration platform pod to be deployed, wherein the dynamic resource definition includes resource usage criteria; a deployment object creation submodule 204, for creating a deployment object to be used when deploying the container orchestration platform pod; a container orchestration platform pod initialization and deployment submodule 206, for deploying a container orchestration platform pod with one or more pod resources in a computing environment, wherein the initialization application generates one or more pod resource limits; and a runtime resource usage monitoring submodule 208, for monitoring container orchestration platform resource usage. The system includes: a runtime resource usage prediction submodule 210 for predicting upcoming resource usage of container orchestration platform pods using a trained machine learning model, wherein the prediction utilizes at least partially monitored runtime resource usage; a dynamic adjustment submodule 212 for dynamically adjusting pod resource limits among the one or more pod resource limits of the container orchestration platform pods in the computing environment; and a propagation dynamic adjustment submodule 214 for dynamically adjusting one or more other pod resource limits of one or more other container orchestration platform pods in the computing environment based on the dynamic adjustment of the pod resource limits of the container orchestration platform pods. Advantageously, processing is facilitated in container-based computing environments by providing enhanced use of resources. By dynamically adjusting pod resource limits, cost-effective use of computing resources is provided, as well as efficient and dynamic utilization of resources, which supports on-demand resource allocation without unnecessary waste of resources. Furthermore, the disclosed processing saves costs, reduces carbon footprint, and decreases energy consumption by using computing resources in container-based computing environments more effectively. Note that while various submodules are described, pod resource limit adjustment module handling, as disclosed in this article, can use or include additional, fewer, and / or different submodules. Specific submodules may include additional code, including code from other submodules or less. Furthermore, additional and / or other modules may be used. Many variations are possible.
[0058] In one or more embodiments, according to one or more aspects of the present invention, a submodule is used to perform pod resource limit adjustment processing. Figure 3 An example of pod resource limit adjustment processing is described, such as that disclosed herein. In one or more examples, this process is performed by a computer (e.g., computer 101). Figure 1 )) and / or processor or processing circuitry (e.g., Figure 1The process is executed by the processor set 110 (or its processing circuitry). In one example, the code or instructions implementing this process are part of a module such as the pod resource limit adjustment module 200. In other examples, the code may be included in one or more other modules and / or in one or more submodules of one or more other modules. Various options are possible.
[0059] As an example, in computers (e.g., Figure 1 Computer 101), processor (e.g., Figure 1 The pod resource constraint adjustment process 300, performed on the processors of processor set 110 and / or processing circuitry (e.g., processing circuitry of processor set 110), obtains a dynamic resource definition that will be used in dynamically allocating pod resources for a container orchestration platform pod to be deployed 302, wherein the dynamic resource definition includes resource usage guidelines. The pod resource constraint adjustment process 300 also includes creating a deployment object 304 that will be used in deploying the container orchestration platform pod. In one embodiment, creating the deployment object may include, for example, resolving the resource usage guidelines of the dynamic resource definition, and generating one or more pod resource constraints based on the resolved resource usage guidelines of the dynamic resource definition for initializing the container orchestration platform pod.
[0060] In one embodiment, the pod resource limit adjustment process 300 further includes initializing and deploying a container orchestration platform pod 306 with one or more pod resources in the computing environment. Initialization may include one or more pod resource limits generated by the application.
[0061] Furthermore, in one embodiment, the pod resource limit adjustment process 300 includes monitoring the runtime resource usage of container orchestration platform containers 308, and predicting upcoming resource usage of container orchestration platform pods by a trained machine learning model 310. In one example, the prediction may be, for example, for a defined time interval, such as predicting upcoming resource usage requirements in the next x seconds or y minutes.
[0062] Additionally, in one embodiment, the pod resource limit adjustment process 300 includes dynamically adjusting pod resource limits 312 of one or more pod resource limits of a container orchestration platform pod in a computing environment. In one embodiment, the dynamic adjustment is based on monitored runtime resource usage and predicted upcoming resource usage. Furthermore, in one embodiment, the pod resource limit adjustment process 300 may include selectively propagating one or more dynamic adjustments to pod resource limits to one or more other container orchestration platform pods 314. For example, the process may include dynamically adjusting one or more other pod resource limits of one or more other container orchestration platform pods in a computing environment based on the dynamic adjustment of pod resource limits of a container orchestration platform pod.
[0063] In another example of a computing environment, Figure 4 A container-based computing environment 400 is described, which (in one embodiment) may reside on or be similar to the above-described combination. Figure 1 The computing environment 100 is described. The computing environment 400 includes a control system 410 with program code configured to implement one or more aspects of a container orchestration platform disclosed herein and one or more aspects of a pod resource constraint adjustment tool. Furthermore, the computing environment 400 includes one or more nodes 405, and underlying hardware, such as processing units, used by the control system 410 and the one or more nodes 405. Nodes can be virtual or physical machines, and they can be deployed in a farm (e.g., in computer 101). Figure 1 This refers to computing environment 400, and / or other computer devices, and / or in a cloud environment (e.g., public cloud 105, private cloud 106, hybrid cloud environment, and / or other cloud environment). In one example, computing environment 400 employs Kubernetes, etc. ® Platforms for managing pods and containers, and / or platforms for other platforms, and also include components and / or features according to one or more aspects of the present invention.
[0064] As an example, the control system 410 includes a call chain management component 412, which in one embodiment includes a Service Message Configuration Reader (SMCR) to automatically create channels and subscriptions based on the SMCR. Furthermore, in one embodiment, the call chain management component 412 is configured to customize which service change events are subscribed to by different pods. Additionally, the control system 410 includes a service mesh configuration data repository 414 and a message queue server (MQ) 416, which facilitates actions taken by subscribers when an event is received, and actions taken by publishers when a change occurs, i.e., sending a change event to the channel. According to one or more aspects of the invention, the application programming interface (API) server 418 of the control system 410 receives a configuration 401 of an application, which may include dynamic resource definitions used to dynamically allocate pod resources for container orchestration platform pods being deployed. In one or more aspects, the dynamic resource definitions include resource usage guidelines, which, for example, facilitate the dynamic adjustment of pod resource limits for one or more container orchestration platform pods. Note that this configuration may be saved to a database 420. In one example, database 420 could be an etcd repository, an open-source distributed key-value repository for storing and managing information about distributed systems, and specifically, it could manage configuration data, state data, and Kubernetes data. ® Metadata.
[0065] Metrics server 430 obtains runtime resource usage data for pod 406 and / or its containers via container monitoring agent 407 associated with one or more nodes 405 having one or more pods 406. Metrics server 430 forwards the runtime resource usage data to historical database 432. In one or more embodiments, runtime resource usage data may include runtime data regarding CPU usage, memory usage, etc., of pods and / or containers running on the node. Historical data may be referenced by prediction component 434, which is configured to collect historical resource usage data and current runtime usage data as input for use by a trained machine learning model to predict upcoming resource usage for one or more container orchestration platform pods running on the node. In one implementation, prediction component 434 predicts input request traffic based on the collected resource usage data. Criterion control component 436 or criteria manager accesses configuration information received from data repository 420 and historical usage data from historical database 432. In one or more embodiments, the criterion control component 436 is used to register criteria and criterion processors, determine when resources are set via criteria (i.e., resource limits), and, if so, parse the syntax of the received configuration criteria. In one implementation, criteria can be parsed at runtime, and the processor for the criterion in the criterion data repository 420 can be located and invoked to obtain the configuration value. This process is repeated until all criteria have been parsed and the corresponding values based on the criteria have been determined.
[0066] Vertical Pod Autoscaling (VPA) engine component 440 communicates with prediction component 434 and criterion control component 436 to, for example, facilitate dynamic adjustment of pod resource limits for one or more pods in a computing environment, based on monitored runtime resource usage and predicted upcoming resource usage. In one implementation, VPA engine component 440 communicates with messaging service component 442 to send and receive events to, for example, propagate changes to pod resource limits to one or more other pods, such as those discussed herein. Additionally, VPA engine component 440 deploys any adjustments to pod resource limits via deployment component 450 and, in one embodiment, via replica set component 452. In one embodiment, deployment component 450 further communicates with horizontal pod autoscaler (HPA) component 435. In one or more implementations, vertical pod autoscaler engine component 440 works in conjunction with horizontal pod autoscaler component 435 and, for example, in post-intervention mode, further dynamically adjusts pod resource limits based on the number of instances of pods with specified pod instance limits remaining within a set time period. Alternatively, in the case that the autoscaling mode is advanced, dynamic adjustment of the number of instances of container orchestration platform pods is performed before autoscaling. For example, in this mode, dynamic pod resource limit settings are applied first, and the horizontal pod autoscaler is not applied until it is detected that the pod resource limits have not changed significantly over a period of time. Note that, at this point, according to one or more aspects disclosed herein, the horizontal pod autoscaler launches and deploys additional container orchestration platform pod instances in the compute environment, and the vertical pod autoscaler engine component controls the dynamic adjustment of one or more pod resource limits of one or more container orchestration platform pods in the compute environment.
[0067] As an example, Figure 5 This illustrates one embodiment of how pod resource limit adjustment processing, as disclosed herein, occurs. It is assumed that there are multiple deployed services, including those described above. Figure 1 and 4The described computing environments 100 and 400 are based on a container-based architecture. The container-based computing environment 400' runs services on pod 510. The container-based computing environment 400' includes a control system 410 (e.g., described above in conjunction with Figure 4) and multiple pods 510, 520 running on one or more nodes of the computing environment. In this example, pod 510 provides, for example, a first service, such as a travel-related service, and the first service invokes one or more second services, such as second services running on multiple other container orchestration platform pods 520. In this example, the travel service provider 510 can receive a request 501 and, as part of processing the request, access one or more other services running on one or more other pods related to the travel service being provided. In one example, where the first service requires high runtime resource usage, Figure 4 The VPA engine component 440 can automatically scale up, for example, one or more pod resource limits for pod 510 to handle a large number of requests, and is based on the call chain management component 412 ( Figure 4 In one example, one or more other container orchestration platforms, such as pod 520, which at least partially support pod 510, can also scale up the resource limits of one or more pods.
[0068] In one or more implementations, processing within a container-based computing environment is facilitated in three distinct phases, according to one or more aspects disclosed herein. Phase 1 is the preparation phase, where processing initializes all necessary setup and configuration deployments, message queues, and establishes monitoring to collect resource usage data for archiving as historical data. Phase 2 is the resource determination and / or prediction phase, where the VPA engine dynamically adjusts resource usage limits based on inputs from the criterion control component, the prediction component, propagated events from the messaging service, and / or real-time metrics from the runtime. Phase 3 includes VPA engine result propagation, where the VPA engine propagates resource limit adjustment events to subsequent services, such as multiple other container orchestration platform pods (e.g., ...). Figure 5 The example shows a second service running on pod 520. In this way, the resource usage limits of the secondary service will also be dynamically adjusted, thus facilitating on-demand resource allocation across pod levels.
[0069] refer to Figure 4 In one embodiment, preparation phase 1 includes a call chain management component 412 periodically synchronizing call chain microservices based on a service mesh configuration 414. The call chain management component 412 establishes a message queue (MQ) based on the call chain information. In one embodiment, the end user or end-user system provides dynamic resource definitions, for example... Figure 4The YAML 401 is used to facilitate the deployment of resources, where one or more criteria are defined based on the criteria data stored in the definition data store 420 (or criteria registry).
[0070] Criterion control component 436 parses applicable criteria and, in one implementation, may maintain a cache for each deployment. In one embodiment, criterion control component 436 includes a parser, a semantic analysis component, a criterion processor, a criterion registry, and a criterion optimization component. In one or more implementations, the parser may include basic syntax checking to ensure the input is in a valid format and may create a syntax tree based on the input. The semantic analysis component may be a runtime component that performs semantic analysis, identifies and extracts criteria such as max(avg(30 days), avg(2 weeks)), and looks up the corresponding processor in the criterion registry. The criterion processor includes a built-in processor and a user-defined processor that facilitate the determination of the results of criteria, for example, input by a system user. The criterion registry manages the registration of criteria and allows users to customize criteria in the processor. In one embodiment, the criterion registry is part of a definition data repository 420. The criterion optimization component may perform criterion optimization based on compiler principles. For example, max(avg(30 days), avg(30 days)) may be optimized to avg(30 days). When a deployment object is created, the guidelines control component resolves the dynamic resource definition of the pod and generates one or more initial pod resource limits for pod initialization and deployment. Figure 6 An embodiment of the recipe control component workflow 600 is described in the document.
[0071] refer to Figure 6 Here, we assume that the dynamic resource definition includes resource usage guidelines, such as `avg(last(30 days))` or `100Mi`. This is an example guideline that can be used for initialization. The guideline is an expression composed of multiple functions and operations. Figure 6In one embodiment, the workflow begins at 601, where semantic analysis is performed on the resource definition received at the control system to determine whether any criteria are used 602. If "no," the processing ends 603. Assuming the resource definition includes criteria, the processing parses the criteria 604, which may include invoking a criteria processor component 606 to parse the criteria. For example, criteria may be obtained from a criteria registry or data repository 420, and the criteria processor component may reference historical data 432 to obtain the data needed to determine one or more limit values 610 using the criteria. Once the limit values are determined, the pod deployment object 612 can be updated using the determined initial pod resource limits. In one or more embodiments, the criteria control component may perform the steps of referencing the criteria database, parsing the criteria, and invoking the criteria processor multiple times to fully parse a particular criterion. For example, for the resource definition example above, the criteria control component first performs these steps on the "last" item and then on the "avg" item. Depending on the criterion, a series of operations may be applied to obtain the expression value of the criterion if necessary. Note that in the case of initialization, if historical data is insufficient, a default value may be specified by the criteria control component. For example, using the resource definition mentioned above, 100Mi can be used as the initial value for a pod resource limit where there is not enough resource usage data to dynamically determine that limit value.
[0072] During deployment, the metrics server is 430 ( Figure 4 It facilitates monitoring the real-time resource usage of deployed pods or objects and maintains historical data storage.432 It is used for runtime metrics, such as the CPU and / or memory usage of container orchestration platform pods.
[0073] In Phase 2 processing, actual resource usage is determined, and upcoming resource usage is predicted. For example, during runtime, the criteria control component 436 interfaces with the VPA engine component to provide historical values of deployed resource usage in order to determine dynamic resource usage limits. For example, in one instance, the dynamic resource definition of application configuration 401 could be as follows: Figure 4 As specified in the documentation, in this example configuration, the resource usage guideline avg(last(30 days)) ×2 can be used to help dynamically adjust the pod resource limits(s) of the container orchestration platform pod(s). Note that in this configuration and guideline example, pod resource usage is determined at runtime based on historical data, because the guideline includes the last function (last(30 days)), meaning that data from the last 30 days is used only in one example.
[0074] The prediction component 434 periodically monitors historical resource usage data and uses, at least in part, the monitored runtime resource usage to predict upcoming resource usage for container orchestration platform pods. Based on the predicted upcoming resource usage, predictive alerts can be generated (e.g., if the predicted usage exceeds a threshold) and provided to the VPA engine component.
[0075] Figures 7A-7C An embodiment of the workflow and environment associated with prediction component 434 and VPA engine component 440 is described. Figure 7A The image shows a container-based computing environment 700, such as... Figure 4 The container-based computing environment 400, in one embodiment, includes a predictive model implementation 710 comprising both a model training phase and a model testing phase. As shown, the predictive model implementation component 710 receives resource metrics 701 such as CPU and / or memory usage, any custom metrics 702 such as queries per second (QPS / RT), and historical data usage metrics 703. During the model training phase, historical resource usage data 712 is fed into one or more machine learning models 714, which facilitate the training of a predictive machine learning model 716 to predict upcoming resource usage for container orchestration platform pods, as discussed herein. During the model testing phase, new runtime resource usage data 720 is applied to the trained predictive machine learning model 716 to generate one or more predictions 722, which can be the basis for one or more actions, such as dynamically adjusting pod resource limits for one or more pods of the container orchestration platform pods. As shown in Figure 7, the predicted upcoming resource usage can be applied to, for example... Figure 4 The VPA engine component 440 of the control system 410 enables dynamic deployment 450 of one or more pod resource limit adjustments, wherein one or more pod resource limit adjustments are propagated to one or more pods 406 of the container-based computing environment.
[0076] As a further explanation, Figure 7B An embodiment of a computing environment is described, which may incorporate or implement one or more aspects of embodiments of the present invention. In one or more implementations, the control system 410 may be implemented as part of the computing environment, such as in the above-described incorporation. Figure 1The described computing environment 100. The control system 410 includes and / or utilizes one or more computing resources executing program code 740, which implements one or more aspects of, for example, modules or facilities disclosed herein, and includes a cognitive engine or agent 742 that trains and / or utilizes one or more machine learning models 716 such as those described herein. Data is obtained from data sources 712, 720, such as historical resource usage data, runtime resource usage data, or other data associated with generating predictive models for dynamic pod resource adjustment according to one or more aspects disclosed herein, and the cognitive agent 742 uses said data to train model 716 to, for example, predict upcoming resource usage for container orchestration platform pods, in order to dynamically adjust pod resource limits of one or more pod resource limits of container orchestration platform pods, and / or take other relevant actions 722, etc., based on the machine learning model, to implement the disclosed dynamic pod resource limit adjustment workflow. In its implementation, the control system 410 may include or utilize one or more networks for interfacing with various aspects of one or more computing resources, one or more data sources 712, 720 for providing data, and one or more components, systems, etc., for receiving the outputs, actions, etc. 722 of one or more machine learning models 716 to facilitate the execution of one or more system operations. As an example, the network may be, for example, a telecommunications network, a local area network (LAN), a wide area network (WAN) such as the Internet, or a combination thereof, and may include wired, wireless, fiber optic connections, etc. The network may include one or more wired and / or wireless networks capable of receiving and transmitting data, including training data for the machine learning models, and the output solutions, suggestions, actions, etc., of the machine learning models, as described herein.
[0077] In one or more implementations, the computing resources accommodate and / or execute program code 740 configured to perform methods according to one or more aspects of the present invention. As an example, the computing resources may be resources implemented in a computing system. Furthermore, for illustrative purposes only, Figure 7B The computing resources described herein are depicted as a single computing system. This is a non-limiting example of an implementation. In one or more other implementations, one or more aspects of the processing discussed herein may be implemented, at least in part, across multiple separate computing resources or systems (such as one or more computing systems in the cloud-hosted environment, as an example). Note that at this point, although this document describes... Figure 4 The container-based computing environment 400 is part of the control system 410, but one or more aspects of machine learning model training and / or testing can be performed by a computing system separate from the control system 410. Many variations are possible.
[0078] In one embodiment, program code 740 executes a cognitive engine or agent 742 that includes and trains one or more models 716. The model can be trained using training data, which may include various types of data depending on the model and the data source. In one or more embodiments, program code 740, executing on one or more computing resources, applies one or more algorithms of the cognitive agent 742 to generate and train a model, which the program code then uses to determine, for example, a value for upcoming resource usage of a container orchestration platform pod, where the prediction uses, for example, monitored runtime resource usage of the pod. This value can be used to dynamically adjust pod resource limits in one or more pod resource limits of the container orchestration platform pod in the computing environment, as described herein. During an initialization or learning phase, program code 740 uses the acquired training data to train one or more machine learning models 716, which in one or more embodiments may include historical usage data or other data to be used by an artificial intelligence system workflow to, for example, generate predictions of upcoming resource usage of a container orchestration platform pod as described herein.
[0079] In one or more embodiments, the data used to train the model may include various types of data, such as heterogeneous data generated from one or more data sources and / or data stored in one or more databases accessible by computing resources. In embodiments of the invention, program code may perform data analysis to generate data structures, including algorithms used by the program code to predict and / or perform actions. As is known, machine learning-based modeling solves problems that cannot be solved solely by numerical means. In one example, program code extracts features / attributes from training data, which may be stored in memory or one or more databases. The extracted features may be used to develop a predictor function h(x), also known as a hypothesis, which the program code uses as a model. When identifying the machine learning model 716, various techniques can be used to select features (elements, patterns, attributes, etc.), including but not limited to diffusion mapping, principal component analysis, recursive feature elimination (a powerful method for feature selection), and / or random forests, to select attributes relevant to a particular model. The program code may utilize one or more algorithms to train the model (e.g., algorithms utilized by the program code), including providing weights for conclusions, such that the program code can train any predictor or performance function included in the model. The conclusions can be evaluated using quality metrics. By selecting different training datasets, the program code trains the model to identify and weight various attributes (e.g., features, patterns) that are relevant to the model's performance enhancement.
[0080] In one or more embodiments, program code executing on one or more processors utilizes existing cognitive analytics tools or agents (now known or developed in the future) to tune a model based on data obtained from one or more data sources. In one or more embodiments, the program code may interface with application programming interfaces (APIs) to perform cognitive analytics on the obtained data. Specifically, in one or more embodiments, some APIs include cognitive agents (e.g., learning agents), which include one or more programs, including but not limited to natural language classifiers, retrieval and ranking services that can reveal the most relevant information from a collection of documents, concept / visual insights, trade-off analysis, document transformation, and / or relation extraction. In one embodiment, one or more programs utilize one or more of a natural language classifier, a retrieval and ranking API, and a trade-off analysis API to analyze data obtained by the program code across various sources.
[0081] In one or more embodiments of the present invention, program code can utilize one or more neural networks to analyze training data and / or collected data to generate an operational machine learning model. A neural network is a programming paradigm that enables computers to learn from observed data. This learning is known as deep learning, which is a set of techniques used for learning in neural networks. Neural networks (including modular neural networks) are capable of pattern (e.g., state) recognition with speed, accuracy, and efficiency when datasets are reciprocal and scalable (including across distributed networks, including but not limited to cloud computing systems). Modern neural networks are tools for nonlinear statistical data modeling. They are typically used to model complex relationships between inputs and outputs, or to identify patterns (e.g., states) in data (i.e., neural networks are nonlinear statistical data modeling or decision-making tools). Typically, program code utilizing neural networks can model complex relationships between inputs and outputs and identified patterns in the data. Due to the speed and efficiency of neural networks, especially when parsing multiple complex datasets, neural networks and deep learning offer solutions to many problems in multi-source processing. In embodiments of the present invention, this program code can be used to implement machine learning models such as those described herein.
[0082] As a further example, Figure 7C Depicting one or more aspects of the invention Figure 4 Another embodiment of the prediction workflow of prediction component 434. The workflow begins by collecting historical resource usage data of container orchestration platform pods and preprocessing and cleaning the historical data 750 by removing any missing or outlier values. For example, based on Figure 4 According to Application Specification 401, for the use of the Central Processing Unit (CPU), the processing can refer to the historical resource usage data of the past 30 days.
[0083] exist Figure 7C In the workflow, the cleaned historical data of resource usage is divided into training and testing datasets, and the training dataset 752 is used to train a machine learning model (e.g., a linear regression model). For example, the process may remove any rows with missing values, or any values that deviate from the mean by more than, for example, 3 standard deviations. In one embodiment, the process may split the historical usage data such that 70% is used for the training dataset and 30% for the testing dataset to facilitate training the linear regression model.
[0084] In the 7C workflow, the performance of the linear regression model is evaluated on a test dataset of historical resource usage data. For example, the R-squared value can be determined, which measures how well the model fits the data.
[0085] The system processes and collects runtime or real-time data on pod resource usage, and preprocesses real-time data by removing any missing or outlier values. For example, in one embodiment, runtime data on CPU usage of a container orchestration platform pod is obtained. For example, in one embodiment, the data may be collected and stored every minute. Furthermore, in one embodiment, any rows with missing values may be removed, and / or any values that deviate from the mean by more than, for example, 3 standard deviations.
[0086] Then, a trained linear regression model (i.e., a trained machine learning model) can be used to predict possible upcoming values of pod resource usage based on preprocessed real-time data.758 For example, a trained machine learning model can be used to predict memory and CPU usage values for the next upcoming minute based on current CPU usage values. Predicted values of pod resource usage can be monitored in real time, and adjustments can be made to the trained machine learning model (e.g., a linear regression model) and / or the pod configuration (i.e., resource limits for one or more pods on a container orchestration platform).760
[0087] Continue with the resource determination phase. Figure 4 The VPA engine component 440 of the control system 410 receives results from both the criterion control component 436 and the prediction component 434, and initiates any adjustments based on these results. For example, based on the results, the VPA engine component can initiate an update to the definition of deployment objects, such as for pod configuration limits. For example, the VPA engine component can dynamically change pod memory limits as needed, such as from 100Mi to 150Mi. In one or more embodiments, the VPA engine component can also work with the horizontal pod autoscaling component 435 to dynamically control pod-level resource usage, as described herein.
[0088] In one or more implementations, if the VPA engine component receives any event from the prediction component 434 or the messaging service 442 component, it will retrieve real-time measurement data from the measurement server 430. Upon receiving this event, the VPA engine component can check the autoscaling mode and, if the current state requires intervention, can make a final decision to scale up or down the pod resource limits, and also notify the deployment component 450 to adjust the resource limits appropriately.
[0089] As described above, in one or more embodiments, the third phase includes the propagation of vertical pod autoscaler results. Assume that VPA engine component 440 decides to scale a service (e.g., in...). Figure 5 The resource limits of service A running on pod 510 can be adjusted by the VPA engine components after the resource limits are changed. This can be sent to a message queue via the message service (publisher) 442 to inform other services (e.g., in...). Figure 5 (The second service on pod 520 in the example) notifies service A of resource changes.
[0090] The corresponding VPA engine component receives a message from the messaging service (subscriber) indicating that a previous service (service A) has changed its resource limits and invokes call chain management to find subsequent services, such as service B. In one embodiment, the VPA engine takes action by calculating the new resource limit value for service B and using Phase 2 processing for service B to propagate the change. If intervention from the VPA engine component is required, it will also notify the deployment component to adjust the resource limits of one or more pods for service B.
[0091] Figure 8 This is an example workflow for publishing and subscribing to restriction event changes. The message service component publishes the restriction change 800 to message service subscribers who have subscribed to the relevant resource restriction change via bus 810. Subscribing to the restriction change 820 allows customization of which service change events will be subscribed to, such as which services are relevant, and actions to take when a restriction event change is received. When a restriction change occurs, the publish restriction change 800 facility can publish and send the change event to the channel or bus 810. For example, a 20% increase in the memory limit can be indicated in the change event payload 802, which can be propagated to subscriber 804 by default.
[0092] As described, in one or more implementations, dynamic adjustments to the pod resource limits of a container orchestration platform pod can be propagated to dynamically adjust one or more pod resource limits of one or more other container orchestration platform pods within a compute environment. Dynamic changes to pod resource limits can work in conjunction with horizontal pod autoscaling components and can be further based on the autoscaling mode of the container orchestration platform pods. For example, a configurable post-intervention mode can be implemented where the control system intervenes when the HPA (Hyperscale Application) is divided by the existing number of pod instances for a set time period—that is, when the number of instances of the container orchestration platform pod remains at a specified pod instance limit within the set time period. For example, the HPA component can be defined to ensure an expected memory utilization of 80%, a minimum of 3 pod instances, and a maximum of 9 pod instances. When the number of instances remains at 3 or 9 for an extended period, the control system can intervene. For example, if the number of instances is 9 for a long period, this means the memory allocation may be too small. Based on this, the control system can be configured to dynamically increase the memory limit by 20% (or another customized value). After running for a period of time, the process then checks if the number of pod instances has decreased. If not, the memory can continue to increase by 20% until the number of pod instances decreases. Figure 9 An example describing the process is provided.
[0093] like Figure 9 As shown, the process begins at 900, where the auto-scaling mode is detected, and specifically, it is determined whether the mode is a post-intervention mode 902. If "no", the process ends 903. Figure 9 In the workflow, otherwise, in post-intervention mode, the process monitors the HPA minimum and maximum pod instance settings and determines whether instances are at the limit value within a set time period, such as a specified (long) time interval 904. If "No", the process ends 903. Otherwise, the process determines the results from the criterion control manager of the control system for adjusting container resource limits 906, and examines the predicted upcoming resource usage values of the container orchestration platform containers obtained from the prediction component 908 using a trained machine learning model. Additionally, the runtime data usage metrics of the metrics server 430 are examined 910, and it is determined whether to increase or decrease the pod resource limit by, for example, a set percentage, such as 20% 912 (in one example). After increasing or decreasing the limit by the set percentage, the process returns to monitor the HPA settings and determine whether the number of pod instances continues to remain at the limit within a specified time length.
[0094] As a concrete example, suppose HPA is defined as follows: minReplicas: 3 (Min Pod) maxReplicas: 9 (Max Pod) metrics: - type: Resource resource: name: memory target: type: Utilization [[ID=!14]]averageUtilization: 80(80% with expected memory usage rate).
[0095] In addition, in one example, the dynamic resource definition of a pod can be: resources: hpaMode: [[ID=!23]]mode: post-in (use HPA first) [[ID=!25]]increment: 20% (per once) [[ID=!27]]stabilizationWindowMinute: 120 (stable window) upper limits: [[ID=!31]]memory: avg(last(30days))* 2 or 100Mi (Twice the average of the last30 days or 100Mi) [[ID=!33]]max: 300Mi (Max300Mi) lower limits: [[ID=!37]]memory: avg(last(30days)) * 0.5 or 10Mi (Half the average of the last30 days or 10Mi) [[ID=!39]]max: 50Mi (Max 50Mi).
[0096] Note: There are some words in the original text that seem to be in Chinese in the tags. I have translated them as best as possible while keeping the tags intact. If these are not meant to be Chinese, please clarify for a more accurate translation.In the example above, where HPA is defined as ensuring expected storage usage of 80% with a minimum of 3 and a maximum of 9 pod instances, the control system can intervene if the number of instances remains at 3 or 9 for a set time period, such as a set long period (e.g., one hour or longer). For example, a pod instance count remaining at 9 for an extended period suggests that the storage allocation may be too small, and as mentioned above, the storage can be increased by a custom value, such as 20%. After running for a period of time, the system monitors whether the number of pod instances has decreased; if not, it continues to increase the storage limit allocation for the pods until the number of pod instances decreases.
[0097] Figure 10 Another example of the above processing is described. This example assumes, for example, a normal working load using 5 pod instances, each using 230Mi of storage. The default value during POD initialization and deployment could be 100Mi. In this case, the processing begins at 1000 by determining if historical resource usage data for the past 30 days of pods exists (i.e., in the example where resource usage guidelines reference the past 30 days of resource usage). Assuming no such data exists, the default value of 100Mi is used for pod initialization and deployment in the container orchestration platform. Furthermore, the processing determines that the autoscaling mode is post-intervention mode 1006. At the start of runtime 1010, the processing determines the pod instance status for, for example, the last 120-minute time interval 1012, and determines whether the number of pod instances is within either limit, for example, 3 pods or 9 pods in the above example. If not, the processing ends 1013. In the example above, after running for a period of time (e.g., 120 minutes) after deployment, it is determined that the number of pod instances has increased to 9, which is the maximum allowed by the HPA definition in this example. In one embodiment, VPA is enabled to check POD resource usage once the deployed replica count reaches the maximum replica setting. For example, when processing a situation with pod instance limits, such as 9 pod instances, which in this example is the maximum replica setting, the VPA process determines whether the maximum storage usage will be exceeded, for example, by increasing the allocated pod storage limit by 20% 1014. If "yes", the process ends 1013. Otherwise, the pod storage limit is increased by, for example, 20%, and the process then waits for the next evaluation window to repeat the process 1016, where in one embodiment, the process is aborted, and runtime is stopped 1018.
[0098] In another mode, the autoscaling mode can be a pre-emptive mode, in which dynamic pod resource limit settings are applied first, and horizontal pod autoscaling is not applied until indicators such as pod CPU usage or pod memory usage are detected to be under limit and have not changed significantly over a specified time period.
[0099] Those skilled in the art will notice from the description provided herein that, in one aspect, a computer-implemented method is provided to facilitate processing within a computing environment. This computer-implemented method includes deploying a container orchestration platform pod having one or more pod resources within the computing environment. The pod resources have one or more pod resource limits associated with them. Furthermore, the computer-implemented method includes monitoring the runtime resource usage of the container orchestration platform pod and predicting upcoming resource usage of the container orchestration platform pod using a trained machine learning model, the prediction at least in part using the monitored runtime resource usage. Additionally, the computer-implemented method includes dynamically adjusting the pod resource limits of the one or more pod resource limits of the container orchestration platform pod within the computing environment. This dynamic adjustment is based on the monitored runtime resource usage and the predicted upcoming resource usage. Advantageously, processing is facilitated in container-based computing environments by providing enhanced use of resources. By dynamically adjusting pod resource limits, cost-effective use of computing resources, as well as efficient and dynamic utilization of resources, is provided, supporting on-demand resource allocation without unnecessary waste of resources. Furthermore, this process saves costs, reduces carbon footprint, and decreases energy consumption by using computing resources more effectively in container-based computing environments.
[0100] In one embodiment, the computer-implemented method further includes obtaining a trained machine learning model by training a machine learning model on historical resource usage data of the container orchestration platform pod. Additionally, in one embodiment, the computer-implemented method further includes using predicted upcoming resource usage of the container orchestration platform pod while continuing to train the trained machine learning model. In one example, the machine learning model includes a linear regression model. Advantageously, using the trained machine learning model to predict upcoming resource usage of the container orchestration platform pod facilitates dynamic adjustment of pod resource constraints to provide, for example, cost-effective use of computing resources and efficient and dynamic utilization of resources, and supports on-demand resource allocation without resource waste.
[0101] In one implementation, the deployment includes initializing and deploying the container orchestration platform pod. Initialization and deployment include, for example, obtaining a dynamic resource definition that will be used to dynamically allocate pod resources for the container orchestration platform pod being deployed. The dynamic resource definition includes resource usage guidelines. Furthermore, initialization and deployment include creating a deployment object to be used when deploying the container orchestration platform pod. In one embodiment, the creation includes resolving the resource usage guidelines of the dynamic resource definition and, based on the resolved resource usage guidelines, generating one or more initial pod resource limits for initializing the container orchestration platform pod. Additionally, initialization and deployment include initializing and deploying a container orchestration platform pod with one or more container resources in a compute environment, applying the generated one or more initial container resource limits to the deployment of the container orchestration platform pod with one or more container resources. Advantageously, generating one or more initial pod resource limits for initializing the container orchestration platform pod based on the resource usage guidelines of the dynamic resource definition further enhances resource utilization, for example, by setting the initial pod resource limits through resource usage guidelines to efficiently initialize the container orchestration platform pod with appropriate pod resource limits.
[0102] In one implementation, the monitored runtime resource usage includes runtime CPU usage and runtime memory usage of the container orchestration platform pod. Advantageously, dynamically adjusting pod resource constraints based on the runtime CPU and runtime memory usage of the container orchestration platform pod allows for more efficient use of computing resources in the container-based computing environment, thereby saving processing costs, reducing carbon footprint, and decreasing energy consumption.
[0103] In one embodiment, dynamic adjustment includes dynamically increasing the pod resource limit of a container orchestration platform pod during runtime in the computing environment. In another implementation, dynamic adjustment includes dynamically decreasing the pod resource limit of a container orchestration platform pod during runtime in the computing environment. Advantageously, dynamically increasing and / or decreasing the pod resource limit of a container orchestration platform pod during runtime in the computing environment provides cost-effective use of computing resources, as well as efficient and dynamic utilization of resources, and supports on-demand resource allocation without allocating unnecessary resources.
[0104] In one implementation, the dynamic adjustment is also based on the autoscaling mode of the container orchestration platform pods. In one example, when the autoscaling mode is a post-intervention mode, the dynamic adjustment is also based on the number of instances of the container orchestration platform pods remaining within a specified time period of a given pod instance limit. In another example, when the autoscaling mode is a pre-invention mode, the dynamic adjustment occurs before the number of instances of the container orchestration platform pods is automatically scaled. Advantageously, dynamically adjusting pod resource limits based on the autoscaling mode provides the ability to dynamically adjust pod resource limits, post-level pod autoscaling, or pre-level pod autoscaling. This further enhances the efficient and dynamic utilization of resources, for example, to better balance load processing and avoid inefficient system operations such as slow computing system operation.
[0105] In one implementation, the computer-implemented method also includes the dynamic adjustment of pod resource constraints based on container orchestration platform pods, dynamically adjusting one or more pod resources of one or more other container orchestration platform pods within the computing environment. Advantageously, the dynamic adjustment of pod resource constraints to dynamically adjust one or more pod resources or one or more other container orchestration platform pods within the computing environment promotes resource scalability within the computing environment by scaling resources up and / or down across multiple container orchestration platform pods according to the needs within the container-based computing environment. This facilitates processing within the container-based computing environment and enhances resource utilization. Furthermore, the dynamic utilization of resources across multiple container orchestration platform pods better balances the processing load and enhances the efficient operation of the container-based computing environment.
[0106] Advantageously, this paper discloses computer-implemented methods, computer systems, and computer program products for facilitating the control of resource constraints on container orchestration platform pods using dynamic data analytics and machine learning model-based predictions. In one or more aspects, this process may include dynamic resource allocation using trained criteria, dynamic prediction of incoming workloads or business processes, and proactively initiating scaling actions by collecting historical usage data and runtime data as input for predictions using machine learning models for training, as well as dynamically configuring resources for microservice call chains, and introducing different modes of integration with horizontal pod autoscalers, including post-intervention and pre-initiation modes. In one or more implementations, the computer-implemented methods, computer systems, and computer program products disclosed herein effectively avoid over-allocation and waste of resources in container-based computing environments. Furthermore, resource allocation is more demand-based, thereby reducing request accumulation and / or service interruptions caused by resource shortages. Moreover, dynamic pod resource constraint adjustments based on data analytics such as those described herein achieve lower technical barriers and reduced reliance on experts.
[0107] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used herein, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the terms “comprise” (and any form of inclusion, such as “comprises” and “comprising”)), “have” (and any form of having, such as “has” and “having”)), “include” (and any form of inclusion, such as “includes” and “including”)), and “contain” (and any form of containing, such as “contains” and “containing”) are open-ended connecting verbs. Thus, a method or apparatus that “comprises,” “has,” “includes,” or “contains” one or more steps or elements has, but is not limited to, those steps or elements. Similarly, the elements of a method or apparatus that "comprises," "has," "includes," or "contains" one or more features possess, but are not limited to, those features. Furthermore, a device or structure configured in a certain way is configured at least in that manner, but may also be configured in a manner not listed.
[0108] If present, all means or steps plus functional elements in the following claims are intended to include corresponding structures, materials, actions, and equivalents for performing functions in combination with other claimed elements of the particular claim. Descriptions of one or more embodiments have been presented for purposes of illustration and description, but such description is not intended to be exhaustive or limited to the forms disclosed. Many modifications and variations will be apparent to those skilled in the art. The embodiments were chosen and described in order to best explain various aspects and practical applications, and to enable others skilled in the art to understand the various embodiments with various modifications suitable for the particular intended use.
Claims
1. A computer-implemented method of facilitating processing within a computing environment, the computer-implemented method comprising: deploying, in the computing environment, a container orchestration platform pod having one or more pod resources, the one or more pod resources having one or more pod resource limits associated therewith; monitoring runtime resource usage of the container orchestration platform pod; predicting, by a trained machine learning model, upcoming resource usage of the container orchestration platform pod, the predicting using at least in part the monitored runtime resource usage; and dynamically adjusting, in the computing environment, a pod resource limit of the one or more pod resource limits of the container orchestration platform pod, the dynamically adjusting based on the monitored runtime resource usage and the predicted upcoming resource usage.
2. The computer-implemented method of claim 1, further comprising: the trained machine learning model is obtained by training a machine learning model on historical resource usage data of the container orchestration platform pod.
3. The computer-implemented method of any of the preceding claims, further comprising: the predicted upcoming resource usage of the container orchestration platform pod is used in continuing training of the trained machine learning model.
4. The computer-implemented method of any one of the preceding claims, wherein, the machine learning model comprises a linear regression model.
5. The computer-implemented method of any of the preceding claims, wherein, the deploying comprises initializing and deploying the container orchestration platform pod, the initializing and deploying comprising: obtaining a dynamic resource definition to be used in dynamically allocating pod resources for the container orchestration platform pod being deployed, the dynamic resource definition comprising resource usage criteria; creating a deployment object to be used in deploying the container orchestration platform pod, the creating comprising: parsing the resource usage criteria of the dynamic resource definition; and based on parsing the resource usage criteria of the dynamic resource definition, generating one or more initial pod resource limits for initializing the container orchestration platform pod; and initializing and deploying, in the computing environment, the container orchestration platform pod having the one or more pod resources, the initializing applying the generated one or more initial pod resource limits to the deployment of the container orchestration platform pod having the one or more pod resources.
6. The computer-implemented method of any one of the preceding claims, wherein, the monitored runtime resource usage comprises runtime central processing unit usage of the container orchestration platform pod and runtime memory usage of the container orchestration platform pod.
7. The computer-implemented method of any of the preceding claims, wherein, the dynamically adjusting comprises dynamically increasing, in the computing environment, the pod resource limit of the container orchestration platform pod during runtime.
8. The computer-implemented method of any one of the preceding claims, wherein, the dynamically adjusting comprises dynamically decreasing, in the computing environment, the pod resource limit of the container orchestration platform pod during runtime.
9. The computer-implemented method of any of the preceding claims, wherein, the dynamically adjusting is further based on an auto-scaling mode of the container orchestration platform pod.
10. The computer-implemented method of claim 9, wherein, the auto-scaling mode is a post-intervention mode and the dynamically adjusting is further based on a number of instances of the container orchestration platform pod staying at a specified pod instance limit for a set period of time.
11. The computer-implemented method of claim 9, wherein, The auto-scaling mode is a pre-emptive mode, and the dynamically adjusting occurs prior to auto-scaling of a number of instances of the container orchestration platform pod.
12. The computer-implemented method of any one of the preceding claims, further comprising: Based on the dynamic adjustment of the pod resource limits of the container orchestration platform pod, dynamically adjusting the one or more pod resource limits of one or more other container orchestration platform pods in the computing environment.
13. A computer system for facilitating processing within a computing environment, the computer system comprising: a memory; and at least one processor in communication with the memory, wherein the computer system is configured to perform a method comprising: deploying a container orchestration platform pod having one or more pod resources in the computing environment, the one or more pod resources having one or more pod resource limits associated therewith; monitoring runtime resource usage of the container orchestration platform pod; predicting upcoming resource usage of the container orchestration platform pod by a trained machine learning model, the prediction using at least in part the monitored runtime resource usage; and dynamically adjusting a pod resource limit of the one or more pod resource limits of the container orchestration platform pod in the computing environment, the dynamically adjusting based on the monitored runtime resource usage and the predicted upcoming resource usage.
14. The computer system of claim 13, further comprising: The trained machine learning model is obtained by training a machine learning model on historical resource usage data of the container orchestration platform pod.
15. The computer system of any one of claims 13-14, wherein, The deploying comprises initializing and deploying the container orchestration platform pod, the initializing and deploying comprising: obtaining a dynamic resource definition to be used in dynamically allocating pod resources for the container orchestration platform pod being deployed, the dynamic resource definition comprising resource usage criteria; creating a deployment object to be used in deploying the container orchestration platform pod, the creating comprising: parsing the resource usage criteria of the dynamic resource definition; and based on parsing the resource usage criteria of the dynamic resource definition, generating one or more initial pod resource limits for initializing the container orchestration platform pod; and initializing and deploying the container orchestration platform pod having the one or more pod resources in the computing environment, the initializing applying the generated one or more initial pod resource limits to the deployment of the container orchestration platform pod having the one or more pod resources.
16. The computer system of any one of claims 13 to 15, wherein, The dynamically adjusting is further based on an auto-scaling mode of the container orchestration platform pod.
17. The computer system of claim 16, wherein, The auto-scaling mode is a post-intervention mode, and the dynamically adjusting is further based on a number of instances of the container orchestration platform pod staying at a specified pod instance limit for a set period of time.
18. The computer system of claim 16, wherein, The auto-scaling mode is a pre-emptive mode, and the dynamically adjusting occurs prior to auto-scaling of a number of instances of the container orchestration platform pod. The auto-scaling mode is a pre-emptive mode, and the dynamically adjusting occurs prior to auto-scaling of a number of instances of the container orchestration platform pod.
19. The computer system of any one of claims 13 to 18, further comprising: based on the dynamic adjustment of the pod resource limits of the container orchestration platform pod, dynamically adjusting the one or more pod resource limits of one or more other container orchestration platform pods in the computing environment.
20. A computer program product for facilitating processing within a computing environment, the computer program product comprising: one or more computer-readable storage media and program instructions embodied thereon that are readable by a processing circuit to cause the processing circuit to perform a method comprising: deploying a container orchestration platform pod having one or more pod resources with one or more pod resource limits associated therewith in the computing environment; monitoring runtime resource usage of the container orchestration platform pod; predicting upcoming resource usage of the container orchestration platform pod by a trained machine learning model, the prediction using at least in part the monitored runtime resource usage; and dynamically adjusting a pod resource limit of the one or more pod resource limits of the container orchestration platform pod in the computing environment, the dynamic adjustment based on the monitored runtime resource usage and the predicted upcoming resource usage.