Predictive Autoscaler for Tiered Computing Infrastructure
A predictive AI-based autoscaler addresses the challenge of optimizing resource allocation in hybrid cloud environments by using hierarchical business rules and AI models to automate scaling, ensuring efficient resource management and quality of service.
Patent Information
- Application Number
- JP2023526686
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-11-11
- Filing Date
- 2021-10-28
- Publication Date
- 2026-01-14
- Estimated Expiration
- 2041-10-28
AI Technical Summary
Existing autoscalers struggle to optimize resource allocation across hybrid cloud environments, failing to efficiently manage scaling in multiple cloud platforms and maintain quality of service metrics.
A predictive AI-based autoscaler that hierarchically organizes cloud resources, using business rules and AI models to anticipate scaling needs and automate resource adjustments across hybrid cloud platforms.
Enables efficient, automated resource scaling across hybrid cloud environments, optimizing resource utilization and maintaining quality of service by anticipating workload changes.
Smart Images

Figure 0007798450000007 
Figure 0007798450000008 
Figure 0007798450000009
Abstract
Description
[Technical Field]
[0001] The present invention relates generally to the field of adaptive autonomic computing, and more particularly to resource management for the deployment of hierarchical hybrid computing infrastructures comprising private or public or a combination of computer networking platforms. [Background technology]
[0002] In a hybrid computing environment, a company may use public and private computing resources that work in conjunction with each other. For example, a company may run an application on a public cloud infrastructure while storing sensitive client data on-premises. The public cloud infrastructure provides flexibility with regard to storage and processing resources, allowing the allocation of resources to be automatically scaled up or down as workloads fluctuate. Summary of the Invention [Means for solving the problem]
[0003] According to an aspect of the present invention, there is provided a method, computer program product, or system, or combination thereof, for auto resource scaling in a multilevel computing platform, the method, computer program product, or system performing the following operations (not necessarily in the following order): (i) receiving a first workload metric for a first resource of a multilevel computing platform; (ii) predicting a scaling action for the first resource based on a combination of the first workload metric and a predetermined criterion; (iii) inserting the predicted metric into a runtime-modifiable rule set (hereinafter also referred to as a "business rule") for the first resource based on the scaling action; (iv) creating a scaling plan for the first resource based on a combination of the scaling action and the runtime-modifiable rule set; (v) sending the scaling plan to a level of the multilevel computing platform associated with the first resource; and (vi) triggering execution of the scaling plan based on the runtime-modifiable rule set. [Brief explanation of the drawings]
[0004] [Figure 1] FIG. 1 illustrates a cloud computing environment in accordance with at least one embodiment of the present invention. [Figure 2] FIG. 2 illustrates an abstract model layer in accordance with at least one embodiment of the present invention. [Figure 3] FIG. 3 is a block diagram of a system in accordance with at least one embodiment of the present invention. [Figure 4] FIG. 4 is a flow chart diagram illustrating a method performed in accordance with at least one embodiment of the present invention. [Figure 5] FIG. 5 is a block diagram illustrating the machine logic (eg, software) portion of a system in accordance with at least one embodiment of the present invention. [Figure 6] FIG. 6 is a block diagram illustrating a predictive artificial intelligence hybrid cloud autoscaler architecture in accordance with at least one embodiment of the present invention. [Figure 7] FIG. 7 is a block diagram illustrating a cloud-level process in accordance with at least one embodiment of the present invention. [Figure 8] FIG. 8 is a flow chart diagram illustrating a method that may be performed in accordance with at least one embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0005] In some embodiments of the present invention, a cloud configuration, e.g., including public and private clouds, is hierarchically organized with a top level and any number of lower levels. A parent-level cloud receives resource workload metrics from each of one or more child-level clouds, makes predictions about future resource needs at the child levels, and establishes run-time modifiable business rules and scaling plans based on the predictions. The parent-level cloud sends the scaling plans to each of one or more child levels. The parent level automatically triggers scaling plans at the child levels when conditions described in the business rules are met. Resources are automatically scaled up or down as needed to maintain optimal resource utilization.
[0006] The detailed description of the present invention is divided into I. Hardware and Software Environment, II. Exemplary Implementations, III. Further Comments or Implementations or Combinations thereof, and IV. Definitions.
[0007] I. Hardware and Software Environment
[0008] The present invention may be a system, method, or computer program product, or combination thereof, at any level of technical detail that may be integrated. The computer program product may include one or more computer-readable storage media having computer-readable program instructions for causing a processor to perform aspects of the present invention.
[0009] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. The computer-readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or a ridge structure in a groove in which instructions are recorded, or any suitable combination thereof. As used herein, the computer-readable storage medium should not be construed as a transitory signal per se, such as an electric wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse passing through a fiber optic cable), or an electrical signal transmitted over an electrical wire.
[0010] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing device / processing device or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network may be comprised of copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface in each computing device / processing device receives the computer-readable program instructions from the network and transmits the computer-readable program instructions to the respective computing device / processing device for storage in a computer-readable storage medium.
[0011] The computer-readable program instructions for carrying out the operations of the present invention may be either assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for an integrated circuit, or source or object code written in any combination of one or more programming languages, such as object-oriented programming languages, e.g., Smalltalk, C++, etc., or conventional procedural programming languages (e.g., the "C" programming language or similar programming languages). The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, partially on the user's computer as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any kind of network, such as a local area network (LAN) or a wide area network (WAN), or the connection may be to an external computer (e.g., over the Internet using an Internet Service Provider). In some embodiments, electronic circuits, such as programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), may execute computer-readable program instructions by utilizing state information of the computer-readable program instructions to personalize the electronic circuitry to perform aspects of the invention.
[0012] Aspects of the present invention are described herein with reference to flowchart illustrations or block diagrams, or combinations thereof, of methods, apparatus (systems), and computer program products or computer programs according to embodiments of the invention. It will be understood that each block of the flowchart illustrations or block diagrams, or combinations thereof, and combinations of blocks in the flowchart illustrations or block diagrams, or combinations thereof, can be implemented by computer-readable program instructions.
[0013] These computer-readable program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to create a machine such that the instructions, executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more blocks of the flowchart diagrams or block diagrams, or any combination thereof. These computer-readable program instructions can also be stored on a computer-readable storage medium that can direct a computer, programmable data processing apparatus, or other device, or any combination thereof, to function in a particular manner, and a computer-readable storage medium having instructions stored therein comprises an article of manufacture containing instructions that implement one or more specified functional / operational aspects of the flowchart diagrams or block diagrams, or any combination thereof.
[0014] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device such that the instructions, which execute on the computer, other programmable data processing apparatus, or other device, implement the functions / acts identified in one or more blocks of the flowchart diagrams or block diagrams, or a combination thereof, to cause the computer, other programmable apparatus, or other device to perform a series of the above steps to generate a computer-implemented process.
[0015] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products or computer programs according to various embodiments of the present invention. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing one or more specified logical functions. In some alternative implementations, the functions shown in the blocks may occur out of the order shown in the figures. For example, two blocks shown in succession may actually be accomplished as a single step performed simultaneously, substantially simultaneously, partially, or fully in a time-overlapping manner, depending on the functionality involved, or the blocks may be performed in the reverse order. It should be noted that each block of the block diagrams or flowchart diagrams or combinations thereof, and combinations of multiple blocks in the block diagrams or flowchart diagrams or combinations thereof, may be implemented by a special-purpose hardware-based system that performs the specified functions or operations, or may execute a combination of special-purpose hardware and computer instructions.
[0016] Although this disclosure includes detailed descriptions related to cloud computing, it should be understood that implementation of the teachings recited herein is not limited to a cloud computing environment. Rather, embodiments of the present invention can be implemented in conjunction with any other type of computing environment now known or later developed.
[0017] Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal administrative effort or interaction with the service provider. The cloud model can include at least five characteristics, at least three service models, and at least four deployment models.
[0018] The features are as follows:
[0019] On-demand self-service: A cloud consumer can unilaterally provision computing capacity, such as server time and network storage, automatically as needed, without requiring human interaction with the provider of the service.
[0020] Broad network access: Functionality is available over the network and accessed via standard mechanisms that facilitate use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs).
[0021] Resource Pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, and various physical and virtual resources are dynamically allocated and reallocated according to demand. Consumers generally have no control or knowledge of the exact location of the resources provided, but are said to be location-independent in that they may be able to specify a location at a higher level of abstraction (e.g., country, state, or data center).
[0022] Rapid Elasticity: Capabilities can be provisioned quickly and elastically, sometimes automatically, scaled out quickly, released quickly, and scaled in quickly. To the consumer, the capabilities available for provisioning are often unlimited and can be purchased in any quantity at any time.
[0023] Measured Services: Cloud systems automatically control and optimize resource usage by using metering capabilities at some level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both providers and consumers of the services being used.
[0024] The service model is as follows:
[0025] Software as a Service (SaaS): The ability to offer consumers the ability to use a provider's applications running on a cloud infrastructure. The applications are accessible from a variety of client devices through a thin-client interface, such as a web browser (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even individual application functions, with the possible exception of limited user-specific application configuration settings.
[0026] Platform as a Service (PaaS): The capability offered to consumers to deploy consumer-created or acquired applications, created using programming languages and tools supported by the provider, onto a cloud infrastructure. The consumer does not manage or control the underlying cloud infrastructure (e.g., including networks, servers, operating systems, or storage), but does have control over the deployed applications and, in some cases, the application-hosting environment configuration.
[0027] Infrastructure as a Service (IaaS): The capability offered to consumers to provision processing, storage, network, and other basic computing resources on which they can deploy and run any software, which may include operating systems and applications. The consumer does not manage or control the underlying cloud infrastructure, but does have control over the operating systems, storage, deployed applications, and in some cases, limited control over selecting network components (e.g., host firewalls).
[0028] The deployment models are as follows:
[0029] Private Cloud: Cloud infrastructure is operated exclusively for an organization. The cloud infrastructure may be managed by the organization or a third party, and may reside on-premises or off-premises.
[0030] Community Cloud: Cloud infrastructure is shared by several organizations and supports a specific community with common interests (e.g., mission, security requirements, policies, and compliance considerations). The cloud infrastructure may be managed by the organizations or a third party and may reside on-premises or off-premises.
[0031] Public Cloud: Cloud infrastructure is available to the general public or large industry groups and is owned by organizations that sell cloud services.
[0032] Hybrid Cloud: A cloud infrastructure is a blend of two or more clouds (private, community, or public) that remain unique entities but are brought together by standardized or proprietary technologies (e.g., cloud bursting for load balancing between clouds) that enable data and application portability.
[0033] A cloud computing environment is service oriented with a focus on statelessness, low coupling, modularity, and semantic interoperability. At the heart of cloud computing is an infrastructure comprising a network of interconnected nodes.
[0034] Referring now to FIG. 1 , an exemplary cloud computing environment 50 is illustrated. As illustrated, the cloud computing environment 50 includes one or more cloud computing nodes 10 with which local computing devices used by cloud consumers, such as a personal digital assistant (PDA) or mobile phone 54A, a desktop computer 54B, a laptop computer 54C, or an automobile computer system 54N, or any combination thereof, may communicate. The nodes 10 may communicate with each other. They may be physically or virtually grouped in one or more networks (not shown), such as a private cloud, community cloud, public cloud, or hybrid cloud, or any combination thereof, as described herein. This enables the cloud computing environment 50 to provide infrastructure, platform, or software, or any combination thereof, as a service without the cloud consumer having to maintain resources on their local computing device. It is understood that the types of computing devices 54A-54N illustrated are intended for illustrative purposes only, and that the computing nodes 10 and the cloud computing environment 50 communicate with any type of computerized device (e.g., using a web browser) over any type of network or network-addressable connection, or any combination thereof.
[0035] Referring now to Figure 2, one set of functional abstraction layers provided by cloud computing environment 50 (Figure 1) is shown. It should be understood that the components, layers, and functions shown in Figure 2 are intended to be merely exemplary, and that embodiments of the present invention are not limited thereto. As shown, the following layers and corresponding functions are provided:
[0036] The hardware and software layer 60 includes hardware and software components. Examples of hardware components include a mainframe 61, a RISC (Reduced Instruction Set Computer) architecture-based server 62, a server 63, a blade server 64, a storage device 65, and a network and networking components 66. In some embodiments, the software components include network application server software 67 and database software 68.
[0037] The virtualization layer 70 provides an abstraction layer from which the following examples of virtual entities can be provided: virtual servers 71; virtual storage 72; virtual networks 73, including, for example, virtual private networks; virtual applications and operating systems 74; and virtual clients 75.
[0038] In one example, management layer 80 may provide several functions, as described below. Resource provisioning 81 provides dynamic procurement of computing and other resources utilized to execute tasks within the cloud computing environment. Metering and pricing 82 provides cost tracking as resources are utilized within the cloud computing environment and billing or invoicing for the consumption of these resources. In one example, these resources may include application software licenses. Security provides identity verification for cloud consumers and tasks and protection for data and other resources. User portal 83 provides access to the cloud computing environment for consumers and system administrators. Service level management 84 provides allocation and management of cloud computing resources so that required service levels are met. Service level agreement (SLA) planning and fulfillment 85 provides pre-allocation and procurement of cloud computing resources where future requirements are predicted according to SLAs. The predictive autoscaler 86 automatically tracks resource usage at various levels of the hybrid tiered cloud computing platform and causes scaling (up or down) of those resources in response to dynamically changing workload conditions.
[0039] The workload layer 90 provides examples of functions for which the cloud computing environment may be utilized. Examples of workloads and functions that may be provided from this layer include mapping and navigation 91; software development and lifecycle management 92; virtual classroom instruction delivery 93; data analytics processing 94; and transaction processing 95.
[0040] One embodiment of a possible hardware and software environment for the software or method, or combination thereof, according to the present invention will now be described in detail with reference to the accompanying drawings. Figure 3 is a functional block diagram showing various portions of a networked computer system 100, including a cloud management subsystem 102, a hybrid cloud 104, a private cloud 106, a public cloud 108, a communication network 114, an autoscaling server 200, a communication unit 202, a set of processors 204, a set of input / output (I / O) interfaces 206, memory 208, persistent storage 210, a display 212, external devices 214, random access memory (RAM) 230, a cache 232, and an autoscaler program 300.
[0041] The cloud management subsystem 102 is in many respects representative of one or more of the various computer subsystems in the present invention. Accordingly, several portions of the cloud management subsystem 102 will now be described in the following paragraphs.
[0042] The cloud management subsystem 102 can be a laptop computer, a tablet computer, a netbook computer, a personal computer (PC), a desktop computer, a personal digital assistant (PDA), a smartphone, or any programmable electronic device capable of communicating with the client subsystems via a communications network 114. The autoscaler program 300 is a collection of machine-readable instructions or data and combinations thereof used to create, manage, and control certain software functions that will be discussed in detail below as part of an exemplary embodiment of the detailed description of the present invention.
[0043] The cloud management subsystem 102 can communicate with other computer subsystems via a communications network 114. The communications network 114 can be, for example, a local area network (LAN), a wide area network (WAN), such as the Internet, or a combination thereof, and can include wired, wireless, or fiber optic connections. In general, the communications network 114 can be any combination of connections and protocols that will support communication between server and client subsystems.
[0044] The cloud management subsystem 102 is shown as a block diagram with many double-headed arrows. These arrows (without individual reference numbers) represent a communications fabric that provides communication between the various components of the cloud management subsystem 102. This communications fabric can be implemented with any architecture designed to pass data or control information, or a combination thereof, between processors (e.g., microprocessors, communications and network processors, etc.), system memory, peripheral devices, and any other hardware components in the system. For example, the communications fabric can be implemented, at least in part, with one or more buses.
[0045] Memory 208 and persistent storage 210 are computer-readable storage media. Generally, memory 208 can encompass any suitable volatile or non-volatile computer-readable storage media. It is further noted that, now or in the near future, or a combination thereof, (i) external device 214 may be able to provide some or all of the memory for cloud management subsystem 102, or (ii) devices external to cloud management subsystem 102 may be able to provide memory for cloud management subsystem 102.
[0046] The autoscaler program 300 is typically stored in persistent storage 210 for access and / or execution by one or more of the respective computer processor sets 204 via one or more of the memories 208. Persistent storage 210: (i) is at least more persistent than signals in transmission, (ii) stores the program (including its soft logic or data or a combination thereof) on a tangible medium (e.g., magnetic or optical domain), and (iii) is substantially less persistent than persistent storage. Alternatively, data storage may be more persistent or permanent than the type of storage provided by persistent storage 210, or a combination thereof.
[0047] Autoscaler program 300 may include both machine-readable instructions or entity data (i.e., the type of data stored in a database), or a combination thereof, and executable instructions or entity data, or a combination thereof. In this particular embodiment, persistent storage 210 includes a magnetic hard disk drive. To name a few possible variations, persistent storage 210 may include a solid-state hard drive, a semiconductor storage device, read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, or any other computer-readable storage medium capable of storing program instructions or digital information.
[0048] The media used by persistent storage 210 may also be removable. For example, a removable hard disk may be used for persistent storage 210. Other examples include optical and magnetic disks, thumb drives, and smart cards that are inserted into a drive for transfer onto another computer-readable storage medium that is also part of persistent storage 210.
[0049] In these examples, communications unit 202 provides for communication with other data processing systems or devices external to cloud management subsystem 102. In these embodiments, communications unit 202 comprises one or more network interface cards. Communications unit 202 may provide communication through the use of either or both physical and wireless communications links. Any software modules described herein may be downloaded to a persistent storage device (e.g., persistent storage 210) through a communications unit (e.g., communications unit 202).
[0050] The I / O interface set 206 allows for the input and output of data to and from other devices that may be locally connected in data communication with the autoscaling server 200. For example, the I / O interface set 206 provides a connection to external devices 214. The external devices 214 will typically include devices such as a keyboard, a keypad, a touchscreen, or any other suitable input device, or a combination thereof. The external devices 214 may also include portable computer-readable storage media such as thumb drives, portable optical or magnetic disks, and memory cards. Software and data used to practice embodiments of the present invention, such as the autoscaler program 300, may be stored on such portable computer-readable storage media. In these embodiments, the associated software may (or may not) be loaded, in whole or in part, into persistent storage 210 via the I / O interface set 206. The I / O interface set 206 is also in data communication with a display 212.
[0051] Display 212 provides a mechanism for displaying data to a user and may be, for example, a computer monitor or a smartphone display screen.
[0052] The programs described herein are identified based on the application for which they are implemented in a particular embodiment of the invention. However, it should be understood that any specific program name herein is used merely for convenience, and therefore the invention should not be limited to use with only any specific application identified or implied or identified and implied by such name.
[0053] The description of various embodiments of the present invention has been presented for illustrative purposes and is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terms used in this specification have been selected to best explain the principles of the embodiments, practical applications, or technical improvements over technologies found in the market, or to enable those skilled in the art to understand the embodiments disclosed herein.
[0054] II. Exemplary Implementations
[0055] Figure 4 shows a flowchart 250 illustrating a method in accordance with the present invention. Figure 5 shows an autoscaler program 300 for performing at least some of the method operations of flowchart 250. The method and associated software will now be described in the following paragraphs with extensive reference to Figure 4 (for the method operation blocks) and Figure 5 (for the software blocks). One physical location where the autoscaler program 300 of Figure 5 may be stored is persistent storage 210 (see Figure 3).
[0056] Processing begins at operation S255, where the predictive AI autoscaler module 302 of the autoscaler program 300 receives workload metrics for resources running on a given level of the hierarchical computing platform. In some embodiments, the workload metrics correspond to processor utilization, memory usage, storage usage, network bandwidth usage, arrival rates, interarrival times, response times, throughput, or service load patterns, or a combination thereof, or the like.
[0057] Processing continues to operation S260, where predictive AI autoscaler module 302 of autoscaler program 300 predicts a scaling action for a given level of a computing platform. The predicted scaling action may include, for example, increasing or decreasing storage space, memory, or other resources allocated to processes running on the given level of the computing platform. In some embodiments, predictive AI autoscaler module 302 predicts the scaling action based on past performance experience, such as, but not limited to, workload variations observed with respect to time of day, day of the week, or product lifecycle for the application type running on the given level of the computing platform.
[0058] Processing continues to operation S265, where the business rules module 306 of the autoscaler program 300 determines the business rules based at least in part on the predicted scaling behavior. a set of business rules Generate business rules set Examples are provided below in the detailed description of the invention under the subheading "Examples of Business Rules" under the heading "Further Comments or Embodiments or Combinations Thereof."
[0059] Processing continues to operation S270, where the predictive AI autoscaler module 302 of the autoscaler program 300 creates a scaling plan based on the predicted scaling actions determined in operation S260. The scaling plan includes a set of scaling actions to be executed when the scaling plan is triggered. a detailed set of actions Includes.
[0060] Processing continues to operation S275, where the automation management module 304 of the autoscaler program 300 sends the scaling plan to a predetermined level of computing platform, where the scaling plan is held in a readiness state and awaits a signal to trigger activation of the plan.
[0061] Processing continues to operation S280, where the automation management module 304 of the autoscaler program 300 determines to trigger the scaling plan and, as a result, sends a signal to the computing platform at the given level that triggers implementation of the scaling plan.
[0062] III. Further Comments or Embodiments or Combinations Thereof
[0063] Some embodiments of the present invention may recognize one or more of the following facts, potential problems, or potential areas for improvement, or a combination thereof, with respect to the current state of conventional autoscalers: (i) they may be able to scale only resources deployed on a provider's platform; (ii) in hybrid cloud deployments, it may be difficult to optimize each metric in each of the clouds as well as overall end-to-end and quality of service (QoS) metrics; (iii) they may have difficulty accommodating infrastructure changes across multiple cloud platforms in addition to configuration changes specific to certain resources and applications; (iv) they focus on adapting to a single cloud provider (as opposed to a large-scale hybrid cloud with multiple cloud providers); or (v) they need to constantly monitor system metrics and respond to those changes at runtime, or any combination of (i)-(v).
[0064] Some embodiments of the present invention may include one or more of the following features, characteristics, or advantages, or combinations thereof: (i) a business rule engine (BRE) that makes auto-scaling decisions in a manner that is easier for cloud operators; (ii) an artificial intelligence (AI) predictive model that predicts future scaling needs and communicates the future scaling needs to the BRE; or (iii) an automation manager (hereinafter sometimes referred to as the "cloud automation manager") that works with the BRE and uses infrastructure as code (IaC) templates to efficiently scale hybrid cloud services; or any combination thereof.
[0065] In some embodiments, the autoscaler uses the BRE to enable simple and sophisticated rules for making scaling decisions in a hybrid cloud platform. The use of business rules simplifies autoscaling requirements and makes them easier for cloud operators to understand and implement. The cloud operator can use simple rules to make scaling decisions based on different system performance or workload metrics, or a combination thereof, for different services in a hybrid cloud environment. set (a set of simple rules) Examples of system performance or workload metrics or combinations thereof include CPU utilization, memory utilization, storage utilization, network bandwidth utilization, arrival rate, inter-arrival time, response time, throughput or service load pattern or combinations thereof, etc.
[0066] While some embodiments of the present invention are directed to hybrid cloud environments including tiered private and public cloud platforms, it should be understood that some embodiments are directed to non-hybrid public or private cloud infrastructures and other networked computing platforms, such as those including non-cloud platforms or other tiered computing platforms or combinations thereof.
[0067] In some embodiments, an AI predictive model predicts future changes in service workloads and cloud platforms based on historical system data. The AI predictive model communicates the need for additional scaling to support future changes to the BRE. Some embodiments organize the BRE and AI predictive model hierarchically, where (i) deployments located on each cloud each maintain local performance metrics; and (ii) deployments at a higher level manage all of the clouds, whether public or private, and maintain end-to-end metrics and cost optimization.
[0068] In some embodiments, a cloud automation manager uses infrastructure as code (IaC) templates to apply scaling decisions (or predicted scaling behaviors) to deploy additional (or excess) resources to the hybrid cloud platform. The cloud automation manager automates and simplifies the deployment or release of different types of resources for multiple services in public and private cloud platforms within a hybrid cloud environment.
[0069] 6 illustrates a process and system architecture according to some embodiments of the present invention, including a predictive artificial intelligence (AI) autoscaler 620; a Level 1 predictive AI model 621; a Level 1 business rules engine (Level 1 BRE 622); business rules 623; a hybrid cloud 630; a public cloud 631; a private cloud 632; managed services 641; a Level 2 predictive AI model 642; and a Level 2 business rules engine 643. The hybrid cloud 630 may include any number of instances of the public cloud 631 and any number of instances of the private cloud 632. The public cloud 631 and each private cloud 632 may each include any number of managed services 641.
[0070] In some embodiments, the predictive AI autoscaler is hierarchical and distributed, where (i) an autoscaler (not shown) associated with each cloud level (e.g., hybrid cloud 630 including public cloud 631 and private cloud 632) maintains metrics local to the associated cloud level, and (ii) a predictive AI autoscaler 620 (at a higher cloud level in the cloud hierarchy) maintains optimization across hybrid cloud 630. Cloud automation manager 624 (sometimes referred to herein as cloud automation manager) manages and deploys services (e.g., any number of instances of managed service 641) across both public cloud 631 and private cloud 632. Level 1 predictive AI model 621 of predictive AI autoscaler 620 receives streamed metrics from instances of managed service 641.
[0071] The Level 1 predictive AI model 621 analyzes the different types of workload mix to make intelligent and predictive scaling decisions. Based on the streamed metric input data, the Level 1 predictive AI model 621 makes scaling decisions and generates predictive metrics for future times and inserts the predictive metrics into existing business rules that underlie the scaling of one or more instances of managed services 641 in hybrid cloud 630. Alternatively, in some embodiments, the predictive AI model 621 generates predictive metrics and inserts them into existing business rules for future times based on the nature of the predictive metrics or the conditions under which the metrics are determined, or a combination thereof.
[0072] Business rules can be changed and modified at runtime. The stateful nature of the Level 1 BRE 622 enables historical and temporal analysis. In the case of predictive scaling, the Level 1 predictive AI model 621 writes (generates) a set of business rules that trigger one or more corresponding business rules in response to certain changes in workload conditions. This enables the Level 1 predictive AI model 621 to predict and implement preemptive scaling changes to support predicted future scaling needs. In addition to business rule modifications by the Level 1 predictive AI model 621, cloud operators can directly modify business rules 623 to address current and future scaling needs.
[0073] In some embodiments, the predictive AI autoscaler 620 operates at the highest level of the cloud hierarchy. The predictive AI autoscaler 620 includes a Level 1 predictive AI model 621, a Level 1 BRE 622, and a cloud automation manager 624. The Level 1 predictive AI model 621 receives streamed metrics from the hybrid cloud 630, more specifically from the public cloud 631 and the private cloud 632 (arrow "1"). The Level 1 predictive AI model 621 modifies autoscaling rules and passes the rules to the Level 1 BRE 622 (arrow "2"). The Level 1 BRE 622 develops one or more plans (e.g., scaling plans) for configuration changes for the hybrid cloud 630 in response to a combination of the modified autoscaling rules and business rules 623. The Level 1 BRE 622 passes the one or more (cloud configuration change) plans to the cloud automation manager 624, which initiates execution of the plans (arrow "3"). Cloud automation manager 624 applies the plan (arrow "4") by passing the plan to a managed service (e.g., managed service 641) associated with public cloud 631 or private cloud 632 or a combination thereof.
[0074] At a lower level, e.g., Level 2 in public cloud 631, a Level 2 autoscaler includes a Level 2 predictive AI model 642 and a Level 2 business rules engine 643. The Level 2 autoscaler maintains local metrics, e.g., CPU utilization thresholds, and streams the local metrics to the Level 1 predictive AI model 621 of the predictive AI autoscaler 620, as described in the previous paragraph, completing a continuous feedback and control loop whose goals include maintaining the overall quality of service (QoS) goals of the hybrid cloud 630.
[0075] In some embodiments of the present invention, the business rules setis the conditional and consequential when-then rule Set (a set of conditional and consequential ''when - then'' rules) A business rules engine (e.g., BRE 622) manages and executes business rules in a runtime environment, which includes a rule definition, which means that when a condition occurs, then execute a result or action. The business rules engine continuously fires rules every x seconds (where, in some embodiments, x is a user-selected time value). The business rules engine collects performance metrics from services or predictive AI models, or a combination thereof, and then uses those metrics to make autoscaling decisions.
[0076] Business Rule Example
[0077] Then, when the rule's conditions are met, the service a set of services High-level business rules expressed in pseudocode for scaling set is as follows:
[0078] Business Rule 1: Prepare a Scale-Up Timer
[0079] [Table 1]
[0080] Business Rule 1 determines whether the service's average CPU utilization exceeds a threshold. In this example, the threshold is 40%. CPU utilization is an average measure of the service collected over a user-selected period (e.g., the last 30 seconds). The rule then inserts a "ScaleUp" state to begin the scaling process.
[0081] Business Rule 2: Prepare to Scale Up
[0082] [Table 2]
[0083] After an additional 30 seconds (in the "Scale Up" state), Business Rule 2 determines whether CPU utilization continues to exceed the threshold (40% in this case). This additional 30-second period prevents unnecessary scaling actions in response to outliers or spikes in utilization. If utilization continues to exceed 40% (if the condition is met), the rule changes the state value from "Scale Up" to "start scaling." In some embodiments, utilization during any given period is measured in various ways, such as (i) the average utilization during the period, (ii) the utilization that remains above the threshold for the entire period, (iii) the utilization that remains above the threshold for a certain percentage of the period, or (iv) any other method, numerical, operational, statistical, or other, or combination thereof, deemed appropriate for a given implementation.
[0084] Business Rule 3: Scale Up
[0085] [Table 3]
[0086] Business rule 3 determines whether a "scale up" state exists and whether its value is "completed." If the "scale up" state = "completed" (condition met), the rule calls the cloud automation manager (CAM) to deploy a new instance for scaling.
[0087] The following code example shows the coding corresponding to the pseudocode example above written in Drools, an open source business rules engine developed by Red Hat. (Note: "Drools" or "Red Hat" or "Drools" and "Red Hat" may be subject to trademark rights in various jurisdictions around the world and are used herein only in connection with products or services appropriately designated by such trademarks, to the extent such trademark rights exist.)
[0088] Business rule 1: Adjust the scale-up timer
[0089] [Table 4] / / Translation: Checks if CPU average is above 40% after 30 seconds and if a scale-up condition exists. / / Inserts the scale-up state into the rule engine memory
[0090] Business Rule 1 checks a "Metric" object that contains the performance indicator of the service (in this case, the average CPU utilization). Once the rule is satisfied, the BRE initializes a "State" object (State.NOTRUN, a binary value).
[0091] Business Rule 2: Adjust for scale-up
[0092] [Table 5] (Translation of / / : Check if CPU average is still above 40% after 30 seconds) / / State values are changed and updated in memory to begin scaling up
[0093] Business Rule 2 uses a temporal (time-related) characteristic. In the "when" statement, "this" represents a state, and "before[30s] metric" determines if the state existed more than 30 seconds before the current Metric object. If the condition is met, the binary value of the state in the BRE memory is updated to State.FINISHED.
[0094] Business Rule 3: Scale Up
[0095] [Table 6] / / Checks if the scale-up state is in memory and complete, i.e., from the previous rule. ( / / : Calls the CAM API) Get the current number of deployed resources ( / / translation: Increase the current resource count by one) Plan a deployment by sending an API request to the cloud automation manager.
[0096] In Business Rule 3, the "when" statement determines if the ScaleUp state has changed to "State.FINISHED" and then proceeds. "CamJson" contains details of the current deployment of the associated service on the cloud automation manager. The BRE calls the CAM API through the CamTemplateAPI and increases the number of resources (web_replica) for the service through a JavaScript Object Notation (JSON) file (acme.json). The BRE then submits the acme.json file through a "plan and apply" request using ".ModifyInstance" and ".ApplyInstance," which initiates the process of deploying additional resources on the cloud.
[0097] If the CPU utilization value decreases below 40% after 30 seconds, the BRE removes the state. This process is the same for "scale up" and "scale down." Because the business rules are easy to write and flexible, users can use the predictive AI model to write rules for various scaling use cases based on CPU utilization or other metrics.
[0098] Note: In the above code example, "nnn.nnn.nnn.nnn" represents an Internet Protocol (IP) address.
[0099] In some embodiments, scale-down behavior is warranted. Consider a transaction processing system that requires a large amount of memory to handle periods of high demand and a smaller amount of memory during periods of low demand. When the period of high demand ends and demand falls below a threshold for a predetermined length of time, the business rules respond in a manner similar to the scale-up scenario in the example above, but instead scale down the memory allocated to the transaction processing system.
[0100] Block diagram 700 of Figure 7 illustrates an autoscaler process according to some embodiments of the present invention. Autoscaler 704 includes a predictive artificial intelligence (AI) model 706 and a business rules engine (BRE 707) at the cloud level, which may be a private or public cloud or any other networked computing platform. Some embodiments of the present invention include a cloud hierarchy organized into multiple levels, with each cloud associated with a corresponding autoscaler. With respect to block diagram 700, autoscaler 704 is associated with a mid-level cloud 702 located between a lower-level cloud (not shown) and a higher-level cloud (not shown).
[0101] In some embodiments, the cloud hierarchy includes, in a nested fashion, a level 3 cloud (e.g., cloud 702), which may be public or private. Level 3 cloud 702 streams metrics local to cloud 702 to a level 2 predictive AI model (e.g., predictive AI model 642, FIG. 6), which in turn streams level 3 metrics, along with level 2 local metrics, to a level 1 predictive AI model (e.g., predictive AI model 621, FIG. 6).
[0102] Managed services 705 streams (arrow "1") cloud 702 local performance metrics to a predictive AI model autoscaler (e.g., Level 1 predictive AI model 621, see FIG. 6) in a higher-level cloud (not shown) and to predictive AI model 706 (arrow "2"), which predicts scaling decisions based on metrics such as CPU, storage, memory, and network utilization.
[0103] The predictive AI model 706 generates (or modifies, or generates and modifies) business rules local to the cloud 702 and sends (arrow "3") the generated or modified business rules to the BRE 707. In response to receiving the business rules, the BRE 707 modifies application-level configuration files and settings in accordance with the business rules to maintain optimal system performance.
[0104] The BRE 707 sends application-level configuration changes to the underlying cloud (arrow "4"). Using an application programming interface (API) provided by the underlying cloud automation manager, the BRE 707 initiates a scaling change plan (arrow "5"). In some embodiments, the API is coded as a representational state transfer API (REST API).
[0105] In some embodiments, a cloud automation manager (CAM, e.g., cloud automation manager 624 (see FIGS. 6 and 8 )) automates service deployments on various cloud providers in an infrastructure-as-code (IaC) environment. A cloud operator, project team, or automated system writes a high-level description of the application's deployment details (IaC script). The CAM executes the IaC script and automatically deploys the necessary infrastructure and services based on the IaC script. The CAM defines, provisions, and manages the deployment of services on public or private clouds. For example, if there are updates to the deployment or configuration values change, the CAM automates the changes. The BRE calls an API and initializes new scaling details, and the CAM applies the changes on the hybrid cloud. This process loops continuously as performance metrics stream to the autoscaler.
[0106] 8 is a hybrid flowchart diagram illustrating an autoscaler process 800 according to some embodiments of the present invention. The autoscaler process includes components, information flows, and operations. The components include a predictive AI model 706; a cloud operator 802; and a cloud automation manager 624. The information flows include multiple incoming metrics 801 and requirements 804. The operational operations include decisions 806, 808, and 812; operation 810 (business rules engine); and operation 816 (adaptive environment changes).
[0107] The process begins with the predictive AI model 706 analyzing the plurality of incoming metrics 801 to determine whether a future scaling decision is needed. If a future scaling decision is not needed (decision 808, "no" branch), the process returns to the predictive AI model 706. If a future scaling decision is needed (decision 808, "yes" branch), the process proceeds to operation 810, where the predictive AI model 706 generates predictive metrics for future times and inserts the predictive metrics into existing business rules based on a combination of the plurality of incoming metrics 801 and requirements 804. Additionally, if new requirements require changes in the hybrid cloud environment, the cloud operator 802 may generate business rules to meet the new requirements.
[0108] Processing continues to decision 812 where, if a business rule is not triggered (decision 812, "no" branch), processing returns to the predictive AI model 706. If a business rule is triggered (decision 812, "yes" branch), the cloud automation manager 624 invokes a representational state transfer (REST) API to implement the business rule. Processing continues to operation 816 where, via the REST API, the cloud automation manager 624 effects adaptive environment changes.
[0109] As services are modified or new services are deployed, those metrics are streamed to the predictive AI model to continue the process.
[0110] In some embodiments, requirements 804 may trigger a cloud operator 802 (which may be an automated system or a human operator) to determine whether a current scaling decision is necessary. If a current scaling decision is not necessary (decision 806, "no" branch), processing returns to the cloud operator 802. If a current scaling decision is necessary (decision 806, "yes" branch), the predictive AI model 706 generates predictive metrics for future times and inserts the predictive metrics into existing business rules (act 814) based on a combination of multiple incoming metrics 801 and requirements 804. The business rules that incorporate the predictive metrics flow to a business rules engine 810. Processing then proceeds to decision 812, as described above.
[0111] In some embodiments of the present invention, a computer-implemented method for artificial intelligence-enabled predictive autoscaler as a service for hybrid cloud deployments includes the steps of: (i) configuring a first autoscaler at each cloud level that maintains simple local metrics and a second autoscaler at a higher level that maintains global optimization; set (ii) streaming metrics from a given service; set (iii) receiving the received metrics using predetermined criteria to determine predictive scaling decisions. set(iv) generating business rules using predictive scaling decisions to scale services in the hybrid cloud platform, where the rules can be changed and modified during runtime; (v) generating a plan for changes to the hybrid cloud environment; (vi) applying the plan by managing and deploying services in both public and private clouds in the hybrid cloud platform with an automation manager in response to receiving the plan; (vii) streaming performance metrics to a second autoscaler for each cloud level that includes at least one managed service; (viii) configuring each first autoscaler to perform a predictive scaling operation. (ix) receiving performance metrics streamed into the predictive artificial intelligence model of the autoscaling engine; (ix) generating predicted scaling decisions using simple metrics, such as metrics including processor utilization and disk utilization; (x) generating local business rules for autoscaling, where the rules can be changed and modified during runtime; (xi) modifying application-level configuration files and settings by the business rules engine; (xii) generating a local plan for changes to the local cloud environment; and (xiii) applying the local plan by managing and deploying services at the respective cloud level by an automation manager in response to receiving the local plan.
[0112] IV. Definition
[0113] The term "the present invention": The term "the present invention" should not be taken as an absolute indication that the subject matter described by the term "the present invention" is covered by either the claims at the time of filing or the claims ultimately issued after patent prosecution. On the contrary, the term "the present invention" is used to help the reader get a general feel for what disclosures herein are believed to be potentially new, and this understanding, as indicated by the use of the term "the present invention," is provisional and tentative and is subject to change during the course of patent prosecution as relevant information develops and claims may be amended.
[0114] The term "embodiment": For the term "embodiment", see the definition of "present invention" above. A similar caution applies to the term "embodiment".
[0115] The word "and / or": The word "and / or" is inclusive or, for example, A, B, "and / or" C means that at least one of A or B or C is true and applicable.
[0116] The word "including / include / includes" means "including but not necessarily limited to," unless expressly stated otherwise.
[0117] The term "user / subscriber": The term "user / subscriber" includes, but is not necessarily limited to, (i) a single human being; (ii) an artificially intelligent entity that has sufficient intelligence to act as a user or subscriber; (iii) a group of related users or subscribers; or combinations thereof.
[0118] The term "data communications": The term "data communications" refers to any type of data communication method now known or later developed, including wireless communications, wired communications, and communications paths having wireless and wired portions; data communications may be, but are not necessarily limited to, (i) direct data communications; (ii) indirect data communications; or (iii) data communications in which the format, packetization, medium, encryption, or protocol of the data communications, or any combination thereof, remains constant throughout the course of the data communications; or any combination thereof.
[0119] The words "receive / provide / transmit / input / output / report": Unless expressly stated otherwise, the words "receive / provide / transmit / input / output / report" should not be construed as implying (i) any particular degree of directness in the relationship between their objects and subjects; or (ii) the absence of any intervening components, acts, or things, or any combination thereof, between their objects and subjects.
[0120] The term "substantially without human intervention": The term "substantially without human intervention" refers to a process that can occur automatically (often through the operation of machine logic, e.g., software) with little or no human input. Examples that encompass the term "substantially without human intervention" include: (i) a computer is performing a complex process and a grid power outage causes a human to switch the computer to another power source, allowing the process to continue uninterrupted; (ii) a computer is about to perform a resource-intensive process and a human confirms that the resource-intensive process should actually occur (in this case, the independently considered confirmation process involves substantial human intervention, but the resource-intensive process does not involve any substantial human intervention, despite a simple "yes / no" confirmation that must be made by a human); and (iii) a computer uses machine logic to make a critical decision (e.g., a decision to ground all flights in anticipation of bad weather), but the computer must obtain a "yes / no" confirmation from a human before implementing the critical decision.
[0121] The term "automatically": The term "automatically" means without any human intervention.
[0122] The term "module / sub-module": The term "module / sub-module" means any set of hardware, firmware, or software, or combination thereof, operatively operating to perform a certain function, whether the module is (i) in a single local vicinity; (ii) distributed over a wide area; (iii) in a single vicinity within a larger portion of software code; (iv) located within a single portion of software code; (v) located within a single storage device, memory, or medium; (vi) mechanically connected; (vii) electrically connected; or (viii) connected in data communication; or any combination thereof.
[0123] The term "computer": The term "computer" refers to any device or combination thereof that has significant data processing or machine-readable instruction capabilities, including, but not limited to, desktop computers, mainframe computers, notebook computers, field-programmable gate array (FPGA)-based devices, smartphones, personal digital assistants (PDAs), body-worn or insertable computers, embedded device computers, and application-specific integrated circuit (ASIC)-based devices.
Claims
1. 1. A computer-implemented method for auto-resource scaling in a multi-level computing platform, the multi-level computing platform comprising a hybrid cloud platform including multiple cloud levels structured in a hierarchical manner, wherein each cloud level is selected from the group consisting of a private cloud computing platform and a public cloud computing platform; The method comprises: receiving a first workload metric for a first resource of a multilevel computing platform, the first resource being selected from the group consisting of memory, storage, network bandwidth, and processor utilization, and the first workload metric being selected from the group consisting of processor utilization, memory utilization, storage utilization, network bandwidth utilization, arrival rate, inter-arrival time, response time, throughput, and service load pattern; predicting a scaling action with respect to the first resource based on the received first workload metric, the scaling action being to increase or decrease storage space, memory, or other resource allocated to processes running at a given level of the computing platform; inserting, for the first resource, a predicted metric into a runtime modifiable business rule based on the predicted scaling behavior, the predicted metric being a metric for a future time; creating a scaling plan for the first resource based on a combination of the predicted scaling behavior and the runtime-modifiable business rules, the scaling plan including a set of specific steps to be performed when the scaling plan is executed; transmitting the scaling plan to a level of the multi-level computing platform associated with the first resource; and Triggering execution of the scaling plan based on mutable business rules at the runtime. The method comprising:
2. 2. The method of claim 1, wherein predicting the scaling behavior is performed by predicting scaling behavior for a given level of the multi-level computing platform using a predictive AI autoscaler module.
3. 2. The method of claim 1 , wherein predicting the scaling behavior is performed using the predictive AI autoscaler module based on observed workload variations with respect to time of day, day of the week, or product lifecycle for an application type running on the given level of the computing platform.
4. receiving a second workload metric for a second resource of the multilevel computing platform, the second workload metric being selected from the group consisting of processor utilization, memory utilization, storage utilization, network bandwidth utilization, arrival rate, inter-arrival time, response time, throughput, and service load pattern; making a second predicted scaling decision based on the received second workload metric; creating runtime modifiable auto-scaling business rules local to each level of the multi-level computing platform; modifying application-level configuration files and application-level settings according to the created business rules; creating a plan for changing the configuration local to each level of the multi-level computing platform; and Deploying and managing services operating on each of the levels of the multi-level computing platform based on the created plan. The method of claim 1 further comprising:
5. The method described in claim 1, wherein the set of business rules includes a set of conditional and consequential "when-then" rules, each of which means that when a condition occurs, the next result or action is to be performed.
6. triggering execution of the scaling plan, determining whether the first workload metric exceeds a threshold value written into the business rule for a predetermined period of time; and deploying the scaling plan in response to determining that the first workload metric exceeds the threshold for the predetermined period of time. The method of claim 1 further comprising:
7. triggering execution of the scaling plan, determining whether a threshold value written into the business rule exceeds the first workload metric for a predetermined period of time; and deploying the scaling plan in response to determining that the threshold exceeds the first workload metric for the predetermined period of time. The method of claim 1 further comprising:
8. 1. A computer program for automatic resource scaling in a multi-level computing platform, the multi-level computing platform comprising a hybrid cloud platform including a plurality of cloud levels constructed in a hierarchical manner, wherein each cloud level is selected from the group consisting of a private cloud computing platform and a public cloud computing platform; receiving a first workload metric for a first resource of a multilevel computing platform, the first resource being selected from the group consisting of memory, storage, network bandwidth, and processor utilization, and the first workload metric being selected from the group consisting of processor utilization, memory utilization, storage utilization, network bandwidth utilization, arrival rate, inter-arrival time, response time, throughput, and service load pattern; predicting a scaling action with respect to the first resource based on the received first workload metric, the scaling action being to increase or decrease storage space, memory, or other resource allocated to processes running at a given level of the computing platform; inserting a predicted metric into a runtime modifiable business rule for the first resource based on the predicted scaling behavior, the predicted metric being a metric for a future time; creating a scaling plan for the first resource based on a combination of the predicted scaling behavior and the runtime-modifiable business rules, the scaling plan including a set of specific steps to be performed when the scaling plan is executed; transmitting the scaling plan to a level of the multi-level computing platform associated with the first resource; and Triggering execution of the scaling plan based on mutable business rules at the runtime. The computer program causing one or more processors to execute each step of the method comprising:
9. 9. The computer program product of claim 8, wherein predicting the scaling behavior is performed by predicting scaling behavior for a given level of the multi-level computing platform using a predictive AI autoscaler module.
10. 9. The computer program product of claim 8, wherein predicting the scaling behavior is performed using the predictive AI autoscaler module based on observed workload variations with respect to time of day, day of the week, or product lifecycle for an application type running on the given level of the computing platform.
11. receiving a second workload metric for a second resource of the multilevel computing platform, the second workload metric being selected from the group consisting of processor utilization, memory utilization, storage utilization, network bandwidth utilization, arrival rate, inter-arrival time, response time, throughput, and service load pattern; making a second predicted scaling decision based on the received second workload metric; creating runtime modifiable auto-scaling business rules local to each level of the multi-level computing platform; modifying application-level configuration files and application-level settings according to the created business rules; creating a plan for changing the configuration local to each level of the multi-level computing platform; and Deploying and managing services operating on each of the levels of the multi-level computing platform based on the created plan. The computer program product of claim 8 , further causing one or more processors to perform each step of a method including:
12. A computer program as described in claim 8, wherein the set of business rules includes a set of conditional and consequential "when-then" rules, each of which means that when a condition occurs, the next result or action is to be performed.
13. triggering execution of the scaling plan, determining whether the first workload metric exceeds a threshold value written into the business rule for a predetermined period of time; and deploying the scaling plan in response to determining that the first workload metric exceeds the threshold for the predetermined period of time. The computer program of claim 8 further comprising:
14. triggering execution of the scaling plan, determining whether a threshold value written into the business rule exceeds the first workload metric for a predetermined period of time; and deploying the scaling plan in response to determining that the threshold exceeds the first workload metric for the predetermined period of time. The computer program of claim 8 further comprising:
15. 1. A computer system for automatic resource scaling in a multi-level computing platform, the multi-level computing platform comprising a hybrid cloud platform including multiple cloud levels constructed in a hierarchical manner, wherein each cloud level is selected from the group consisting of a private cloud computing platform and a public cloud computing platform; The computer system a processor set; and one or more computer-readable storage media It is equipped with wherein the set of processors is structured, arranged, connected, or programmed, or any combination thereof, to execute a computer program according to any one of claims 8 to 14, stored on the one or more computer-readable storage media.
Citation Information
Patent Citations
Resource monitoring method and apparatus
JP2009259005A
Cloud-based application resource manager
JP2014527221A
Resource allocation optimization system and method
JP2019082801A
Predictive Asset Optimization for Computer Resources
JP2020504382A