Pre-built configured AI proxy for cloud services

By using pre-built configurations generated by AI agents, utilization data, and project metadata to optimize cloud service resource allocation, the problem of excessive resource throttling or idleness in cloud services is solved, achieving efficient resource use and cost control.

CN121970029APending Publication Date: 2026-05-01MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-24
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Cloud service customers face the problem of excessive throttling or excessive idleness when configuring computing resources, resulting in poor performance or waste of resources. Existing methods rely on historical performance data for adjustment, but this is inefficient.

Method used

By employing pre-built configurations of artificial intelligence (AI) agents, leveraging prior utilization data and project metadata, pre-built configurations of computing resources are generated through hierarchical and target coding models. This approach optimizes resource allocation by balancing cost and performance, thereby reducing excessive throttling and idleness.

Benefits of technology

It effectively reduces the waste of computing resources, improves resource utilization efficiency, and achieves a balance between performance and cost by dynamically adjusting resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121970029A_ABST
    Figure CN121970029A_ABST
Patent Text Reader

Abstract

An example solution provides an artificial intelligence (AI) proxy for a pre-build configuration of a cloud service in order to enable initial build of computing resources (e.g., in the cloud service) to minimize the likelihood of excessive throttling or idling. Examples utilize previously existing utilization data and item metadata to identify similar use cases. Utilization data includes capacity information and resource consumption information (e.g., throttling and idling) for previously existing computing resources, and project metadata includes information for hierarchical classification to identify similar resources. A pre-built configuration is generated for resources of a customer that may adjust the pre-built configuration based on the customer's preferences for cost and performance balance points.
Need to check novelty before this filing date? Find Prior Art

Description

Pre-built and configured AI agents for cloud services Background Technology

[0001] The availability of public cloud services has facilitated access to a wide range of data services with diverse data analytics requirements, such as SQL / NoSQL databases, streaming, machine learning (ML), business insight analytics, and more. However, with a large number of options exposed, the complexity of configuring an optimal deployment increases significantly. Translating use cases into cloud service resource capacity configuration requirements is challenging. If a cloud service customer configures resources too low (e.g., too few processors), throttling occurs when resources are overwhelmed, damaging performance. If resources are configured generously, they experience idle (unused capacity), meaning the cloud service customer is paying for unnecessary capacity.

[0002] Customers can tailor resources based on performance history cycles (e.g., using support tickets to increase or remove capacity), but this approach requires collecting performance history during cycles of potential throttling or excessive idleness. In other words, customers suffer from poor performance or wasted money until the efficient level of required resource capacity is calculated. Summary of the Invention

[0003] The disclosed examples are described in detail below with reference to the accompanying drawings. The following summary of the invention is provided to illustrate some examples disclosed herein.

[0004] The example solution provides a pre-built configuration of an artificial intelligence (AI) agent for cloud services. The example receives previously existing utilization data and project metadata, where the utilization data includes capacity and resource consumption information for the previously existing computing resources, and where the project metadata includes information for hierarchically classifying the previously existing computing resources; it uses the utilization data and project metadata to create a capacity prediction model, which generates a pre-built configuration for a first computing resource; it uses the capacity prediction model to generate the pre-built configuration for the first computing resource; and it adjusts the pre-built configuration using a selected cost and performance balance point and historical data from the previously existing project. The capacity prediction model can take different forms based on the available metadata: a hierarchical model and a target-encoded model. Attached Figure Description

[0005] The disclosed example is described in detail below with reference to the accompanying drawings:

[0006] Figure 1 illustrates an example architecture that advantageously provides a pre-built configuration of an artificial intelligence (AI) agent for cloud services;

[0007] Figure 1A shows an alternative presentation of the architecture in Figure 1, illustrating both pre-built configurations and optimized cloud services;

[0008] Figure 2A illustrates how an example architecture such as Figure 1 can prevent excessive throttling of computing resources;

[0009] Figure 2B illustrates how an example architecture such as Figure 1 can prevent excessive idle computing resources;

[0010] Figure 2C illustrates a scenario where computing resources have achieved an optimal balance between throttling and idleness, which can be achieved using an example architecture such as that in Figure 1.

[0011] Figure 3 illustrates an exemplary workflow used by an example architecture such as that in Figure 1;

[0012] Figures 4A and 4B illustrate different options for filtering or encoding metadata, which can be used in example architectures such as Figure 1;

[0013] Figure 5 illustrates an exemplary hierarchy of one form of capacity prediction model, which can be used in example architectures such as Figure 1;

[0014] Figure 6 illustrates another exemplary workflow used by an example architecture such as that in Figure 1;

[0015] Figure 7 illustrates the cost and performance analyzer loop, which can be used in example architectures such as Figure 1;

[0016] Figure 8 shows an exemplary UI as seen by a viewer using an example architecture such as that in Figure 1;

[0017] Figure 9 illustrates an example feedback flow, which can be used in example architectures such as Figure 1;

[0018] Figures 10 and 11 show flowcharts illustrating exemplary operations that can be performed when using an architecture such as that of Figure 1; and

[0019] Figure 12 shows a block diagram of an example computing device suitable for implementing some of the various examples disclosed herein.

[0020] In all the accompanying drawings, the corresponding reference numerals indicate the corresponding parts. Detailed Implementation

[0021] Various aspects of this disclosure provide a pre-built configuration of an artificial intelligence (AI) agent for cloud services to minimize the possibility of excessive throttling or idleness during the initial construction of computing resources (e.g., in cloud services). Examples leverage previously existing utilization data and project metadata to identify similar use cases. Utilization data includes capacity and resource consumption information (e.g., throttling and idleness) for previously existing computing resources, and project metadata includes information for hierarchical classification, enabling the identification of similar projects and resources. A pre-built configuration is generated for a customer's resources, which the customer can tailor based on their preferences for a cost-performance balance point. Based on the available metadata, capacity prediction models employing different forms are used: a hierarchical model and a target-encoded model.

[0022] Various aspects of this disclosure reduce the count of computing resources used by customers of cloud services by providing pre-built configurations that reduce the possibility of excessive idleness. Various aspects of this disclosure further improve the performance of computing resources (including underlying devices) used by customers of cloud services by providing pre-built configurations that reduce the possibility of excessive throttling. This is achieved, at least in part, by using utilization data and project metadata for previously existing computing resources to create a capacity prediction model for generating pre-built configurations for the first computing resource. Therefore, various aspects of this disclosure address problems specific to the computing domain.

[0023] Various examples are described in detail with reference to the accompanying drawings. Preferably, the same reference numerals are used in all the drawings to denote the same or similar parts. References to specific examples and implementations throughout this disclosure are provided for illustrative purposes only and are not intended to limit all examples unless otherwise indicated.

[0024] Figure 1 illustrates an example architecture 100 that advantageously provides a pre-built configuration for an AI agent (trained model 110) for cloud services. Specifically, the trained model 110 (AI agent) generates a pre-built configuration 140 for configuring computing resources 144 to perform operations to generate output data 154 from input data 152, while minimizing the possibility of excessive throttling (e.g., avoiding under-configuration) and minimizing the possibility of excessive idleness (e.g., avoiding over-configuration).

[0025] Figures 2A through 2C illustrate throttling, idling, and optimization at the balance point between the target throttling ratio and the target idle ratio. Figure 2A shows excessive throttling of computing resources that can be avoided by using the example of architecture 100. Plot 200a shows curve 202a plotting processor utilization (axis 204) against time (axis 206). Curve 202a reaches its maximum and stabilizes during throttling events 210a, 210b, 210c, 210d, 210e, 210f, and 210g. Throttling events 210a through 210g occur when the demand for processor performance exceeds a certain threshold (such as 95%) of the maximum available capacity due to insufficient configuration for the current task (e.g., too few processor cores or too slow processors). This can lead to a decline in overall resource performance, as experienced by customers. Some data processing or retrieval jobs may take longer than expected (e.g., longer than without throttling) or be canceled entirely.

[0026] While processor utilization is plotted, other performance curves, such as memory utilization and storage device usage, can also reflect inadequate configuration. In some cases, high memory utilization drives the use of slower swap space to alleviate memory pressure, and memory usage can lead to failure when there isn't enough space in the configured persistent storage to hold the data. In cloud service configurations, virtual machines (VMs) can be used, meaning that processor cores, memory, and storage devices are all virtualized. In some examples, error ratios are used instead of throttling as a metric to indicate performance degradation due to inadequate configuration.

[0027] Figure 2B illustrates excessive idleness of computing resources, which can also be avoided by using the example of architecture 100. Plot 200b shows curve 202b plotting processor utilization (axis 204) against time (axis 206). Curve 202b shows a large gap between its typical maximum value and the maximum capacity of the resource. This is conceptually labeled as idle 212, although a more detailed definition of idleness is required.

[0028] In some examples, the following time series is used to define average slack. average ): Equation (1) where c is the capacity and slack is the instantaneous idle time at time sample t.

[0029] While processor utilization is plotted, other performance curves can also be used to reflect overprovisioning, such as memory utilization and storage usage, which may not be used in cases of overprovisioning.

[0030] Figure 2C illustrates a scenario where computing resources have achieved an optimal balance (“optimization”) between throttling and idleness using an example of architecture 100. Plot 200c shows curve 202c plotting processor utilization (axis 204) against time (axis 206). Curve 202c shows a target throttling ratio 214 for a period less than throttling events 210a to 210g, and a target idle ratio 216, theoretically smaller than idleness 212, with a smaller gap between its typical maximum and the maximum capacity of the resources.

[0031] In some cases, customers may be more sensitive to throttling than to idleness, so hard constraints can be set for throttling, and the capacity used to achieve the closest expected average idleness can be used. In some examples, the capacity c (given by k in Equation (2)) of the pre-built configuration 140 is optimized for a target idle ratio of 216 (Figure 1) and the target throttling ratio 214 (given by k in Equation (2)) is optimized. The given information is obtained by selecting c from a set of available capacity configurations C, where the available capacity configuration C makes the target idle ratio 216 and the average idle (slack) of equation (1) equal. average Minimize the difference between ). The probability P of being throttled for this capacity P(c) is lower than the target throttling ratio 214. This is shown as: Equation (2)

[0032] In some examples, the target throttling ratio 214 can be set to 0 (zero), and the target idle ratio 216 can be set to 50%. Other values ​​for the target throttling ratio (e.g., non-zero values) and other values ​​for the target idle ratio 216 can be used based on customer preferences.

[0033] It may be necessary to optimize the capacity of several computing resources individually, such as processor count (and speed), amount of memory, and storage space (capacity). For example, if computing resources are used for database applications, processor count and speed performance may be lower relative to memory and storage devices, while still providing acceptable performance, rather than if computing resources are used for heavy computations on relatively small amounts of data.

[0034] Returning to Figure 1, the previously existing computing resources 102 are computing resources that have been built and used by earlier customers and therefore have a performance history. This performance history can be used to generate a pre-built configuration 140 for a new customer project, such that the pre-built configuration 140 is optimized for a target throttling ratio 214 and a target idle ratio 216 (i.e., satisfying equation (2)). The previously existing computing resources 102 include previously existing computing resources 102a, 102b, and 102c.

[0035] As will be described below, the previously existing computing resource 102b is used to explain to the client (in the user interface (UI) 800) how the trained model (AI agent) generates the pre-built configuration 140. Interpretability may not be used in some examples, although it can be used to train architecture 100 for some new clients or other users.

[0036] Historical data 104 is collected from the activity of previously existing computing resources 102 and includes customer metadata 104a, resource solution history 104b, and resource health and utilization history 104c. Resource solution history 104b contains information such as processor counts and speeds, memory quantity, and storage capacity for previously existing computing resources 102, and provides a set of available configurations C for equation (2).

[0037] Resource health and utilization history 104c contains throttling and idle information for each resource in the previously existing computing resource 102, such as information collected on a per-second basis and processed for statistical characteristics, which can be stored more efficiently. Resource health and utilization history 104c includes information about the owners of the resources in the previously existing computing resource 102, such as industry and segment.

[0038] Customer metadata 104a includes information such as the industry in which the customer operates, departments within the customer's organization (referred to as "resource groups"), which may each have their own projects, and other data that characterizes specific projects so that other projects can be identified by other customers who may be more similar or less similar. For example, two customers in the food service industry may have similar requirements for cloud resources, although differences within the department may lead to differences in similarity. For example, a single entity in the food service industry may have a transportation department and a sales department, which have significantly different needs. For example, for the transportation department, delivery and route planning functions may have a critical need to avoid cost-cutting (or other performance degradation), while the marketing department may have a lower sensitivity to cost-cutting and a higher sensitivity to cost (e.g., a higher need to reduce idle time).

[0039] The generation of hierarchical similarity based on the resolution of this data collection improves the reliability of pre-built configuration 140. In some examples, at least some portions of the historical data can be anonymized. In some examples, historical data 104 has a history of thousands of projects, with daily updates (or more frequently for some data), indexed by customers, subscriptions, and resource groups (departments), and hierarchically categorized by providing types such as burstable (development), general (small production), and memory-optimized (large production) . To better match this dimension of project performance when generating pre-built configuration 140, projects can be clustered by the prevalence of workload spikes.

[0040] Trainer 106 uses historical data 104 to train training model 110 to generate capacity prediction model 130, which in turn produces pre-built configuration 140. Capacity prediction model 130 can take different forms based on available metadata: hierarchical model and target encoding model. This will be explained in conjunction with Figures 4A and 4B.

[0041] Pre-built configuration 140 provides customer-specific recommendations for configuring cloud service resources (such as VMs). In some examples, pre-built configuration 140 specifies the number of processors (and speed), the amount of memory, and the storage capacity. In some examples, training model 110 includes a single model, including a machine learning (ML) model. As used herein, AI includes, but is not limited to, ML. In some examples, training model 110 includes three different models, capacity model 112, workload prediction model 114, and balancing model 116, which will be described below. In some examples, capacity model 112 and workload prediction model 114 are combined into a single model (or ML model).

[0042] In some examples, trainer 106 executes the training that is currently being carried out by training model 110 (e.g., continuing to train one or more of the capacity model 112, workload prediction model 114, and balancing model 116). For example, after generating pre-built configuration 140 and builder 142 builds compute resource 144 based on pre-built configuration 140, compute resource 144 can begin execution within cloud execution environment 150. This initiates the development history for compute resource 144.

[0043] For example, customer feedback 160 in the form of Customer Reported Events (CRIs) and support tickets is provided to adjuster 162 that adjusts the capacity of computing resource 144. Customer feedback and capacity (and capacity change events) are added to historical data 104. Furthermore, utilization data such as workload (including spike information), throttling, and idle time are included in utilization data 164 for computing resource 144. This is also added to historical data 104, such as resource health and usage history 104c. These additions to historical data 104 provide new training material for trainer 106 for further training of model 110.

[0044] The trained model uses source data 120, which includes utilization data 122 extracted from historical data 104, project metadata 124, and project historical data 126. Utilization data 122 includes capacity information, resource consumption information, and workload information for previously existing computing resources. Capacity information includes processor counts for each resource in the previously existing computing resources 102, and sometimes processor speed, memory amount, and / or storage capacity; these values ​​may vary over time as these resources are scaled up or down. In some examples, each processor in the processor count includes a virtual core (vcore). In some examples, resource consumption information is idle and / or throttling information, and workload information, as described with respect to Figures 2A and 2B. In some examples, telemetry of utilization for each resource is provided at one-minute intervals.

[0045] Project metadata 124 includes information for hierarchically categorizing previously existing computing resources 102, such as metadata tags (e.g., classification identifiers for customers or resources). Examples include software version, localization tags (e.g., the region or country where the resource resides), or development / testing / production tags. Both resource-specific tags (e.g., development / testing / production) and broader customer-related tags (e.g., industry and other segmented data) can be used to allow for intelligent recommendations for both existing and new customers. In some cases, a hierarchy of metadata is utilized, starting from the subscription identifier (ID) and extending to wide segmented tags such as industry names (e.g., food and beverage or food service, manufacturing, consumer electronics).

[0046] Resource-specific tags (such as software version and development / small production / large production) can be extracted from the same source as capacity and utilization data. Customer metadata (e.g., subscriptions and resource groups) can be inferred from the resource ID path in the utilization table or extracted from customer subscription metadata. For consistency, customer data can be anonymized and processed.

[0047] Project history data 126 includes requested changes or reported events to previously existing computing resources 102, and includes customer satisfaction signal 602 (Figure 6). Project history data 126 may include anything indicating cost sensitivity (e.g., customers prefer cheaper offerings and are willing to make slight performance clicks to reduce costs) and performance sensitivity (e.g., customers prefer higher-performance offerings and are willing to pay more to avoid cost-cutting). Examples include support tickets and manual scaling actions on resources (e.g., using adjuster 162). Multiple CRIs can be used or labeled as cost-sensitive or performance-sensitive using keyword search or Large Language Model (LLM) to extract representations of CRIs. This is shown in Figure 7 below.

[0048] Some existing customers may use multiple computing resources, for example, with different departments (resource groups, such as transportation and marketing), and different departments may each have their own profiles. In some examples, a customer satisfaction signal 602 (see Figure 6) from an existing customer is propagated to the project profile via a weighted addition. For example, if a customer complains about the performance of a resource used by one of their departments, the signal will have a full impact on the profile of the project used by that department, but a reduced (e.g., lower-weighted) impact on the profiles of projects for other departments of the same customer.

[0049] The trained model 110 uses three phases, each providing a capacity recommendation. Phase 1 is capacity optimization, which calculates the ideal capacity using most or all of the previously existing computing resources 102 that have already been used in training. This is illustrated by capacity model 112, which generates capacity optimization phase 132 in capacity prediction model 130. Phase 2 is workload prediction, which recommends the optimal capacity of resources for new requests (e.g., pre-built recommendations), starting from capacity optimization phase 132 as the foundation. This is illustrated by workload prediction model 114, which generates workload prediction phase 134 in capacity prediction model 130. Regardless of how accurate workload prediction phase 134 is, different customers may have different preferences. This situation is addressed by balancing model 116.

[0050] Phase 3 is balancing or personalization, which adjusts (e.g., recalibrates) the recommendations calculated in Phase 2 (e.g., workload forecasting phase 134) based on the customer's cost and performance preferences. This is illustrated by the balancing model 116 that produces adjustment phase 136. In some examples, adjustment phase 136 is considered as a third phase within capacity forecasting model 130, while in other examples, capacity forecasting model 130 has only two phases (capacity optimization phase 132 and workload forecasting phase 134), and adjustment phase 136 falls outside of capacity forecasting model 130.

[0051] These three stages can be used together or independently. For example, stage 2 is required for a pre-built configuration, and in some examples, stages 1 and 3 are optional. Using all three stages together to generate a pre-built configuration 140 can be viewed as a two-stage approach: using stages 1 and 2 as the first stage to generate an initial pre-built configuration 140a, and then, in the second stage, personalizing the pre-built configuration 140 to a modified pre-built configuration 140b upon receiving customer input regarding preferences after seeing the initial pre-built configuration 140 (in UI 800, as described below with reference to Figure 8). Furthermore, stage 3 can be used to adjust the computing resource 144 after it has been built and execution has begun, for example, by providing input to the adjuster 162.

[0052] Capacity model 112 generates capacity optimization phase 132 by evaluating the relationship between resource workload and capacity, identifying opportunities for cost savings through scaling down, or for performance gains through scaling up. This phase can be considered as identifying the goodness of fit of a given resource capacity to the workload. For example, a workload that typically requires three processors (e.g., processor cores or vcores) can be recommended to scale up from two processors to four processors to adjust the resource, thereby achieving performance improvements, and can facilitate pre-built recommendations for four processors.

[0053] Optimization requires two inputs: resource utilization and resource capacity. In some cases, capacity may not change much over time, while workload typically changes over short time spans. Therefore, capacity and workload information may be recorded at different frequencies. Some examples involve operating on aggregate values ​​such as peak resource usage or average unused resource capacity. However, when resources undergo throttling, the true workload is unobservable. These are called censored workloads. Alternative optimization methods are employed for these cases. For example, censored workloads are correctly sized under the assumption that throttling a capacity will not occur in the next larger capacity selection.

[0054] The workload prediction model 114 generates the workload prediction phase 134 by utilizing the capacity optimization phase 132 and metadata describing new customers and new requests for resources (e.g., what will become computing resources 144). This task can be described as follows: given a vector M describing a customer with type O and the resources it requests, a function Y=f(M, O) is defined to recommend the optimal resource capacity (e.g., pre-built configuration 140).

[0055] The requested provision type O corresponds to burstable (e.g., development), general purpose (e.g., small production), and memory-optimized (e.g., large production), and is shown as a UI input in Figure 8. In some examples, different sets of possible capacities may exist for each provision type. Some examples use provision types to stratify predictions and constrain configurations to ensure effective capacity for a given tier. For example, a burstable configuration may be provided using a single processor, while general-purpose and memory-optimized provisioning may have a minimum of two processors.

[0056] This phase is powered by metadata tags, which are discrete attributes that can take any value. The range of metadata tags can range from software version to resource URI path information (resource groups, subscriptions, etc.) to customer segmentation data (e.g., industry name). Metadata tags are typically associated with underlying workloads (such as test / development / production tags) and are useful for associating similar workloads and for predicting typical workloads for new resource requests.

[0057] The balancing model 116 generates the adjustment phase 136 by personalizing the workload forecasting phase 134 based on the new customer's cost and performance preferences. The balancing model 116 can use various sources, such as subscription metadata, resource metadata, customer interactions, and VM telemetry. The balancing model 116 can learn each customer's (or a customer's department's) cost and performance preferences by evaluating historical customer interactions related to resource configuration, scaling actions, and performance CRIs.

[0058] Figure 1A shows alternative architecture 100a. Unless otherwise specified or impractical, references to architecture 100 thereafter also refer to architecture 100a. Historical utilization data 104 of architecture 100a is used for training using trainer 106 and trainer 106a as described above. Trainer 106 provides training to build a capacity prediction model 130 that determines a pre-built configuration 140, while trainer 106a provides training to build an optimized model 130a that determines an optimized configuration 140. The pre-built configuration 140 is used for the initial build of computational resources 144 (via builder 142), attempting to obtain an appropriate size for the client at the outset.

[0059] However, optimization is employed when the pre-built configuration 140 is not ideal, and whenever customer needs (e.g., workload) change. Even if the pre-built configuration 140 is initially ideal, it may not remain ideal over time. Therefore, optimization model 130a determines an optimized configuration 140 for adjusting the size / capacity of computing resources 144, which is used by adjuster 162.

[0060] As shown, the new customer metadata 104d is associated with the new project that builds computing resources 144 and is used in the generation of the capacity prediction model 130. However, customer metadata 104d is added to historical data 104 for use in improving future new projects.

[0061] Computational resource 144 generates online resource 146, which, over time, produces new resource solution history 104e and resource health and utilization history 104f. Resource solution history 104e and resource health and utilization history 104f are used for optimization, for example, in the generation of optimization model 130a. Resource solution history 104b and resource health and utilization history 104c are also added to historical data 104 for use in improving future new projects.

[0062] Figure 3 illustrates an exemplary workflow 300 used by an example of architecture 100. In training pipeline 302, trainer 106 uses historical data 104 to train trained model 110. A client creates a database in operation 304 and accesses trained model 110 using API endpoint 306. Trained model 110 generates pre-built configuration 140, which is passed to initial configuration operation 308, in which builder 142 builds compute resources 144 based on pre-built configuration 140.

[0063] Figures 4A and 4B illustrate different options for filtering or encoding metadata. An example of architecture 100 can use the hierarchical filter 400a (or hierarchical configuration side) of Figure 4A as the underlying model, or alternatively, the target encoder 400b (or target encoding configuration side) of Figure 4B as the underlying model. The choice between hierarchical filter 400a and target encoder 400b depends on the nature of the item metadata 124. If the item metadata 124 is sufficiently hierarchical (see Figure 5), hierarchical filter 400a is used, and the capacity prediction model 130 takes the form of a hierarchical model. However, if insufficient hierarchical levels exist, target encoder 400b is used, and the capacity prediction model 130 takes the form of a target encoding model.

[0064] Figure 4A provides a conceptual illustration of hierarchical filter 400a, which includes a top-level hierarchy 402a, hierarchy 402b, hierarchy 402c, hierarchy 402d, and hierarchy 402e. Hierarchical filter 400a employs a heuristic approach that leverages the inherent hierarchical structure of some metadata features to recommend capacity based on previously existing computing resources for similar clients. The approach uses several stages: (1) computing the hierarchical structure of metadata features, (2) categorizing each observed previously existing computing resource into buckets along this hierarchy, and (3) using the filled buckets to recommend capacity for new computing resources (e.g., computing resource 144).

[0065] Some examples compute pairwise hierarchical relationships between features to construct a hierarchy graph. In this method, pairwise entropy is calculated, and uncertainty reduction is derived from this entropy, followed by the computation of a table of hierarchy strength values. This table can include moderately strict hierarchy relationships, as well as some near-strict and some weak hierarchy relationships. To ensure that only meaningful relationships are included in the final hierarchy, a minimum pairwise entropy threshold is used, where all values ​​below the threshold are set to zero.

[0066] To compute the hierarchical chain, a weighted directed acyclic graph (DAG) is constructed using a threshold table as the adjacency matrix. That is, the edges pointing from one node to another in the table values ​​between nodes are non-zero (they may be set to zero due to the threshold). The most granular feature in the hierarchy is the node with the highest out-degree in the DAG (i.e., the row with the least non-zero value in the table). To extract the remaining elements of the hierarchical chain, the DAG is traversed through the neighbors of each consecutive node with the highest out-degree. The chain terminates at the first node reached outside 0 degrees (i.e., the coarsest feature).

[0067] Next, each previously existing compute resource is categorized into buckets along the computed hierarchy. In some examples, other categorization schemes may also be used. The optimal (e.g., optimized) capacity (e.g., c in equation (2)) for a new compute resource is found by first selecting the most granular metadata feature with non-empty buckets, using filled buckets. Next, the percentage of the target capacity distribution of the buckets is computed as a capacity recommendation. Some examples use the 50th percentile (e.g., the median). Because the final recommendation is a simple percentile of the existing workloads in the bucket, it can be interpreted by the hierarchy level used to identify similar resources and the list of existing compute resources in that bucket and their capacities. This is shown in Figure 8 below.

[0068] Figure 4B provides a conceptual illustration of a target encoder 400b, which includes a top-level encoder 404 and encoder branches 406a, 406b, and 406c in a common feed structure. The target encoder 400b learns arbitrary relations within the metadata. The target encoding method encodes categorical, potentially high-cardinality attributes as real numbers by using a grouped set of regression targets, thus implementing classic ML algorithms such as random forests and gradient boosting trees. This tends to outperform other encoding methods (e.g., one-hot encoding), especially in the case of tree-based methods.

[0069] This method can be implemented in two stages: (1) For each attribute, target value, and value, an aggregation function such as the mean or percentage points is mapped. This value can represent the average number of processors selected for all customers sharing the same value for that attribute. (2) After transforming the features in the data using a function (e.g., features learned in a hierarchical model), regression is used to estimate the target value given the attribute. For interpretability (see Figure 8), the average target value for each bucket used as input to the predictive model is shown.

[0070] Figure 5 illustrates an exemplary hierarchy 500 of a capacity forecasting model 130 (when the capacity forecasting model takes the form of a hierarchical model) with example values. Hierarchy 500 has a segment name 502a as the top level, an industry name 502b as the next level, a vertical name 502c as the next level, a customer name 502d as the next level, a subscription ID 502e as the next level, and a resource group 502f (e.g., department) as the lowest level shown. Other hierarchical arrangements are also possible, with different names and different numbers of levels.

[0071] Figure 6 illustrates an exemplary workflow 600 used by an example of architecture 100. A trained model 110 generates a pre-built configuration 140 for constructing computing resource 144. During the execution of computing resource 144, the customer provides a customer satisfaction signal 602, such as a requested change (scaling action) to computing resource 144 or a reported event (e.g., CRI). In operation 604, the customer satisfaction signal 602 is used to update the project profile 708 (see Figure 7), and this new information is used in the ongoing training of the trained model 110.

[0072] Figure 7 illustrates a cost and performance analyzer loop 700 that can be used in an example of architecture 100. Customer selection 702 and resource telemetry 704 (e.g., utilization data, idle data, and throttling data) provided via UI 800 (shown in Figure 8) are provided to profile initializer 706, which creates a project profile 708 that is persistently stored in profile repository 710. A customer satisfaction signal 602, including CRI 712, is used to update project profile 708 (as mentioned with respect to Figure 6). This is achieved by providing customer satisfaction signal 602 to profile updater 714. Profile updater 714 performs updates to project profile 708 in profile repository 710.

[0073] CRI 712 is sent to profile updater 714 via CRI tagger 716 to be marked as cost-sensitive or performance-sensitive. In some examples, profile updater 714 uses LLM to extract a representation of each CRI.

[0074] Figure 8 illustrates UI 800. The subscription number window 802 displays the subscription number for the current project (e.g., the project for which pre-build configuration 140 is being generated). The offer selection window 804 will offer options displayed as small production 804a (e.g., general), large production 804b (e.g., memory optimized), and development 804c (e.g., burstable). As shown, the build target selection 806 is currently set to large production 804b, but can be changed by the customer after the customer sees the pre-build configuration 140 displayed in the model prediction window 820.

[0075] The cost-performance balance point selection slider 808 allows the customer to indicate the selected cost-performance balance point 810, although other input schemes can also be used. The user clicks the "Generate" button 812 to generate a pre-built configuration 140. The trained model 110 generates the pre-built configuration 140 at this point, and the UI 800 displays the pre-built configuration 140 in the model prediction window 820. As shown, the model prediction window 820 displays the processor count 820a, memory amount 820b, and storage capacity 820c for the pre-built configuration 140. Some examples may display different or additional information.

[0076] Once the customer changes the build target selection 806 in the selection window 804, and / or changes the selected cost and performance balance point 810 on the cost and performance balance point selection slider 808, and clicks the "Generate" button 812 to generate a new version of the pre-build configuration 140, the model interpretability window 830 displays information 830a about the previously existing computing resources 102b used to generate the pre-build configuration 140, as well as at least a portion of the hierarchy 500—in those examples that provide interpretability.

[0077] Use the interpretability settings selection 816 within the interpretability settings selection window 814 to adjust the amount and type of information displayed in the model interpretability window 830. As shown, the interpretability settings selection window 814 displays two options, showing instances of them in bucket selection 814a and a histogram in logarithmic scale selection 814b. In some examples, alternative options may be provided in place of the displayed options or in addition to the displayed options. The information window 818 provides the customer with additional information about other options the customer may wish to try or use.

[0078] Figure 9 illustrates an example feedback flow 900 that can be used as in the example of architecture 100. The customer creates a database in operation 304 and initial configuration operation 308, in which builder 142 builds compute resource 144 based on pre-built configuration 140, as described above with respect to Figure 3. In parallel, the new database is a new project that generates a new project profile in project profile 708, as described with respect to Figure 7. A change 902 in the use of compute resource 144 triggers a scaling action (e.g., scaling up or scaling down) by adjuster 162.

[0079] Figure 10 shows a flowchart 1000 illustrating exemplary operations that can be performed by architecture 100. In some examples, the operations described for flowchart 1000 are performed by computing device 1200 of Figure 12. Flowchart 1000 begins in operation 1002 by compiling tenant data as historical data 104 (e.g., including customer metadata 104a, resource solution history 104b, and resource health and utilization history 104c) into training data for use in training trained model 110.

[0080] In operation 1004, trainer 106 trains trained model 110, which in some examples includes: training capacity model 112 to perform capacity sizing adjustments for capacity prediction model 130 based at least on utilization data 122; training workload prediction model 114 to perform workload prediction for capacity prediction model 130 based at least on project metadata 124; and training balancing model 116 to perform capacity sizing adjustments for capacity prediction model 130 based at least on project historical data 126.

[0081] Operation 1006 presents UI 800 to the customer to receive customer preferences in the form of choices within UI 800. This includes initial build target selection 806 and the selected cost and performance balance point 810. However, the customer can change the build target selection 806 and / or the selected cost and performance balance point 810 at a later time.

[0082] Alternatively or additionally, either or both of the initial build target selection 806 and the cost-performance trade-off point 810 can be automatically inferred via one or more ML models described herein, without requiring any UI input from the customer. For example, for the initial build target selection 806, one of the "burstable, general, memory-optimized" options can be automatically selected as an internal policy or attribute for the new resource based on the ML model. In contrast to the customer inputting and managing the initial build target selection 806 via UI 800, the underlying infrastructure applies and transitions between these internal policies.

[0083] Similarly, the ML model predicts the cost and performance balance point using past scaling action and event data from the customer (e.g., CRI). The ML model is applied to new resources, and the ML model is adjusted as part of the feedback loop via any manually scaled data from the new resources and / or via CRI data.

[0084] In operation 1008, the trained model 110 receives customer metadata (e.g., customer metadata 104d). In operation 1010, the trained model 110 uses utilization data 122 and project metadata 124 to create a capacity prediction model 130, and the capacity prediction model 130 generates a pre-built configuration 140 for computing resources 144. This includes, in operation 1012, optimizing the pre-built configuration 140 to minimize the expected throttling idleness by achieving a target idle ratio 216 and a target throttling ratio 214.

[0085] Operation 1014 displays at least a portion of the pre-built configuration 140 in UI 800 (e.g., processor count, memory amount, and / or storage capacity). Operation 1016 displays at least a portion of the hierarchy 500 of the capacity prediction model 130 in UI 800. Operation 1018 receives interpretability setting selection 816 via UI 800. Operation 1020 displays information in UI 800 regarding the previously existing computing resources 102b used in generating the capacity prediction model 130.

[0086] In decision operation 1022, the customer determines whether the pre-built configuration 140 is acceptable to the customer's needs. If not, in operation 1024, the customer has the option to perform manual adjustments. When the customer is satisfied with the pre-built configuration 140, flowchart 1000 continues. In operation 1026, builder 142 builds compute resource 144 based on the pre-built configuration 140. Computation resource 144 begins execution in operation 1028, thereby receiving input data 152 to generate output data 154 and receiving online resource data (e.g., new resource solution history 104e and resource health and utilization history 104f).

[0087] Operation 1030 performs online optimization to adjust computing resources 144 to a more optimal capacity using tuner 162. This includes minimizing operation throttling and idle time while adjusting computing resources 144 based on a selected cost and performance balance point 810.

[0088] Operation 1036 adds (e.g., further compiles) historical data 104, including customer metadata 104a, resource solution history 104b, and resource health and utilization history 104c, so that historical data 104 now includes pre-built configuration 140 and utilization data 164 for compute resource 144. In operation 1038, trainer 106 also uses customer metadata 104a and resource solution history 104b to train capacity model 112, workload prediction model 114, and / or balancing model 116.

[0089] Figure 11 shows a flowchart 1100 illustrating exemplary operations that can be performed by architecture 100. In some examples, the operations described for flowchart 1100 are performed by computing device 1200 of Figure 12. Flowchart 1100 begins with operation 1102, which includes receiving previously existing utilization data and item metadata, wherein the utilization data includes capacity information and resource consumption information for previously existing computing resources, and wherein the item metadata includes information for hierarchically classifying the previously existing computing resources.

[0090] Operation 1104 involves using utilization data and project metadata to create a capacity prediction model for generating a pre-built configuration for the first computing resource. Operation 1106 involves using the capacity prediction model to generate a pre-built configuration for the first computing resource. Operation 1108 involves adjusting the pre-built configuration using a selected cost and performance trade-off point and existing project history data. Additional examples

[0091] An exemplary system includes: a processor; and a computer-readable medium storing instructions, which, when executed by the processor, are configured to: receive previously existing utilization data and project metadata, wherein the utilization data includes capacity information and resource consumption information for a previously existing computing resource, and wherein the project metadata includes information for hierarchically classifying the previously existing computing resources; use the utilization data and project metadata to create a capacity prediction model for generating a pre-built configuration for a first computing resource; use the capacity prediction model to generate a pre-built configuration for the first computing resource; and adjust the pre-built configuration using a selected cost and performance balance point and previously existing project history data.

[0092] An example computer-implemented method includes: receiving previously existing utilization data and project metadata, wherein the utilization data includes capacity information and resource consumption information for a previously existing computing resource, and wherein the project metadata includes information for hierarchically classifying the previously existing computing resources; using the utilization data and project metadata to create a capacity prediction model for generating a pre-built configuration for a first computing resource; using the capacity prediction model to generate a pre-built configuration for the first computing resource, the pre-built configuration including processor counts, memory quantity, and / or storage capacity for the first computing resource; and adjusting the pre-built configuration using a selected cost and performance trade-off point and previously existing project history data.

[0093] One or more example computer storage devices have computer-executable instructions stored thereon that, when executed by a computer, cause the computer to perform operations including: receiving previously existing utilization data and project metadata, wherein the utilization data includes capacity information and resource consumption information for a previously existing computing resource, and wherein the project metadata includes information for hierarchically classifying the previously existing computing resource; using the utilization data and project metadata to create a capacity prediction model for generating a pre-built configuration for a first computing resource; using the capacity prediction model to generate the pre-built configuration for the first computing resource; adjusting the pre-built configuration using a selected cost and performance balance point and previously existing project historical data; and building the first computing resource according to the pre-built configuration.

[0094] Alternatively, or in addition to other examples described herein, examples include any combination of the following: - a pre-built configuration including processor counts, memory quantity, and / or storage capacity for a first computing resource; - building the first computing resource according to the pre-built configuration; - receiving input data from the first computing resource; - generating output data using the first computing resource, at least based on the input data; - capacity information including processor counts, memory quantity, and / or storage capacity for a previously existing computing resource; - resource consumption information including idle and throttling information for a previously existing computing resource; - project history data including requested changes or reported events for a previously existing computing resource; - training a first model, at least based on the utilization data, to perform capacity sizing for a capacity prediction model; - training a second model, at least based on the project metadata, to perform workload prediction for a capacity prediction model; - training a third model, at least based on the project history data, to perform capacity sizing for a capacity prediction model; - presenting a UI; - receiving an initial build target selection through the UI, wherein generating the pre-built configuration for the first computing resource includes at least based on... The process includes: - Selecting an initial build target to generate a pre-built configuration for a first computing resource; - Receiving the selected cost and performance trade-off point through the UI; - Displaying at least a portion of the pre-built configuration in the UI; - Displaying at least a portion of the hierarchy of the capacity prediction model in the UI; - Displaying information for a previously existing computing resource used to generate the capacity prediction model in the UI; - Generating the pre-built configuration by optimizing it for a target idle ratio and a target throttling ratio; - Adjusting the pre-built configuration by optimizing it for a target idle ratio and a target throttling ratio; - Each of the first, second, and third multi-models includes an ML model; - Each of the first, second, and third multi-models is within a fourth model; - Each processor in the processor count includes a virtual core; - Receiving interpretability setting selections through the UI; - Compiling the configuration history and computing resource history; - Using the configuration history and computing resource history to further train the first, second, and / or third models; - The configuration history includes the pre-built configuration; and - The computing resource history includes utilization data for the first computing resource.

[0095] While various aspects of this disclosure have been described with reference to various examples having their associated operations, those skilled in the art will understand that combinations of operations from any number of different examples are also within the scope of various aspects of the invention. Example operating environment

[0096] Figure 12 is a block diagram of an example computing device 1200 (e.g., a computer storage device) used to implement the various aspects disclosed herein, and is generally designated as computing device 1200. In some examples, one or more computing devices 1200 are provided for provisioning computing solutions. In some examples, one or more computing devices 1200 are provided as cloud computing solutions. In some examples, a combination of provisioning and cloud computing solutions is used. Computing device 1200 is merely one example of a suitable computing environment and is not intended to impose any limitation on the scope or functionality of the examples disclosed herein, whether used alone or as part of a larger set.

[0097] The computing device 1200 should not be construed as having any dependency or requirement on any of the components / modules shown or a combination thereof. The examples disclosed herein can be described in the general context of computer code or machine-usable instructions (including computer-executable instructions such as program components) executed by a computer or other machine, such as a personal data assistant or other handheld device. Typically, a program component, including routines, programs, objects, components, data structures, etc., refers to code that performs a specific task or implements a specific abstract data type. The disclosed examples can be practiced in a variety of system configurations, including personal computers, laptops, smartphones, mobile tablets, handheld devices, consumer electronics, dedicated computing devices, etc. The disclosed examples can also be practiced in distributed computing environments when the task is performed by a remote processing device linked via a communication network.

[0098] Computing device 1200 includes a bus 1210 that directly or indirectly couples to the following devices: computer storage memory 1212, one or more processors 1214, one or more presentation components 1216, input / output (I / O) ports 1218, I / O components 1220, power supply 1222, and network components 1224. Although computing device 1200 is depicted as appearing as a single device, multiple computing devices 1200 can work together and share the depicted device resources. For example, memory 1212 can be distributed across multiple devices, and processors(s) 1214 can be housed in different devices.

[0099] Bus 1210 represents a bus that may be one or more buses (e.g., an address bus, a data bus, or a combination thereof). Although lines are used to show the various boxes in FIG12 for clarity, alternative representations may be used to complete the depiction of the various components. For example, presentation components such as display devices are I / O components in some examples, and some examples of processors have their own memory. No distinction is made between categories such as “workstation,” “server,” “laptop,” “handheld device,” etc., because all of these are considered to be within the scope of FIG12, and the term “computing device” is used herein. Memory 1212 may take the form of computer storage media referenced below and is operatively provided for computing device 1200 to store computer-readable instructions, data structures, program modules, and other data. In some examples, memory 1212 stores one or more of an operating system, a general-purpose application platform, or other program modules and program data. Thus, memory 1212 is capable of storing and accessing data 1212a and instructions 1212b, which can be executed by processor 1214 and configured to perform the various operations disclosed herein. Therefore, computing device 1200 includes a computer storage device having computer-executable instructions 1212b stored thereon.

[0100] In some examples, memory 1212 includes computer storage media. Memory 1212 may include any number of memories associated with or accessible by computing device 1200. Memory 1212 may be internal to computing device 1200 (as shown in Figure 12), external to computing device 1200 (not shown), or both (not shown). Additionally or alternatively, memory 1212 may be distributed across multiple computing devices 1200, for example, in a virtualized environment where instruction processing is performed on multiple computing devices 1200. For the purposes of this disclosure, “computer storage media,” “computer storage memory,” “memory,” and “memory device” are synonymous terms for memory 1212, and none of these terms includes a carrier or propagation signaling.

[0101] The (multiple) processors 1214 may include any number of processing units that read data from various entities, such as memory 1212 or I / O components 1220. Specifically, the (multiple) processors 1214 are programmed to execute computer-executable instructions for implementing various aspects of this disclosure. These instructions may be executed by a processor, multiple processors within computing device 1200, or a processor external to client computing device 1200. In some examples, the (multiple) processors 1214 are programmed to execute instructions, such as those shown in the flowcharts discussed below and depicted in the accompanying drawings. Furthermore, in some examples, the (multiple) processors 1214 represent an implementation of analog technology for performing the operations described herein. For example, the operations may be performed by analog client computing device 1200 and / or digital client computing device 1200. A presentation component 1216 presents data indications to a user or other device. Exemplary presentation components include display devices, speakers, printing components, vibration components, etc. Those skilled in the art will understand and recognize that computer data can be presented in a variety of ways, such as visually in a graphical user interface (GUI), audibly through speakers, wirelessly between computing devices 1200, via a wired connection, or otherwise. I / O port 1218 allows computing device 1200 to be logically coupled to other devices including I / O components 1220, some of which may be built-in. Example I / O components 1220 include, for example, but not limited to, microphones, joysticks, game controllers, disc-shaped satellite antennas, scanners, printers, wireless devices, etc.

[0102] Computing device 1200 can operate in a networked environment via network component 1224 through a logical connection to one or more remote computers. In some examples, network component 1224 includes a network interface card and / or computer-executable instructions (e.g., a driver) for operating the network interface card. Communication between computing device 1200 and other devices can occur via any wired or wireless connection, using any protocol or mechanism. In some examples, network component 1224 is operable to transmit data over public, private, or hybrid (public and private) networks using a transport protocol that utilizes short-range communication technologies (e.g., near-field communication, Bluetooth). TM Between wireless devices (such as brand communications, etc.) or combinations thereof. Network component 1224 communicates with remote resource 1228 (e.g., cloud resource) via network 1230 on wireless communication link 1226 and / or wired communication link 1226a. Various examples of communication links 1226 and 1226a include wireless connections, wired connections, and / or dedicated links, and in some examples, are at least partially routed via the Internet.

[0103] While described in conjunction with example computing device 1200, the examples of this disclosure can be implemented using many other general-purpose or special-purpose computing system environments, configurations, or devices. Examples of well-known computing systems, environments, and / or configurations that may be applicable to various aspects of this disclosure include, but are not limited to, smartphones, mobile tablets, mobile computing devices, personal computers, server computers, handheld or laptop devices, multiprocessor systems, game consoles, microprocessor-based systems, set-top boxes, programmable consumer electronics, mobile phones, mobile computing and / or communication devices of wearable or accessory form factors (e.g., watches, glasses, headsets, or headphones), network PCs, minicomputers, mainframes, distributed computing environments of any of the systems or devices included above, virtual reality (VR) devices, augmented reality (AR) devices, mixed reality devices, holographic devices, etc. Such systems or devices can accept input from users in any manner, including from input devices such as keyboards or pointing devices, via gesture input, proximity input (such as by hover), and / or via voice input.

[0104] Examples of this disclosure can be described in the general context of computer-executable instructions (such as program modules) that are executed by one or more computers or other devices as software, firmware, hardware, or a combination thereof. Computer-executable instructions can be organized into one or more computer-executable components or modules. Typically, program modules include, but are not limited to, routines, programs, objects, components, and data structures that perform a particular task or implement a particular abstract data type. Various aspects of this disclosure can be implemented using any number and organization of such components or modules. For example, various aspects of this disclosure are not limited to the specific computer-executable instructions or specific components or modules shown in the accompanying drawings and described herein. Other examples of this disclosure may include different computer-executable instructions or components having more or fewer functions than those shown and described herein. In examples involving general-purpose computers, various aspects of this disclosure, when configured to execute the instructions described herein, transform a general-purpose computer into a special-purpose computing device.

[0105] By way of example and not limitation, computer-readable media include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable memory implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, etc. Computer storage media are tangible and mutually exclusive with communication media. Computer storage media are implemented in hardware and do not include carrier waves and propagating signals. Computer storage media used in this disclosure do not include signals. Exemplary computer storage media include hard disks, flash drives, solid-state memory, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, optical disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage devices, magnetic tape cassettes, magnetic tape, disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information for access by a computing device. In contrast, communication media typically embody computer-readable instructions, data structures, program modules, etc., in the form of modulated data signals (such as carrier waves or other transmission mechanisms), and include any information transmission medium.

[0106] The execution or order of operations shown and described herein is not essential and may be performed in different orders in various examples. For example, it is contemplated that a particular operation may be performed before, simultaneously with, or after another operation within the scope of various aspects of the invention. When introducing elements of various aspects of this disclosure or its examples, the articles “a,” “an,” “the,” and “described” are intended to indicate the presence of one or more elements. The terms “comprising,” “including,” and “having” are intended to be inclusive and mean that additional elements may be present in addition to the listed elements. The term “exemplary” is intended to mean “an example of…”. The phrase “one or more of the following: A, B, and C” means “at least one of A and / or at least one of B and / or at least one of C.”

[0107] Various aspects of this disclosure have been described in detail, and it will be apparent that modifications and variations are possible without departing from the scope of the aspects of this disclosure as defined in the appended claims. Since various changes can be made to the above constructions, products, and methods without departing from the scope of the aspects of this disclosure, everything contained in the foregoing description and shown in the accompanying drawings should be interpreted as illustrative rather than restrictive.

Claims

1. A system comprising: Processor (1214); as well as A computer-readable medium (1212) storing instructions (1212b) that, when executed by the processor, operate to: receive (1102) previously existing utilization data (122) and project metadata (124), wherein the utilization data includes capacity information and resource consumption information for a previously existing computing resource (102), and wherein the project metadata includes information for hierarchically classifying the previously existing computing resource; use the utilization data and project metadata to create (1104) a capacity prediction model (130) for generating a pre-built configuration (140) for a first computing resource (144); use the capacity prediction model to generate (1106) the pre-built configuration for the first computing resource; and adjust (1108) the pre-built configuration using a selected cost and performance balance point (810) and previously existing project history data (126).

2. The system of claim 1, wherein the instructions further operate to: construct the first computing resource according to the pre-built configuration, the pre-built configuration including processor count, memory amount, and / or storage capacity for the first computing resource.

3. The system of claim 2, wherein the instructions further operate to: receive input data from the first computing resource; and generate output data using the first computing resource, at least based on the input data.

4. The system according to any one of claims 1 to 3, wherein the capacity information includes processor counts, memory quantity, and / or storage capacity for the pre-existing computing resources; wherein the resource consumption information includes idle information and throttling information for the pre-existing computing resources; and wherein the project history data includes requested changes or reported events for the pre-existing computing resources.

5. The system according to any one of claims 1 to 4, wherein the instructions further operate to: train a first model at least based on the utilization data to perform capacity sizing adjustments for the capacity prediction model; A second model is trained based at least on the project metadata to perform workload prediction against the capacity prediction model; And at least based on the project's historical data, a third model is trained to perform capacity size adjustments against the capacity prediction model.

6. The system according to any one of claims 1 to 5, wherein the instructions further operate to: present a user interface (UI); receive an initial build target selection via the UI, wherein generating the pre-build configuration for the first computing resource includes generating the pre-build configuration for the first computing resource based at least on the initial build target selection; receive the selected cost and performance balance point via the UI; and display at least a portion of the pre-build configuration in the UI.

7. The system of claim 6, wherein the instructions further operate to: display at least a portion of the hierarchy of the capacity prediction model in the UI; and display information in the UI for a previously existing computing resource used to generate the capacity prediction model.

8. A computer-implemented method, comprising: Receive (1102) previously existing utilization data (122) and project metadata (124), wherein the utilization data includes capacity information and resource consumption information for a previously existing computing resource (102), and wherein the project metadata includes information for hierarchically classifying the previously existing computing resource; use the utilization data and project metadata to create (1104) a capacity prediction model for generating a pre-built configuration (140) for a first computing resource (144); use the capacity prediction model to generate (1106) the pre-built configuration for the first computing resource, the pre-built configuration including processor count (820a), memory quantity (820b), and / or storage capacity (820c) for the first computing resource; and adjust (1108) the pre-built configuration using a selected cost and performance balance point (810) and previously existing project history data (126).

9. The computer-implemented method according to claim 8, further comprising: The first computing resource is constructed according to the pre-built configuration; The input data is received by the first computing resource; And at least based on the input data, the first computing resources are used to generate output data.

10. The computer-implemented method of claim 8 or 9, wherein the capacity information includes processor counts, memory amounts, and / or storage capacity for the previously existing computing resources; wherein the resource consumption information includes idle information and throttling information for the previously existing computing resources; and wherein the project history data includes requested changes or reported events for the previously existing computing resources.

11. The computer-implemented method according to any one of claims 8 to 10, further comprising: The first model is trained at least based on the utilization data to perform capacity sizing adjustments for the capacity prediction model; A second model is trained based at least on the project's historical data to perform workload prediction against the capacity prediction model; And at least based on the project's historical data, a third model is trained to perform capacity size adjustments against the capacity prediction model.

12. The computer-implemented method according to any one of claims 8 to 11, further comprising: Present the user interface (UI); The UI is used to receive an initial build target selection, wherein generating the pre-build configuration for the first resource includes generating the pre-build configuration for the first computing resource based at least on the initial build target selection; the UI is used to receive the selected cost and performance balance point; and at least a portion of the pre-build configuration is displayed in the UI.

13. The computer-implemented method according to claim 12, further comprising: The UI displays at least a portion of the hierarchy of the capacity prediction model; The UI also displays information about previously existing computing resources used to generate the capacity prediction model.

14. The computer-implemented method according to any one of claims 8 to 13, wherein generating the pre-built configuration and adjusting the pre-built configuration each include optimizing the pre-built configuration for a target idle ratio and a target throttling ratio.

15. A computer storage device having computer-executable instructions stored thereon, the computer-executable instructions, when executed by a computer, causing the computer to perform operations, the operations including: Receive (1102, 1008) previously existing utilization data (122) and project metadata (124), wherein the utilization data includes capacity information and resource consumption information for the previously existing computing resource (102), and wherein the project metadata includes information for hierarchically classifying the previously existing computing resource; use the utilization data and project metadata to create (1104, 1010) a capacity prediction model (130), the capacity prediction model being used to generate a pre-built configuration (140) for a first computing resource (144). The capacity prediction model is used to generate (1106, 1010) the pre-built configuration for the first computing resource; the pre-built configuration is adjusted (1108, 1024, 1032) using the selected cost and performance balance point (810) and previously existing project history data (126); and the first computing resource is constructed (1026) according to the pre-built configuration, the pre-built resource including the processor count (820a), memory amount (820b), and / or storage capacity (820c) for the first computing resource.

16. The computer storage device of claim 15, wherein the operation further comprises: The input data is received by the first computing resource; And at least based on the input data, the first computing resources are used to generate output data.

17. The computer storage device according to claim 15 or 16, wherein the operation further comprises: The capacity information includes processor counts, memory quantity, and / or storage capacity for the previously existing computing resources; The resource consumption information includes idle and throttling information for the previously existing computing resources; and the project history information includes requested changes or reported events for the previously existing computing resources.

18. The computer storage device according to any one of claims 15 to 17, wherein the operation further comprises: The first model is trained at least based on the utilization data to perform capacity sizing adjustments for the capacity prediction model; A second model is trained based at least on the project metadata to perform workload prediction against the capacity prediction model; And at least based on the project's historical data, a third model is trained to perform capacity size adjustments against the capacity prediction model.

19. The computer storage device according to any one of claims 15 to 18, wherein the operation further comprises: Present the user interface (UI); The UI is used to receive an initial build target selection, wherein generating the pre-build configuration for the first computing resource includes generating the pre-build configuration for the first computing resource based at least on the initial build target selection; the UI is used to receive the selected cost and performance balance point; and at least a portion of the pre-build configuration is displayed in the UI.

20. The computer storage device of claim 19, wherein the operation further comprises: The UI is used to receive interpretable setting selections; The UI displays at least a portion of the hierarchy of the capacity prediction model; The UI also displays information about previously existing computing resources used to generate the capacity prediction model.