Machine Learning Capacity Planning for Extreme Demand Forecasts

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing capacity planning methods struggle to accurately forecast computing resource needs for extreme scenarios, leading to inefficiencies due to over- or under-allocation, which can result in increased costs or poor user experience.

Innovation Solution

A machine learning model utilizing Auto Regressive Integrated Moving Average (ARIMA) with Quantile Regression and Monte-Carlo simulations is employed to predict computing resource demands, enabling proactive and reactive resource allocation based on specified quality of service thresholds.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional capacity planning methods are used, then resource allocation is simplified, but forecasting accuracy for extreme scenarios deteriorates

Engineering Contradiction:
Improveforecasting accuracyVSAvoidmodel complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the forecasting problem by focusing on extreme percentiles (e.g., 95th, 99th) separately from average demand. This segmentation allows the model to specifically target and improve accuracy for extreme scenarios without being constrained by traditional aggregate forecasting methods, directly addressing the need for better extreme scenario prediction.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary actions by proactively allocating resources based on forecasted extreme demand percentiles before the actual demand occurs. This advance planning enables organizations to prepare adequate capacity for extreme scenarios rather than reacting after demand exceeds capacity, improving forecasting effectiveness for rare but critical events.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If capacity is over-allocated to ensure availability, then service reliability improves, but resource waste increases

Engineering Contradiction:
Improveservice availabilityVSAvoidresource waste
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The patent applies partial action by allocating resources to meet specific percentile thresholds (e.g., 95th percentile) rather than planning for absolute maximum demand. This partial allocation strategy ensures adequate service availability for extreme scenarios while avoiding the excessive resource allocation that would occur if organizations planned for every possible outlier event, thus reducing waste while maintaining reliability.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent changes the parameter from which capacity is planned - shifting from average demand or maximum historical demand to extreme percentile forecasts. This parameter change enables more precise resource allocation that matches actual extreme scenario needs, improving the balance between availability and waste reduction by targeting specific demand levels rather than using blanket allocation rules.

Inventive Principle:
Principle #35Parameter changes

3Loss of energy

If capacity is under-allocated to reduce costs, then resource efficiency improves, but service quality deteriorates during peak demand

Engineering Contradiction:
Improveresource efficiencyVSAvoidservice quality
Core Design Contradiction:
Loss of energyVSReliability

Solution Approach 1:

The patent performs preliminary action by forecasting and allocating resources for extreme demand scenarios before they occur. This advance preparation ensures that adequate capacity is available during peak demand periods while avoiding the need for continuous over-allocation, thus maintaining service quality during critical periods while improving overall resource efficiency through targeted rather than blanket allocation.

Inventive Principle:
Principle #10Preliminary action

4Measurement precision

If traditional forecasting models are used, then model simplicity is maintained, but ability to predict extreme demand percentiles deteriorates

Engineering Contradiction:
Improveextreme percentile predictionVSAvoidforecasting model
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the forecasting approach by specifically targeting extreme percentile predictions rather than attempting to model the entire demand distribution. This segmentation allows the use of specialized techniques focused on tail events, improving extreme percentile prediction accuracy without requiring a complete overhaul of the forecasting infrastructure.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by concentrating forecasting efforts and computational resources on predicting extreme percentiles rather than attempting to accurately forecast the entire demand spectrum. This focused approach improves extreme scenario prediction while avoiding the excessive complexity that would result from trying to model all demand characteristics with equal precision.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12393855B1Capacity planning using machine learning
Publication Date: 2025.08.19 AMAZON TECH INC
  • US12393855B1 patent drawing
  • US12393855B1 patent drawing
  • US12393855B1 patent drawing

AI summary

Systems, devices, and methods are provided for training and/or inferencing capacity planning using a machine learning model. A first time series may be provided as an input to a machine learning model, which may be an Auto Regressive Integrated Moving Average (ARIMA)-based forecasting model. The machine learning model may be trained solve a conditional maximum likelihood problem by performing quantile regression. The machine learning model may forecast one or more innovations using Monte-Carlo simulations. The machine learning model may generate, as an output, a value that corresponds to an amount of computing resources that is predicted, over a second time series, to be sufficient to satisfy a threshold level of availability or quality.