Multi-Tenant Service Scaling Units for Accurate Capacity Planning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Accurately determining the number of hosts to add for a given workload increase in large-scale service-oriented architecture (SOA) applications with multi-tenant resources is challenging due to complex multi-tenancy configurations, leading to potential resource wastage or poor user experience.

Innovation Solution

A scalability management service (SMS) identifies resource scaling units (RSUs) based on multi-tenancy information, defining sets of uniform-configuration host groups and load balancers to predict workload changes and adjust resources accordingly.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of energy

If multi-tenant hosts are used to process workloads of multiple constituent services, then resource utilization is improved, but capacity planning and scaling decisions become more complex and less accurate

Engineering Contradiction:
Improveresource utilizationVSAvoidcapacity planning complexity
Core Design Contradiction:
Loss of energyVSDevice complexity

Solution Approach 1:

The patent segments the complex multi-tenant environment into individual constituent service workloads, analyzing each service's resource consumption patterns separately. This allows for granular capacity planning by breaking down the aggregate multi-tenant workload into manageable service-level units that can be independently scaled based on their specific requirements.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements feedback mechanisms that monitor actual workload patterns and resource usage of constituent services running on multi-tenant hosts. This feedback data is used to continuously refine capacity planning models and scaling decisions, improving accuracy over time by learning from real-world performance data and adjusting resource allocation accordingly.

Inventive Principle:
Principle #23Feedback

2Adaptability or versatility

If multi-tenant load balancers distribute service requests across numerous hosts, then system flexibility is improved, but accurate capacity planning becomes non-trivial

Engineering Contradiction:
Improvesystem flexibilityVSAvoidcapacity planning accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary layer between the load balancer and the constituent services that tracks and attributes service requests to their source services. This intermediary mechanism enables precise measurement of workload distribution patterns while maintaining the flexibility of multi-tenant load balancing, providing the data needed for accurate capacity planning without restricting system adaptability.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If hosts are added to handle anticipated workload increases, then service capacity is improved, but resource wastage may occur due to inaccurate scaling decisions

Engineering Contradiction:
Improveservice capacityVSAvoidresource wastage
Core Design Contradiction:
ProductivityVSLoss of substance

Solution Approach 1:

The patent performs preliminary analysis of workload patterns and growth trends before making scaling decisions. By forecasting future workload requirements based on historical data and constituent service performance metrics, the system can provision resources in advance with greater accuracy, ensuring capacity is available when needed while minimizing premature or excessive resource allocation that would lead to wastage.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12579004B1Capacity planning and scaling of service-oriented applications implemented using multi-tenant resources
Publication Date: 2026.03.17 AMAZON TECH INC
  • US12579004B1 patent drawing
  • US12579004B1 patent drawing
  • US12579004B1 patent drawing

AI summary

A request to perform scalability analysis of a constituent service of an application at a first deployment environment is received. The constituent service is deployed to a plurality of host groups of the deployment environment. A data set which indicates a collection of front end request routers configured to distribute service requests among hosts of a collection of deployment environments including the first deployment environment is obtained. Analysis of the data set is used to identify a resource scaling unit which indicates a set of host groups, including a particular host group of the plurality of host groups, such that individual hosts within the set of host groups have the same estimated workload level as other hosts within the set of host groups. A representation of the resource scaling unit is stored.