Predicting Microservices for Serverless Cold-Start Latency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Serverless microservices experience significant latency due to the 'cold-start' problem, where each microservice must scale from zero, leading to cumulative response latencies in microservice-based applications.

Innovation Solution

A predictive model is built using tracing data from historical requests to proactively scale only the microservices required for an incoming request, reducing the cold-start time by dynamically predicting and scaling necessary microservices before they are invoked.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Use of energy by moving object

If serverless microservices scale from zero for each incoming request, then resource efficiency is improved, but response latency increases due to cold-start problem

Engineering Contradiction:
Improveresource efficiencyVSAvoidresponse latency
Core Design Contradiction:
Use of energy by moving objectVSLoss of time

Solution Approach 1:

The system performs preliminary scaling actions by proactively launching microservice containers before they are actually needed to handle requests. The prediction module analyzes historical tracing data and request patterns to anticipate which microservices will be required, initiating their startup process in advance. This preliminary action reduces the cold-start latency when requests arrive, while maintaining resource efficiency by only scaling services that are predicted to be needed based on the predictive model.

Inventive Principle:
Principle #10Preliminary action

2Loss of time

If all microservices are scaled up proactively, then response latency is reduced, but resource usage and costs increase

Engineering Contradiction:
Improvecold-start latencyVSAvoidresource usage
Core Design Contradiction:
Loss of timeVSUse of energy by moving object

Solution Approach 1:

The system applies local quality by differentiating the scaling behavior of individual microservices based on specific request characteristics and historical patterns. Rather than uniformly scaling all microservices, the prediction module analyzes the incoming request attributes and selectively identifies which specific microservices are likely to be required. This targeted approach ensures that only the necessary microservices are scaled up proactively, optimizing the balance between reducing cold-start latency and minimizing resource consumption.

Inventive Principle:
Principle #3Local quality

3Use of energy by moving object

If microservices are scaled on-demand without prediction, then resource efficiency is maintained, but cumulative response latency increases

Engineering Contradiction:
Improveresource efficiencyVSAvoidresponse time
Core Design Contradiction:
Use of energy by moving objectVSProductivity

Solution Approach 1:

The system implements feedback by continuously monitoring and analyzing historical tracing data from processed requests to build and refine the predictive model. The model learns from past request patterns, microservice invocation sequences, and performance metrics to improve its predictions over time. This feedback mechanism enables the system to make increasingly accurate predictions about which microservices will be needed for incoming requests, allowing for more effective proactive scaling decisions that improve response time while maintaining resource efficiency.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11683391B2Predicting microservices required for incoming requests
Publication Date: 2023.06.20 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11683391B2 patent drawing
  • US11683391B2 patent drawing
  • US11683391B2 patent drawing

AI summary

A method, system, and computer program product for predicting microservices required for incoming requests for reducing the start latency of serverless microservices. The method may include obtaining tracing data of microservices of an application for historical requests processed by the application. The method may also include grouping the tracing data based on common request attributes. The method may also include aggregating each group into rules relating the common request attributes to lists of microservices. The method may also include building a predictive model formed of the rules for processing incoming requests to obtain a list of predicted microservices required for the incoming request based on attributes of the incoming request.