Site Reliability Engineering Service Level Indicators

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Modern software architectures, often hosted on third-party infrastructure, face challenges in measuring and ensuring user experience reliability, leading to difficulties in defining and maintaining service level objectives and agreements.

Innovation Solution

A method and system for site reliability engineering that tracks user experience through service level indicators, calculates their values, and generates alerts when thresholds are breached, allowing for proactive management of service performance and reliability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If software services are hosted on third-party infrastructure off premises, then service delivery flexibility is improved, but measurement and management of user experience reliability becomes more difficult

Engineering Contradiction:
Improveservice delivery flexibilityVSAvoiduser experience reliability measurement
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent introduces service level indicators (SLIs) and monitoring systems as intermediaries between the third-party infrastructure and the business users. These SLIs act as measurable proxies that bridge the gap between off-premises service delivery and on-premises reliability measurement, enabling indirect but effective monitoring of user experience through standardized metrics.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements continuous feedback loops by tracking SLI values against service level objectives (SLOs) and generating alerts when thresholds are breached. This feedback mechanism enables real-time monitoring and management of user experience reliability despite the services being hosted externally, allowing proactive response to performance degradation.

Inventive Principle:
Principle #23Feedback

2Reliability

If service level indicators are continuously monitored and tracked, then user experience reliability is improved, but system complexity and resource consumption increase

Engineering Contradiction:
Improveuser experience reliabilityVSAvoidmonitoring system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent transforms the complex problem of monitoring entire user experiences into monitoring specific standardized parameters (SLIs) such as latency, throughput, and availability. By changing the monitoring approach from holistic to parameter-based, the system achieves reliable user experience measurement through a manageable set of key metrics rather than comprehensive system observation.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The monitoring system is segmented into distinct components: SLI definition, data collection, calculation, threshold comparison, and alerting. This segmentation allows each component to be independently managed and optimized, reducing overall system complexity while maintaining comprehensive monitoring capability through modular architecture.

Inventive Principle:
Principle #1Segmentation

3Productivity

If multiple service level indicators are calculated and monitored, then service performance management is improved, but computational overhead and processing time increase

Engineering Contradiction:
Improveservice performance managementVSAvoidprocessing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system calculates only the necessary SLIs required for service level management rather than computing all possible performance metrics. By applying partial action - focusing on critical indicators like latency, throughput, and availability - the system achieves effective performance management with reduced computational overhead compared to comprehensive metric analysis.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12160351B2Systems and methods for site reliability engineering
Publication Date: 2024.12.03 JPMORGAN CHASE BANK NA
  • US12160351B2 patent drawing
  • US12160351B2 patent drawing
  • US12160351B2 patent drawing

AI summary

In some aspects, the techniques described herein relate to a method including: defining a user experience based on one or more software services provided by a platform; defining a service level indicator based on service metrics data; tracking an execution of the user experience, wherein the tracking includes recording of metadata that is output by each software service, wherein the metadata is defined as a parameter of the service level indicator; calculating a value of the services level indicator; and determining whether the value of the service level indicator is lower than a threshold value of a service level objective associated with the service level indicator.