Monitoring Service Exploration for Cloud Incident Localization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing cloud computing systems struggle to efficiently and accurately identify and mitigate service incidents and outages due to vague, sparse, or noisy incident tickets, leading to delayed resolutions and manual errors.

Innovation Solution

A service incident resolution system that utilizes service models to correlate monitoring signals with outage tickets, determine relevant monitoring incident tickets, and direct issues to appropriate mitigation teams, enhancing accuracy and efficiency by filtering out noise and leveraging real-time data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If existing systems rely on monitoring service reports to identify service incidents, then incident detection coverage is provided, but identification accuracy deteriorates due to vague, sparse, or noisy incident tickets

Engineering Contradiction:
Improveincident identification accuracyVSAvoidinformation quality in incident tickets
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent introduces an intermediary system that sits between the monitoring service reports and the incident identification process. This intermediary uses service models to correlate monitoring signals with outage tickets, filtering and enriching the information to produce accurate incident identifications without relying directly on the noisy original reports

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the manual/mechanical process of reviewing incident tickets with an automated system using service models and signal correlation. The system automatically processes monitoring data, correlates signals with outages, and identifies incidents without human intervention, eliminating the limitations of manual analysis of vague tickets

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If systems process all monitoring incident tickets, then comprehensive coverage is achieved, but processing efficiency deteriorates due to the sheer volume of incidents

Engineering Contradiction:
Improveincident processing efficiencyVSAvoidnumber of incident tickets
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent extracts only the relevant information from the large volume of incident tickets by using service models to identify which monitoring signals are actually correlated with service outages. Instead of processing all tickets equally, the system extracts and focuses on the subset of tickets that contain genuine incident information, dramatically improving processing efficiency

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies partial action by selectively processing only those incident tickets that meet specific correlation criteria with monitoring signals. Rather than exhaustively analyzing every ticket in the system, the method performs targeted analysis on relevant subsets, achieving efficient incident identification without the overhead of complete processing

Inventive Principle:
Principle #16Partial or excessive action

3Speed

If manual review of incident tickets is performed, then detailed analysis is possible, but response time deteriorates leading to slow mitigation

Engineering Contradiction:
Improveincident response speedVSAvoidtime to mitigate outages
Core Design Contradiction:
SpeedVSLoss of time

Solution Approach 1:

The patent implements self-service by enabling the system to automatically identify, correlate, and flag incidents without requiring manual review. The service models autonomously process monitoring signals, correlate them with outage patterns, and generate incident identifications, allowing the system to serve itself in the incident detection process and eliminate delays associated with human intervention

Inventive Principle:
Principle #25Self-service

4Reliability

If systems ignore numerous incidents to focus on critical ones, then processing load is reduced, but reliability deteriorates as critical incidents may be missed

Engineering Contradiction:
Improvecritical incident detection reliabilityVSAvoidincident filtering complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent uses feedback mechanisms where the service models continuously learn from the correlation between monitoring signals and actual service outages. The system refines its understanding of what constitutes a critical incident based on observed patterns, improving its ability to reliably identify important incidents while automatically filtering noise, without requiring complex manual filtering rules

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12463876B2Utilizing monitoring service exploration to improve service incident mitigation and localization
Publication Date: 2025.11.04 MICROSOFT TECHNOLOGY LICENSING LLC
  • US12463876B2 patent drawing
  • US12463876B2 patent drawing
  • US12463876B2 patent drawing

AI summary

The disclosure relates to utilizing a service incident resolution system to determine and mitigate service incidents in a cloud computing system. For example, based on identifying an outage ticket (e.g., a customer-impacting incident ticket), the service incident resolution system identifies additional context of the outage by detecting a number of relevant monitoring signals. For instance, the service incident resolution system utilizes various monitoring signals and service models to determine monitoring signals that are relevant to the outage ticket by efficiently selecting relevant monitor signals and filtering out noisy signals. In this way, vaguely reported outages are supplemented with rich information that enable these outages to be resolved more quickly. Additionally, the service incident resolution system may utilize service-based models to efficiently send a report of an outage to a service or mitigation team that is well-equipped to quickly address the outage.