Monitoring Service Exploration for Cloud Incident Localization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing cloud computing systems struggle to efficiently and accurately identify and mitigate service incidents and outages due to vague, sparse, or noisy incident tickets, leading to delayed resolutions and manual errors.
Innovation Solution
A service incident resolution system that utilizes service models to correlate monitoring signals with outage tickets, determine relevant monitoring incident tickets, and direct issues to appropriate mitigation teams, enhancing accuracy and efficiency by filtering out noise and leveraging real-time data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If existing systems rely on monitoring service reports to identify service incidents, then incident detection coverage is provided, but identification accuracy deteriorates due to vague, sparse, or noisy incident tickets
Solution Approach 1:
The patent introduces an intermediary system that sits between the monitoring service reports and the incident identification process. This intermediary uses service models to correlate monitoring signals with outage tickets, filtering and enriching the information to produce accurate incident identifications without relying directly on the noisy original reports
Solution Approach 2:
The patent replaces the manual/mechanical process of reviewing incident tickets with an automated system using service models and signal correlation. The system automatically processes monitoring data, correlates signals with outages, and identifies incidents without human intervention, eliminating the limitations of manual analysis of vague tickets
2Productivity
If systems process all monitoring incident tickets, then comprehensive coverage is achieved, but processing efficiency deteriorates due to the sheer volume of incidents
Solution Approach 1:
The patent extracts only the relevant information from the large volume of incident tickets by using service models to identify which monitoring signals are actually correlated with service outages. Instead of processing all tickets equally, the system extracts and focuses on the subset of tickets that contain genuine incident information, dramatically improving processing efficiency
Solution Approach 2:
The patent applies partial action by selectively processing only those incident tickets that meet specific correlation criteria with monitoring signals. Rather than exhaustively analyzing every ticket in the system, the method performs targeted analysis on relevant subsets, achieving efficient incident identification without the overhead of complete processing
3Speed
If manual review of incident tickets is performed, then detailed analysis is possible, but response time deteriorates leading to slow mitigation
Solution Approach 1:
The patent implements self-service by enabling the system to automatically identify, correlate, and flag incidents without requiring manual review. The service models autonomously process monitoring signals, correlate them with outage patterns, and generate incident identifications, allowing the system to serve itself in the incident detection process and eliminate delays associated with human intervention
4Reliability
If systems ignore numerous incidents to focus on critical ones, then processing load is reduced, but reliability deteriorates as critical incidents may be missed
Solution Approach 1:
The patent uses feedback mechanisms where the service models continuously learn from the correlation between monitoring signals and actual service outages. The system refines its understanding of what constitutes a critical incident based on observed patterns, improving its ability to reliably identify important incidents while automatically filtering noise, without requiring complex manual filtering rules
Data Source
AI summary
The disclosure relates to utilizing a service incident resolution system to determine and mitigate service incidents in a cloud computing system. For example, based on identifying an outage ticket (e.g., a customer-impacting incident ticket), the service incident resolution system identifies additional context of the outage by detecting a number of relevant monitoring signals. For instance, the service incident resolution system utilizes various monitoring signals and service models to determine monitoring signals that are relevant to the outage ticket by efficiently selecting relevant monitor signals and filtering out noisy signals. In this way, vaguely reported outages are supplemented with rich information that enable these outages to be resolved more quickly. Additionally, the service incident resolution system may utilize service-based models to efficiently send a report of an outage to a service or mitigation team that is well-equipped to quickly address the outage.


