ML-Based Incident Routing for Cloud Service Misclassification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for incident routing in cloud environments are inefficient and prone to misrouting, leading to prolonged service-level effects and resource wastage, as they rely on human prediction and are not optimized for quick or efficient resolution.

Innovation Solution

The implementation of team-specific scouts using machine learning models to evaluate incident descriptions and predict the capability of associated teams to resolve incidents, with a scout master aggregating predictions for accurate routing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If human prediction is used for incident routing, then incidents can be routed to teams, but misrouting occurs and resolution efficiency decreases

Engineering Contradiction:
Improverouting accuracyVSAvoidresolution efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent replaces the mechanical system of human prediction and manual incident routing with an automated machine learning-based classification system. The ML model analyzes incident data and automatically determines the appropriate routing, eliminating human error and subjectivity while improving both routing accuracy and resolution efficiency.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The incident routing system performs self-service by automatically classifying and routing incidents without requiring human intervention. The machine learning model independently evaluates incident characteristics and determines optimal routing decisions, freeing human operators from manual classification tasks while maintaining high accuracy.

Inventive Principle:
Principle #25Self-service

2Speed

If incidents are routed quickly without proper classification, then resolution speed may improve, but misrouting increases and resources are wasted

Engineering Contradiction:
Improveincident routing speedVSAvoidresource wastage
Core Design Contradiction:
SpeedVSLoss of energy

Solution Approach 1:

The system performs preliminary classification and routing decisions using machine learning before incidents are assigned to teams. By pre-analyzing incident characteristics and predicting the most appropriate routing, the system ensures both speed and accuracy from the outset, preventing resource wastage while maintaining rapid response times.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The incident classification system incorporates feedback mechanisms where routing outcomes and resolution results are fed back into the machine learning model. This continuous feedback loop improves the model's accuracy over time, ensuring that quick routing decisions remain both fast and accurate, thereby preventing resource wastage.

Inventive Principle:
Principle #23Feedback

3Loss of time

If manual incident classification is used, then routing decisions can be made, but time consumption increases and service-level objectives are not met

Engineering Contradiction:
Improveclassification timeVSAvoidservice-level objective compliance
Core Design Contradiction:
Loss of timeVSReliability

Solution Approach 1:

The patent replaces manual incident classification with an automated machine learning-based classification system. This substitution eliminates the time-consuming nature of human classification while maintaining or improving accuracy, thereby reducing classification time and ensuring compliance with service-level objectives.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The machine learning classification system operates continuously and automatically, processing incidents as they occur without interruption. This continuous operation eliminates the delays inherent in manual classification processes, ensuring that incidents are routed quickly and consistently, thereby meeting service-level objectives.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentEP4091110B1Systems and methods for distributed incident classification and routing
Publication Date: 2023.11.15 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP4091110B1 patent drawingFigure 1A
  • EP4091110B1 patent drawingFigure 1B
  • EP4091110B1 patent drawingFigure 2

AI summary

Aspects of the present disclosure relate to incident routing in a cloud environment. In an example, cloud provider teams utilize a scout framework to build a team-specific scout based on that team's expertise. In examples, an incident is detected and a description is sent to each team-specific scout. Each team-specific scout uses the incident description and the scout specifications provided by the team to identify, access, and process monitoring data from cloud components relevant to the incident. Each team-specific scout utilizes one or more machine learning models to evaluate the monitoring data and generate an incident-classification prediction about whether the team is responsible for resolving the incident. In examples, a scout master receives predictions from each of the team-specific scouts and compares the predictions to determine to which team an incident should be routed.