Automated Datacenter Network Failure Mitigation Pipeline

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Diagnosing and repairing network failures in datacenter networks is time-consuming due to the variability of failure sources, including hardware components, software bugs, and configuration errors, often requiring manual intervention and third-party assistance, leading to prolonged recovery times.

Innovation Solution

An automated system that monitors network state data to detect failures, determines a set of suspected components, and executes mitigation actions through a planner and plan executor, allowing for automated trial-and-error approaches to alleviate network failures without precise diagnosis or repair, utilizing a pipeline comprising a failure detector, aggregator, planner, impact estimator, and plan executor.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual diagnosis and repair procedures are used, then operators can identify and fix root causes, but failure recovery time is prolonged

Engineering Contradiction:
Improvefailure recoveryVSAvoidfailure recovery time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system enables self-service failure mitigation by automatically detecting failures, determining affected components, executing mitigation actions, and verifying resolution without requiring manual operator intervention. The automated pipeline performs the entire failure recovery process independently, transforming a manual operation into an autonomous system that resolves issues on its own.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary actions by proactively monitoring network state data continuously and immediately detecting failures as they occur. The automated pipeline prepares and executes mitigation actions before manual operators can intervene, reducing the window of vulnerability and minimizing recovery time through pre-positioned detection and response capabilities.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If automated tools are used to localize failures, then diagnosis speed improves, but manual intervention is still required for root cause diagnosis and repair

Engineering Contradiction:
Improvefailure diagnosis speedVSAvoidfailure repair automation
Core Design Contradiction:
ProductivityVSExtent of automation

Solution Approach 1:

The system extends automation from merely localizing failures to fully autonomously executing repair actions. The automated pipeline not only identifies affected components through monitoring and analysis but also independently performs mitigation actions such as restarting services, reconfiguring network elements, and verifying resolution, eliminating the need for manual repair intervention.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The automated pipeline acts as an intermediary between failure detection and resolution by orchestrating the entire recovery process. It mediates between the raw failure data and the necessary repair actions, translating detected issues into structured mitigation plans and executing them automatically, thereby bridging the gap between monitoring and repair without human intervention.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If third-party vendor assistance is sought, then specialized expertise is available, but failure recovery time increases

Engineering Contradiction:
Improvefailure resolution qualityVSAvoidfailure recovery time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system eliminates dependency on third-party vendors by embedding specialized failure resolution capabilities directly into the automated pipeline. It performs complex diagnostic analysis, root cause identification, and mitigation action execution autonomously, replacing the need to outsource failure resolution to external vendor support teams.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The automated pipeline provides universal failure resolution capabilities that handle diverse failure types and complex network configurations without requiring external vendor expertise. The system integrates multiple functions including monitoring, diagnosis, mitigation strategy generation, and action execution into a single autonomous platform that can resolve failures independently.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Measurement precision

If comprehensive failure detection and monitoring is implemented, then failure detection accuracy improves, but system complexity increases

Engineering Contradiction:
Improvefailure detection accuracyVSAvoidmonitoring system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The monitoring system is segmented into modular components that independently monitor specific network elements and failure conditions. Each segment processes and analyzes data for its designated scope, then feeds results to the automated pipeline. This segmentation enables comprehensive monitoring of complex networks while maintaining manageable system architecture through functional decomposition.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The automated pipeline serves as a universal platform that handles multiple functions including data collection, failure detection, root cause analysis, mitigation strategy generation, and action execution. By consolidating these functions into a single multi-functional system, the patent reduces overall complexity compared to having separate specialized systems for each function.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10075327B2Automated datacenter network failure mitigation
Publication Date: 2018.09.11 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10075327B2 patent drawing
  • US10075327B2 patent drawing
  • US10075327B2 patent drawing

AI summary

The subject disclosure is directed towards a technology that automatically mitigates datacenter failures, instead of relying on human intervention to diagnose and repair the network. Via a mitigation pipeline, when a network failure is detected, a candidate set of components that are likely to be the cause of the failure is identified, with mitigation actions iteratively targeting each component to attempt to alleviate the problem. The impact to the network is estimated to ensure that the redundancy present in the network will be able to handle the mitigation action without adverse disruption to the network.