Virtual Instance Network Failure Detection via Segmented Monitoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing datacenter virtualization systems lack a reliable and fast mechanism for detecting failures in host and VM-level networking components, leading to network outages and inadequate remediation workflows, especially in high-availability clusters.

Innovation Solution

The system proactively monitors multiple networking routes by periodic pinging or testing, using fault monitors to detect failures and initiate remediation actions such as live migration or fault tolerance, ensuring uninterrupted availability of virtual computing instances.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If existing systems monitor only host level failures and storage component failures, then the monitoring scope is limited and simple, but network component failures cannot be detected reliably and quickly

Engineering Contradiction:
Improvefailure detection reliabilityVSAvoidmonitoring system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the monitoring scope into three distinct levels: host level monitoring (via host monitors), virtual machine level monitoring (via VM monitors), and network route level monitoring (via route monitors). Each monitor type specializes in detecting specific failure modes at its respective level, enabling comprehensive and reliable failure detection across the entire virtualized infrastructure without requiring a single complex monitoring system.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces heartbeat monitors as intermediary components that facilitate communication and failure detection between VCIs and external systems. These heartbeat monitors send periodic heartbeat signals through network routes to detect connectivity failures, acting as intermediaries that can identify network-level issues without requiring direct access to the underlying hardware or virtualization layer.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Speed

If existing systems lack fast detection mechanism for network component failures, then the response time is slow, but implementing comprehensive monitoring increases system complexity

Engineering Contradiction:
Improvefailure detection speedVSAvoidmonitoring mechanism complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The patent implements preliminary action by having multiple monitors (host monitors, VM monitors, route monitors) continuously and proactively monitor their respective components before failures occur. These monitors continuously check the status of network routes, hosts, and VMs, enabling immediate detection of failures as soon as they occur, rather than waiting for symptoms to manifest or for periodic comprehensive checks.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent establishes feedback loops where monitors continuously report status information about network routes, hosts, and VMs to a central management system. When a monitor detects a failure condition (such as a network route becoming unavailable), it immediately sends a failure notification, creating a real-time feedback mechanism that enables rapid response to failures without requiring complex manual intervention or analysis.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10367711B2Protecting virtual computing instances from network failures
Publication Date: 2019.07.30 VMWARE INC
  • US10367711B2 patent drawing
  • US10367711B2 patent drawing
  • US10367711B2 patent drawing

AI summary

The subject matter described herein provides virtual computing instance (VCI) component protection against networking failures in a datacenter cluster. Networking routes at the host level, VCI level, and application level are monitored for connectivity. Failures are communicated to a primary host or to a datacenter virtualization infrastructure that initiates policy-based remediation, such as moving affected VCIs to another host in the cluster that has all the necessary networking routes functional.