Non-Preemptive Logical Router Failure Handling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional failure handling in logical routers often results in disruption and overheads due to preemptive modes, which can adversely impact the availability and performance of stateful services.

Innovation Solution

Implementing a non-preemptive mode in logical router failure handling, where the standby routing component continues to operate as active after the recovery of the primary, avoiding the need for switching and reducing disruption, overheads, and latency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If preemptive failure handling mode is used in logical routers, then the primary routing component can take over after recovery, but this causes traffic loss and service disruption due to mandatory switching

Engineering Contradiction:
Improveservice availabilityVSAvoidtraffic loss
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent inverts the traditional preemptive failure handling approach by implementing a non-preemptive mode where the standby routing component does not automatically take over upon primary component recovery. Instead of forcing a switchover, the system allows the standby component to remain active, thereby avoiding the traffic loss and disruption associated with mandatory switching while maintaining service availability

Inventive Principle:
Principle #13The other way round (Inversion)

2Reliability

If preemptive failure handling mode is used in logical routers, then failure recovery is ensured through automatic switching, but this increases system overhead and latency due to frequent state transitions

Engineering Contradiction:
Improvefailure recoveryVSAvoidswitching overhead
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-configuring the standby routing component with all necessary routing information and state data before failure occurs. This allows the standby component to immediately assume the active role without requiring time-consuming state synchronization or configuration transfers during the switching process, thereby reducing switching overhead and latency while ensuring reliable failure recovery

Inventive Principle:
Principle #10Preliminary action

3Object-affected harmful factors

If non-preemptive failure handling mode is used in logical routers, then traffic loss is reduced by avoiding unnecessary switching, but the primary routing component may not resume service after recovery

Engineering Contradiction:
Improvetraffic lossVSAvoidservice resumption
Core Design Contradiction:
Object-affected harmful factorsVSReliability

Solution Approach 1:

The patent implements feedback mechanisms where the standby routing component continuously monitors the health and operational status of the primary routing component. When the primary component recovers, the standby component receives feedback about its current state and can make an informed decision about whether to maintain its active role or transition the primary component back to service, thereby balancing traffic loss reduction with service resumption capability

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10447581B2Failure handling at logical routers according to a non-preemptive mode
Publication Date: 2019.10.15 VMWARE INC
  • US10447581B2 patent drawing
  • US10447581B2 patent drawing
  • US10447581B2 patent drawing

AI summary

Example methods are provided to handle failure at one or more logical routers according to a non-preemptive mode. The method may include in response to detecting, by a first routing component operating in a standby state, a failure associated with a second routing component operating in an active state, generating a control message that includes a non-preemptive code to instruct the second routing component not to operate in the active state after a recovery from the failure, sending the control message to the second routing component, and performing a state transition from the standby state to the active state. The method may also include in response to detecting, by the first routing component operating in the active state, network traffic during the failure or after the recovery of the second routing component, forwarding the network traffic from the first network to the second network, or from the second network to the first network.