Hitless Reset for Network Device System Control Devices

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Network devices with system control devices (SCDs) are susceptible to single event upsets (SEUs), which can cause unexpected state changes and disrupt network operations, requiring a method to perform updates and resets without impacting network traffic.

Innovation Solution

A hitless repair method where the SCD agent determines if the SCD and network device support hitless reset, performs a pre-hitless repair action set, and executes a hitless reset, ensuring agents are gracefully shut down and restarted to maintain continuous network traffic processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a traditional reset method is used to update or repair the SCD, then the SCD can be reset or updated, but network traffic processing is disrupted

Engineering Contradiction:
ImproveSCD update reliabilityVSAvoidnetwork traffic processing continuity
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system separates the SCD reset function from the network traffic processing function by introducing an intermediary component (such as a switch fabric or bypass path) that allows traffic to continue flowing while the SCD is reset independently. This segmentation enables the control plane (SCD) and data plane (traffic processing) to operate independently, resolving the contradiction between updating the SCD and maintaining network traffic continuity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

An intermediary component is introduced between the SCD and the network traffic processing path. This intermediary acts as a mediator that can redirect traffic around the SCD during reset operations or provide alternative pathways for traffic to flow while the SCD is being updated or repaired, thus maintaining productivity while allowing reliability improvements.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If the SCD is reset to repair SEU damage, then the SEU impact is mitigated, but network operations are interrupted

Engineering Contradiction:
ImproveSCD operation reliabilityVSAvoidnetwork operation interruption time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-configuring alternative traffic paths or standby SCD components before an SEU occurs. When an SEU is detected, the pre-configured alternative paths are immediately activated, allowing the affected SCD to be reset without causing network operation interruptions. This preliminary preparation eliminates the need for time-consuming reconfiguration during actual reset operations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system maintains continuous network traffic processing by implementing hot-standby SCD components or bypass paths that can immediately take over when the primary SCD needs resetting due to SEU. The useful action of network traffic processing continues uninterrupted while the SCD is reset, achieving both reliability improvement and minimal operational disruption.

Inventive Principle:
Principle #20Continuity of useful action

3Productivity

If a hitless repair method is implemented, then network traffic continuity is maintained, but system complexity increases

Engineering Contradiction:
Improvenetwork traffic processing continuityVSAvoidhitless repair system complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system implements self-service mechanisms where the SCD or associated control logic automatically detects SEUs, initiates reset procedures, and switches to alternative paths without requiring external intervention or complex manual configuration. This automation reduces the operational complexity of managing hitless repair capabilities while maintaining network traffic continuity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system uses discardable temporary configurations or state information that can be quickly discarded and recovered during SCD reset operations. By using temporary, easily replaceable configuration data structures and state representations, the system enables rapid reset and recovery without requiring complex persistent state management, thus reducing overall system complexity while maintaining continuity.

Inventive Principle:
Principle #34Discarding and recovering

Data Source

PatentUS10846179B2Hitless repair for network device components
Publication Date: 2020.11.24 ARISTA NETWORKS INC
  • US10846179B2 patent drawing
  • US10846179B2 patent drawing
  • US10846179B2 patent drawing

AI summary

Methods, systems, and computer readable mediums for hitless repair. Hitless repair may include making a first determination, by a system control device (SCD) agent of a network device, that a SCD of the network device has experienced an error and/or is to be updated; making a second determination, by the SCD agent, that the SCD and the network device support the hitless repair; performing, by the SCD agent, a pre-hitless repair action set; and performing, by the SCD agent and after completing the pre-hitless repair action set, a post-hitless repair action set, including a hitless reset of the SCD.