I/O Management Node Failover via Switch Reconfiguration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Managing input/output (I/O) devices in large, distributed networks is challenging due to the difficulty in ensuring high availability and real-time management across various geographic locations, particularly in industries like manufacturing, energy, and transportation, where consistent data measurement and system health monitoring are critical.
Innovation Solution
Implementing a method with multiple I/O management nodes, including active and standby nodes, that communicate through switches to ensure high availability by quickly identifying failures through frequent health check communications and configuring data paths to switch over to standby nodes in case of primary node failures.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If multiple I/O management nodes are deployed to ensure high availability, then system reliability is improved, but device complexity increases
Solution Approach 1:
The system segments I/O management functionality across multiple independent nodes (primary and standby), where each node handles a portion of the management workload. This segmentation enables failover capability while distributing complexity across separate units rather than concentrating it in a single complex system.
Solution Approach 2:
The patent changes the operational state parameter of the standby node from inactive to actively monitoring through health check communications. By transitioning the standby node to a state where it actively pings the primary node and can rapidly take over, the system achieves high availability without requiring complex active-active load balancing mechanisms.
2Loss of time
If health check communications occur at high frequency to quickly identify failures, then failure detection speed is improved, but use of energy increases
Solution Approach 1:
The health check mechanism uses periodic ping communications at defined intervals between standby and primary nodes. This periodic action enables timely failure detection while allowing energy consumption to be controlled through adjustable interval timing, rather than requiring continuous monitoring.
Solution Approach 2:
The system implements feedback through health check responses where the primary node confirms its operational status to the standby node. This feedback mechanism enables rapid failure detection through simple presence/absence of responses rather than complex continuous monitoring, reducing energy requirements while maintaining fast detection capability.
Data Source
AI summary
Disclosed herein are enhancements for operating an input/output (I/O) management cluster with end I/O devices. In one implementation, a method of operating an I/O cluster includes, in a first I/O management node of the I/O management cluster, executing a first application to manage data for an I/O device communicatively coupled via at least one switch to the first I/O management node. The method further provides identifying a failure in the first I/O management node related to processing the data for the I/O device and, in response to the failure, configuring the at least one switch to communicate the data for the I/O device with a second I/O management node of the I/O management cluster. The method also includes, in the second I/O management node and after configuring the at least one switch, executing a second application to manage the data for the I/O device.


