Monitoring Device for Fault-Tolerant System Crash Prevention
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Fault-tolerant systems using the lockstep scheme face challenges in preventing system crashes or degradation due to inconsistencies and faults in hardware components, particularly when external devices like flash memory devices suffer faults, leading to potential separation of healthy processor systems from the fault-tolerant system.
Innovation Solution
A monitoring device is introduced that reads data from a storage area in an accessory device connected to the processor system, compares it with reference data, and separates the processor system from the fault-tolerant system when discrepancies are detected, thereby preventing system crashes by quickly identifying and isolating faulty components.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data comparison is performed between processor systems to detect faults, then system reliability is improved, but false separation of healthy systems may occur due to external noise or temporary abnormalities
Solution Approach 1:
The monitoring device performs preliminary monitoring of accessory devices (storage devices, I/O devices) connected to processor systems before faults affect the processor systems themselves. By detecting abnormalities in accessory devices early through continuous data comparison and status monitoring, the system can isolate faulty accessory devices before they cause processor system failures or false lockstep losses, thereby preventing unnecessary separation of healthy processor systems and maintaining system availability while ensuring reliable fault detection
Solution Approach 2:
The monitoring device acts as an intermediary between accessory devices and processor systems. It monitors the accessory devices independently and provides fault information to the controller, which then makes separation decisions. This intermediary role allows the system to distinguish between faults in accessory devices and faults in processor systems, preventing false separation of healthy processor systems due to external noise or temporary abnormalities in accessory devices
2Reliability
If accessory devices are monitored to detect faults early, then system crash prevention is improved, but device complexity increases due to additional monitoring components
Solution Approach 1:
The monitoring device is designed to monitor multiple types of accessory devices (storage devices, I/O devices) connected to multiple processor systems using a unified monitoring mechanism. This multi-functional approach allows the system to detect faults in various accessory devices through a single monitoring component, preventing system crashes without requiring separate monitoring systems for each device type, thereby limiting the increase in device complexity
Solution Approach 2:
The monitoring device utilizes data already present in the accessory devices (stored data, I/O data) and the existing controller infrastructure to perform monitoring functions. By leveraging existing system resources and data, the monitoring device can detect faults without requiring additional complex hardware or external monitoring systems, thus achieving improved crash prevention while minimizing increases in overall system complexity
Data Source
AI summary
A monitoring device is mounted in each of a plurality of operational systems constituting a fault-tolerant system. The plurality of operational systems have an identical configuration including a processor system. The monitoring device includes a processor. The processor executes instruction to read data from a predetermined storage area in a memory of an accessory device to be monitored, connected to the processor system. The processor further executes instruction to compare the read data with reference data held in advance. The processor further executes instruction to separate the processor system connected to the accessory device to be monitored from the fault-tolerant system when the read data is different from the reference data.


