Redundant Node Controllers for Multiprocessor Failure Recovery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In multiprocessor systems, system stop times due to node controller failures are prolonged due to the need for extensive settings changes in address decoders and routing tables, which often fail to meet constraints for continuous operation without rebooting.

Innovation Solution

A multiprocessor system design with redundant node controllers, each equipped with unique identifiers and request control sections, registers, and routing tables, allowing for efficient rerouting of requests during failures without requiring extensive system changes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If redundant node controllers are provided and extensive settings changes are made in address decoders and routing tables to handle failures, then system reliability is improved, but system stop time increases and productivity deteriorates

Engineering Contradiction:
Improvesystem reliabilityVSAvoidsystem stop time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-configuring multiple node controllers (first and second node controllers) with identical functional capabilities before any failure occurs. Each node controller is pre-equipped with address decoder settings and routing table information for all possible communication scenarios. When a failure is detected, the system can immediately switch to the pre-configured redundant node controller without requiring time-consuming settings changes, thus resolving the contradiction between maintaining high reliability and minimizing system stop time.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If redundant node controllers are provided and extensive settings changes are made in address decoders and routing tables to handle failures, then system reliability is improved, but device complexity increases

Engineering Contradiction:
Improvesystem reliabilityVSAvoiddevice complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies merging by combining the functional capabilities of multiple node controllers into a unified redundant architecture. The first and second node controllers are designed with identical functional blocks (address decoders, routing tables, communication interfaces), allowing them to be treated as interchangeable units. This standardization reduces the overall system complexity compared to having heterogeneous failure handling mechanisms, while still providing robust redundancy for high reliability.

Inventive Principle:
Principle #5Merging (Combining)

3Reliability

If extensive settings changes are made in address decoders and routing tables to handle node controller failures, then system reliability is improved, but ease of operation deteriorates

Engineering Contradiction:
Improvesystem reliabilityVSAvoidease of operation
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent applies self-service by enabling the node controller system to automatically detect failures and perform settings changes without requiring manual intervention. The system includes failure detection mechanisms that automatically trigger the switching process between redundant node controllers. The address decoders and routing tables are automatically reconfigured by the system itself when a failure occurs, eliminating the need for operators to manually adjust settings, thus improving ease of operation while maintaining high reliability.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS8051325B2Multiprocessor system and failure recovering system
Publication Date: 2011.11.01 NEC PLATFROMS LTD
  • US8051325B2 patent drawing
  • US8051325B2 patent drawing
  • US8051325B2 patent drawing

AI summary

A multiprocessor system includes a plurality of nodes, each of which includes a plurality of processors, a plurality of memories, and first and second node controllers. Unique identifiers are assigned to all the components. Each of the first and second node controllers includes: each of first and second request control sections configured to determine the identifier of a transmission destination of a request based on a memory address of an access destination of the request; each of first and second registers configured to hold in the first request control section, the identifier of the transmission destination of the request; a first routing table configured to specify one of the first request control section and the second request control section as an output destination of the request based on the identifier held by the first register, the identifier held by the second register, the identifier of the transmission destination of the request, when receiving the request, and a second routing table configured to specify a signal line for the identifier of the transmission destination of the request based on the identifier of the transmission destination which is determined by the first request control section or the second request control section, to transmit the request.