Cell Boundary Fault Detection in Parallel Processing Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In parallel processing computer systems, the redundancy designed for fault tolerance complicates the detection of faulty nodes and connections along cell surfaces, making it difficult to pinpoint and service faulty components due to the large number of nodes and alternative data paths.
Innovation Solution
A method is implemented to detect nodal faults by communicating between nodes on adjacent boundary surfaces, checking for errors, and verifying latency and bandwidth specifications, using features like error detection registers and interrupt delivery to identify hardware failures and correct miscabled or misconfigured hardware.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If redundancy is added to enable fault tolerance in parallel processing systems, then system reliability is improved, but the complexity of detecting and locating faulty nodes increases
Solution Approach 1:
The patent divides the fault detection process into segmented phases: initialization phase where nodes are organized into groups and assigned test patterns, execution phase where tests are performed in parallel, and analysis phase where results are processed. This segmentation allows systematic management of complexity in detecting faulty nodes within redundant parallel processing systems.
Solution Approach 2:
The patent implements feedback mechanisms where nodes receive test patterns, execute tests, and return results to controlling nodes. The controlling nodes analyze these feedback results to identify faulty nodes. This feedback loop enables reliable detection of faulty nodes while managing complexity through structured information flow and parallel processing of test results.
2Productivity
If the number of nodes is increased to improve processing capability, then productivity is improved, but the difficulty of pinpointing faulty components increases
Solution Approach 1:
The patent applies local quality by assigning unique test patterns to specific groups of nodes and using localized test sequences that target particular regions. Each node or group of nodes receives customized test patterns based on its position and function, enabling precise identification of faulty components even in large-scale parallel systems with many nodes.
Solution Approach 2:
The patent introduces additional dimensions to fault detection by organizing nodes into multi-dimensional groups and using multi-phase test sequences. Tests are executed in parallel across different dimensions, and results are analyzed systematically to pinpoint faults. This dimensional organization transforms the complexity of locating faults in large node arrays into a manageable multi-step process.
3Ease of manufacture
If physical cabling is used to connect nodes, then ease of manufacture and assembly is improved, but susceptibility to cable damage and configuration errors increases
Solution Approach 1:
The patent performs preliminary testing and configuration verification during the initialization phase before normal operation begins. Nodes exchange identification information, verify connection integrity, and establish test patterns in advance. This preliminary action detects cable damage and configuration errors before they affect productivity, maintaining reliability while preserving the assembly ease of physical cabling.
Data Source
AI summary
A method determines a nodal fault along the boundary, or face, of a computing cell. Nodes on adjacent cell boundaries communicate with each other, and the communications are analyzed to determine if a node or connection is faulty.


