Reconfigurable Processor Self-Healing via Redundant Dataflow Configurations
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems with reconfigurable processors face challenges in efficiently managing and healing defective components in dataflow graphs, which affects the overall performance of these systems.
Innovation Solution
An intelligent redundancy management framework (IRMF) is introduced to self-heal configurable data units in a reconfigurable data processor by identifying and replacing defective components with healthy alternatives, either statically or dynamically, to ensure optimal system performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If defective components are not managed, then system complexity is reduced, but system reliability deteriorates
Solution Approach 1:
The system performs preliminary actions by pre-identifying defective components through built-in self-test (BIST) circuits during manufacturing or initial operation. Configuration files are prepared in advance with alternative routing paths and healthy component mappings, so when a component fails, the system can immediately switch to pre-planned alternatives without complex real-time decision-making, thus improving reliability while managing complexity.
Solution Approach 2:
The reconfigurable data processor implements self-service through automated defect detection and healing mechanisms. The system uses onboard test circuits to automatically identify defective components and configures alternative paths without external intervention. This self-healing capability improves reliability while avoiding the complexity of manual defect management, as the system manages its own redundancy automatically.
2Adaptability or versatility
If component healing is performed dynamically, then system adaptability is improved, but processing time increases
Solution Approach 1:
The system performs preliminary configuration of alternative paths and healthy component mappings before failures occur. Configuration files contain pre-computed alternative routing information, allowing the system to quickly switch to backup paths when components fail, rather than computing new configurations in real-time. This reduces healing time while maintaining adaptability through pre-planned alternatives.
Solution Approach 2:
The system implements dynamic reconfiguration capabilities that allow it to adapt to component failures in real-time. The reconfigurable architecture can dynamically switch between different operational modes and alternative component paths based on detected failures. This dynamic adaptation improves versatility while the optimization of the healing process minimizes time loss through efficient configuration switching.
3Speed
If static configuration is used, then processing speed is maintained, but system reliability decreases
Solution Approach 1:
The system performs preliminary identification of defective components and prepares alternative configuration paths in advance, but maintains the primary static configuration for normal operation to preserve processing speed. When failures are detected, the pre-prepared alternative configurations are activated without requiring complex real-time reconfiguration, thus maintaining speed while improving reliability through failover capability.
Solution Approach 2:
The system changes configuration parameters dynamically only when needed, rather than continuously reconfiguring. The static configuration parameters are optimized for normal high-speed operation, while backup parameter sets are prepared for fault tolerance. This selective parameter changing maintains processing speed during normal operation while enabling reliability improvements when failures occur.
Data Source
AI summary
A data processing system comprises a coarse-grained reconfigurable (CGR) processor including an array of reconfigurable units which are configured to execute a dataflow graph. The system further includes a compiler coupled to provide a configuration file including a configuration for a set of components in a plurality of components in the array of reconfigurable units. The configuration file is coupled to configure the set of components using the configuration. An intelligent redundancy management framework (IRMF) checks health of the configuration and identify the configuration as defective if a component in the set of component is defective. further performs a healing operation for the defective configuration by replacing the defective configuration with an alternate configuration using a different set of all healthy components.


