Multi-Core Mutual Inter-Checking for Hardware Fault Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional fault detection solutions for automotive system-on-chip circuits face challenges in efficiently and flexibly performing in-field functional safety testing, especially under changing operating conditions, due to the need for dedicated hardware and the risk of error-prone software-based self-tests on faulty hardware, which can lead to performance degradation and limited test coverage.
Innovation Solution
An adaptive hardware fault detection method where multiple processor cores execute software-based self-tests concurrently and perform mutual inter-core checking, allowing for flexible and low-cost detection of hardware faults without requiring dedicated hardware fault detection or recovery mechanisms, by synchronizing cores to execute software-based self-test programs and cross-checking results.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If dedicated test hardware components are used to detect faults, then fault detection capability is improved, but device cost and complexity increase
Solution Approach 1:
The patent makes existing processor cores perform multiple functions: they execute both normal computational tasks and self-test programs. The cores use their existing computational units, memory interfaces, and I/O capabilities for both production work and fault detection, eliminating the need for dedicated test hardware while maintaining comprehensive fault detection capability
Solution Approach 2:
The processor cores perform self-diagnosis by executing self-test programs that monitor their own operational parameters. Each core acts as its own test device, detecting faults in its computational units, memory interfaces, and I/O peripherals without external intervention, thereby reducing system complexity while maintaining reliability
2Reliability
If software-based self-test programs are run sequentially on each processor core, then fault detection is achieved, but system performance deteriorates due to suspension of normal operations
Solution Approach 1:
The patent enables fault detection to occur continuously during normal operations rather than requiring suspension. Self-test programs are executed in the background concurrently with computational tasks, and operational parameters are monitored in real-time, ensuring both continuous productivity and continuous fault detection capability
Solution Approach 2:
The system performs preliminary fault detection during idle periods or low-utilization windows by executing comprehensive self-test programs. This allows thorough testing to be conducted in advance during periods when performance impact is minimal, while maintaining rapid fault detection capability during high-utilization periods through continuous monitoring
3Adaptability or versatility
If a faulty processor core executes self-test programs on itself, then self-testing capability is maintained, but test accuracy decreases due to error-prone checking on faulty hardware
Solution Approach 1:
The patent introduces a healthy processor core as an intermediary to verify the self-test results of potentially faulty cores. The monitoring core independently checks the test outcomes and operational parameters reported by the core-under-test, providing an external validation mechanism that prevents faulty cores from masking their own defects through corrupted self-assessment
Solution Approach 2:
Instead of relying on a core to assess its own health accurately, the system inverts the assessment direction by having other healthy cores evaluate the test results and operational status of each core. This cross-validation approach ensures that no single core can corrupt the fault detection process, thereby maintaining high measurement precision
4Ease of manufacture
If single-core self-testing is performed, then implementation simplicity is maintained, but test coverage is limited for complex cache coherence logic
Solution Approach 1:
The patent segments the testing function across multiple processor cores, where each core is responsible for monitoring specific aspects of system operation and cache coherence. This distribution of testing responsibilities among multiple cores enables comprehensive coverage of complex multi-core interactions while maintaining modular implementation that builds upon simple single-core self-testing foundations
Data Source
AI summary
A method, apparatus, article of manufacture, and system are provided for detecting hardware faults on a multi-core integrated circuit device by executing runtime software-based self-test code concurrently on multiple processor cores to generate a first set of self-test results from a first processor core and a second set of self-test results from a second processor core; performing mutual inter-core checking of the self-test results by using the first processor core to check the second set of self-test results from the second processor core while simultaneously using the second processor core to check the first set of self-test results from the first processor core; and then using the second processor core to immediately execute a recovery mechanism for the first processor core if comparison of the first set of self-test results against reference test results indicates there is a hardware failure at the first processor core.


