Multi-Core Mutual Inter-Checking for Hardware Fault Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional fault detection solutions for automotive system-on-chip circuits face challenges in efficiently and flexibly performing in-field functional safety testing, especially under changing operating conditions, due to the need for dedicated hardware and the risk of error-prone software-based self-tests on faulty hardware, which can lead to performance degradation and limited test coverage.

Innovation Solution

An adaptive hardware fault detection method where multiple processor cores execute software-based self-tests concurrently and perform mutual inter-core checking, allowing for flexible and low-cost detection of hardware faults without requiring dedicated hardware fault detection or recovery mechanisms, by synchronizing cores to execute software-based self-test programs and cross-checking results.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If dedicated test hardware components are used to detect faults, then fault detection capability is improved, but device cost and complexity increase

Engineering Contradiction:
Improvefault detection capabilityVSAvoiddevice complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent makes existing processor cores perform multiple functions: they execute both normal computational tasks and self-test programs. The cores use their existing computational units, memory interfaces, and I/O capabilities for both production work and fault detection, eliminating the need for dedicated test hardware while maintaining comprehensive fault detection capability

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The processor cores perform self-diagnosis by executing self-test programs that monitor their own operational parameters. Each core acts as its own test device, detecting faults in its computational units, memory interfaces, and I/O peripherals without external intervention, thereby reducing system complexity while maintaining reliability

Inventive Principle:
Principle #25Self-service

2Reliability

If software-based self-test programs are run sequentially on each processor core, then fault detection is achieved, but system performance deteriorates due to suspension of normal operations

Engineering Contradiction:
Improvefault detectionVSAvoidsystem performance
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent enables fault detection to occur continuously during normal operations rather than requiring suspension. Self-test programs are executed in the background concurrently with computational tasks, and operational parameters are monitored in real-time, ensuring both continuous productivity and continuous fault detection capability

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The system performs preliminary fault detection during idle periods or low-utilization windows by executing comprehensive self-test programs. This allows thorough testing to be conducted in advance during periods when performance impact is minimal, while maintaining rapid fault detection capability during high-utilization periods through continuous monitoring

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If a faulty processor core executes self-test programs on itself, then self-testing capability is maintained, but test accuracy decreases due to error-prone checking on faulty hardware

Engineering Contradiction:
Improveself-testing capabilityVSAvoidtest accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent introduces a healthy processor core as an intermediary to verify the self-test results of potentially faulty cores. The monitoring core independently checks the test outcomes and operational parameters reported by the core-under-test, providing an external validation mechanism that prevents faulty cores from masking their own defects through corrupted self-assessment

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

Instead of relying on a core to assess its own health accurately, the system inverts the assessment direction by having other healthy cores evaluate the test results and operational status of each core. This cross-validation approach ensures that no single core can corrupt the fault detection process, thereby maintaining high measurement precision

Inventive Principle:
Principle #13The other way round (Inversion)

4Ease of manufacture

If single-core self-testing is performed, then implementation simplicity is maintained, but test coverage is limited for complex cache coherence logic

Engineering Contradiction:
Improveimplementation simplicityVSAvoidtest coverage
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent segments the testing function across multiple processor cores, where each core is responsible for monitoring specific aspects of system operation and cache coherence. This distribution of testing responsibilities among multiple cores enables comprehensive coverage of complex multi-core interactions while maintaining modular implementation that builds upon simple single-core self-testing foundations

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10628275B2Runtime software-based self-test with mutual inter-core checking
Publication Date: 2020.04.21 NXP BV
  • US10628275B2 patent drawing
  • US10628275B2 patent drawing
  • US10628275B2 patent drawing

AI summary

A method, apparatus, article of manufacture, and system are provided for detecting hardware faults on a multi-core integrated circuit device by executing runtime software-based self-test code concurrently on multiple processor cores to generate a first set of self-test results from a first processor core and a second set of self-test results from a second processor core; performing mutual inter-core checking of the self-test results by using the first processor core to check the second set of self-test results from the second processor core while simultaneously using the second processor core to check the first set of self-test results from the first processor core; and then using the second processor core to immediately execute a recovery mechanism for the first processor core if comparison of the first set of self-test results against reference test results indicates there is a hardware failure at the first processor core.