Error Blocking Unit for Fault Containment in Multi-CPU Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In multi-CPU or multi-core microcontrollers, existing fault containment methods fail to prevent error propagation effectively, leading to system instability and lack of defined error recovery mechanisms in automotive electronic control units (ECUs), which are critical for ensuring system availability and safety.

Innovation Solution

A system and method that utilize a delayed lockstep CPU configuration with an error blocking unit to delay and block write transactions from a first CPU until an error is detected by a second CPU, preventing corruption of system resources and ensuring integrity by comparing output signals from both CPUs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If fault containment methods are implemented in multi-CPU microcontrollers, then system reliability is improved, but device complexity increases

Engineering Contradiction:
Improvesystem reliabilityVSAvoiddevice complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system divides the multi-CPU microcontroller into separate functional domains, with each CPU operating independently within its own protected memory space. Error blocking units are inserted between CPUs and shared resources to segment error propagation paths, allowing faults to be contained within specific domains while maintaining overall system reliability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Error blocking units serve as intermediary components between CPUs and shared system resources. These units monitor and control access to shared memory and peripherals, blocking erroneous transactions before they can corrupt system state. The intermediary mechanism adds reliability without requiring complete system redesign.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If error detection mechanisms are added to prevent error propagation, then system safety is improved, but manufacturing precision requirements increase

Engineering Contradiction:
Improvesystem safetyVSAvoidmanufacturing precision
Core Design Contradiction:
ReliabilityVSManufacturing precision

Solution Approach 1:

Error blocking units are configured with pre-defined blocking rules and criteria before system operation. These units proactively prevent potential error propagation by monitoring transaction patterns and blocking suspicious operations before they can cause system corruption. The preliminary configuration approach reduces the need for high-precision real-time detection, lowering manufacturing tolerance requirements.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If delayed lockstep CPU configuration is used to detect errors, then system availability is improved, but device complexity increases

Engineering Contradiction:
Improvesystem availabilityVSAvoiddevice complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The error blocking functionality is merged with the existing CPU interconnect architecture and memory interface structures. Rather than adding completely separate error detection systems, the patent integrates error monitoring and blocking capabilities into the existing data paths, reducing overall device complexity while maintaining system availability.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS9417946B2Method and system for fault containment
Publication Date: 2016.08.16 INFINEON TECHNOLOGIES AG
  • US9417946B2 patent drawing
  • US9417946B2 patent drawing
  • US9417946B2 patent drawing

AI summary

Embodiments relate to systems and methods for error containment in a system comprising detecting an error by processing an input signal by multiple processing units, and delaying at least one output signal of a processing unit to enable, in case an error has been detected, modifying at least one output signal of the processing unit that would cause propagation of the error through the system.