Parallel Hardware Diagnostics for Data Processing Pipeline Faults

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing hardware fault detection methods, such as error correction coding and software diagnostic methods, are ineffective in detecting transient and permanent faults in data processing pipelines, leading to performance degradation and undetected faults, especially in complex systems like autonomous vehicles.

Innovation Solution

A hardware diagnostic circuit is integrated into a System on a Chip (SoC) that includes a safety monitor and simulation system to perform diagnostic operations, using error detection hardware and software diagnostic methods like PFSD to detect faults, with enhanced observability through structural and workload analysis and hardware-based error detection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If software diagnostic methods are used to detect faults in processing logic, then some transient and permanent faults can be detected, but system performance deteriorates because computing resources are tied up performing diagnostic operations

Engineering Contradiction:
Improvefault detection capabilityVSAvoidsystem performance
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent replaces software-based diagnostic methods with hardware-based diagnostic circuits integrated into the processing logic. This substitution allows fault detection to occur in parallel with normal operations, eliminating the performance penalty associated with software diagnostics while maintaining high detection coverage for both transient and permanent faults

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces intermediate diagnostic circuits and test data paths that act as mediators between the processing logic and fault detection mechanisms. These intermediaries enable continuous monitoring of internal signals without disrupting the main data flow, allowing simultaneous functional operation and diagnostic evaluation

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If error correction coding is used to detect transient faults in data storage, then transient faults can be detected, but it is not effective for detecting faults in processing logic

Engineering Contradiction:
Improvetransient fault detectionVSAvoidapplicability to processing logic
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent designs diagnostic circuits with universal applicability that can detect both transient and permanent faults across different types of processing logic. The diagnostic mechanism is configured to monitor various signal types and processing stages, making it adaptable to diverse computational operations while maintaining effective fault detection

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent employs parameter changes in the diagnostic approach by switching between different detection modes and test configurations. The system adjusts diagnostic parameters such as sampling rates, test patterns, and monitoring points to optimize detection effectiveness for different fault types and processing contexts

Inventive Principle:
Principle #35Parameter changes

3Reliability

If diagnostic operations are performed instead of functional operations, then fault detection improves, but runtime bottlenecks increase due to tied up computing resources

Engineering Contradiction:
Improvefault detection rateVSAvoidruntime bottlenecks
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements continuous fault detection through dedicated hardware circuits that operate in parallel with functional operations. This allows the diagnostic process to continue uninterrupted alongside normal computing tasks, eliminating the need to pause or slow down functional operations for fault detection while maintaining continuous monitoring capability

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS12417137B2Detecting hardware faults in data processing pipelines
Publication Date: 2025.09.16 NVIDIA CORP
  • US12417137B2 patent drawing
  • US12417137B2 patent drawing
  • US12417137B2 patent drawing

AI summary

In various examples, a system comprising at least one circuit to detect whether a fault has occurred during performance of an operation by the at least one circuit. In at least some embodiments, the at least one circuit generates error detecting values and determines a fault has occurred when the error detecting values do not match predetermined error detecting data.