PCIe Endpoint Failure Detection via Host Software Scanning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing PCIe error detection solutions are inadequate for detecting endpoint device failures, particularly when devices are located behind a switch or experiencing software-related issues, as they rely on Link Training and Status State Machine (LTSSM) which may not reach the host CPU, leading to undetected failures.

Innovation Solution

A method and system that involves host software detecting connected PCIe endpoint devices, scanning for advanced status reporting capabilities, programming specific values in memory control words to trigger status updates, and monitoring for failure indications in the endpoint response register, enabling detection of endpoint device failures even when LTSSM is ineffective.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If LTSSM hardware is used for error detection, then basic link failures are detected, but failures behind switches or software-related failures are not detected

Engineering Contradiction:
Improveerror detection capabilityVSAvoiddetection coverage for various failure types
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent implements preliminary action by having the host software proactively scan the extended configuration space of PCIe endpoint devices to detect advanced status reporting capabilities before failures occur. The host software programs predetermined values into the endpoint response register and root complex request register to establish a monitoring mechanism in advance, enabling detection of failures including those behind switches or caused by software states, thereby resolving the limitation of passive LTSSM-based detection.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If comprehensive error detection is implemented, then detection coverage is improved, but system complexity increases

Engineering Contradiction:
Improvefailure detection accuracyVSAvoidhost software complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies universality by designing a multi-functional monitoring mechanism where the host software performs multiple functions: scanning extended configuration space for capability detection, programming predetermined values into registers, triggering status updates, and monitoring for failures. This unified approach using existing PCIe configuration space registers eliminates the need for separate dedicated hardware for each detection function, thereby achieving comprehensive detection coverage without proportionally increasing system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If advanced status reporting is enabled, then detection of endpoint failures is improved, but compatibility with existing PCIe systems is reduced

Engineering Contradiction:
Improveendpoint failure detectionVSAvoidcompatibility with existing systems
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent implements partial action by selectively enabling advanced status reporting only for PCIe endpoint devices that support it. The host software scans the extended configuration space to identify devices with advanced status reporting capability and applies the enhanced monitoring mechanism only to those devices. This approach maintains compatibility with existing PCIe systems that do not support the advanced feature while providing improved failure detection where applicable, thus resolving the contradiction between enhancement and compatibility.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentEP3851964A1Method and system to detect failure in pcie endpoint devices
Publication Date: 2021.07.21 NXP USA INC
  • EP3851964A1 patent drawingFigure 1
  • EP3851964A1 patent drawingFigure 2
  • EP3851964A1 patent drawingFigure 3

AI summary

A method, system, apparatus, and architecture are provided for detecting failure of a PCIe endpoint device by scanning an extended configuration space for each connected PCIe endpoint device to detect a first PCIe endpoint device that supports advance status reporting, and then by programming a first predetermined value and a second predetermined value, respectively, into an endpoint response register and a root complex request register of a dedicated memory control word in the extended configuration space for the first PCIe endpoint device, where the second predetermined value signals a request to the first PCIe endpoint device to update the endpoint response register of the dedicated memory control word with a new status value so that, after a minimum specified delay, a report that the first PCIe endpoint device has failed may be generated in response to detecting that the first predetermined value is stored in the endpoint response register.