Lane Error Counter View for Faulty Lane Isolation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Hyperscalers face challenges in accurately diagnosing and disambiguating transient versus persistent errors in data center systems, leading to unscheduled reboots and significant revenue loss, due to the limitations of existing error detection and correction mechanisms that fail to provide timely and precise fault isolation and prediction.

Innovation Solution

Implementing a lane-based normalized historical error counter view that records error occurrences per lane, using ECC or CRC syndromes combined with hardware logic to maintain error counts, providing a historical error map and probabilistic bad lane identification, applicable to high-speed I/O protocols like DDR, UPI, and PCIe.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing error detection and correction mechanisms are used, then errors can be detected, but fault isolation and disambiguation of transient versus persistent errors are inaccurate and timely

Engineering Contradiction:
Improvefault isolation accuracyVSAvoiderror diagnosis time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments error tracking by creating separate error counters for each physical lane instead of using a single aggregate error counter. This segmentation allows precise identification of which specific lane is experiencing errors, enabling accurate fault isolation without requiring time-consuming analysis of mixed error data from all lanes.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds a temporal dimension to error detection by implementing historical error counters that track errors over time. This allows the system to distinguish between transient errors (temporary fluctuations) and persistent errors (consistent failures) by analyzing error patterns across multiple time periods, significantly improving diagnostic accuracy.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Device complexity

If aggregate error counting is used, then error detection is simple, but lane-specific fault identification is impossible

Engineering Contradiction:
Improveerror counting structureVSAvoidlane identification accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent divides the error counting function into multiple lane-specific counters, where each counter tracks errors for a specific physical lane. This segmentation maintains relative structural simplicity while enabling precise lane identification, as each counter independently records errors for its associated lane without requiring complex analysis.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates multiple copies of the error counter structure, one for each physical lane. This copying approach allows the system to maintain a simple, replicated error tracking mechanism for each lane, making fault identification straightforward by simply reading the appropriate lane-specific counter rather than analyzing complex aggregated data.

Inventive Principle:
Principle #26Copying

3Reliability

If comprehensive error tracking is implemented, then fault diagnosis capability is improved, but software complexity increases

Engineering Contradiction:
Improvefault diagnosis capabilityVSAvoidsoftware complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements self-service error tracking by having the hardware automatically maintain lane-specific error counters and historical error data. This eliminates the need for complex software to manually track and analyze errors from multiple sources, as the hardware itself performs the comprehensive tracking and organizes the data in a readily interpretable format.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent provides direct feedback through lane-specific error counters that immediately reflect the error status of each physical lane. This feedback mechanism enables simple software to quickly determine which lane is faulty by reading the counter values, avoiding complex diagnostic algorithms while maintaining high reliability in fault identification.

Inventive Principle:
Principle #23Feedback

4Reliability

If historical error data is collected per lane, then predictive failure analysis is enabled, but data storage requirements increase

Engineering Contradiction:
Improvepredictive failure analysis capabilityVSAvoiderror data volume
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential error information needed for predictive analysis by maintaining historical counters that record the presence or absence of errors over time periods. Rather than storing complete error logs with all details, the system extracts and retains only the critical temporal pattern data necessary for predicting future failures, significantly reducing storage requirements.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the parameter representation from detailed error logs to normalized error counts over time periods. By transforming the error data into a compact temporal pattern representation (e.g., error counts per time period rather than individual error events), the system enables predictive failure analysis while minimizing the quantity of stored data.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12541416B2Lane based normalized historical error counter view for faulty lane isolation and disambiguation of transient versus persistent errors
Publication Date: 2026.02.03 INTEL CORP
  • US12541416B2 patent drawing
  • US12541416B2 patent drawing
  • US12541416B2 patent drawing

AI summary

Methods and apparatus relating to lane based normalized historical error counter view for faulty lane isolation and disambiguation of transient versus persistent errors are described. In an embodiment, a plurality of storage entries store error information to be detected at one or more physical lanes of an interface. Faulty lane detection logic circuitry determines which of the one or more physical lanes is faulty or more likely to be faulty based at least in part on the stored error information for the one or more physical lanes of the interface. The stored error information comprises historical error details for the one or more physical lanes of the interface. Other embodiments are also disclosed and claimed.