Management Controller Crash Data Collection via I3C Link
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current crash data harvesting methods in computing systems, particularly using the PECI protocol state machine, face challenges in accessing crash data from CPU core/uncore IP blocks and companion dice due to failures in internal and external sideband networks, leading to incomplete data recovery and prolonged reset times, which can result in data loss and extended downtime.
Innovation Solution
Implementing a management controller that utilizes a two-wire I3C communication link to interface with primary and secondary PECI protocol state machines, enabling direct access to sticky registers on companion dice through integrated and discrete companion dice, and employing a demoted warm reset to preserve error information, thereby ensuring data integrity and reducing recovery time.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If PECI protocol state machine is used to access crash data from CPU and companion dice, then crash data can be harvested, but access may fail due to sideband network failures leading to incomplete data recovery
Solution Approach 1:
The patent introduces an I3C interface as an intermediary communication path between the management controller and companion dice. This alternative interface bypasses the failed PECI sideband network, allowing crash data to be retrieved through a different communication channel when the original path is unavailable.
Solution Approach 2:
The system dynamically switches between different communication interfaces (PECI and I3C) based on availability. The management controller can adaptively select the functional interface for data retrieval, making the crash data harvesting process resilient to interface failures.
2Productivity
If traditional reset methods are used after catastrophic error, then system can recover, but error information may be lost and downtime is extended
Solution Approach 1:
The patent performs crash data harvesting before executing the reset operation. By collecting error information from sticky registers and other sources prior to the demoted warm reset, the system preserves diagnostic data that would otherwise be lost, enabling faster recovery without sacrificing debugging capability.
Solution Approach 2:
The system prepares alternative data collection mechanisms (I3C interface) beforehand to ensure crash data can be retrieved even if the primary interface fails during the crash event, cushioning against potential data loss scenarios.
Data Source
AI summary
Examples include techniques to collect crash data for a computing system following a catastrophic error. Examples include a management controller gathering error information from components of a computing system that includes a central processing unit (CPU) coupled with one or more companion dice following the catastrophic error. The management controller to gather the error information via a communication link coupled between the management controller, the CPU and the one or more companion dice.


