Crashlog Unit Hang Detection Microprocessor Data Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for detecting and addressing system hangs in microprocessor systems are inefficient, often resulting in data loss and requiring complex out-of-band data flows or physical probes that are impractical for large-scale systems, leading to difficulties in identifying and resolving hardware and software errors.
Innovation Solution
Implementing a crashlog unit with processing logic to detect hangs and collect data from core, uncore, and controller hub registers, allowing for independent and integrated data collection that persists even during system resets, using existing in-band data paths and enhanced serial peripheral interface messaging for efficient data retrieval.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If out-of-band data flows or physical probes are used to detect system hangs, then hang detection capability is improved, but device complexity and ease of operation deteriorate
Solution Approach 1:
The patent extracts the hang detection and data collection functionality into a dedicated crashlog unit that operates independently from the main processor. This unit can collect crash logs from multiple sources (core, uncore, controller hub) and store them in shared memory, allowing detection without requiring complex out-of-band data flows or physical probes during system hangs.
Solution Approach 2:
The crashlog unit serves multiple functions: detecting hangs, collecting crash logs from different system regions, storing data in shared memory, and supporting both in-band and out-of-band retrieval paths. This multi-functional design improves hang detection capability while avoiding the need for separate complex detection systems.
2Loss of information
If out-of-band data flows are used for data collection, then data recovery capability is improved, but ease of operation and device complexity worsen
Solution Approach 1:
The patent introduces shared memory as an intermediary storage location between the crashlog unit and various retrieval paths. Crash logs are collected and stored in this shared memory buffer, which can then be accessed through multiple paths including out-of-band interfaces. This intermediary approach enables data recovery without requiring direct complex out-of-band data flows from the crash site.
3Measurement precision
If complex out-of-band data flows are implemented, then hang detection accuracy is improved, but ease of manufacture and device complexity worsen
Solution Approach 1:
The crashlog unit is designed to automatically detect hangs and collect crash logs without requiring external physical probes or complex out-of-band instrumentation. The unit autonomously monitors for hang conditions, collects relevant data from system registers, and stores it in shared memory, enabling accurate hang detection while simplifying manufacturing by eliminating the need for complex external detection infrastructure.
4Device complexity
If in-band data paths are used for data retrieval, then device complexity is reduced, but reliability during hangs worsens
Solution Approach 1:
The patent segments the data retrieval paths into multiple independent channels: in-band paths for normal operation and out-of-band paths for hang recovery. The crashlog unit can select appropriate retrieval paths based on system state, allowing in-band paths to be used during normal operation (reducing complexity) while providing out-of-band alternatives when hangs occur (maintaining reliability).
Data Source
AI summary
Embodiment of this disclosure provides a mechanism to support hang detection and data recovery in microprocessor systems. In one embodiment, a processing device comprising a processing core and a crashlog unit operatively coupled to the core is provided. An indication of an unresponsive state in an execution of a pending instruction by the core is received. Responsive to receiving the indication, a crash log comprising data from registers of at least one of: a core region, a non-core region and a controller hub associated with the processing device is produced. Thereupon, the crash log is stored in a shared memory of a power management controller (PMC) associated with the controller hub.


