Error Recovery Between Independent Processors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies face challenges in coordinating error information recovery between independently operable processors, particularly in handling fatal errors and maintaining consistent frameworks for error recovery across multiple processors, leading to unreliability and loss of debugging information.
Innovation Solution
A method and apparatus for controlled recovery of error information between independently operable processors, utilizing a PCIe bus interface and GPIO signals to detect and manage crash events, perform error recovery procedures, and ensure successful completion before resetting processors, thereby preserving error information and coordinating reset conditions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If error recovery procedures are performed between independently operable processors, then error information can be collected and preserved, but the complexity of coordinating reset conditions and ensuring consistent recovery frameworks increases
Solution Approach 1:
The patent introduces a coordinated error recovery mechanism that acts as an intermediary between independently operable processors. When a crash event occurs, the system establishes a consistent recovery framework that manages the interaction between processors, ensuring error information is preserved without requiring complex direct coordination between each processor pair. The recovery mechanism mediates the reset conditions and information transfer, simplifying the overall coordination complexity while maintaining reliability.
2Reliability
If processors are reset immediately upon crash detection, then system stability is restored, but error information and debugging data are lost
Solution Approach 1:
The patent implements preliminary error information preservation actions before processor reset. When a crash event is detected, the system first collects and preserves error information from the crashed processor's memory and registers, then establishes recovery procedures, and only after these preliminary actions are complete does the reset occur. This ensures debugging information is captured before the processor state changes, preventing information loss while still restoring system stability.
3Loss of information
If error recovery procedures are delayed to preserve information, then debugging data is maintained, but system downtime and recovery time increase
Solution Approach 1:
The patent enables the crashed processor to perform self-service error information preservation. The processor that experiences the crash autonomously preserves its own error information (register states, memory contents, stack traces) before being reset, without requiring extensive external intervention or prolonged recovery procedures. This self-service approach minimizes the time the system remains in a degraded state while ensuring complete error information is captured for debugging.
Data Source
AI summary
Methods and apparatus for controlled recovery of error information between two (or more) independently operable processors. The present disclosure provides solutions that preserve error information in the event of a fatal error, coordinate reset conditions between independently operable processors, and implement consistent frameworks for error information recovery across a range of potential fatal errors. In one exemplary embodiment, an applications processor (AP) and baseband processor (BB) implement an abort handler and power down handler sequence which enables error recovery over a wide range of crash scenarios. In one variant, assertion of signals between the AP and the BB enables the AP to reset the BB only after error recovery procedures have successfully completed.


