Error Recovery Between Independent Processors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies face challenges in coordinating error information recovery between independently operable processors, particularly in handling fatal errors and maintaining consistent frameworks for error recovery across multiple processors, leading to unreliability and loss of debugging information.

Innovation Solution

A method and apparatus for controlled recovery of error information between independently operable processors, utilizing a PCIe bus interface and GPIO signals to detect and manage crash events, perform error recovery procedures, and ensure successful completion before resetting processors, thereby preserving error information and coordinating reset conditions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If error recovery procedures are performed between independently operable processors, then error information can be collected and preserved, but the complexity of coordinating reset conditions and ensuring consistent recovery frameworks increases

Engineering Contradiction:
Improveerror information recovery reliabilityVSAvoidcoordination complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces a coordinated error recovery mechanism that acts as an intermediary between independently operable processors. When a crash event occurs, the system establishes a consistent recovery framework that manages the interaction between processors, ensuring error information is preserved without requiring complex direct coordination between each processor pair. The recovery mechanism mediates the reset conditions and information transfer, simplifying the overall coordination complexity while maintaining reliability.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If processors are reset immediately upon crash detection, then system stability is restored, but error information and debugging data are lost

Engineering Contradiction:
Improvesystem stabilityVSAvoiderror information loss
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent implements preliminary error information preservation actions before processor reset. When a crash event is detected, the system first collects and preserves error information from the crashed processor's memory and registers, then establishes recovery procedures, and only after these preliminary actions are complete does the reset occur. This ensures debugging information is captured before the processor state changes, preventing information loss while still restoring system stability.

Inventive Principle:
Principle #10Preliminary action

3Loss of information

If error recovery procedures are delayed to preserve information, then debugging data is maintained, but system downtime and recovery time increase

Engineering Contradiction:
Improveerror information preservationVSAvoidrecovery time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent enables the crashed processor to perform self-service error information preservation. The processor that experiences the crash autonomously preserves its own error information (register states, memory contents, stack traces) before being reset, without requiring extensive external intervention or prolonged recovery procedures. This self-service approach minimizes the time the system remains in a degraded state while ensuring complete error information is captured for debugging.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS9842036B2Methods and apparatus for controlled recovery of error information between independently operable processors
Publication Date: 2017.12.12 APPLE INC
  • US9842036B2 patent drawing
  • US9842036B2 patent drawing
  • US9842036B2 patent drawing

AI summary

Methods and apparatus for controlled recovery of error information between two (or more) independently operable processors. The present disclosure provides solutions that preserve error information in the event of a fatal error, coordinate reset conditions between independently operable processors, and implement consistent frameworks for error information recovery across a range of potential fatal errors. In one exemplary embodiment, an applications processor (AP) and baseband processor (BB) implement an abort handler and power down handler sequence which enables error recovery over a wide range of crash scenarios. In one variant, assertion of signals between the AP and the BB enables the AP to reset the BB only after error recovery procedures have successfully completed.