Machine Check Architecture Partitioning for Cross-Partition Error Handling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

As computing systems become more complex, the increasing bit error rate due to signal integrity issues and decreasing operating voltage complicates error handling across multiple semiconductor chips, making communication between separate components in distributed storage systems less straightforward, necessitating efficient methods for error management.

Innovation Solution

A computing system architecture that employs multiple partitions with distinct machine check architectures, utilizing a host processor to manage errors and an address translation unit to bridge communication between partitions, allowing for effective error detection, reporting, and handling across different hardware components.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If the computing system complexity increases to support multiple applications and distributed storage, then the computing capability and data bandwidth increase, but the bit error rate increases due to signal integrity issues and decreasing operating voltage

Engineering Contradiction:
Improvecomputing capabilityVSAvoidbit error rate
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system is divided into multiple partitions, each with its own machine check architecture. This segmentation allows error handling to be localized and managed independently in each partition, preventing error propagation across the entire system while maintaining overall computing capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

An address translation unit acts as an intermediary between partitions with different machine check architectures. It enables communication and error information exchange between partitions without requiring direct integration, thus maintaining reliability while supporting system complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If separate components such as motherboard and peripheral device cards are used to increase hardware topology options, then the adaptability and computing capability increase, but error handling communication becomes less straightforward

Engineering Contradiction:
Improvehardware topologyVSAvoiderror handling communication
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The machine check architecture is designed to be universal across different hardware components and configurations. Each partition implements the same MCA framework, allowing error handling to work consistently regardless of the specific hardware topology or peripheral devices connected.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The address translation unit serves as a mediator that handles communication between different hardware components and partitions. It translates error information and communication protocols between different machine check architectures, simplifying error handling across diverse hardware topologies.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Use of energy by moving object

If operating voltage decreases to reduce power consumption, then energy efficiency improves, but signal integrity deteriorates and noise margin decreases

Engineering Contradiction:
Improvepower consumptionVSAvoidsignal integrity
Core Design Contradiction:
Use of energy by moving objectVSReliability

Solution Approach 1:

The machine check architecture implements error detection and correction mechanisms in advance, cushioning against the increased bit error rates that result from low-voltage operation. Error checking is performed before errors can propagate and cause system failures.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

Solution Approach 2:

The system implements feedback mechanisms through machine check architecture that continuously monitor for errors caused by low-voltage operation. When errors are detected, the system can take corrective actions such as retrying operations or activating error correction protocols.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12072756B2Scalable machine check architecture
Publication Date: 2024.08.27 ADVANCED MICRO DEVICES INC
  • US12072756B2 patent drawing
  • US12072756B2 patent drawing
  • US12072756B2 patent drawing

AI summary

An apparatus and method for supporting communication during error handling in a computing system. A computing system includes a first partition and a second partition, each capable of performing error management based on a respective machine check architecture (MCA). When a host processor in the first partition detects an error that requires information from processor cores of the second partition, the host processor generates an access request with a target address pointing to a storage location in a memory of the second partition, not the first partition. When the host processor receives the requested error log information from the second partition, the host processor completes processing of the error. To support the host processor in generating the target address for the access request, during an earlier bootup operation, the second partition communicates the hardware topology of the second partition to the host processor.