CPU Management Engine Crash Dump Circuit for Error Logging

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Computer systems without a baseboard management controller (BMC) face challenges in error recording from CPU failures, leading to increased downtime due to the inability to analyze errors, as they require specialized knowledge and protocols for error logging and pose a security risk.

Innovation Solution

A dedicated crash dump hardware circuit is introduced, which includes a programmable device coupled to the CPU to receive error signals, request and store error data in a non-volatile memory, eliminating the need for a BMC by enabling error logging directly within the CPU management engine.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a baseboard management controller (BMC) is used to store error data from CPU failures, then error recording capability is improved, but device complexity and cost increase

Engineering Contradiction:
Improveerror recording capabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts the error data storage function from the complex BMC system and implements it directly in the CPU's management engine. The management engine collects error data from the catastrophic error event signal and stores it in dedicated crash dump memory, eliminating the need for BMC involvement in error logging while maintaining full error recording capability

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The management engine within the CPU is given the additional function of error data collection and storage. This multi-functional approach allows the same processor component to handle both computational tasks and error logging, removing the need for a separate BMC dedicated solely to error recording

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If a baseboard management controller (BMC) is used for error logging, then error data storage is improved, but security risks increase

Engineering Contradiction:
Improveerror logging capabilityVSAvoidsecurity risk
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The error logging function is extracted from the BMC, which is a separate processor unit with potential security vulnerabilities. By implementing error logging within the CPU's management engine, the system eliminates the security risks associated with BMC access while preserving error logging capabilities

Inventive Principle:
Principle #2Taking out (Extraction)

3Reliability

If a baseboard management controller (BMC) is used for error recording, then error analysis capability is improved, but cost increases

Engineering Contradiction:
Improveerror analysis capabilityVSAvoidsystem cost
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The management engine within the CPU performs both error data collection and storage functions that would otherwise require a separate BMC. This consolidation eliminates the need for additional processor units and reduces overall system cost while maintaining error analysis capability

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Device complexity

If no baseboard management controller (BMC) is used, then device complexity is reduced, but error recording capability is lost

Engineering Contradiction:
Improvesystem simplicityVSAvoiderror recording capability
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The CPU's management engine performs error data collection and storage autonomously without requiring external BMC assistance. The catastrophic error event signal triggers the management engine to automatically collect error data from model-specific registers and store it in crash dump memory, enabling the system to serve its own error logging needs

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11360839B1Systems and methods for storing error data from a crash dump in a computer system
Publication Date: 2022.06.14 QUANTA COMPUTER INC
  • US11360839B1 patent drawing
  • US11360839B1 patent drawing
  • US11360839B1 patent drawing

AI summary

A system and method for logging error data from a central processing unit on a computer system using a dedicated crash dump device, is disclosed. The central processing unit has a management engine. The central processing unit sends an error signal. The dedicated crash dump device is coupled to the central processing unit to receive the error signal. A storage device is coupled to the crash dump device. The crash dump device sends a request to the central processing unit for error data. The crash dump device receives error data from the central processing unit. The crash dump device stores the error data in the storage device.