Multi-Socket CPU Overcurrent Protection with Graceful Workload Migration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional over current protection (OCP) mechanisms in server systems cause abrupt shutdowns, leading to poor service level agreements (SLAs) due to system unavailability, data loss, and error propagation, as they fail to manage excessive current draws effectively without disrupting the entire system.

Innovation Solution

The Application Aware Graceful Over Current Protection (AAGOCP) system employs a configurable policy-based approach that includes telemetry thresholds and system management interrupt (SMI) capabilities, allowing for isolated actions such as frequency throttling, deactivation of specific CPUs, and workload migration, rather than a complete system shutdown, thereby mitigating the impact of over current conditions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional over current protection (OCP) mechanisms are used to protect hardware components from excessive current draw, then hardware safety is improved, but system availability deteriorates due to abrupt shutdowns

Engineering Contradiction:
Improvehardware safetyVSAvoidsystem availability
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent segments the system into multiple independent sockets, allowing OCP to be applied at the individual socket level rather than system-wide. When an overcurrent condition is detected on one socket, only that specific socket is affected while other sockets continue operating normally, thus maintaining system availability while protecting hardware

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamic response strategies that adapt to application requirements. The system can choose between graceful degradation (throttling CPU frequency) or isolation (deactivating specific CPUs) based on real-time conditions and application awareness, allowing flexible balancing between hardware protection and system availability

Inventive Principle:
Principle #15Dynamics

2Reliability

If abrupt shutdown is implemented to protect against over current conditions, then hardware damage is prevented, but data loss increases

Engineering Contradiction:
Improvehardware protectionVSAvoiddata loss
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent implements preliminary actions by notifying the operating system and applications of impending OCP events before actual protection is enforced. This allows applications to perform cleanup operations, save data, and gracefully terminate processes, preventing data loss while maintaining hardware protection

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system provides a warning period between detecting overcurrent conditions and enforcing protection, creating a cushion that allows safe shutdown procedures to be executed. This intermediate state enables data preservation while ensuring hardware will be protected if the condition persists

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

3Reliability

If complete system shutdown is used for over current protection, then hardware safety is ensured, but service level agreement (SLA) experience deteriorates

Engineering Contradiction:
Improvehardware safetyVSAvoidservice level agreement experience
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent applies segmentation by isolating OCP impact to specific sockets rather than the entire system. Affected sockets can be deactivated or throttled while other sockets continue to serve workloads, maintaining SLA compliance and system usability while ensuring hardware safety

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes operational parameters dynamically by adjusting CPU frequency or deactivating specific processors based on OCP conditions. This allows the system to maintain safety while adapting performance parameters to preserve service quality and user experience

Inventive Principle:
Principle #35Parameter changes

4Reliability

If traditional OCP mechanisms are implemented, then hardware protection is achieved, but error propagation increases due to abrupt shutdowns

Engineering Contradiction:
Improvehardware protectionVSAvoiderror propagation
Core Design Contradiction:
ReliabilityVSObject-generated harmful factors

Solution Approach 1:

The patent extracts the error source by isolating and containing it to specific affected sockets. By deactivating or throttling only the problematic sockets rather than shutting down the entire system, the patent prevents error propagation to other system components while maintaining hardware protection

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS12112194B2Application aware graceful over current protection for multi-socket platforms
Publication Date: 2024.10.08 INTEL CORP
  • US12112194B2 patent drawing
  • US12112194B2 patent drawing
  • US12112194B2 patent drawing

AI summary

Systems, apparatuses and methods may provide for technology that detects an over current condition associated with a voltage regulator in a computing system, identifies a configurable over current protection policy associated with the voltage regulator, and automatically takes a protective action based on the configurable over current protection policy. In one example, the protective action includes one or more of a frequency throttle of a processor coupled to the voltage regulator in isolation from one or more additional processors in the computing system, a deactivation of the processor in isolation from the one or more additional processors, an issuance of a virtual machine monitor notification, an issuance of a data center fleet manager notification, or an initiation of a migration of a workload from the processor to at least one of the additional processor(s).