Multi-Socket CPU Overcurrent Protection with Graceful Workload Migration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional over current protection (OCP) mechanisms in server systems cause abrupt shutdowns, leading to poor service level agreements (SLAs) due to system unavailability, data loss, and error propagation, as they fail to manage excessive current draws effectively without disrupting the entire system.
Innovation Solution
The Application Aware Graceful Over Current Protection (AAGOCP) system employs a configurable policy-based approach that includes telemetry thresholds and system management interrupt (SMI) capabilities, allowing for isolated actions such as frequency throttling, deactivation of specific CPUs, and workload migration, rather than a complete system shutdown, thereby mitigating the impact of over current conditions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional over current protection (OCP) mechanisms are used to protect hardware components from excessive current draw, then hardware safety is improved, but system availability deteriorates due to abrupt shutdowns
Solution Approach 1:
The patent segments the system into multiple independent sockets, allowing OCP to be applied at the individual socket level rather than system-wide. When an overcurrent condition is detected on one socket, only that specific socket is affected while other sockets continue operating normally, thus maintaining system availability while protecting hardware
Solution Approach 2:
The patent implements dynamic response strategies that adapt to application requirements. The system can choose between graceful degradation (throttling CPU frequency) or isolation (deactivating specific CPUs) based on real-time conditions and application awareness, allowing flexible balancing between hardware protection and system availability
2Reliability
If abrupt shutdown is implemented to protect against over current conditions, then hardware damage is prevented, but data loss increases
Solution Approach 1:
The patent implements preliminary actions by notifying the operating system and applications of impending OCP events before actual protection is enforced. This allows applications to perform cleanup operations, save data, and gracefully terminate processes, preventing data loss while maintaining hardware protection
Solution Approach 2:
The system provides a warning period between detecting overcurrent conditions and enforcing protection, creating a cushion that allows safe shutdown procedures to be executed. This intermediate state enables data preservation while ensuring hardware will be protected if the condition persists
3Reliability
If complete system shutdown is used for over current protection, then hardware safety is ensured, but service level agreement (SLA) experience deteriorates
Solution Approach 1:
The patent applies segmentation by isolating OCP impact to specific sockets rather than the entire system. Affected sockets can be deactivated or throttled while other sockets continue to serve workloads, maintaining SLA compliance and system usability while ensuring hardware safety
Solution Approach 2:
The patent changes operational parameters dynamically by adjusting CPU frequency or deactivating specific processors based on OCP conditions. This allows the system to maintain safety while adapting performance parameters to preserve service quality and user experience
4Reliability
If traditional OCP mechanisms are implemented, then hardware protection is achieved, but error propagation increases due to abrupt shutdowns
Solution Approach 1:
The patent extracts the error source by isolating and containing it to specific affected sockets. By deactivating or throttling only the problematic sockets rather than shutting down the entire system, the patent prevents error propagation to other system components while maintaining hardware protection
Data Source
AI summary
Systems, apparatuses and methods may provide for technology that detects an over current condition associated with a voltage regulator in a computing system, identifies a configurable over current protection policy associated with the voltage regulator, and automatically takes a protective action based on the configurable over current protection policy. In one example, the protective action includes one or more of a frequency throttle of a processor coupled to the voltage regulator in isolation from one or more additional processors in the computing system, a deactivation of the processor in isolation from the one or more additional processors, an issuance of a virtual machine monitor notification, an issuance of a data center fleet manager notification, or an initiation of a migration of a workload from the processor to at least one of the additional processor(s).


