Processor Core Adaptive Voltage-Frequency Scaling for Silent Data Corruption

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Silent data corruptions occur due to timing errors propagating through computing devices, leading to incorrect data without immediate detection, compromising dependability and reliability.

Innovation Solution

Implementing adaptive voltage-frequency scaling (AVFS) via in-situ monitors like razor flops across processor cores and caches, with a system management unit that adjusts voltage and frequency to mitigate timing errors, using telemetry data and predictive simulations to optimize performance and power consumption.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If voltage and frequency are increased to improve processing speed, then productivity increases, but timing errors occur causing silent data corruptions

Engineering Contradiction:
Improveprocessing speedVSAvoiddata integrity
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system dynamically adjusts voltage and frequency based on real-time operational conditions and error rates. The computing device transitions from static voltage-frequency settings to adaptive scaling, where parameters are continuously modified to maintain optimal performance while preventing timing errors that cause silent data corruptions.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The invention changes the voltage and frequency parameters adaptively based on detected error patterns. When silent data corruptions are detected, the system modifies these electrical parameters to stabilize timing-critical operations, thereby preventing future errors while maintaining acceptable productivity levels.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If voltage is increased to stabilize timing and prevent errors, then reliability improves, but power consumption increases

Engineering Contradiction:
Improvetiming stabilityVSAvoidpower consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

Instead of maintaining a constantly high voltage for timing stability, the system dynamically adjusts voltage levels based on actual error rates and operational needs. Voltage is increased only when and where timing errors are detected, and reduced when conditions permit, thereby achieving reliability improvements with minimized power consumption impact.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The adaptive voltage-frequency scaling is applied selectively to specific circuits or components that exhibit timing errors, rather than uniformly increasing voltage across the entire system. This localized approach stabilizes critical timing paths while avoiding unnecessary power consumption in non-problematic areas.

Inventive Principle:
Principle #3Local quality

3Productivity

If frequency is increased to improve performance, then productivity increases, but timing errors propagate causing data corruptions

Engineering Contradiction:
Improveprocessing throughputVSAvoidtiming accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system implements dynamic frequency scaling that adjusts operating frequency based on real-time timing error detection. When timing errors are detected, frequency is reduced to eliminate error propagation; when operations are stable, frequency can be increased to maximize throughput, creating a dynamic balance between productivity and timing accuracy.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12399772B2Devices, systems, and methods for detecting and mitigating silent data corruptions via adaptive voltage-frequency scaling
Publication Date: 2025.08.26 ADVANCED MICRO DEVICES INC
  • US12399772B2 patent drawing
  • US12399772B2 patent drawing
  • US12399772B2 patent drawing

AI summary

An exemplary computing device includes a plurality of circuits and/or a plurality of in-situ monitors configured to generate outputs that indicate one or more operating conditions of the circuits. The computing device also includes a system management unit configured to detect a potentially faulty voltage-to-frequency ratio implemented by one of the circuits based at least in part on one or more of the outputs. The system management unit is also configured to modify the potentially faulty voltage-to-frequency ratio based at least in part on one or more of the outputs. Various other devices, systems, and methods are also disclosed.