Processor Error Handling via Dynamic Operating Point Adjustment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Despite extensive testing and on-chip temperature and frequency control, processor errors still occur within the specified operating range, and as processors age, errors become more frequent near the upper limits of their operating range, necessitating better techniques for error reduction and handling.
Innovation Solution
A processor system with an on-chip controller that detects errors and adjusts operating points by changing the clock and power signals to transition to a new operating point, allowing retry of the operation, and dynamically updates the operating range to prevent future errors, thereby ensuring robustness and consistency with current capabilities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If the processor operates at higher frequency and power to achieve better performance, then productivity is improved, but processor errors occur more frequently especially as the processor ages
Solution Approach 1:
The processor dynamically adjusts its operating point by changing frequency and voltage levels in response to detected errors. The system transitions from static operating parameters to dynamic adaptation, allowing the processor to move between different operating states based on real-time error detection and environmental conditions.
Solution Approach 2:
The system changes physical parameters (frequency, voltage, temperature) to resolve errors. When an error is detected, the processor modifies its operating parameters by transitioning to a different operating point with adjusted frequency and voltage levels, thereby eliminating the error condition while maintaining operational continuity.
2Productivity
If the processor operates near the upper limits of its operating range to maximize performance, then productivity is improved, but errors become more frequent as the processor ages
Solution Approach 1:
The system implements feedback by monitoring for processor errors and using this information to adjust operating parameters. Error detection triggers a response where the processor modifies its operating point, creating a closed-loop control system that continuously adapts to maintain reliability while optimizing performance.
Solution Approach 2:
The processor performs self-diagnosis and self-correction by detecting its own errors and autonomously adjusting its operating parameters without external intervention. The system serves itself by identifying problematic operating conditions and independently transitioning to corrected operating states.
3Reliability
If extensive testing is performed to ensure processor reliability across operating points, then reliability is improved, but device complexity and testing time increase
Solution Approach 1:
The system performs preliminary error detection and operating point adjustment before critical failures occur. By continuously monitoring for errors and proactively adjusting operating parameters, the system prevents catastrophic failures rather than relying solely on extensive pre-release testing to identify all potential issues.
Data Source
AI summary
A processor comprises a processor core and a controller. The processor core has an execution unit configured to execute instructions and to attempt to perform at least one operation in executing one of the instructions. The processor core is configured to detect a processor error associated with the at least one operation. The controller is configured to change an operating point of the processor core in response to a detection of the processor error such that the processor core operates at a new operating point, and the processor core is configured to retry the at least one operation while the processor core is operating at the new operating point.


