Processor Core Resource Adjustment for Abnormal Event Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Computing systems face degraded performance due to abnormal events like hardware or software failures, which require additional resources for diagnosis and recovery, leading to increased workload and potential SLA violations.
Innovation Solution
A resource management module detects abnormal events and dynamically adjusts processor settings to increase computing resources, such as adding cores or altering frequencies, to enhance processing capacity and facilitate quicker recovery and workload completion.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If computing resources are increased to handle abnormal events, then recovery speed improves, but system complexity increases
Solution Approach 1:
The system dynamically adjusts processor settings based on operational conditions. When abnormal events are detected, the resource management module modifies processor frequencies and activates additional cores temporarily. This dynamic adjustment allows the system to have high recovery capacity when needed while maintaining low complexity during normal operation, resolving the contradiction between recovery speed and system complexity.
2Productivity
If processor frequency is increased to complete abnormal event processing, then processing capacity improves, but energy consumption increases
Solution Approach 1:
The system uses periodic monitoring to detect abnormal events and applies processor frequency adjustments only during these specific periods when abnormal events occur. During normal operation, processors run at standard frequencies with normal energy consumption. This periodic intervention approach allows high processing capacity during abnormal events while maintaining acceptable energy consumption during normal operation.
3Reliability
If computing resources are allocated to handle abnormal events, then service level agreement compliance improves, but resource availability for other partitions decreases
Solution Approach 1:
The resource management module applies computing resources locally and selectively to specific partitions experiencing abnormal events. Instead of globally reducing resources across all partitions, the system identifies the affected partition and allocates additional processor capacity specifically to that partition. This localized approach ensures SLA compliance for the affected partition while minimizing impact on resource availability for other partitions.
Data Source
AI summary
According to one or more embodiments of the present invention, a computer-implemented method includes detecting an abnormal event in operation of a first partition from a plurality of partitions of a computer server, the first partition being associated with a set of processors of the computer server and with a set of computing resources of the computer server. The method further includes in response, determining the set of processors associated with the first partition. The method further includes adjusting one or more settings of the set of processors to increase the set of computing resources associated with the first partition to complete the abnormal event.


