Server RAS Configuration for Crash Recovery and Fault Isolation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Server faults, particularly crashing faults during starting and operation, affect performance and operational efficiency, and existing RAS functions do not adequately address these issues, leading to potential server crashes.
Innovation Solution
A method and apparatus for configuring server maintainability that calculates CPU utilization rates before and after crashes, determines fault components, and switches server configuration modes based on service migration states to isolate faults and reduce crash rates.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the server operates with default RAS configuration parameters, then the server can continue operating with fault components, but server performance and operational efficiency deteriorate and may lead to server crashing
Solution Approach 1:
The patent implements dynamic adjustment of RAS configuration parameters based on real-time service migration states. The system monitors CPU utilization rates before and after crashes, determines service migration states, and dynamically switches between different RAS configuration modes (default mode, service migration mode, and fault isolation mode) to optimize both reliability and operational efficiency under different operating conditions
Solution Approach 2:
The patent changes RAS configuration parameters based on detected service migration states. When a service migration state is detected (indicated by CPU utilization rate changes), the system modifies RAS parameters such as fault isolation levels and service migration policies to prevent crashes while maintaining operational efficiency
2Reliability
If the server continues operating with fault components using default RAS configuration, then availability is maintained, but performance deteriorates and crash risk increases
Solution Approach 1:
The patent implements a feedback mechanism that continuously monitors CPU utilization rates before and after crashes to detect service migration states. This feedback information is used to adjust RAS configuration parameters in real-time, enabling the system to respond to changing conditions and prevent crashes while maintaining availability
Solution Approach 2:
The system automatically detects service migration states through CPU utilization monitoring and self-adjusts RAS configuration parameters without requiring external intervention. The server performs self-diagnosis and self-configuration to maintain availability while minimizing crash risk
Data Source
AI summary
A method and apparatus for configuring server maintainability, an electronic device and a non-transitory readable storage medium are provided by embodiments of the present application. The method includes: calculating a first utilization rate of a central processing unit, in response to a server starting and operating normally; determining a fault component, in response to the server crashing and restarting; calculating a second utilization rate of the central processing unit; determining a service migration state based on the first utilization rate and the second utilization rate; switching a server configuration mode according to the service migration state; and isolating the fault component in the server configuration mode. In the embodiments of the present application, by determining whether a service of a client has been migrated, and based on whether the service has been migrated, different server configuration modes are activated, the crash rate of the server may be reduced.


