Server RAS Configuration for Crash Recovery and Fault Isolation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Server faults, particularly crashing faults during starting and operation, affect performance and operational efficiency, and existing RAS functions do not adequately address these issues, leading to potential server crashes.

Innovation Solution

A method and apparatus for configuring server maintainability that calculates CPU utilization rates before and after crashes, determines fault components, and switches server configuration modes based on service migration states to isolate faults and reduce crash rates.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If the server operates with default RAS configuration parameters, then the server can continue operating with fault components, but server performance and operational efficiency deteriorate and may lead to server crashing

Engineering Contradiction:
Improveserver reliabilityVSAvoidoperational efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent implements dynamic adjustment of RAS configuration parameters based on real-time service migration states. The system monitors CPU utilization rates before and after crashes, determines service migration states, and dynamically switches between different RAS configuration modes (default mode, service migration mode, and fault isolation mode) to optimize both reliability and operational efficiency under different operating conditions

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes RAS configuration parameters based on detected service migration states. When a service migration state is detected (indicated by CPU utilization rate changes), the system modifies RAS parameters such as fault isolation levels and service migration policies to prevent crashes while maintaining operational efficiency

Inventive Principle:
Principle #35Parameter changes

2Reliability

If the server continues operating with fault components using default RAS configuration, then availability is maintained, but performance deteriorates and crash risk increases

Engineering Contradiction:
Improveserver availabilityVSAvoidcrash risk
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent implements a feedback mechanism that continuously monitors CPU utilization rates before and after crashes to detect service migration states. This feedback information is used to adjust RAS configuration parameters in real-time, enabling the system to respond to changing conditions and prevent crashes while maintaining availability

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system automatically detects service migration states through CPU utilization monitoring and self-adjusts RAS configuration parameters without requiring external intervention. The server performs self-diagnosis and self-configuration to maintain availability while minimizing crash risk

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250225023A1Server maintainability configuration method and apparatus, electronic device and storage medium
Publication Date: 2025.07.10 INSPUR SUZHOU INTELLIGENT TECH CO LTD
  • US20250225023A1 patent drawing
  • US20250225023A1 patent drawing
  • US20250225023A1 patent drawing

AI summary

A method and apparatus for configuring server maintainability, an electronic device and a non-transitory readable storage medium are provided by embodiments of the present application. The method includes: calculating a first utilization rate of a central processing unit, in response to a server starting and operating normally; determining a fault component, in response to the server crashing and restarting; calculating a second utilization rate of the central processing unit; determining a service migration state based on the first utilization rate and the second utilization rate; switching a server configuration mode according to the service migration state; and isolating the fault component in the server configuration mode. In the embodiments of the present application, by determining whether a service of a client has been migrated, and based on whether the service has been migrated, different server configuration modes are activated, the crash rate of the server may be reduced.