Server Recovery via Dynamic A2SU Reconfiguration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional large servers with a chip-to-chip topology connecting storage control (SC) and CPU chips do not allow all CPU chips to operate after an SC chip failure, leading to system instability and potential data loss.

Innovation Solution

Implementing a computer-implemented method that configures an Address-to-SC unit (A2SU) in each CPU chip based on the number of valid SC chips, pausing system operations, reconfiguring the A2SU when a change in SC chips is detected, and resuming operations once the reconfiguration is complete, ensuring seamless operation even with failed SC chips.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a conventional chip-to-chip topology is used to connect SC and CPU chips, then the system structure is simple, but the system cannot operate after an SC chip fails

Engineering Contradiction:
Improvesystem operation continuityVSAvoidtopology structure
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements a dynamic reconfiguration mechanism where the system automatically adapts its topology when an SC chip fails. The A2SU in each CPU chip is reconfigured to update the mapping between memory addresses and remaining valid SC chips, allowing the system to dynamically adjust to the changed hardware state and maintain operation continuity.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the operational parameters of the A2SU by clearing and reconfiguring the address-to-SC chip mappings when an SC chip failure is detected. This parameter change allows the system to redirect memory access requests to the remaining valid SC chips, enabling continued operation despite the failure.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If the system continues operation without reconfiguration after SC chip failure, then productivity is maintained, but data loss or system crashes occur

Engineering Contradiction:
Improvesystem operationVSAvoiddata integrity
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent performs preliminary detection of SC chip validity and proactively reconfigures the A2SU before actual data loss or system crash can occur. By detecting the failure and reconfiguring the address mappings in advance, the system prevents subsequent operations from accessing failed hardware, thereby maintaining both productivity and data integrity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements a feedback mechanism where the system continuously monitors the validity of SC chips and automatically triggers reconfiguration when a failure is detected. This closed-loop feedback ensures that the system adapts to hardware changes in real-time, maintaining reliable operation without data loss.

Inventive Principle:
Principle #23Feedback

3Reliability

If the A2SU is reconfigured immediately upon detecting SC chip change, then reliability is improved, but system operation is interrupted

Engineering Contradiction:
Improvesystem recoveryVSAvoidsystem downtime
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements a pause-resume mechanism that temporarily suspends system operation during the critical reconfiguration period and then quickly resumes operation once reconfiguration is complete. This approach minimizes the interruption time while ensuring that the reconfiguration is properly completed, balancing reliability improvement with minimal productivity loss.

Inventive Principle:
Principle #21Skipping (Rushing through)

Data Source

PatentUS11886345B2Server recovery from a change in storage control chip
Publication Date: 2024.01.30 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11886345B2 patent drawing
  • US11886345B2 patent drawing
  • US11886345B2 patent drawing

AI summary

Configuring an address-to-SC unit (A2SU) of each of a plurality of CPU chips based on a number of valid SC chips in the computer system is disclosed. The A2SU is configured to correlate each of a plurality of memory addresses with a respective one of the valid SC chips. In response to detecting a change in the number of valid SC chips, pausing operation of the computer system including operation of a cache of each of the plurality of CPU chips; while operation of the computer system is paused, reconfiguring the A2SU in each of the plurality of CPU chips based on the change in the number of valid SC chips; and in response to reconfiguring the A2SU, resuming operation of the computer system.