CPU Fault Diagnosis Line for Partitioned Server Boards

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The complexity of CPU fault indication signal circuit topologies in high-performance servers increases circuit complexity and cabling difficulty, making it challenging to efficiently diagnose CPU faults and adapt to changing physical partitioning configurations.

Innovation Solution

A fault diagnosis system that includes a control unit, pull-up units, and switches on the node board, which forms a fault diagnosis line to detect CPU faults by pulling up fault indication signals and determining if they are below a diagnosis threshold, allowing for fault synchronization without involving the management board, thus simplifying the circuit and reducing cabling complexity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If CPU fault indication signals are separately transmitted to the server management board by using a backplane, then fault detection capability is improved, but circuit complexity and cabling difficulty increase

Engineering Contradiction:
Improvefault detection capabilityVSAvoidcircuit complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the fault indication signal transmission by implementing separate fault indication lines for each CPU on the node board, with dedicated pull-up resistors and switches for each CPU. This segmentation allows independent fault detection for each CPU while maintaining manageable circuit complexity through modular design.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary mechanism by using a dedicated fault indication signal line that connects CPUs to the management board through controlled switches and pull-up resistors. This intermediary structure enables reliable fault detection without requiring complex direct connections or signal convergence circuits on the management board.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If CPU fault indication signals are separately transmitted to the server management board by using a backplane, then fault detection capability is improved, but cabling difficulty increases

Engineering Contradiction:
Improvefault detection capabilityVSAvoidcabling difficulty
Core Design Contradiction:
ReliabilityVSEase of manufacture

Solution Approach 1:

The fault indication system is segmented into independent units for each CPU, with each unit having its own dedicated signal line, pull-up resistor, and switch. This segmentation simplifies the cabling architecture by eliminating the need for complex signal convergence and reduces the difficulty of manufacturing and assembly.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts the fault indication functionality from the main signal transmission path by implementing dedicated fault indication lines separate from data and control signals. This extraction simplifies the overall cabling requirements and makes the system easier to manufacture and maintain.

Inventive Principle:
Principle #2Taking out (Extraction)

3Reliability

If all isolation circuits and level shift circuits are placed on the server management board, then fault detection completeness is improved, but device complexity increases

Engineering Contradiction:
Improvefault detection completenessVSAvoidmanagement board circuit complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the isolation and level shifting functions by implementing them at the node board level for each CPU fault indication signal, rather than consolidating all such circuits on the management board. This segmentation distributes the circuit complexity across multiple boards, reducing the burden on the management board while maintaining complete fault detection capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts the isolation and level shifting circuits from the management board and relocates them to the node board, where they can service individual CPU fault indication signals. This extraction reduces the management board's circuit complexity while preserving complete fault detection functionality.

Inventive Principle:
Principle #2Taking out (Extraction)

4Adaptability or versatility

If the fault diagnosis system adapts to changing physical partitioning configurations, then system versatility is improved, but control complexity increases

Engineering Contradiction:
Improvephysical partitioning adaptationVSAvoidcontrol complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements dynamic adaptability by using controllable switches that can be programmed to connect or disconnect fault indication lines based on the current physical partitioning configuration. This dynamic control allows the system to adapt to changing configurations without requiring complex rewiring or hardware changes.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system adapts to different physical partitioning configurations by changing the state parameters of the switches (on/off positions) rather than changing the physical circuit topology. This parameter-based adaptation simplifies control complexity while maintaining system versatility.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP3835903B1Fault diagnosis system and server
Publication Date: 2023.01.04 XFUSION DIGITAL TECH CO LTD
  • EP3835903B1 patent drawingFigure 1
  • EP3835903B1 patent drawingFigure 2
  • EP3835903B1 patent drawingFigure 3

AI summary

A fault diagnosis system is disclosed, including: a control unit (102), a first management board (1011), a first pull-up unit (1031), a second pull-up unit (1032), a first pull-up switch (1051), a second pull-up switch (1052), and at least one central processing unit (CPU 1, ..., CPU 8), where the first pull-up unit (1031) is electrically connected to the first pull-up switch (1051), the second pull-up unit (1032) is electrically connected to the second pull-up switch (1052), the control unit (102) is electrically connected to the first management board (1052), and the control unit (102) is electrically connected to the first pull-up switch (1051) and the second pull-up switch (1032) separately, the control unit (102) is configured to receive physical partitioning information sent by the first management board (1011), the first pull-up unit (1031) and the second pull-up unit (1032) are configured to pull up a fault indication signal of a fault diagnosis line to obtain a target signal, the first management board (1011) is configured to detect whether a level of the target signal is lower than a diagnosis threshold, and when the level of the target signal is lower than the diagnosis threshold, determine that a faulty central processing unit (CPU 1, ..., CPU 8) exists in the at least one central processing unit (CPU 1, ..., CPU 8).