Server Memory Fault Detection and Bit-Redundancy Repair

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The frequent failure of server memories leads to low utilization rates due to direct replacement without effective repair, contributing significantly to server downtime and maintenance costs.

Innovation Solution

A memory processing method that includes determining a detection strategy based on memory hardware parameters, detecting faulty memories, and repairing them by replacing fault bits with redundant bits, utilizing different detection strategies to enhance accuracy and efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If faulty memories are directly replaced, then server stability is maintained, but memory utilization rate decreases

Engineering Contradiction:
Improveserver stabilityVSAvoidmemory utilization rate
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent applies preliminary action by detecting and identifying faulty memories before they cause server breakdowns. The system continuously monitors memory health parameters and detects potential failures early, allowing for proactive repair operations that prevent complete memory failure while maximizing memory utilization.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements discarding and recovering by repairing faulty memory modules rather than immediately replacing them. The repair process recovers usable functionality from failed memories through bit-level correction and redundancy activation, extending their service life and improving overall memory utilization rates.

Inventive Principle:
Principle #34Discarding and recovering

2Measurement precision

If comprehensive memory detection is performed, then detection accuracy improves, but detection time increases

Engineering Contradiction:
Improvedetection accuracyVSAvoiddetection time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies local quality by implementing selective detection strategies tailored to different memory modules based on their individual health parameters. Rather than uniformly detecting all memories, the system identifies at-risk modules through continuous monitoring and applies comprehensive detection only to those needing it, optimizing the balance between accuracy and time.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent implements dynamics by adapting detection intensity based on real-time memory health status. The system dynamically adjusts detection strategies from basic monitoring to comprehensive testing based on detected anomalies, memory age, and failure history, allowing flexible optimization of detection accuracy versus time consumption.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12609176B2Memory processing method based on a server and apparatus, processor and electronic device
Publication Date: 2026.04.21 INSPUR SUZHOU INTELLIGENT TECH CO LTD
  • US12609176B2 patent drawing
  • US12609176B2 patent drawing
  • US12609176B2 patent drawing

AI summary

Disclosed are a memory processing method based on a server and apparatus, a processor, and an electronic device. Said method comprises: determining, according to a target flag bit stored in a baseboard management controller, whether to perform a detecting and repairing operation on memories of a target server; when it is determined to perform a detecting and repairing operation on the memories of the target server, acquiring memory hardware parameters of a plurality of memories in the target server, and determining a detection strategy for each memory according to the memory hardware parameters; and detecting each memory according to the detection strategy for each memory to determine a faulty memory, and repairing the faulty memory. The present disclosure solves the technical problem in the related art that when a memory of a server fails, the faulty memory is directly replaced, resulting in a low utilization rate of the memory.