Proactive Memory Poison Recovery in Cloud Hosts

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Cloud computing systems face significant downtime and data loss due to uncorrectable memory errors, which can propagate and affect large numbers of virtual machines and applications, leading to poor customer experience.

Innovation Solution

The implementation of a proactive detection and containment system in a cloud computing environment that scans for memory errors, generates machine check exceptions, and isolates poisoned memory pages to prevent error propagation during live migration of virtual machines.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If proactive memory scanning and error detection is implemented, then memory error detection capability is improved, but system complexity increases

Engineering Contradiction:
Improvememory error detection capabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements preliminary action by continuously scanning memory in the background to detect errors before they affect VM operations. The scanner proactively identifies corrupted memory pages and generates MCEs ahead of time, allowing the system to prevent error propagation rather than reacting after failures occur.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary mechanism (the scanner and MCE handling layer) between the physical memory and the VMs. This intermediary detects memory errors, generates appropriate exceptions, and manages poisoned page isolation, thereby protecting VMs from direct exposure to memory faults while adding controlled complexity at the management layer.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If poisoned memory pages are isolated to prevent error propagation, then system reliability is improved, but memory access performance deteriorates

Engineering Contradiction:
Improveerror containment capabilityVSAvoidmemory access performance
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent applies segmentation by isolating only the specific poisoned memory pages that contain errors, rather than taking down the entire memory system or host. This granular approach contains errors to minimal affected areas, allowing the rest of the memory system to continue operating at full performance.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements local quality by applying different access characteristics to different memory pages. Poisoned pages are marked with special attributes (e.g., MAP_POPULATE, MADV_DONTFORK) that restrict access, while healthy pages maintain normal access patterns. This ensures that error containment does not unnecessarily impact performance of unaffected memory regions.

Inventive Principle:
Principle #3Local quality

3Measurement precision

If continuous memory scanning is performed, then error detection speed is improved, but computational overhead increases

Engineering Contradiction:
Improveerror detection speedVSAvoidcomputational overhead
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent implements periodic action by performing memory scans at scheduled intervals rather than continuously monitoring every memory access in real-time. This periodic scanning approach maintains adequate error detection capability while significantly reducing the computational overhead and energy consumption compared to continuous real-time monitoring.

Inventive Principle:
Principle #19Periodic action

Data Source

PatentUS12271253B2Memory error prevention by proactive memory poison recovery
Publication Date: 2025.04.08 GOOGLE LLC
  • US12271253B2 patent drawing
  • US12271253B2 patent drawing
  • US12271253B2 patent drawing

AI summary

The disclosed technology provides techniques, systems, and apparatus for proactively detecting, containing, and recovering from uncorrectable memory errors in distributed computing environment. An aspect of the disclosed technology includes scanning, by a scanner of a host machine, memory of the host machine for errors. After the scanner detects an error, the scanner may generate an error notification. The scanner may transmit the error notification to one or more processors of the host machine to implement mitigation techniques.