Proactive Hardware Error Screening via Sensor-Based Health Monitoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for testing and screening hardware errors in computer processing systems are invasive, disruptive, and not suitable for continuous proactive error detection, leading to potential hardware failures and reduced quality of service.

Innovation Solution

A proactive hardware error screening system that uses a software-based scheduler to continuously monitor and assess the health status of computational units without interrupting normal operations, utilizing existing sensors to detect impending failures and allocate tasks to healthy units, thereby preventing mission-critical tasks from being affected by hardware failures.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If offline screening tests are performed to detect hardware errors, then hardware reliability is improved, but system availability and quality of service deteriorate due to computational units being taken offline

Engineering Contradiction:
Improvehardware reliabilityVSAvoidsystem availability
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system performs preliminary health assessments by analyzing sensor data (temperature, voltage, timing drift) before hardware failures occur. This allows proactive identification of at-risk computational units and enables preventive task migration, avoiding the need to take units offline for testing while maintaining system availability

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary health assessment system that uses existing sensors and telemetry data to indirectly evaluate hardware health without direct intervention. This intermediary layer provides failure predictions and enables preventive actions, resolving the contradiction between reliable detection and continuous operation

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If computational units are taken offline for error screening, then hardware errors are detected, but quality of service and system performance deteriorate

Engineering Contradiction:
Improveerror detection accuracyVSAvoidquality of service
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system enables computational units to self-report their health status through existing sensors and telemetry mechanisms. This self-service approach provides continuous health information without requiring external testing intervention, maintaining ease of operation while improving error detection capability

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

An intermediary health assessment module aggregates and analyzes data from multiple sensors (temperature, voltage, timing) to provide comprehensive error detection. This intermediary layer achieves high measurement precision through multi-parameter analysis while remaining transparent to ongoing operations, preserving quality of service

Inventive Principle:
Principle #24Intermediary (Mediator)

3Device complexity

If existing sensors are utilized for health monitoring, then additional hardware complexity is reduced, but measurement precision for early failure detection may be insufficient

Engineering Contradiction:
Improvehardware complexityVSAvoidfailure prediction accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent merges data from multiple existing sensors (temperature sensors, voltage regulators, timing drift measurements) into a unified health assessment model. This combination approach compensates for individual sensor limitations and achieves high measurement precision without adding dedicated failure detection hardware, maintaining low device complexity

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

Existing sensors designed for other purposes (temperature monitoring, voltage regulation, timing) are repurposed for health assessment and failure prediction. This multi-functionality approach maximizes measurement precision using existing hardware resources, avoiding additional complexity while improving failure detection capability

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250004896A1Method and apparatus to proactively screen hardware errors of a computer processing system
Publication Date: 2025.01.02 INTEL CORP
  • US20250004896A1 patent drawing
  • US20250004896A1 patent drawing
  • US20250004896A1 patent drawing

AI summary

Methods and apparatus to implement proactive hardware error screening are disclosed. In one embodiment, a computer processing system includes a plurality of computational units to execute tasks for one or more applications; a plurality of sensors collects measurement data of the plurality of computational units, to collect measurement data of the plurality of computational units; a data structure indicating hardware health statuses of the plurality of computational units determined based on the measurement data is stored in a storage; and the plurality of computational units is scheduled to perform task execution on the computer processing system for the one or more applications based on the hardware health statuses of the plurality of computational units indicated in the data structure, wherein a first computational unit is excluded from the task execution when a corresponding first hardware health status of the first computational unit indicates an impending hardware failure.