Proactive Hardware Error Screening via Sensor-Based Health Monitoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for testing and screening hardware errors in computer processing systems are invasive, disruptive, and not suitable for continuous proactive error detection, leading to potential hardware failures and reduced quality of service.
Innovation Solution
A proactive hardware error screening system that uses a software-based scheduler to continuously monitor and assess the health status of computational units without interrupting normal operations, utilizing existing sensors to detect impending failures and allocate tasks to healthy units, thereby preventing mission-critical tasks from being affected by hardware failures.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If offline screening tests are performed to detect hardware errors, then hardware reliability is improved, but system availability and quality of service deteriorate due to computational units being taken offline
Solution Approach 1:
The system performs preliminary health assessments by analyzing sensor data (temperature, voltage, timing drift) before hardware failures occur. This allows proactive identification of at-risk computational units and enables preventive task migration, avoiding the need to take units offline for testing while maintaining system availability
Solution Approach 2:
The patent introduces an intermediary health assessment system that uses existing sensors and telemetry data to indirectly evaluate hardware health without direct intervention. This intermediary layer provides failure predictions and enables preventive actions, resolving the contradiction between reliable detection and continuous operation
2Measurement precision
If computational units are taken offline for error screening, then hardware errors are detected, but quality of service and system performance deteriorate
Solution Approach 1:
The system enables computational units to self-report their health status through existing sensors and telemetry mechanisms. This self-service approach provides continuous health information without requiring external testing intervention, maintaining ease of operation while improving error detection capability
Solution Approach 2:
An intermediary health assessment module aggregates and analyzes data from multiple sensors (temperature, voltage, timing) to provide comprehensive error detection. This intermediary layer achieves high measurement precision through multi-parameter analysis while remaining transparent to ongoing operations, preserving quality of service
3Device complexity
If existing sensors are utilized for health monitoring, then additional hardware complexity is reduced, but measurement precision for early failure detection may be insufficient
Solution Approach 1:
The patent merges data from multiple existing sensors (temperature sensors, voltage regulators, timing drift measurements) into a unified health assessment model. This combination approach compensates for individual sensor limitations and achieves high measurement precision without adding dedicated failure detection hardware, maintaining low device complexity
Solution Approach 2:
Existing sensors designed for other purposes (temperature monitoring, voltage regulation, timing) are repurposed for health assessment and failure prediction. This multi-functionality approach maximizes measurement precision using existing hardware resources, avoiding additional complexity while improving failure detection capability
Data Source
AI summary
Methods and apparatus to implement proactive hardware error screening are disclosed. In one embodiment, a computer processing system includes a plurality of computational units to execute tasks for one or more applications; a plurality of sensors collects measurement data of the plurality of computational units, to collect measurement data of the plurality of computational units; a data structure indicating hardware health statuses of the plurality of computational units determined based on the measurement data is stored in a storage; and the plurality of computational units is scheduled to perform task execution on the computer processing system for the one or more applications based on the hardware health statuses of the plurality of computational units indicated in the data structure, wherein a first computational unit is excluded from the task execution when a corresponding first hardware health status of the first computational unit indicates an impending hardware failure.


