Consolidated Failure Reporting for Partitioned Server Hardware
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional server systems face challenges in accurately managing hardware failures when using physical partitioning, leading to redundant failure reports and increased costs due to the need for additional software and hardware resources, which complicates the management of shared and monopolized resources.
Innovation Solution
A system that divides hardware resources into physical partitions, with each partition capable of operating independently, includes a failure notification mechanism to detect and report failures accurately, eliminating redundant reports by determining if a failure is shared across partitions and creating a common failure report, thus improving reliability without requiring additional dedicated devices or increasing manufacturing costs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If physical partitioning function is activated to divide hardware resources into multiple physical partitions, then resource flexibility and independent operation of partition blocks are improved, but redundant failure reports are generated when failures occur in shared hardware resources
Solution Approach 1:
The patent merges failure information from multiple physical partitions that share the same hardware resource into a single consolidated failure report. When a failure occurs in a shared hardware resource, the failure information is collected from all affected physical partitions and combined into one report, eliminating redundant notifications while maintaining comprehensive failure coverage.
2Reliability
If conventional failure management is implemented in partitioned systems, then system reliability is maintained, but additional software and hardware resources are required increasing manufacturing costs
Solution Approach 1:
The patent implements a universal failure management mechanism where the same failure information collection and consolidation process handles both shared and non-shared hardware resources across all partition blocks. This multi-functional approach maintains system reliability without requiring separate dedicated failure management systems for different scenarios, thereby reducing manufacturing costs.
Solution Approach 2:
The failure management system uses existing hardware resource management structures and firmware components to collect and process failure information, rather than requiring additional dedicated hardware or software resources. The system leverages existing partition management infrastructure to provide failure management services, reducing overall system cost while maintaining reliability.
3Measurement precision
If each physical partition independently manages hardware failures, then accurate failure detection is achieved, but management complexity increases due to multiple independent failure reports
Solution Approach 1:
The patent segments the failure management process into two distinct stages: (1) independent failure detection at the physical partition level, and (2) centralized failure information consolidation at the system level. This segmentation allows each partition to accurately detect failures independently while the consolidation stage reduces management complexity by aggregating reports and eliminating duplicates.
Solution Approach 2:
The patent introduces an intermediary failure information consolidation mechanism that acts as a mediator between individual physical partition failure detectors and the overall system failure management. This intermediary collects failure information from multiple partitions, consolidates redundant reports, and presents a unified failure status, thereby reducing management complexity while preserving detection accuracy.
Data Source
AI summary
An information processing apparatus includes partitioning mode information retaining section, hardware resource management information retaining section, failure notifying section, operation mode detecting section, shared hardware resource judging section, common failure report creating section for creating, if operation in the partitioning mode is detected and a hardware resource in which the failure occurrence has been detected is judged to be a shared resource, a common failure report on the basis of the detection of failure occurrence output by the notifying sections of physical partitions that share the shared hardware resource. That can avoid excessive report of a failure occurred even at a shared hardware resource, making it possible to grasp the accurate number of failure occurrence and to manufacture in a low cost.


