Partitioned I/O Fabric Flush Control for Shared Resource Errors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing server computers with partitioned processors face challenges in managing error handling for shared resources, necessitating effective stability and isolation of partitioned components.
Innovation Solution
Implementing a system with a baseboard management controller (BMC) and security controller to manage error handling, using cryptographic verification and secure boot processes to ensure system integrity, and employing a security chip to mediate firmware updates and control system resets, while maintaining logical and physical isolation between host components and security cards.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If processors are partitioned for independent control and management, then flexibility of server computers is improved, but error handling management for shared resources becomes more complex
Solution Approach 1:
A hypervisor is introduced as an intermediary layer between the physical hardware and partitioned virtual machines. The hypervisor manages shared resources uniformly across multiple partitions, abstracting the complexity of error handling from individual partition management. When errors occur in shared resources, the hypervisor mediates the error response, determining which partitions are affected and coordinating the appropriate error handling actions, thereby simplifying the overall error management architecture despite the presence of multiple partitions.
2Productivity
If shared resources are used between host partitions, then resource utilization efficiency is improved, but system stability and isolation are compromised
Solution Approach 1:
The system segments error handling into partition-specific and shared resource-specific components. Each partition maintains its own error handling policies and isolation boundaries, while shared resources have dedicated error management mechanisms. This segmentation allows efficient resource sharing while maintaining stability through clear separation of error domains. The hypervisor enforces these segmentation boundaries, ensuring that errors in one partition or shared resource do not propagate uncontrollably to other partitions.
3Reliability
If configuration is limited to manufacturer design, then system reliability is improved, but flexibility of server computers deteriorates
Solution Approach 1:
The system transitions from static manufacturer-defined configurations to dynamic runtime reconfigurability. Virtual machines can be created, migrated, and deleted dynamically through the hypervisor, allowing the server to adapt its configuration based on workload requirements. This dynamic capability is built upon a reliable foundation of standardized interfaces and controlled access mechanisms that ensure system stability even as configurations change. The hypervisor provides controlled dynamic reconfiguration while maintaining underlying system reliability through consistent error handling and resource management protocols.
Data Source
AI summary
Computing systems and associated methods are described for managing error conditions associated with shared resources in a partitioned computing system. In some examples, a flush signal is generated by a controller of an expansion card or other component of the partitioned computing system responsive to detecting a communication issue between the component and a first host partition. Responsive to the flush signal, an Input/Output (I/O) fabric of the component terminates active memory transactions for the first host partition and transmits a success status signal associated with the active memory transactions to an initiator of the active memory transactions. The controller may then reset a port coupled to the first host partition to prepare the component for re-establishing a link with the first host partition.


