Redundant Fabric Access for Storage Engine Compute Nodes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current storage systems with individualized fabric access for compute nodes face performance bottlenecks and reduced scalability when attempting to access memory across nodes, especially in scenarios where host adapters fail, leading to increased latency and reduced system availability.
Innovation Solution
Implementing a storage engine with redundant fabric access, where each compute node has access to both its own and the other compute node's fabric adapter, enabling native memory access operations and multi-initiator capabilities to simplify and accelerate atomic operations across nodes, thereby improving system availability and scalability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If individualized fabric access is implemented for compute nodes, then system structure is simplified, but memory access latency increases and system availability decreases
Solution Approach 1:
Each compute node is configured with distinct fabric access characteristics: primary fabric adapter for normal operations, secondary fabric adapter for failover. This local differentiation allows optimized access paths while maintaining simplicity in the overall fabric topology.
Solution Approach 2:
The system pre-configures redundant fabric adapters and establishes failover pathways before failures occur. Compute nodes maintain connectedness to both primary and secondary fabric adapters, with readiness to switch access paths, thereby eliminating latency during actual memory access operations.
2Device complexity
If individualized fabric access is implemented for compute nodes, then device complexity is reduced, but system availability decreases due to host adapter failures
Solution Approach 1:
The system implements redundant fabric adapters as a protective measure against host adapter failures. Each compute node maintains connectivity through both primary and secondary fabric adapters, cushioning the system against single-point failures and ensuring continuous availability.
Solution Approach 2:
The system dynamically changes the operational state of fabric adapters based on failure conditions. Upon detecting host adapter failure, the system transitions from single-adapter mode to dual-adapter mode, altering the connectivity parameters to maintain system availability while preserving structural simplicity.
3Productivity
If redundant fabric access is implemented with multi-initiator capabilities, then memory access efficiency improves, but device complexity increases
Solution Approach 1:
Fabric adapters are designed with multi-initiator capabilities, allowing them to serve multiple compute nodes simultaneously. This universal functionality enables efficient memory access operations across the fabric while consolidating adapter resources, thereby improving productivity without proportionally increasing device complexity.
Solution Approach 2:
The system merges the functionality of multiple fabric adapters into a unified redundant access architecture. Compute nodes leverage both primary and secondary fabric adapters as interconnected resources, combining their capabilities to achieve efficient memory access while maintaining a manageable level of device complexity through shared fabric infrastructure.
4Reliability
If redundant fabric adapters are implemented, then failure tolerance increases, but system complexity increases
Solution Approach 1:
Redundancy is implemented locally at each compute node through configuration rather than through additional physical hardware. Each node is programmed to utilize both primary and secondary fabric adapters, achieving failure tolerance through software-defined redundancy that minimizes the increase in physical device complexity.
Solution Approach 2:
The system creates a logical copy of fabric adapter functionality through the secondary adapter configuration. Instead of duplicating entire adapter assemblies, the system replicates the essential fabric access capabilities through configured pathways, achieving failure tolerance while controlling the increase in system complexity.
Data Source
AI summary
A storage system includes a storage engine having a first compute node, a second compute node, a first fabric adapter, and a second fabric adapter, the first compute node having a first memory and the second compute node having a second memory. The first compute node is connected to both the first and second fabric adapters, and the second compute node is connected to both the second and first fabric adapters. Both fabric adapters are configured to perform atomic operations on a memory of its respective compute node, and each fabric adapter contains a multi-initiating module configured to enable both the first compute node and the second compute node to initiate memory access operations on its respective memory.


