InfiniBand LID Reassignment for Non-Disruptive Topology Expansion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In XtremIO storage systems, expanding from a single brick configuration to a multi-brick system with InfiniBand switches requires a non-disruptive method to change Local Identifier (LID) assignments to prevent routing issues and disconnections, as duplicate LID assignments can occur when merging separate networks into a single fabric.
Innovation Solution
The method involves rebooting storage controllers, waiting for one to become the OpenSM master, changing cache files for LID assignments, restarting the system manager, and reconnecting with new LID assignments, allowing for the addition of new switches without service disruption by ensuring persistent and consistent LID changes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If storage controllers are connected to different switches and switches are interconnected to form a multi-brick system, then system expandability is improved, but routing issues and disconnections occur due to duplicate LID assignments
Solution Approach 1:
The patent applies preliminary action by detecting potential LID assignment conflicts before they cause network failures. The system proactively identifies duplicate LID assignments that would occur when connecting storage controllers to switches, and pre-resolves these conflicts by reassigning LIDs before the actual connection is made, preventing routing issues and disconnections from occurring
Solution Approach 2:
The patent introduces an intermediary mechanism in the form of a LID assignment management system that acts as a mediator between storage controllers and switches. This intermediary detects LID conflicts, coordinates LID reassignments across multiple controllers, and ensures consistent LID allocation throughout the multi-brick system, thereby preventing duplicate LID assignments and maintaining network reliability
2Reliability
If LID assignments are changed to prevent duplicate assignments in switched topology, then routing issues are prevented, but network disruption occurs during the change process
Solution Approach 1:
The patent applies preliminary action by preparing LID assignment changes in advance and coordinating them across multiple storage controllers before implementation. The system detects potential conflicts, plans the reassignment sequence, and executes changes in a coordinated manner that minimizes network disruption, allowing LID changes to be made without causing significant network unavailability
Solution Approach 2:
The patent maintains continuity of useful action by ensuring that LID assignment changes are performed in a coordinated sequence that keeps the network operational. The system orchestrates LID reassignments across controllers and switches in a manner that maintains network connectivity throughout the transition, preventing complete network disruption while routing correctness is being established
3Reliability
If storage controllers are rebooted and system manager is restarted to apply new LID assignments, then consistent LID changes are achieved, but service disruption occurs during the reboot process
Solution Approach 1:
The patent applies preliminary action by coordinating reboot schedules and LID assignment applications across storage controllers and the system manager before actual reboots occur. The system plans the sequence of reboots and LID changes in advance, ensuring that controllers are rebooted in an order that minimizes service disruption and that LID assignments are applied consistently across all components
Solution Approach 2:
The patent implements periodic action by scheduling reboots and LID assignment updates in a staged, periodic manner rather than all at once. The system performs LID changes and reboots in sequential phases across different controllers, allowing each phase to complete and stabilize before proceeding to the next, thereby reducing overall service downtime while achieving consistent LID assignments system-wide
Data Source
AI summary
Moving from a back-to-back topology to a switched topology in an InfiniBand network includes, prior to connecting a switch for a first storage controller in the network and during reboot of the first storage controller, waiting for a second storage controller in the network to become master, and upon the second storage controller becoming master, changing cache files for local ports on the first storage controller regarding adjacent ports' LID assignments. An aspect further includes restarting a system manager for the first storage controller, connecting the first storage controller to the system with new LID assignments provided by changed files on first storage controller, and upon the first storage controller becoming active, rebooting the second storage controller, changing the LID assignments in the active storage controller, and adding new switches to the system.


