Network Controller Abort Mechanism for Sequential Firmware Updates
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Network device upgrades in groups often result in significant downtime and delayed troubleshooting due to simultaneous rebooting, which can lead to prolonged disruptions and delays in identifying and rectifying faults during the update process.
Innovation Solution
A network controller initiates and manages a sequential group update process, allowing for user-controlled or condition-based abortion of the update process, removing subsequent devices from the update list and rolling back firmware to prevent further disruptions and ensure all devices return to a functional state.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If all network devices in a group are upgraded simultaneously, then the upgrade process is completed efficiently, but network downtime increases and availability is disrupted
Solution Approach 1:
The patent divides the group of network devices into multiple subsets and upgrades them sequentially rather than simultaneously. The controller manages multiple upgrade processes for different subsets, allowing some devices to remain operational while others are upgraded, thus maintaining network availability while completing the upgrade process.
2Reliability
If sequential updates are applied to network devices, then network availability is maintained, but the update process takes longer and productivity decreases
Solution Approach 1:
The patent implements a dynamic upgrade approach where the controller can adjust the number and composition of device subsets being upgraded based on network conditions and requirements. The system can transition between more parallel updates when network capacity allows and more sequential updates when availability is prioritized, optimizing both productivity and reliability dynamically.
3Productivity
If the upgrade process continues until completion, then all devices are updated, but troubleshooting delays increase when faults occur
Solution Approach 1:
The patent implements preliminary abort capabilities that allow the upgrade process to be stopped at any point before completion. When a fault is detected or suspected, the system can abort the ongoing upgrade process and roll back to the previous stable state, enabling immediate troubleshooting without waiting for the entire upgrade process to complete, thus reducing troubleshooting delays.
4Reliability
If multiple retry attempts are made for failed device updates, then update reliability improves, but the total update time increases significantly
Solution Approach 1:
The patent applies a limited retry strategy where each device subset is attempted a specific number of times (e.g., twice) before moving on to the next subset. This partial action approach ensures that critical failures are addressed through retries while preventing excessive delays from prolonged retry attempts on all devices, balancing update reliability with acceptable total update duration.
Data Source
AI summary
Examples of the present disclosure relate to updating network devices belonging to a group of network devices. In one aspect, a network controller coupled to the network devices of the group of network access devices, responsive to a first command, initiates a group update process for the network devices of the group is to update the network devices of the group sequentially according to an ordered list. Responsive to a second command during the group update process while a firmware image of a particular network device is updated, the network controller aborts the group update process for the network devices of the group. Aborting the group update process comprises removing a first subset of network devices subsequent to the particular network device in the ordered list from the ordered list such that the firmware image of the first subset of network devices will not be updated and rolling back the firmware image of the particular network device.


