Server Farm Patching System for Expedited Failure Resolution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Server farm patching systems face disruptions during peak hours due to software regressions that cause increased processing cycles and memory usage, leading to delayed resolution of performance failures and strain on limited computing resources.
Innovation Solution
The system allows developers to override off-peak patching schedules by expediting the installation of specific patches that resolve performance failures, preventing intermediate builds from being installed to reduce resource demand and network traffic, thereby minimizing disruptions and resource utilization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If software patches are installed chronologically during off-peak hours, then service disruption is minimized, but resolution of performance failures is delayed
Solution Approach 1:
The patching system dynamically adjusts the patching schedule based on detected performance failures. When a performance failure is detected, the system transitions from adhering to the off-peak schedule to expediting the relevant patch installation, thereby resolving the contradiction between maintaining service continuity and reducing patch resolution time.
Solution Approach 2:
The system changes the timing parameter of patch installation from fixed off-peak hours to variable timing based on performance monitoring. When performance failures are detected, the patch installation time parameter is adjusted to prioritize expedited installation, resolving the contradiction between service continuity and resolution time.
2Stability of the object's composition
If intermediate builds are installed in sequence, then system stability is maintained, but computing resources are unnecessarily consumed
Solution Approach 1:
The system extracts and skips the installation of intermediate builds that are not needed to resolve the detected performance failure. By identifying the specific patch required and bypassing unnecessary intermediate builds, the system maintains stability while reducing computing resource consumption.
Solution Approach 2:
Instead of installing all intermediate builds in sequence (excessive action), the system installs only the specific build containing the patch needed to resolve the performance failure (partial action), thereby reducing unnecessary resource consumption while maintaining system stability.
3Loss of time
If patches are expedited during peak hours, then performance failures are resolved faster, but service disruption increases
Solution Approach 1:
The system applies different patching strategies to different server farms based on their specific performance failure status. Affected server farms receive expedited patches while unaffected server farms continue with normal off-peak scheduling, thereby resolving performance failures faster without causing widespread service disruption.
Solution Approach 2:
The patching system segments the server farm population into affected and unaffected groups based on performance monitoring. This segmentation allows expedited patching to be applied locally to affected servers while maintaining normal scheduling for others, reducing overall service disruption while resolving critical issues faster.
Data Source
AI summary
A system to reduce strain on server farm computing resources by over-riding “off-peak” patching schedules in response to performance failures occurring on a server farm. Embodiments disclosed herein determine a patching schedule for causing builds of patches to be sequentially installed on server farms during an off-peak usage time-range. Responsive to a performance failure occurring on the server farm, embodiments disclosed herein identify a particular patch that is designed to resolve the performance failure. Then, the patching schedule is over-ridden to expedite an out-of-sequence installation of whichever build is first to include the particular patch. Because resolution of the performance failure is expedited, the impact of the performance failure on the computing resources of the server farm is reduced as compared to existing server farm patching systems.


