Automated Grid Compute Node Cleaning via Snapshot Comparison
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Grid compute nodes become inefficient and unusable over time due to residual grid-enabled software applications that are not properly removed after usage, leading to increased disk space usage, port occupation, and memory issues.
Innovation Solution
An automated method for cleaning grid compute nodes involves taking initial snapshots, comparing current states to predefined criteria, and adjusting nodes by rebooting, restarting grid containers, or installing new disk images to ensure compliance with system criteria, which includes file and port management, disk space, and memory usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If grid-enabled software applications are deployed and used on grid compute nodes, then computing tasks can be executed, but residual application parts remain on the nodes causing performance degradation over time
Solution Approach 1:
The system takes an initial snapshot of the compute node's state before deploying applications, and periodically compares current snapshots against this baseline. This preliminary comparison mechanism enables early detection of residual application parts before they cause significant performance degradation, allowing proactive cleaning actions.
Solution Approach 2:
The system continuously monitors the compute node state by taking snapshots and comparing them against the initial state and predefined criteria. This feedback loop identifies when application removal is needed and triggers appropriate cleaning actions, ensuring node performance is maintained without manual intervention.
2Ease of operation
If manual application removal is performed, then some cleaning can be achieved, but it is error-prone and incomplete without restarting the node
Solution Approach 1:
The system automatically performs the cleaning operation by comparing snapshots against criteria and executing appropriate removal actions without requiring manual intervention. This self-service approach eliminates human error while maintaining thoroughness through automated verification against predefined completeness criteria.
Solution Approach 2:
The patent replaces manual mechanical operations with automated software-based snapshot comparison and analysis. Instead of manual inspection and removal, the system uses computational comparison of data snapshots to identify and remove residual application parts, achieving both ease of operation and cleaning completeness.
3Reliability
If the node is rebooted to remove applications, then complete removal can be achieved, but service interruption occurs
Solution Approach 1:
The system attempts partial removal actions first by comparing snapshots and identifying removable application parts, only resorting to complete node reboot if necessary. This partial action approach minimizes service interruption while maintaining the option of complete removal when required.
Solution Approach 2:
The system changes the state parameters by taking snapshots at different points in time and comparing them to identify changes. By monitoring parameter changes in the node state, the system can determine when application removal is needed and what actions are required, minimizing unnecessary reboots.
4Ease of operation
If application developers manually remove applications, then some cleaning is achieved, but they often forget to remove all parts leading to gradual node degradation
Solution Approach 1:
The system provides continuous feedback by comparing current node state against the initial state and predefined criteria. This feedback mechanism automatically detects when application parts remain and triggers cleaning actions, eliminating the need for developer memory and preventing gradual node degradation.
Solution Approach 2:
The system performs self-service cleaning by automatically comparing snapshots, identifying residual application parts, and executing removal actions without requiring developer intervention. This eliminates human error and ensures complete removal of application parts.
Data Source
AI summary
A method includes, in a network of interconnected grid compute nodes, storing system criteria for a first grid compute node, storing an initial snapshot of the first grid compute node, comparing a current snapshot of the first grid compute node with the initial snapshoot to identify parts of the current snapshot that do not meet the criteria, and adjusting the first compute node to meet the criteria.


