Dynamic Service Management in Administration Clusters
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
High-performance computing clusters face scalability issues due to rapid growth in equipment numbers outpacing processing capacity, leading to limitations in administrative tools, hardware, and operating system parameters, which hinder high availability and efficient maintenance.
Innovation Solution
A method for dynamic service management in administration clusters that includes obtaining configuration data for virtual machines, synchronizing data between file servers and virtual machines, and dynamically redistributing services across processing stations to ensure high availability and facilitate maintenance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Power
If the number of equipment in the cluster is increased to enhance computing capacity, then the processing capability is improved, but the scalability of administrative tools deteriorates due to increased complexity and resource consumption
Solution Approach 1:
The patent segments administrative services into multiple independent service nodes distributed across the cluster. Each service node manages specific administrative functions, dividing the overall administrative complexity into manageable units that can be independently scaled and maintained.
Solution Approach 2:
The patent implements dynamic service placement and migration mechanisms that allow administrative services to be automatically redistributed across service nodes based on real-time cluster state. This dynamic adaptation enables the system to maintain scalability as equipment numbers change.
2Ease of operation
If static service distribution is used to simplify administration, then ease of operation is improved, but service availability deteriorates when failures occur
Solution Approach 1:
The patent transitions from static to dynamic service distribution, enabling automatic service migration when failures occur. The system continuously monitors service health and dynamically redistributes affected services to healthy nodes, maintaining high availability while preserving operational simplicity through automation.
Solution Approach 2:
The patent implements feedback mechanisms that monitor service status and trigger automatic migration actions when failures are detected. This closed-loop system maintains service availability by continuously adapting to failure conditions without requiring manual intervention.
3Device complexity
If centralized administration is implemented to reduce complexity, then device complexity is reduced, but processing latency increases due to single point of congestion
Solution Approach 1:
The patent segments the centralized administration function into multiple distributed service nodes. Each service node handles administrative tasks locally, reducing the distance data must travel and eliminating the single point of congestion, thereby reducing processing latency while maintaining manageable complexity through functional segmentation.
4Power
If hardware capacity is increased to handle more equipment, then processing capability is improved, but input/output flow limitations worsen due to hardware constraints
Solution Approach 1:
The patent segments administrative processing across multiple service nodes, distributing the input/output load rather than concentrating it at a single point. This segmentation enables the system to handle increased equipment numbers by spreading hardware constraints across multiple nodes rather than being limited by a single hardware bottleneck.
Data Source
Figure 1
Figure 2~3
Figure 4~6
AI summary
The invention is aimed in particular at the dynamic management of services in an administration cluster comprising at least one central administration node and processing stations adapted for implementing services targeting several calculation nodes, the central administration node distributing an execution of services in processing stations. At least one service is implemented via at least one virtual machine executed by a processing station. If an anomaly is detected during the monitoring (435) of an execution parameter of the virtual machine executed by a first processing station, the virtual machine is stopped (450) and restarted (455) in a second processing station, the restarting of the virtual machine being at least partially based on configuration data for the virtual machine.