Auto-Clustering of HPC Services for Easier Cluster Deployment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge of making high performance computing (HPC) accessible to non-expert users with limited IT budgets and capabilities, particularly in transitioning from monolithic workstation-based platforms to HPC environments, is significant due to the complexity and cost associated with traditional HPC systems.
Innovation Solution
A system and method for automating the deployment of clusters of clusterable services, utilizing a controller to manage and configure compute, storage, and networking resources, including templates and scheduling, to provide seamless scaling and efficient resource management across GPU clusters and mixed hardware environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Power
If traditional HPC systems are deployed, then high performance computing capability is achieved, but system complexity and cost increase significantly
Solution Approach 1:
The patent segments HPC functionality into discrete, modular services that can be independently deployed and managed. Each service represents a specific computing function that can be clustered and scaled separately, reducing overall system complexity while maintaining high computing capability.
Solution Approach 2:
The patent creates a universal service deployment framework that can accommodate multiple types of HPC workloads and applications through a common infrastructure. This multi-functional approach allows diverse computing tasks to be handled by the same system architecture, reducing complexity.
2Power
If traditional HPC systems are deployed, then high performance computing capability is achieved, but ease of use deteriorates for non-expert users
Solution Approach 1:
The patent implements self-service mechanisms where the system automatically performs service discovery, clustering, and configuration without requiring expert user intervention. The automated service deployment framework handles complex setup tasks independently, making HPC accessible to non-expert users while maintaining high computing capability.
Solution Approach 2:
The patent performs preliminary configuration and setup actions automatically during system initialization. Service templates and clustering parameters are pre-configured, allowing users to deploy HPC capabilities with minimal effort while the system handles complex configuration details in advance.
3Ease of manufacture
If automated service clustering is implemented, then ease of deployment is improved, but service coordination complexity increases
Solution Approach 1:
The patent introduces an intermediary service registry that mediates between individual services and the clustering infrastructure. This registry automatically manages service discovery, registration, and coordination, simplifying deployment while handling the complexity of service interactions through a standardized intermediary layer.
Solution Approach 2:
The patent implements feedback mechanisms where services automatically report their status and capabilities to the clustering system. This feedback loop enables automated coordination and dynamic adjustment of service clusters, improving ease of deployment while managing coordination complexity through real-time system awareness.
Data Source
AI summary
A system can be configured to automatically deploy clusters of clusterable services. For example, controller can deploy a plurality of copies of an application, and these applications can interdepend on each other. The controller can also configure a scheduler to manage (which may include load balancing) these applications. A service template used by the controller can include clustering rules, and these clustering rules can tell the controller how to connect those services. The clustering rules can be a set of logic instructions and/or templates that provide for the deployment of a service to a plurality of resources. Coupling instructions in the clustering rules define the coordination and interaction of separately booked physical and/or virtual resources and set up dependencies. The clustering rules define the use of information to scale up or scale down resources being used by a service.


