Cluster Resource Allocation via Service Assurance Manager
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In cluster computing environments, existing resource scheduling solutions fail to dynamically manage hardware resources effectively, leading to resource imbalances and unsatisfied service level agreements due to static allocation methods, which do not account for real-time needs of applications, resulting in poor user experience.
Innovation Solution
Implementing a node-level service assurance technique with a Service Assurance Manager (SAM) agent and master that dynamically allocates CPU, memory, and I/O resources across computing nodes to ensure compliance with derived node-level service level agreements, derived from application-level agreements, thereby ensuring proximate compliance with overall application-level service agreements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If static resource allocation is used, then resource distribution is simple to manage, but service level agreement compliance deteriorates due to inability to meet real-time application needs
Solution Approach 1:
The patent implements dynamic resource allocation where the resource allocation module continuously monitors application performance metrics and adjusts resource distribution in real-time based on actual application needs and current system state, transforming static allocation into a dynamic adaptive system that maintains SLA compliance
Solution Approach 2:
The system employs feedback mechanisms by monitoring application performance metrics and using this information to adjust resource allocation decisions, creating a closed-loop control system that responds to actual application behavior and ensures service level agreement compliance
2Reliability
If operating system or node level resource isolation is implemented, then service level agreement guarantee improves, but resource sharing capability deteriorates due to coarse granularity
Solution Approach 1:
The patent segments resource management into multiple hierarchical levels: cluster-level resource pool, node-level allocation, and application-level distribution. This multi-level segmentation enables fine-grained control at each level while maintaining overall system coordination, allowing both SLA guarantees and flexible sharing
Solution Approach 2:
The system applies different resource management strategies at different levels: global resource pooling at cluster level, isolation mechanisms at node level, and dynamic allocation at application level. Each level has optimized quality characteristics suited to its specific requirements, enabling both guarantee and sharing
3Productivity
If multiple applications share hardware resources without isolation, then resource utilization improves, but resource imbalance occurs leading to poor user experience
Solution Approach 1:
The patent creates a universal resource pool at the cluster level that serves multiple applications simultaneously, allowing resources to be dynamically allocated to different applications based on their current needs. This multi-functional resource pool maintains high utilization while preventing any single application from monopolizing resources
Solution Approach 2:
The resource allocation module acts as an intermediary between the shared hardware resources and multiple applications, mediating resource access and distribution. This intermediary ensures fair and balanced resource allocation among competing applications, preventing resource imbalance while maintaining high utilization
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Apparatuses, methods and storage medium associated with cluster computing are disclosed herein. In embodiments, a server of a computing cluster may include memory. input/output resources, and one or more processors to operate one of a plurality of application slaves of an application master; wherein the other application slaves are operated on other servers, which, together with the server, are members of the computing cluster. The server may further include a service assurance manager agent to manage allocation of the one or more processors, the memory and the input/output resources to the application slave, to assure compliance with a node level service level agreement, derived from an application level service level agreement, to contribute to proximate assurance of compliance with the application level service agreement; wherein the application level service agreement specifies the aggregate service level to be jointly provided by the application master and slaves. Other embodiments may be described or claimed.