Cluster Metrics Access for Agile Scheduling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Large-scale data processing systems face challenges in dynamically balancing loads and detecting failures in clusters due to resource overloads, congestion, and machine failures, requiring timely and responsive adjustments in computational and communicational scheduling.

Innovation Solution

The implementation of a cluster-wide operational metrics access system using L4 protocols for subscribing and de-subscribing metrics, allowing nodes to access and share metrics across the cluster, with timestamped updates to detect disruptions and adjust resource allocation, and a dedicated communication channel for low-latency metric delivery.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional data processing tools are used for big data, then system complexity is reduced, but processing capability and speed become inadequate

Engineering Contradiction:
Improvedata processing capabilityVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the monolithic data processing system into distributed microservices running on multiple container instances across a cluster. Each microservice handles specific data processing tasks independently, enabling the system to process big data through parallel execution while maintaining manageable complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from single-machine processing to multi-dimensional distributed processing across a cluster of machines. By adding spatial distribution as a new dimension, the system achieves enhanced processing capability through parallel computation while managing complexity through standardized communication protocols and orchestration layers.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If clusters grow larger in size, then processing capability increases, but timeliness of failure detection and load balancing decreases

Engineering Contradiction:
Improveprocessing capabilityVSAvoidfailure detection time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent implements continuous feedback mechanisms through health check endpoints and metrics collection that monitor cluster node status in real-time. This feedback loop enables timely detection of failures and imbalances even as clusters scale, allowing the orchestration system to reactively adjust task allocation and proactively rebalance loads before performance degradation occurs.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent employs preliminary actions through pre-configured health checks, metrics subscriptions, and proactive rebalancing mechanisms that detect and respond to failures before they impact processing capability. By establishing monitoring and response protocols in advance, the system maintains timely failure detection regardless of cluster size.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If clusters grow larger in size, then processing capability increases, but responsiveness of scheduling adjustments decreases

Engineering Contradiction:
Improveprocessing capabilityVSAvoidscheduling responsiveness
Core Design Contradiction:
ProductivityVSSpeed

Solution Approach 1:

The patent implements dynamic scheduling through the Kubernetes orchestration system that continuously monitors cluster conditions and adjusts task allocation in real-time. The system dynamically rebalances loads across nodes based on current resource availability and task priorities, maintaining scheduling responsiveness even in large clusters through event-driven architecture and incremental updates.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent employs preliminary scheduling decisions through pre-computed task placement strategies and resource allocation policies that are established before runtime. These preliminary configurations enable rapid scheduling responses by reducing the computational overhead during actual task allocation, allowing the system to maintain responsiveness as cluster size increases.

Inventive Principle:
Principle #10Preliminary action

4Productivity

If more distributed hosts and internetworking machinery are coordinated, then processing capability increases, but coordination complexity increases

Engineering Contradiction:
Improveprocessing capabilityVSAvoidcoordination complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements a universal orchestration platform (Kubernetes) that provides multi-functional coordination across diverse distributed hosts and internetworking machinery. This universal system handles scheduling, resource management, fault tolerance, and load balancing through standardized APIs and protocols, enabling coordination of large clusters while managing complexity through abstraction and reuse of common coordination patterns.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10547527B2Apparatus and methods for implementing cluster-wide operational metrics access for coordinated agile scheduling
Publication Date: 2020.01.28 INTEL CORP
  • US10547527B2 patent drawing
  • US10547527B2 patent drawing
  • US10547527B2 patent drawing

AI summary

Apparatus, methods, and system for implementing cluster-wide operational metrics access for coordinated agile scheduling. One embodiment of the apparatus includes a memory to store instructions; a processing circuitry to execute instructions; and an interface circuitry. The interface circuitry to provide metrics associated with the apparatus to one or more subscriber nodes or network components in a managed cluster and to subscribe, via a metrics subscription request, to receive from one or more publisher nodes or network components in the managed cluster, metrics associated with the one or more publisher nodes or network components. The metrics to be stored in a dedicated location of the memory. The provision and subscription of metrics may be made using new protocols added to Layer 4 or transport layer of a network communication model and/or over a dedicated communication channel. The dedicated communication channel may be of low bandwidth with fixed priority and deterministic latency.