Distributed Job Scheduler Using Pervasive State Estimation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing distributed compute systems, such as Hadoop, face challenges in precisely estimating server states due to volatile performance caused by factors like hardware issues and VM migrations, leading to resource congestion and reduced productivity, and require a framework that updates beliefs with minimal production loss.

Innovation Solution

A decision-theoretic approach extending the Pervasive Diagnosis framework to continuous domains, where server states are modeled as random variables, allowing for autonomous and efficient state estimation and task scheduling that maximizes immediate and future production by using a cost function that balances production utility and information gain.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the scheduler uses a static model of compute node capabilities, then the scheduling process is simple and fast, but the model becomes inaccurate over time due to volatile performance from hardware issues and VM migrations

Engineering Contradiction:
Improvemodel accuracyVSAvoidscheduling complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system implements feedback by continuously monitoring state variables from compute nodes and using pervasive state estimation to update the capability model dynamically. The scheduler receives feedback about actual node performance and adjusts its beliefs about node capabilities accordingly, resolving the contradiction between model accuracy and scheduling simplicity.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The capability model transitions from static to dynamic through continuous updating based on observed state variables. The system adapts the model of compute node capabilities over time to reflect changing conditions such as hardware issues and VM migrations, maintaining accuracy without requiring complete re-scheduling.

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If the scheduler frequently updates the capability model by running diagnostic tasks, then the model accuracy improves, but production throughput decreases due to resource congestion

Engineering Contradiction:
Improvestate estimation accuracyVSAvoidcluster throughput
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

Instead of running comprehensive diagnostic tasks on all nodes, the system performs partial updates only where needed based on detected discrepancies. The pervasive state estimation selectively updates capability models for specific compute nodes rather than进行全面 diagnostics, reducing the impact on production throughput while maintaining necessary accuracy.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system uses production workloads themselves to gather state information needed for model updates. By leveraging existing task executions and their observed performance, the scheduler obtains diagnostic information without requiring separate diagnostic tasks, thus maintaining throughput while improving accuracy.

Inventive Principle:
Principle #25Self-service

3Ease of operation

If the scheduler assumes fixed compute node capabilities, then resource allocation is straightforward, but resource congestion occurs when node performance changes due to hardware issues or VM migrations

Engineering Contradiction:
Improvescheduling easeVSAvoidperformance reliability
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The system makes capability models dynamic by continuously updating them based on observed state variables from compute nodes. This allows the scheduler to adapt to changing node performance due to hardware issues or VM migrations while maintaining relatively simple scheduling operations through automated belief updates.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The scheduler implements feedback loops that monitor compute node performance and automatically update capability models. This feedback mechanism detects when node capabilities change and adjusts resource allocation accordingly, maintaining reliability without significantly complicating the scheduling process.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS9934071B2Job scheduler for distributed systems using pervasive state estimation with modeling of capabilities of compute nodes
Publication Date: 2018.04.03 GENESEE VALLEY INNOVATIONS LLC
  • US9934071B2 patent drawing
  • US9934071B2 patent drawing
  • US9934071B2 patent drawing

AI summary

The following relates generally to computer system efficiency improvements. Broadly, systems and methods are disclosed that improve efficiency in a cluster of nodes by efficient processing of tasks among nodes in a cluster of nodes. Initially, tasks may be scheduled on the nodes in the cluster of nodes. Following that, state information may be received, and a determination may be made as to if tasks should be rescheduled.