Dynamic Task Mapping for Heterogeneous Parallel Nodes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing job schedulers and task mappers in parallel computing systems do not account for node variability, leading to sub-optimal performance and power consumption due to assumptions of homogeneous nodes, resulting in poor task scheduling and mapping decisions.
Innovation Solution
A cluster agent monitors and tracks physical and functional sensory data from nodes to calculate variability metrics, using this information for dynamic task scheduling and mapping, and adjusts node parameters through DVFS or other reconfigurations to maximize performance and power efficiency while reducing variability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If existing job schedulers and task mappers assume nodes are homogeneous for task scheduling, then the scheduling process is simplified, but performance and power consumption become sub-optimal
Solution Approach 1:
The patent applies parameter changes by transitioning from static homogeneous node assumptions to dynamic heterogeneous node characterization. The system captures physical and functional sensory data to create time-varying variability metrics for each node, then uses these metrics to dynamically adjust task scheduling decisions. This resolves the contradiction by accepting increased scheduling complexity in exchange for optimal performance utilization of heterogeneous nodes.
Solution Approach 2:
The patent implements dynamics by making the task mapping process adaptive rather than static. The cluster agent continuously monitors node variability metrics and dynamically remaps tasks based on current node performance states. This dynamic approach allows the system to exploit temporal variations in node performance, resolving the contradiction between scheduling simplicity and performance optimality.
2Device complexity
If existing task mappers do not account for node variability, then the mapping algorithm is simpler, but task execution efficiency deteriorates
Solution Approach 1:
The patent applies preliminary action by capturing and storing physical and functional sensory data from nodes before task mapping decisions are made. The cluster agent pre-characterizes node variability metrics and maintains this information for use during task scheduling. This preliminary characterization enables informed mapping decisions that reduce task execution time without excessive algorithmic complexity.
Solution Approach 2:
The patent implements feedback by continuously monitoring node performance metrics and using this information to adjust task mapping decisions. The cluster agent receives feedback from node variability measurements and dynamically remaps tasks to optimize execution efficiency. This feedback loop resolves the contradiction by enabling adaptive mapping that reduces execution time while maintaining manageable algorithmic complexity.
3Speed
If nodes operate at maximum performance, then critical tasks can be completed faster, but power consumption increases and variability among nodes increases
Solution Approach 1:
The patent applies local quality by assigning different operational states to different nodes based on their individual variability metrics and the requirements of specific tasks. Rather than uniformly maximizing all nodes, the system selectively operates nodes at different performance levels matched to task criticality. This resolves the contradiction by enabling fast execution of critical tasks on suitable nodes while conserving power on less critical workloads.
Solution Approach 2:
The patent implements parameter changes by dynamically adjusting node operating parameters based on captured sensory data and variability metrics. The system changes voltage, frequency, or other operational parameters to match task requirements, allowing critical tasks to run at maximum speed while non-critical tasks use lower power states. This resolves the speed-power contradiction through adaptive parameter optimization.
4Productivity
If nodes are reconfigured to reduce variability, then workload partitioning efficiency improves, but device reconfiguration complexity increases
Solution Approach 1:
The patent applies self-service by enabling nodes to autonomously adjust their operational parameters based on captured variability metrics and cluster-wide workload conditions. The cluster agent facilitates this self-adjustment by providing guidance signals, but the actual reconfiguration is performed by the nodes themselves. This resolves the contradiction by distributing reconfiguration complexity across individual nodes rather than requiring centralized complex control.
Solution Approach 2:
The patent implements feedback by using captured node variability metrics to drive reconfiguration decisions. The system monitors node performance characteristics and uses this feedback to automatically adjust node parameters for reduced variability. This feedback-driven approach improves workload execution efficiency while keeping reconfiguration complexity manageable through automated adaptation rather than manual intervention.
Data Source
AI summary
Systems, apparatuses, and methods for managing variations among nodes in parallel system frameworks. Sensor and performance data associated with the nodes of a multi-node cluster may be monitored to detect variations among the nodes. A variability metric may be calculated for each node of the cluster based on the sensor and performance data associated with the node. The variability metrics may then be used by a mapper to efficiently map tasks of a parallel application to the nodes of the cluster. In one embodiment, the mapper may assign the critical tasks of the parallel application to the nodes with the lowest variability metrics. In another embodiment, the hardware of the nodes may be reconfigured so as to reduce the node-to-node variability.


