Massively Parallel Node Sampling via Round-Robin Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional performance assessment techniques for massively parallel computer systems suffer from significant overhead, leading to compromised sampling rates and longer analysis times, which can hinder system performance and increase costs due to the burden on nodes during data collection.
Innovation Solution
A method that identifies nodes performing similar work and samples their performance data at different times, distributing the overhead evenly across nodes using a round-robin technique, allowing for efficient performance sampling with reduced impact on the system.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional performance assessment techniques sample all nodes simultaneously, then comprehensive performance data is collected, but system overhead increases significantly and sampling rates are compromised
Solution Approach 1:
The patent divides the set of all nodes into multiple subsets and samples each subset at different times rather than sampling all nodes simultaneously. This segmentation approach maintains comprehensive data collection while reducing the overhead burden on any single node at any given moment, thereby preserving system throughput.
Solution Approach 2:
The patent implements periodic sampling of different node subsets at different time intervals. By rotating through different subsets in a periodic manner, the system collects comprehensive performance data across all nodes while ensuring that no single node is overwhelmed by simultaneous sampling operations, thus maintaining system productivity.
2Loss of information
If performance sampling is performed on all nodes at once, then complete system profile is obtained, but analysis time increases and node burden increases
Solution Approach 1:
By segmenting nodes into subsets and sampling them at different times, the patent reduces the time required for data collection while maintaining information completeness. The segmented approach allows parallel processing of sampling operations across subsets, significantly reducing total analysis time compared to sequential sampling of all nodes.
Solution Approach 2:
The patent performs preliminary grouping of nodes into subsets based on their work similarity before sampling. This preliminary action enables efficient organization of sampling operations, allowing the system to collect complete performance information faster by preparing the sampling structure in advance rather than processing nodes individually.
3Measurement precision
If high sampling rates are maintained, then accurate performance measurement is achieved, but overhead on nodes increases and system performance is compromised
Solution Approach 1:
The patent segments the sampling process across different time intervals for different node subsets. This allows high sampling rates to be maintained for each subset without overwhelming any single node, as the sampling burden is distributed across subsets rather than concentrated on all nodes simultaneously. This resolves the contradiction by maintaining measurement precision while reducing node overhead.
Data Source
AI summary
An apparatus, program product and method sample at different times nodes that are performing similar work. Performance data associated with first and second node subsets performing the similar work are sampled at different times, e.g., in a round-robin fashion, and in accordance with a given sampling rate. The performance data is analyzed. Nodes whose performance suffers as a result of a sampling operation may be identified and removed from a subsequent operation.


