Graph-Based Call-Stack Prefetching for HPC Power Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
High-Performance Computing (HPC) systems face significant power consumption issues due to idle computing resources during loading phases, where existing prefetching methods are inefficient as they load data globally without granularity, leading to unnecessary data loading.
Innovation Solution
A method for prefetching data in HPC systems based on a graph representation of call-stacks, predicting the next call-stack and corresponding Input/Output request, and prefetching the required data to minimize unnecessary loading.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If global prefetching is performed over the whole region, then data availability is improved, but unnecessary data is loaded consuming excessive power and resources
Solution Approach 1:
The patent segments the prefetching process by analyzing call-stack patterns to identify specific data regions that will be accessed, rather than prefetching the entire memory region. This segmentation allows the system to prefetch only the necessary portions of data, reducing energy consumption while maintaining data availability.
Solution Approach 2:
The patent applies local quality by making the prefetching behavior adaptive to local access patterns. By monitoring and analyzing call-stack sequences, the system identifies locally accessed data regions and prefetches only those specific areas, rather than applying a uniform global prefetching strategy across the entire memory space.
2Measurement precision
If source code analysis is performed to detect data requirements, then prefetching precision is improved, but computing resource consumption increases significantly
Solution Approach 1:
The patent implements self-service by having the system automatically monitor and analyze its own call-stack patterns during execution. Instead of requiring external source code analysis tools, the system uses its runtime execution information to identify access patterns and trigger prefetching operations autonomously, reducing the need for additional computing resources.
Solution Approach 2:
The patent applies preliminary action by performing call-stack analysis and identifying prefetch candidates during idle periods or in advance of actual data access. This allows the system to prepare prefetch operations without interfering with the main computation workload, thereby maintaining high prefetching precision while minimizing impact on computing resource availability.
3Loss of information
If call-stack analysis is performed using grammatical models, then loading patterns are identified, but model size becomes large and assumptions limit applicability
Solution Approach 1:
The patent extracts only the essential features from call-stack analysis - specifically the sequence of function calls and their associated memory access patterns - rather than building complete grammatical models. By extracting and storing only the relevant call-stack sequences and their temporal patterns, the system achieves effective loading pattern detection with significantly reduced model size.
Solution Approach 2:
The patent changes the parameters of the analysis model by focusing on temporal sequences of call-stack frames rather than comprehensive grammatical structures. Instead of modeling all possible syntactic relationships, the system parameters are simplified to track the order and frequency of function calls, reducing model complexity while maintaining the ability to detect loading patterns.
Data Source
AI summary
A computer implemented method for prefetching data related to an application executed by a node of a High-Performance Computing system while said node is running an application. The prefetching is carried out based on a call-stack and corresponding Input/Output request predicted by using a graph.

