Near Memory Miss Prediction for Latency Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current computing system memory architectures face performance and power consumption issues due to frequent cache misses in near memory, which lead to slower data retrieval from far memory, and existing hit-miss prediction techniques are not scalable for larger memory-side caches.
Innovation Solution
A miss predictor is implemented that tracks near memory misses using a prediction table, bypassing entry allocations for hits to maintain a smaller and more scalable structure, allowing for parallel access to near and far memory when a valid page address is found, thereby improving prediction accuracy and reducing latency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If hit-miss prediction techniques are used for small host processor-side caches, then prediction accuracy is improved, but scalability to larger memory-side caches deteriorates due to predictor size and overhead
Solution Approach 1:
The prediction system is segmented into a lightweight miss predictor component that operates independently within the memory controller, separating the prediction function from the main cache structure. This allows the prediction mechanism to scale independently from cache size without proportionally increasing overhead.
Solution Approach 2:
A prediction table serving as an intermediary structure is introduced between the access request and the near memory. This table stores only missed page addresses and uses validation bits to indicate predicted misses, acting as a mediator that enables scalable prediction without requiring complex predictor structures.
2Productivity
If frequent misses in near memory occur, then data retrieval from far memory increases, but performance and power consumption deteriorate
Solution Approach 1:
The system performs preliminary action by predicting near memory misses before actual access occurs. When a miss is predicted, the far memory is accessed in parallel with the near memory, so that when the near memory eventually returns data, it is already available in the far memory, eliminating the performance penalty of actual misses.
Solution Approach 2:
The system converts the harmful effect of near memory misses into a benefit by using the prediction table to identify missed page addresses. These previously harmful misses become useful information that triggers parallel far memory access, transforming performance degradation into performance optimization.
3Reliability
If a prediction table tracking all page addresses is implemented, then prediction coverage is improved, but device complexity increases
Solution Approach 1:
The invention extracts only the essential information needed for prediction - specifically missed page addresses and validation bits - from a complete cache tracking system. By taking out only what is necessary rather than tracking all cache states, the prediction table achieves effective coverage with minimal structural complexity.
Solution Approach 2:
The prediction table implements partial action by tracking only missed addresses rather than all cache accesses. This selective tracking provides sufficient prediction coverage for the critical case (misses) while avoiding the complexity of comprehensive cache state monitoring.
Data Source
AI summary
Systems, apparatuses and methods may provide for technology to maintain a prediction table that tracks missed page addresses with respect to a first memory. If an access request does not correspond to any valid page addresses in the prediction table, the access request may be sent to the first memory. If the access request corresponds to a valid page address in the prediction table, the access request may be sent to the first memory and a second memory in parallel, wherein the first memory is associated with a shorter access time than the second memory.


