A method for merging GPU level 2 cache miss requests based on data locality

By introducing a request merger between the L2 cache and the DRAM controller, the redundant DRAM access problem caused by L2 cache misses in the GPGPU architecture is solved, thereby reducing the number of DRAM accesses and memory access latency, and improving system performance.

CN122412337APending Publication Date: 2026-07-17EAST CHINA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
EAST CHINA UNIV OF TECH
Filing Date
2026-04-20
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

In the existing GPGPU architecture, redundant DRAM access issues caused by L2 cache misses lead to increased DRAM bandwidth usage and memory access latency, becoming a bottleneck in system performance.

Method used

A request merger is introduced between the L2 cache and the DRAM controller. By merging control logic, merging metadata tables, request buffer queues, and data distribution units, the online merging of missed requests is achieved, reducing the number of requests entering DRAM.

Benefits of technology

It effectively reduces the number of DRAM accesses, improves bandwidth utilization, reduces memory access latency, enhances the spatial locality utilization of the cache, and has low hardware overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122412337A_ABST
    Figure CN122412337A_ABST
Patent Text Reader

Abstract

本发明公开了一种基于数据局部性的GPGPU二级缓存未命中请求合并方法。此方法将请求合并器设置于GPU的L2缓存写回队列与DRAM访问延迟队列之间,包括合并控制逻辑、合并元数据表、请求缓冲队列和数据分发单元四个核心模块。合并元数据表以哈希表形式实现,每个条目记录缓存块地址、原始请求指针列表、状态标识、创建时间及访问掩码,配置维护64个活跃条目,总硬件开销仅约2.34KB。本发明通过在L2缓存未命中请求到达DRAM控制器之前进行在线合并,从源头减少冗余DRAM访问次数,降低访存延迟,提升GPGPU系统整体性能,同时以极低的硬件代价实现显著的性能增益。
Need to check novelty before this filing date? Find Prior Art