Far-end memory data caching method based on shared cache
By building a shared cache architecture based on two-stage on-chip networks in multi-chip GPUs, and combining the reconfigurable GPU L1 cache and request router, the problems of reduced effective cache capacity, poor scalability and insufficient flexibility in the existing technology are solved, and higher memory access performance and better scalability are achieved.
Patent Information
- Application Number
- CN202510119715.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-06
AI Technical Summary
In existing multi-chip GPUs, remote data caching methods such as L2 cache and L1.5 cache on the processor unit side are facing the problems of reduced effective cache capacity, poor scalability and insufficient flexibility caused by data duplication, and it is difficult to effectively apply in larger-scale multi-chip GPUs in the future.
Using the remote memory data cache method based on shared cache, by building a shared cache architecture based on two-level on-chip networks, combining the reconfigurable GPU L1 cache and request router, cache sharing and flexible configuration between processor units are realized, reducing data copy and cross-chip data access.
It improves the memory access performance and scalability of multi-chip GPUs, and flexibly configures the L1 cache of processor units. It is suitable for multi-chip GPUs containing more chips in the future, avoiding unnecessary performance losses in applications with poor locality.
Smart Images

Figure CN120104518A_ABST
Abstract
Description
Technical field:
[0001] The invention discloses a remote memory data caching method based on shared cache, relates to a multi-chip GPU micro-architecture, in particular to a multi-chip GPU cache architecture, and belongs to the field of computer technology. Background technology:
[0002] GPU is widely used in various applications such as machine learning and high-performance computing due to its large-scale parallel computing capabilities and high energy efficiency. As the demand for computing power continues to grow, multi-chip GPUs have become the mainstream trend of future GPU development. In a typical multi-chip GPU architecture, it usually includes several computing chips and corresponding memories. Each computing chip contains multiple processor units and L2 caches, which are interconnected through an on-chip network. Each processor unit is equipped with a private L1 cache.
[0003] In multi-chip GPUs, the L2 cache is a memory-side cache with a much higher access latency than the L1 cache. The L2 cache of each computing chip only caches data from the memory corresponding to that chip. The latency of the computing unit accessing the memory on the remote chip is much higher than the latency of accessing the memory on the local chip. Therefore, remote memory data access is an important performance bottleneck for multi-chip GPUs. By utilizing the memory access locality of GPU applications, caching remote memory data on the chip is an important way to effectively overcome this bottleneck.
[0004] In existing multi-chip GPUs, remote data caching methods are mainly divided into two methods: processor unit-side L2 cache and L1.5 cache:
[0005] The processor unit-side L2 cache is a different design from the memory-side L2 cache in traditional GPUs. In the processor unit-side L2 cache architecture, the L2 cache is privately owned by the processor unit in each chip, rather than shared by all the processor units in a multi-chip GPU. The L2 cache slice in each chip corresponds to the entire memory address space. When the memory access requests issued by the processor unit in the chip miss the L1 cache, these requests first access the L2 cache inside the chip. If the request hits the processor unit-side L2 cache, cross-chip data access is avoided, thereby reducing data transfer between chips; if it misses, further access to the memory. Compared with the traditional GPU L2 cache architecture, the processor unit-side L2 cache can fully utilize locality within the chip and avoid unnecessary cross-chip data access. However, the processor unit-side L2 cache faces the problem of reduced effective capacity of the L2 cache due to data duplication between L2 caches in different chips, that is, multiple copies of the same data will reduce the effective capacity of the cache and impair cache performance. In addition, the processor unit-side L2 cache relies on a complex consistency protocol. When the number of chips in a multi-chip GPU further increases, the overhead of maintaining cache consistency is unacceptable. Therefore, the scalability of the processor unit-side L2 cache is poor.
[0006] The L1.5 cache architecture is designed to reduce the L2 cache capacity while ensuring a certain number of chip transistors, and add a new cache called L1.5 cache to each chip. The L1.5 cache is located between the L1 cache and the L2 cache, and is only used to cache data from remote memory. Therefore, data access to local memory does not pass through the L1.5 cache. The L1.5 cache can also take advantage of locality to reply to some requests that miss the L1 cache within the chip, thereby reducing cross-chip data access. However, the L1.5 cache needs to reduce the L2 cache capacity when it is designed. This approach lacks flexibility. For applications with large L2 cache capacity requirements, the L1.5 cache architecture design will result in performance degradation. Therefore, the L1.5 cache is greatly restricted in actual GPU applications.
[0007] In summary, in order to solve the inefficiency problem caused by cross-chip data access in multi-chip GPUs, the existing methods are divided into processor unit-side L2 cache and L1.5 cache. However, the processor unit-side L2 cache faces the problem of data inefficiency caused by duplicate data and poor scalability; the L1.5 cache faces the problem of poor flexibility and difficulty in practical application. With the increase in the scale and number of chips in multi-chip GPUs in the future, the above two solutions face more severe challenges. Summary of the invention:
[0008] The invention discloses a remote memory data caching method based on shared cache, which consists of three parts: a shared cache architecture based on a two-level on-chip network; a reconfigurable GPU L1 cache; and a request router.
[0009] The construction method of the shared cache architecture based on two-level on-chip network is as follows: First, in the benchmark multi-chip GPU, the processor units are divided into multiple clusters to avoid all processor units sharing cache with each other, which leads to serious resource contention. Secondly, all processor units are interconnected through two-level on-chip networks. The processor units in the cluster are interconnected through the on-chip network within the cluster, and the on-chip networks in each cluster are interconnected through the on-chip network between clusters. Then, the input queue of the L1 cache in the processor unit is modified so that the request from the current processor unit and the request from the on-chip network are used as the input of the L1 cache through the first-come-first-served policy. Finally, the output queue of the L1 cache is modified so that the request after the memory access is completed can be returned to the current processor unit or returned to other processor units through the on-chip network. By constructing a shared cache architecture based on two-level on-chip networks, a basis is provided for subsequent L1 cache reconfiguration to serve different memory access requests respectively.
[0010] The main design of the reconfigurable GPU L1 cache is as follows: After completing the construction of a shared cache architecture based on a two-level on-chip network, an architecture is obtained in which the L1 caches in the cluster are shared with each other and the L1 caches between clusters can communicate with each other. On this basis, the L1 cache is configured, and the L1 caches of some processor units in the cluster are used as the shared L1 caches in the cluster, and the L1 caches of other processor units in the cluster are used as L1.5 caches. The shared L1 cache in the cluster is the first-level cache that processes requests from the processor units. Each shared L1 cache in the cluster corresponds to a different address space segment, and all shared L1 caches in the cluster cover the entire address space together. Through this mapping method, requests with the same address in different processor units can be mapped to the same cache, thereby reducing the decrease in cache capacity caused by data copies. The L1.5 cache is used to process requests that miss the shared L1 cache in the cluster. Unlike the shared L1 cache in the cluster, which is only shared within the cluster, the L1.5 cache is shared by all processor units in the current chip. The L1.5 cache only caches data from remote memory. Each L1.5 cache in the chip corresponds to a segment of address space, and all L1.5 caches share the entire remote memory address space. When a request hits the L1.5 cache, there is no need to access data across chips, thereby improving the memory access performance of multi-chip GPUs. During chip configuration, the L1 caches of different processor units can be configured as shared L1 caches or L1.5 caches within the cluster by configuring registers.
[0011] The main design purpose of the request router is: since in the shared cache architecture, the request no longer only accesses the L1 cache of the processor unit that issued the request, the request needs to be routed to the correct cache. The request router has two main functions, namely, mapping the request from the processor unit to the appropriate shared L1 cache in the cluster and routing the request that misses the shared L1 cache in the cluster to the appropriate L1.5 cache or directly to the local L2 cache. To this end, the main logic of the request router is divided into two parts. The input of the first part of the logic is the request from the processor unit. First, it passes through the tag selector to select the specified number of bits from the address tag of the request, and then maps the request to the specified shared L1 cache in the cluster according to the value of these bits. The input of the other part of the logic of the request router is the request from the shared L1 cache in the cluster. These requests miss the shared L1 cache in the cluster and need to further access the next level cache. At this time, the request router first determines whether the request accesses the local memory. If it accesses the local memory, the request is sent directly to the local memory; otherwise, the specified bit in the request address tag is first selected, and then the request is mapped to the corresponding L1.5 cache. If the target L1.5 cache is in the current cluster, the request goes through the on-chip network within the cluster. If it is in another cluster, it goes through the two-level on-chip network.
[0012] Under the cache architecture proposed by the present invention, the main access process of the request is as follows: when the memory access data loading unit in the processor unit issues a memory access request, the request first goes to the request router, and the request router maps the request to the shared L1 cache in the target cluster. When the request hits the shared L1 cache in the cluster, the requested data is directly returned to the memory access data loading unit; otherwise, the request is sent to the request router again. In the request router, it is determined whether the request is to access local memory or remote memory. For requests to access local memory, they are sent directly to the local memory subsystem. For requests to access remote memory, the request first accesses the L1.5 cache. If the request hits the L1.5 cache, there is no need for cross-chip data access. If the request does not hit the L1.5 cache, it goes to the remote memory subsystem through the chip-to-chip interconnection.
[0013] Advantages of the present invention include:
[0014] The present invention discloses a remote memory data caching method based on shared cache. Compared with the existing design, the advantages are: the existing remote memory data caching method has poor scalability and flexibility and is difficult to fully utilize locality. The present invention adopts a reconfigurable shared cache method to flexibly configure the L1 cache of the processor unit, and can adopt different configurations for different applications. It has good flexibility, thereby avoiding unnecessary performance loss in applications with poor locality. The present invention has high scalability and flexibility, and is suitable for multi-chip GPUs containing more chips in the future. Description of the drawings:
[0015] Figure 1 Schematic diagram of remote memory data cache architecture based on shared cache.
[0016] Figure 2 Schematic diagram of the address mapping of the shared L1 cache and L1.5 cache within the cluster.
[0017] Figure 3 Request a diagram of the router hardware organization. Specific implementation method:
[0018] The present invention is further described in detail below with reference to the accompanying drawings.
[0019] like Figure 1 As shown, a schematic diagram of a remote memory data cache architecture based on shared cache is shown. Figure 1 The processor units and memory hierarchy in a chip in a multi-chip GPU are shown. The present invention first groups the processor units in the chip, and the processor units are divided into two clusters. The processor units in the cluster are interconnected through the intra-cluster on-chip network, and the on-chip networks in each cluster are interconnected through the inter-cluster on-chip network to achieve full interconnection of all processor units. In each processor unit in the cluster, the L1 cache is configured as a shared L1 cache or L1.5 cache within the cluster by configuring registers. In addition, in each processor unit, a request router is introduced between the memory access data loading unit and the L1 cache to ensure that the request is routed to the correct L1 cache.
[0020] like Figure 2 As shown in the figure, a schematic diagram of the address mapping of the shared L1 cache and L1.5 cache within the cluster. The figure takes a 4-chip multi-chip GPU as an example. There is a memory on each chip, so there are 4 memories in the entire multi-chip GPU. These memories can be accessed by all processor units and jointly store data in the entire address space. The figure shows chip 0 with 6 processor units, which are divided into two clusters. Each cluster has 2 intra-cluster shared L1 caches and 1 L1.5 cache. Figure 2In the example, different textures represent different address ranges. For the shared L1 cache within the cluster, the cache address ranges of cache 0 and cache 1 are different, and they cache all address spaces together. Similarly, cache 3 and cache 4 also cache all address spaces together. For the L1.5 cache, that is, cache 2 and cache 5, since they are on chip 0, they only cache the address space of the remote memory, that is, the memory address space on chips 1-3. Cache 2 and cache 5 also cache different address ranges respectively, and cache the entire remote memory address space together.
[0021] Figure 3 The hardware organization of the request router is shown. The request router is located between the memory access data loading unit and the L1 cache in the processor unit, and is responsible for routing the memory access requests from the memory access data loading unit and the requests from the shared L1 cache in the cluster that miss the L1 cache to the correct shared L1 cache in the cluster and the L1.5 cache respectively. Therefore, the hardware organization of the request router is divided into two parts. The first part is to receive the request from the memory access data loading unit and map the request to the specified shared L1 cache in the cluster. The request first selects the specified bit in the address tag in the tag selector and passes it to the cache mapping unit. In the cache mapping unit, the request is mapped to the shared L1 cache in the cluster according to the request tag. If the shared L1 cache in the cluster is in the current processor unit, it goes directly to the local L1 cache, otherwise it goes to the target L1 cache through the on-chip network in the cluster. The second part receives the requests from the shared L1 cache in the cluster that miss the cache, and first classifies the requests into local memory requests and remote memory requests in the memory request classifier. Local memory requests go directly to the local L2 cache, while remote memory requests first select the specified bit from the address tag through the tag selector, and then the cache mapping unit maps the request to the target L1.5 cache and sends the request to the target L1.5 cache through two levels of on-chip interconnect.
[0022] Finally, it should be noted that the present invention may also have many other application scenarios. Without departing from the spirit and essence of the present invention, technicians familiar with the field can make various corresponding changes and deformations based on the present invention, but these corresponding changes and deformations should all fall within the scope of protection of the present invention.
Claims
1. A remote memory data caching method based on shared cache, characterized in that: It includes the following three parts: a shared cache architecture based on a two-level on-chip network, a reconfigurable GPU L1 cache, and a request router; In a multi-chip GPU architecture, there is no need to add additional cache capacity. By sharing the L1 cache in the processor unit and reconfiguring the L1 cache into a shared L1 cache and L1.5 cache within the cluster, data in the entire address space and remote memory data can be cached respectively, thereby making full use of data locality.
2. The method according to claim 1, characterized in that: In the shared cache architecture based on the two-level on-chip network, the processor units are divided into several clusters to avoid all the processor units sharing cache with each other, which leads to serious resource contention; in addition, all the processor units are interconnected through the two-level on-chip network, the processor units in the cluster are interconnected through the on-chip network within the cluster, and the on-chip networks in each cluster are interconnected through the inter-cluster on-chip network; then, the input queue of the L1 cache in the processor unit is modified so that the request from the current processor unit and the request from the on-chip network are used as the input of the L1 cache through the first-come-first-served strategy; finally, the output queue of the L1 cache is modified so that the request after the memory access is completed can be returned to the current processor unit or returned to other processor units through the on-chip network. By constructing a shared cache architecture based on the two-level on-chip network, a basis is provided for subsequent L1 cache reconfiguration to serve different memory access requests respectively.
3. The method according to claim 1, characterized in that: The reconfigurable L1 cache uses the L1 caches of some processor units in the cluster as the shared L1 caches in the cluster, and the L1 caches of other processor units in the cluster as the L1.5 caches; the shared L1 cache in the cluster is the first-level cache for processing requests from processor units; each shared L1 cache in the cluster corresponds to a different address space segment, and all shared L1 caches in the cluster jointly cover the entire address space; through this mapping method, requests for the same address in different processor units can be mapped to the same cache, thereby reducing the decrease in cache capacity caused by data copies; the L1.5 cache is used to process requests that miss the shared L1 cache in the cluster; unlike the shared L1 cache in the cluster, which is only shared within the cluster, the L1.5 cache is shared by all processor units in the current chip; the L1.5 cache only caches data from remote memory; each L1.5 cache in the chip corresponds to a segment of address space, and all L1.5 caches share the entire address space of the remote memory; When a request hits the L1.5 cache, there is no need to access data across chips, thereby improving the memory access performance of multi-chip GPUs; when configuring the chip, the L1 cache of different processor units can be configured as a shared L1 cache or L1.5 cache within the cluster by configuring registers.
4. The method according to claim 1, characterized in that: The request router has two functions, namely, mapping requests from processor units to appropriate intra-cluster shared L1 caches, and routing requests that miss the intra-cluster shared L1 caches to appropriate L1.5 caches or directly to local L2 caches; To this end, the logic of the request router includes the following two parts. The input of the first part of the logic is the request from the processor unit. It first passes through the label selector to select the specified number of bits from the address label of the request. Then, according to the value of these bits, the request is mapped to the specified shared L1 cache in the cluster. The input of another part of the logic of the request router is the request from the shared L1 cache within the cluster. These requests do not hit the shared L1 cache within the cluster and need to further access the next level cache. At this time, the request router first determines whether the request accesses the local memory. If it accesses the local memory, the request is sent directly to the local memory. Otherwise, first select the specified bit in the request address tag, and then map the request to the corresponding L1.5 cache. If the target L1.5 cache is in the current cluster, the request goes through the on-chip network within the cluster. If it is in other clusters, it goes through two levels of on-chip networks.