Universal graphics processor cache sharing design based on on-chip cache and core separation architecture

By separating and combining the L1 cache of the general graphics processor from the core into a large cache, combining MESI consistency and frequency table management, the problem of L1 cache resource competition is solved, and cache efficiency and processor performance is improved.

CN120371727APending Publication Date: 2025-07-25WUHAN TEXTILE UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410091586.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-23
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

When existing general-purpose graphics processors handle complex applications, L1 cache resource competition leads to inefficiency and fails to fully utilize their performance.

Method used

The on-chip cache and core separation architecture is adopted, and the L1 cache is merged into large caches and divided into exclusive and shared caches. The MESI consistency protocol is used to manage cache access in combination with the frequency table, and different strategies are used to process cache data of different frequencies.

Benefits of technology

It improves the efficiency of L1 cache usage, improves the number of IPCs of the graphics processor and the running speed of complex programs, and avoids data loss caused by cache storms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention discloses a universal graphics processor cache sharing design based on an on-chip cache and core separation architecture, which comprises the following steps: firstly separating each core on a universal graphics processor chip and an L1 cache corresponding to the core, and then synthesizing a plurality of separated L1 caches into a'large cache '; a plurality of large caches are divided into exclusive large caches and shared large caches, each large cache uses MESI to achieve data cache consistency, each core maintains a frequency table, the number of times of access of the cores to cache data and the core in which the cache data are stored are recorded, and meanwhile data consistency is kept. Caches accessed by the core are divided into'low-frequency cache data ', 'medium-frequency cache data' and'high-frequency cache data ', different types of storage strategies are adopted, and finally different access strategies are adopted for different types of cache data accessed by the core, so that the on-chip cache is efficiently utilized, the utilization rate of the cache is improved, and the method is simple and practical.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a cache sharing design for a general - purpose graphics processing unit (GPGPU) based on an on - chip cache and core separation architecture, belonging to the fields of computer architecture and high - performance computing. Background Art

[0002] Nowadays, the application fields of general - purpose graphics processing units are no longer limited to graphics processing and are becoming more and more extensive. The original design purpose of general - purpose graphics processors was to process graphics applications. However, due to their excellent parallel performance, simple pipeline design, and numerous cores, users are running more complex applications on general - purpose graphics processors, such as meteorology, artificial intelligence, big data, cloud computing, blockchain computing, quantitative trading, and other business scenarios. However, the negative impact of this is that these applications cannot fully utilize the performance of general - purpose graphics processors. Some applications only use less than 30% of the computing power of the processor due to limitations such as resource competition.

[0003] One of the most prominent problems is that some applications require multiple cores to process the same data fetched from the main memory for different processing calculations. This causes a copy of the data in the main memory to be cached in the on - chip L1 caches of multiple cores of the processor for use, resulting in the already limited - capacity L1 cache becoming more crowded, thereby reducing the utilization rate of the on - chip L1 cache and ultimately leading to a sharp decline in performance.

[0004] Under the influence of this problem, a new architecture of on - chip cache and core has emerged. The advantage of separating the on - chip cache from the core is that the caches accessed by different cores are no longer limited to one, but can access all the caches on the general - purpose graphics processor. Based on this architecture, the present invention further proposes a new shared cache design to further improve the utilization efficiency of the on - chip L1 cache. Summary of the Invention

[0005] The solution of the present invention is a cache sharing design for a general - purpose graphics processing unit based on an on - chip cache and core separation architecture, and its features include the following steps:

[0006] Step 1: Separate each core on the GPGPU chip from the L1 cache corresponding to that core;

[0007] Step 2: Combine multiple separated L1 caches into a "large cache" (the L1 caches of every 4 cores are combined into one);

[0008] Step 3: Divide multiple "large caches" into "exclusive large caches" and "shared large caches";

[0009] Step 4: Use MESI to achieve data cache consistency for each "large cache";

[0010] Step 5: Each core maintains a frequency table to record the access times of the core to cache data and which core the cached data is stored in, while maintaining data consistency;

[0011] Step 6: Divide the caches accessed by the core into "low-frequency cached data", "medium-frequency cached data", and "high-frequency cached data", and adopt different storage strategies;

[0012] Step 7: Different access strategies are adopted for the core to access different types of cached data.

[0013] The specific steps of Step 1 include:

[0014] Step 1.1: Separate the on-chip L1 cache and the core, and connect them with a high-speed internal network of the chip in the middle.

[0015] The specific steps of Step 2 include:

[0016] Step 2.1: Separate the core and the corresponding L1 cache of the core.

[0017] Step 2.2: Merge the separated multiple caches into a large cache (the L1 caches of every 4 cores are merged into one).

[0018] Step 2.3: It is stipulated that the large cache can only be accessed by the core corresponding to the L1 cache before merging.

[0019] The specific steps of Step 3 include:

[0020] Step 3.1: Set all "large caches" as "exclusive large caches" in the initial stage, and the "exclusive large caches" can only be accessed by the core corresponding to the L1 cache before merging.

[0021] Step 3.2: If the number of "large caches" is greater than ten, set one-tenth of the "large caches" as "shared large caches"; if the number of "large caches" is less than ten (including ten) and greater than five, set one "large cache" as "shared large cache"; if the number of "large caches" is less than five (including five), no "shared large cache" is set.

[0022] Step 3.3: The "shared large cache" loads the cached data with high access frequencies.

[0023] The specific steps of Step 4 include:

[0024] Step 4.1: Use the MESI consistency protocol to ensure the consistency of cached data among multiple "large caches".

[0025] The specific steps of Step 5 include:

[0026] Step 5.1: Each core has a frequency table built-in, and the table entries include the access times of cache lines, the number of the "Exclusive Large Cache" being accessed, and the cache type ("Low-frequency Cache Data", "Medium-frequency Cache Data", "High-frequency Cache Data").

[0027] Step 5.2: When a core accesses the "Exclusive Large Cache", count in the frequency data table and record the number of the "Exclusive Large Cache" being accessed.

[0028] Step 5.3: When a core changes the frequency table of the "Exclusive Large Cache", notify other cores through broadcasting. After the unified data update is completed, update its own frequency table and then access the "Exclusive Large Cache".

[0029] The specific steps of Step Six include:

[0030] Step 6.1: Cache lines with an access frequency lower than 10 times are recorded as "Low-frequency Cache Data"; those with an access frequency higher than 10 times and lower than 50 times are recorded as "Medium-frequency Cache Data"; cache data with an access policy higher than 50 times is recorded as "High-frequency Cache Data".

[0031] Step 6.2: For "Low-frequency Cache Data", it is not cached in any "Large Cache", and only the cache metadata (access times) is recorded in the core's frequency table; when the access times of "Low-frequency Cache Data" reach 10 times, the data is cached in the "Exclusive Large Cache"; when the access times of the cached data reach 50 times, the cached data is moved to the "Shared Large Cache".

[0032] The specific steps of Step Seven include:

[0033] Step 7.1: For "Low-frequency Cache Data", the core directly accesses the main memory; for "Medium-frequency Cache Data", the core obtains data from the corresponding "Exclusive Large Cache"; for "High-frequency Cache Data", the core obtains data from the "Shared Large Cache", and the data requests issued by the core corresponding to the "Shared Large Cache" are directly bypassed to the main memory, that is, the core corresponding to the "Shared Large Cache" does not cache any data.

[0034] The present invention has the following beneficial effects:

[0035] 1. Improve the on-chip L1 cache usage efficiency of general-purpose image processors.

[0036] 2. Increase the IPC number of general-purpose image processors and improve the running speed of complex programs.

[0037] 3. Prevent the "cache storm" caused by complex application scenarios from making the cache unable to effectively store data and generating a large number of cache misses. Description of the Drawings

[0038] Figure 1 For the architecture where the on-chip cache is separated from the core

[0039] Figure 2 For the schematic diagrams of the shared large cache and the exclusive large cache

[0040] Figure 3 For the frequency table architecture

[0041] Figure 4 For the flowchart of multi-core frequency table data synchronization

[0042] Figure 5 For the flowcharts of low-frequency cache access, medium-frequency cache access, and high-frequency cache access Specific implementation method

[0044] The present invention will be further described below in conjunction with the architecture diagrams.

[0045] First, as Figure 1 shown, separate the cache from the core and connect them with the high-speed on-chip network (Interconnect Network 1). When the cache access fails, access the main memory of the general-purpose image processor (Global Memory) through another high-speed on-chip network (Interconnect Network 2). Figure 1 It is the schematic diagram of the architecture described in step 1.1.

[0046] After separating the cache from the core, merge the caches corresponding to every 4 cores into a large cache. If the number of "large caches" is greater than ten, set one-tenth of the "large caches" as "shared large caches". If the number of "large caches" is less than ten (including ten) and greater than five, set one "large cache" as a "shared large cache". If the number of "large caches" is less than five (including five), do not set a "shared large cache", as Figure 2 shown.

[0047] Secondly, design a frequency table on each core. The frequency table architecture is as Figure 3 shown. In the figure, FT is the frequency table located beside the Tag Unit in the pipeline of the general-purpose image processor. Before the Tag Unit is ready to determine whether a memory read instruction hits the cache, update the frequency table first. The frequency table contains the address of the cache line (Cache line Address), the access count of the cache line (Frequency), whether the cache line is a high-frequency cache (IsHigh), and which "large cache" the cache is located in (Location).

[0048] When a core accesses a cache line, it needs to broadcast and update the frequency tables in all cores. The specific process is as Figure 4As shown, when Core 1 is ready to access memory and reaches the Tag Unit through the internal chip pipeline, the Tag Unit prepares to decide whether to cache the access and first updates the local FT (as Figure 4 shown in ①). After the update, it broadcasts a notification to all other cores through the internal chip network (Interconnect Network 1) to let them change their FTs. After receiving the broadcast message, other cores update their FTs (as Figure 4 shown in ②). Finally, Core 1 accesses its "exclusive large cache" or "shared large cache" or directly bypasses the memory access according to the type of cache line (as Figure 4 shown in ③).

[0049] Subsequently, each cache line is classified into "low-frequency cached data", "medium-frequency cached data", and "high-frequency cached data". Different types of cache lines have different storage strategies and different access mechanisms. For "low-frequency cached data", it is not cached in any "large cache", and only the cache metadata (access count) is recorded in the core's frequency table; when the access count of "low-frequency cached data" reaches 10 times, the data is cached in the "exclusive large cache"; when the access count of the cached data reaches 50 times, the cached data is moved to the "shared large cache". At the same time, there are also different access strategies for different types of cache lines, that is, for "low-frequency cached data", the core directly accesses the main memory ( Figure 5 as described in ③); for "medium-frequency cached data", the core obtains the data from the corresponding "exclusive large cache" ( Figure 5 as shown in ①); for "high-frequency cached data", the core obtains the data from the "shared large cache" ( Figure 5 as shown in ②). The data requests issued by the core corresponding to the "shared large cache" are directly bypassed to the main memory, Figure 5 which describes these characteristics.

Claims

1. A cache sharing design for a general - purpose graphics processor based on an on - chip cache and core separation architecture The features include the following steps: Step 1: Separate each core on the GPGPU chip and the L1 cache corresponding to that core; Step 2: Combine multiple separated L1 caches into a "large cache" (the L1 caches of every 4 cores are combined into one); Step 3: Divide multiple "large caches" into "exclusive large caches" and "shared large caches"; Step 4: Use MESI for each "large cache" to achieve data cache consistency; Step 5: Each core maintains a frequency table to record the number of accesses to cache data by the core and in which core the cache data is stored, while maintaining data consistency; Step 6: Divide the caches accessed by the core into "low-frequency cache data", "medium-frequency cache data", and "high-frequency cache data", and adopt different storage strategies; Step 7: The core adopts different access strategies for accessing different types of cache data. For the general-purpose graphics processor cache sharing design according to claim 1, the specific steps of step 2 include: Step 2.1: Separate the core and the L1 cache corresponding to the core. Step 2.2: Combine multiple separated caches into a large cache (the L1 caches of every 4 cores are combined into one). Step 2.3: Specify that the large cache can only be accessed by the core corresponding to the L1 cache before combination. For the general-purpose graphics processor cache sharing design according to claim 1, the specific steps of step 3 include: Step 3.1: Set all "large caches" as "exclusive large caches" in the initial stage, and the "exclusive large caches" can only be accessed by the core corresponding to the L1 cache before combination. Step 3.2: If the number of "large caches" is greater than ten, set one-tenth of the "large caches" as "shared large caches"; if the number of "large caches" is less than ten (including ten) and greater than five, set one "large cache" as "shared large cache"; if the number of "large caches" is less than five (including five), no "shared large cache" is set. Step 3.3: The "shared large cache" loads cache data with a high access frequency. For the general-purpose graphics processor cache sharing design according to claim 1, the specific steps of step 5 include: Step 5.1: Each core has a built-in frequency table, and the table entries include the number of accesses to the cache line, the number of the "exclusive large cache" accessed, and the cache type ("low-frequency cache data", "medium-frequency cache data", "high-frequency cache data"). Step 5.2: When the core accesses the "exclusive large cache", count in the frequency data table and record the number of the "exclusive large cache" accessed. Step 5.3: When the core changes the frequency table of the "exclusive large cache", notify other cores through broadcasting. After the unified data update is completed, update its own frequency table and then access the "exclusive large cache". For the general-purpose graphics processor cache sharing design according to claim 1, the specific steps of step 6 include: Step 6.1: Record the cache lines with an access frequency lower than 10 times as "low-frequency cache data"; record the cache lines with an access frequency higher than 10 times and lower than 50 times as "medium-frequency cache data"; record the cache data with an access strategy higher than 50 times as "high-frequency cache data". Step 6.2: For "low-frequency cached data", it is not cached in any "large cache", and only the cache metadata (access count) is recorded in the core frequency table; when the access count of "low-frequency cached data" reaches 10 times, the data is cached in the "dedicated large cache"; when the access count of the cached data reaches 50 times, the cached data is moved to the "shared large cache". According to the general graphics processor cache sharing design of claim 1, the specific steps of step seven include: Step 7.1: For "low-frequency cached data", the core directly accesses the main memory; for "medium-frequency cached data", the core obtains the data from the corresponding "dedicated large cache"; for "high-frequency cached data", the core obtains the data from the "shared large cache", and the data request issued by the core corresponding to the "shared large cache" is directly bypassed to the main memory, that is, the core corresponding to the "shared large cache" does not cache any data.