Multi-core GPU cache division method, architecture and data access method

Through the topology-aware cache allocation mechanism, the cache division of multi-core GPUs is optimized, which solves the problem of large access delays for computing core and remote video memory, and improves system performance and cache hit rate.

CN120278871AActive Publication Date: 2025-07-08BEIHANG UNIV

Patent Information

Application Number
CN202510763985.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

In multi-core GPU systems, the access delay between the computing core and remote video memory is large, resulting in a bottleneck in performance improvement. The existing architecture based on remote data cache has a low cache hit rate, which affects the overall performance.

Method used

A cache allocation mechanism based on topology awareness is introduced. By obtaining the topology structure of a multi-core GPU, cache division is performed according to the number of hops between cores and cache capacity ratios, and cross-core access performance is optimized.

Benefits of technology

Reduces the probability of cache misses, reduces remote video memory access, improves large-scale image data processing performance, and avoids cache fragmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278871A_ABST
    Figure CN120278871A_ABST
Patent Text Reader

Abstract

The invention provides a multi-core GPU cache division method, architecture and data access method, relates to the technical field of graphics processor design, and aims to realize cross-core access performance system-level optimization and improve the performance of processing large-scale image data by introducing a cache allocation mechanism based on topology awareness. The method comprises the steps that a multi-core-grain GPU topological structure comprising a target computing core grain and at least one associated computing core grain is obtained, a to-be-divided far-end data cache of the target computing core grain is used for caching image data from a far-end video memory, and the to-be-divided far-end data cache is connected with the associated computing core grain through an inter-core-grain interconnection network; if the number of the calculation core particles contained in the multi-core-particle GPU does not exceed a first preset threshold value, the inter-core-particle hop count between each associated calculation core particle and a target calculation core particle is determined based on the topological structure, and then the corresponding capacity of the associated calculation core particles in the to-be-divided cache of the target calculation core particle is determined; and dividing the to-be-divided cache according to the capacity corresponding to each associated calculation core particle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of graphics processing unit design, and particularly to a method for partitioning a multi-die GPU cache, an architecture, and a data access method. Background Art

[0002] With its excellent parallel computing capabilities and high throughput characteristics, the Graphics Processing Unit (GPU) has become the core computing platform for image data processing tasks. However, as the physical limit of transistors is approaching, developing GPUs for large-scale image data processing by relying on process scaling faces various challenges such as yield, cost, and development cycle. Therefore, the multi-die graphics processor architecture integrating multiple computing and storage dies has become the key to promoting the improvement of large-scale image data processing capabilities.

[0003] In such an architecture, multiple graphics processing computing dies (i.e., computing dies) are connected through an inter-die interconnect network to work collaboratively to meet the performance requirements of large-scale image data processing. Each computing die is equipped with local video memory (i.e., storage die), and all video memories are jointly mapped to the entire address space. However, as the number of computing dies increases, the maximum physical distance (the physical distance can be referred to as the hop count) between the computing die and the video memory on other dies significantly increases, resulting in a significant increase in the access latency between dies during image data processing tasks, which has become a bottleneck problem restricting the performance improvement of multi-die GPUs.

[0004] To solve the above problems, existing research has adopted a cache technology based on remote data caches. This technology sets an additional cache (i.e., remote data cache) in each computing die of the multi-die GPU to cache image data from remote video memories and is shared by all computing units in the computing die. When the request of each computing unit misses the L1 cache inside the computing unit, if the requested address is in the address space of the local video memory, the remote data cache is skipped and the local last-level cache and video memory are directly accessed; if the requested address is in the address space of the remote video memory, the remote data cache is first accessed. If the request hits the remote data cache, the data access can be completed inside the computing die. At this time, there is no need to access the remote last-level cache and video memory through other computing dies, thus greatly reducing the frequency of remote data access and improving the access performance of multi-die GPUs.

[0005] However, the current architecture based on remote data caches is relatively simple and has the problem of low cache hit rate when processing large-scale image data, which affects the further improvement of the overall performance of the multi-die GPU system and urgently needs to be optimized and improved through innovative solutions. Summary of the Invention

[0006] The present application provides a method for partitioning a multi-die GPU cache, an architecture, and a data access method. By introducing a cache allocation mechanism based on topology awareness, system-level optimization of cross-die access performance is achieved, and the performance of processing large-scale image data is improved.

[0007] In a first aspect, a method for partitioning a multi-die GPU cache is provided, including: Obtain the topology of the multi-die GPU. The multi-die GPU includes a target computing die for processing image data and at least one associated computing die. The target computing die is the computing die for which cache partitioning operation is to be performed. The associated computing die is connected to the target computing die through an inter-die interconnect network. The remote data cache to be partitioned of the target computing die is used to cache image data from the video memory of at least one associated computing die; If the number of computing dies included in the multi-die GPU does not exceed a first preset threshold, determine the number of inter-die hops between each associated computing die and the target computing die based on the topology; According to the number of inter-die hops corresponding to each associated computing die, determine the cache capacity ratio corresponding to each associated computing die, where the number of inter-die hops corresponding to each associated computing die is positively correlated with the corresponding cache capacity ratio, and the cache capacity ratio is the ratio of the required capacity of the associated computing die to the capacity of the remote data cache to be partitioned; According to the cache capacity ratio corresponding to each associated computing die and the capacity of the remote data cache to be partitioned, determine the capacity corresponding to each associated computing die in the remote data cache to be partitioned; Partition the remote data cache to be partitioned according to the capacity corresponding to each associated computing die.

[0008] In a feasible design, the method includes: If the number of computing dies included in the multi-die GPU exceeds the first preset threshold, perform region partitioning on the multi-die GPU according to the topology to obtain a target region and at least one associated region. The target region is the region to which the target computing die belongs, and each associated region is a region other than the target region in the multi-die GPU. The number of computing dies included in each region is the same, and the maximum distance between the computing dies in each region is less than a second preset threshold; Based on the topology, determine the number of inter-region hops between each region and the target region; Based on a first linear scaling factor and a first global offset constant, adjust the number of inter-region hops corresponding to each region; According to the adjusted number of inter-region hops corresponding to each region, determine the cache capacity ratio corresponding to each region, where the number of inter-region hops corresponding to each region is positively correlated with the corresponding cache capacity ratio; Determine the capacity corresponding to each region in the to-be-partitioned remote data cache according to the cache capacity ratio corresponding to each region and the capacity of the to-be-partitioned remote data cache. Partition the to-be-partitioned remote data cache according to the capacity corresponding to each region. Among them, the associated computing dies in each region share the storage space allocated to the region in the to-be-partitioned remote data cache.

[0009] In a feasible design, after partitioning the to-be-partitioned remote data cache according to the capacity corresponding to each associated computing die, periodically re-partition the capacity of the to-be-partitioned remote data cache. The steps of re-partitioning the capacity of the to-be-partitioned remote data cache after each partitioning cycle include: Detect the first request quantity for the target computing die to access the video memory of each associated computing die in the previous partitioning cycle. Determine the first capacity that needs to be re-partitioned in the to-be-partitioned remote data cache. Adjust the first request quantity corresponding to each associated computing die based on the second linear scaling factor and the second global offset constant. According to the adjusted first request quantity corresponding to each associated computing die, determine the first capacity ratio corresponding to each associated computing die. Among them, the first capacity ratio is the ratio of the capacity that the associated computing die needs to be re-allocated to the first capacity. The adjusted first request quantity corresponding to each associated computing die is positively correlated with the corresponding first capacity ratio. According to the first capacity ratio, the original capacity, and the first capacity corresponding to each associated computing die, determine the new capacity corresponding to each associated computing die. Among them, the original capacity is the capacity corresponding to the associated computing die in the previous partitioning cycle. Re-partition a part of the cache of the to-be-partitioned remote data cache according to the new capacity and the original capacity corresponding to each associated computing die.

[0010] In a feasible design, re-partition a part of the cache of the to-be-partitioned remote data cache according to the new capacity and the original capacity corresponding to each associated computing die, including: Send a memory access latency perception request to each associated computing die. The memory access latency perception request is used to trigger the round-trip communication between the target computing die and the associated computing die. Determine the access latency value corresponding to each associated computing die according to the sending time and receiving time of each memory access latency perception request. Determine the second capacity that needs to be re-partitioned in the to-be-partitioned remote data cache. Adjust the access latency value corresponding to each associated computing die based on the third linear scaling factor and the third global offset constant. Calculate the adjusted access latency value corresponding to each associated computing die according to each association, and determine the second capacity ratio corresponding to each associated computing die, where the second capacity ratio is the ratio of the capacity that needs to be redistributed for the associated computing die to the second capacity, and the adjusted access latency value corresponding to each associated computing die is positively correlated with the corresponding second capacity ratio; Determine the final capacity corresponding to each associated computing die according to the new capacity, the second capacity ratio, and the second capacity corresponding to each associated computing die; Redivide a partial cache of the remote data cache to be partitioned according to the final capacity and the original capacity corresponding to each associated computing die.

[0011] In a feasible design, after partitioning the remote data cache to be partitioned according to the capacity corresponding to each region, periodically re-partition the capacity of the remote data cache to be partitioned. The steps of re-partitioning the capacity of the remote data cache to be partitioned after each partitioning cycle include: Detect the second request quantity for accessing the video memory of each region by the target computing die in the previous partitioning cycle. The video memory of each region includes the video memories of each associated computing die in the region; Determine the third capacity that needs to be re-partitioned in the remote data cache to be partitioned; Adjust the second request quantity corresponding to each region based on the fourth linear scaling factor and the fourth global offset constant; Determine the third capacity ratio corresponding to each region according to the adjusted second request quantity corresponding to each region, where the third capacity ratio is the ratio of the capacity that needs to be redistributed for the region to the third capacity, and the adjusted second request quantity corresponding to each region is positively correlated with the corresponding third capacity ratio; Determine the new capacity corresponding to each region according to the third capacity ratio, the original capacity, and the third capacity corresponding to each region, where the original capacity is the capacity corresponding to the region in the previous partitioning cycle; Redivide a partial cache of the remote data cache to be partitioned according to the new capacity and the original capacity corresponding to each region.

[0012] In a feasible design, re-partitioning a partial cache of the remote data cache to be partitioned according to the new capacity and the original capacity corresponding to each region includes: Send a memory access latency perception request to each associated computing die, and the memory access latency perception request is used to trigger the round-trip communication between the target computing die and the associated computing die; Determine the access latency value corresponding to each associated computing die according to the sending time and the receiving time of each memory access latency perception request; Determine the sum of the access latency values corresponding to each associated computing die in each region as the access latency value corresponding to each region; Determine the fourth capacity that needs to be re-partitioned in the to-be-partitioned remote data cache; Adjust the access latency value corresponding to each region based on the fifth linear scaling factor and the fifth global offset constant; According to the adjusted access latency value corresponding to each region, determine the fourth capacity ratio corresponding to each region, where the fourth capacity ratio is the ratio of the capacity that needs to be re-allocated in the region to the fourth capacity, and the adjusted access latency value corresponding to each region is positively correlated with the corresponding fourth capacity ratio; According to the new capacity, the fourth capacity ratio, and the fourth capacity corresponding to each region, determine the final capacity corresponding to each region; Re-partition the to-be-partitioned remote data cache according to the final capacity and the original capacity corresponding to each region.

[0013] In a feasible design, determining the cache capacity ratio corresponding to each associated computing die according to the number of die-to-die hops corresponding to each associated computing die includes: Adjust the number of die-to-die hops corresponding to each associated computing die based on the sixth linear scaling factor and the sixth global offset constant; According to the adjusted number of die-to-die hops corresponding to each associated computing die, determine the cache capacity ratio corresponding to each associated computing die.

[0014] In a second aspect, a multi-die GPU cache architecture is provided, including: A topology-aware unit for obtaining the topology of a multi-die GPU. The multi-die GPU includes a target computing die for processing image data and at least one associated computing die. The target computing die is the computing die to be operated on for cache partitioning. The associated computing die is connected to the target computing die through a die-to-die interconnect network. The to-be-partitioned remote data cache of the target computing die is used to cache image data from the video memory of at least one associated computing die; The topology-aware unit, for if the number of computing dies included in the multi-die GPU does not exceed a first preset threshold, determining the number of die-to-die hops between each associated computing die and the target computing die based on the topology; A capacity calculation unit for determining the cache capacity ratio corresponding to each associated computing die according to the number of die-to-die hops corresponding to each associated computing die, where the number of die-to-die hops corresponding to each associated computing die is positively correlated with the corresponding cache capacity ratio, and the cache capacity ratio is the ratio of the capacity required by the associated computing die to the capacity of the to-be-partitioned remote data cache; The capacity calculation unit is further configured to determine the capacity corresponding to each associated computing die in the to-be-partitioned remote data cache according to the cache capacity ratio corresponding to each associated computing die and the capacity of the to-be-partitioned remote data cache. The cache control unit is configured to partition the to-be-partitioned remote data cache according to the capacity corresponding to each associated computing die.

[0015] In a feasible design, the to-be-partitioned remote data cache adopts a set-associative mapping structure, including multiple cache groups, each cache group includes multiple cache lines, and each cache line includes a partition identification bit, and the partition identification bit is used to store a video memory identification, and the video memory identification is used to identify the video memory to which the image data cached by the cache line belongs. Wherein, the cache control unit is configured to partition the to-be-partitioned remote data cache according to the capacity corresponding to each associated computing die, including: The cache control unit is configured to determine the video memory identification stored in the partition identification bit of each cache line according to the capacity corresponding to each associated computing die. The cache control unit partitions the to-be-partitioned remote data cache by writing the corresponding video memory identification to the partition identification bit of each cache line.

[0016] In a third aspect, a data access method for a multi-die GPU cache is provided, which is applied to the multi-die GPU cache architecture as described in the foregoing example. The multi-die GPU cache architecture includes a cache control unit and a remote data cache. The remote data cache adopts a set-associative mapping structure, including multiple cache groups, each cache group includes multiple cache lines, and each cache line includes a partition identification bit. The partition identification bit is configured to store a video memory identification or a region identification. The video memory identification is used to identify the video memory to which the image data cached by the cache line belongs, and the region identification is used to identify the region to which the image data cached by the cache line belongs. The method includes: The computing unit of the target computing die generates a request including a data source identification, an access address, and an identification of the computing unit, and the data source identification is a video memory identification or a region identification. If the request misses the L1 cache inside the computing unit and the access address includes the address of the remote video memory, send the request to the cache control unit. The cache control unit parses the request to obtain the data source identification, the access address, and the identification of the computing unit of the request. The cache control unit locates the target cache group according to the access address. The cache control unit performs a request hit judgment on each cache line in the target cache group in the storage order according to the access address and the data source identification. If a hit occurs, the cache control unit extracts the target image data from the target cache line according to the access address, and sends the target image data to the computing unit according to the identification of the computing unit. If there is a miss, the cache control unit sends indication information to the computing unit. The indication information is used to indicate that the target image data is not stored in the target cache group. After receiving the indication information, the computing unit sends a request to the target associated computing die corresponding to the video memory mapped by the access address through the inter-die interconnect network.

[0017] The multi-die GPU cache partitioning method described in the embodiments of the present application breaks through the limitation of the traditional architecture based on remote data caching that does not partition the remote data cache. By introducing a cache allocation mechanism based on topology awareness, it realizes the system-level optimization of the performance of accessing image data across dies. Based on the positive correlation design of the ratio of the number of hops between dies to the cache capacity in the present application, the associated computing dies with a longer access distance (i.e., a higher number of hops) between the dies for processing image data obtain a larger proportion of the cache space, thereby reducing the cache miss probability of this part of the associated computing dies when processing image data, reducing the access to the remote video memory, and improving the performance of processing large-scale image data.

[0018] In addition, by restricting the number of computing dies included in the multi-die GPU not to exceed the first preset threshold, it avoids the fragmentation of the remote data cache for storing image data caused by the over-fine partitioning of the cache due to too many computing dies. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0020] Figure 1 is a schematic diagram of a multi-die GPU architecture provided by an exemplary embodiment of the present application; Figure 2 is a schematic flowchart of a multi-die GPU cache partitioning method provided by an exemplary embodiment of the present application; Figure 3 is another schematic diagram of a multi-die GPU architecture provided by an exemplary embodiment of the present application; Figure 4 is another schematic diagram of a multi-die GPU architecture provided by an exemplary embodiment of the present application; Figure 5 is another schematic diagram of a multi-die GPU architecture provided by an exemplary embodiment of the present application; Figure 6 is another schematic diagram of a multi-die GPU architecture provided by an exemplary embodiment of the present application; Figure 7It is a schematic diagram of a multi-die GPU cache architecture provided by an exemplary embodiment of the present application; Figure 8 It is a schematic diagram of the cache line structure of a multi-die GPU provided by an exemplary embodiment of the present application; Figure 9 It is a schematic diagram of the composition of an access address provided by an exemplary embodiment of the present application. Detailed implementation manners

[0021] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0022] The cache technology based on remote data caching is an important method for reducing the number of times of accessing image data in the remote video memory in a multi-die GPU and improving the memory access performance. The traditional architecture based on remote data caching is relatively simple. The image data from the video memory corresponding to each associated computing die (which can be called the remote video memory) uses the corresponding storage space in the remote data cache in the order of access by the target computing die. However, since the distances from the target computing die to different associated computing dies are different, the costs of requests that miss the address spaces of the video memories of the associated computing dies with smaller distances and those that miss the address spaces of the video memories of the associated computing dies with larger distances are different in the remote data cache of the target computing die.

[0023] The following uses Figure 1 the multi-die GPU shown as an example to illustrate the above problems: Figure 1 The multi-die GPU shown includes 4 computing dies, namely computing die 0, computing die 1, computing die 2, and computing die 3. Each computing die corresponds to a video memory. The video memory corresponding to computing die 0 is video memory 0, the video memory corresponding to computing die 1 is video memory 1, the video memory corresponding to computing die 2 is video memory 2, and the video memory corresponding to computing die 3 is video memory 3. If there is a solid line connection between two computing dies, it means that the two computing dies can communicate directly without passing through other computing dies. That is to say, the access distance (i.e., the number of hops) between the two computing dies is 1. Each computing unit inside the computing die accesses the last-level cache through the on-chip interconnection network. Taking Figure 1 the computing die 3 shown as an example, its remote data cache stores the image data of the address spaces of video memories 0-2. For the target computing die 3, the access distances to video memory 0 and video memory 1 are different. Therefore, in the remote data cache, the costs of requests that miss the address space of video memory 0 and those that miss the address space of video memory 1 are different.

[0024] To improve the overall performance of the graphics processing unit in processing image data, the present application creatively proposes to allocate the capacity of the remote data cache according to the cost of request misses. Based on this, the present application provides a method for partitioning the multi-die GPU cache, as Figure 2 shown, the method includes: S110, obtaining the topological structure of the multi-die GPU.

[0025] Among them, the multi-die GPU includes a target computing die for processing image data and at least one associated computing die. The target computing die is the computing die to be operated on for cache partitioning. The associated computing die is connected to the target computing die through an inter-die interconnect network. The remote data cache to be partitioned of the target computing die is used to cache the image data from the video memory of at least one associated computing die.

[0026] S120, if the number of computing dies included in the multi-die GPU does not exceed a first preset threshold, determining the number of inter-die hops between each associated computing die and the target computing die based on the topological structure.

[0027] To prevent the problem of remote data cache fragmentation caused by the cache being partitioned too finely due to too many computing dies, the present application partitions the multi-die GPU caches with the number exceeding the first preset threshold and not exceeding the first preset threshold in different ways. When the number of computing dies included in the multi-die GPU does not exceed the first preset threshold, methods S120 to S150 are used for partitioning in the remote data cache to be partitioned of the target computing die. Among them, the first preset threshold is set according to actual needs, and the present application does not limit this.

[0028] Exemplarily, it is implemented in the following way. Determining the number of inter-die hops between the target associated computing die and the target computing die based on the topological structure: Obtaining the preset topological structure of the multi-die GPU; Determining the target associated computing die to be subject to routing discovery based on the topological structure; Sending a routing discovery request to the target associated computing die, where the routing discovery request is used to record the number of forwarding times on the path between the target computing die and the target associated computing die; Determining the number of inter-die hops between the target associated computing die and the target computing die as the number of forwarding times.

[0029] As Figure 3As shown in the figure, taking the target computing die as computing die 3 and computing dies 0-2 as associated computing dies as an example, the die-to-die hop counts corresponding to computing dies 0-2 are 2, 1, and 1 respectively. Among them, computing die 3 includes multiple computing units, and each computing unit can access the remote data cache and access the last-level cache through the on-chip interconnect network. The capacity of the remote data cache partition corresponding to video memory 0 in the remote data cache of computing die 3 is greater than the capacities of the remote data cache partitions corresponding to video memory 1 and video memory 2 respectively in the remote data cache of computing die 3. The capacities of the remote data cache partitions corresponding to video memory 1 and video memory 2 in the remote data cache of computing die 3 are the same.

[0030] S130. Determine the cache capacity ratio corresponding to each associated computing die according to the die-to-die hop count corresponding to each associated computing die.

[0031] Among them, the cache capacity ratio is the ratio of the required capacity of the associated computing die to the capacity of the remote data cache to be partitioned, and the die-to-die hop count corresponding to each associated computing die is positively correlated with the corresponding cache capacity ratio.

[0032] Exemplarily, it is implemented in the following manner. Determine the cache capacity ratio corresponding to each associated computing die according to the die-to-die hop count corresponding to each associated computing die: Determine the sum value of the die-to-die hop counts corresponding to each associated computing die; Determine the ratio of the die-to-die hop count corresponding to each associated computing die to the sum value as the cache capacity ratio corresponding to the associated computing die.

[0033] In order to more accurately determine the cache capacity ratio, in a feasible design, it is implemented in the following manner. Determine the cache capacity ratio corresponding to each associated computing die according to the die-to-die hop count corresponding to each associated computing die: Adjust the die-to-die hop count corresponding to each associated computing die based on the sixth linear scaling factor and the sixth global offset constant; Determine the cache capacity ratio corresponding to each associated computing die according to the adjusted die-to-die hop count corresponding to each associated computing die.

[0034] It should be noted that the first linear scaling factor, the second linear scaling factor, the third linear scaling factor, the fourth linear scaling factor, the fifth linear scaling factor, and the sixth linear scaling factor in this application are all used to perform linear scaling on the die-to-die hop count. The first global offset constant, the second global offset constant, the third global offset constant, the fourth global offset constant, the fifth global offset constant, and the sixth global offset constant in this application are all used to perform an overall translation operation on the linearly scaled die-to-die hop count.

[0035] Exemplarily, it is implemented as follows. The adjusted number of hops between dielets corresponding to each associated calculation dielet is calculated, and the cache capacity ratio corresponding to each associated calculation dielet is determined: Determine the sum value of the adjusted number of hops between dielets corresponding to each associated calculation dielet; The ratio of the adjusted number of hops between dielets corresponding to each associated calculation dielet to the sum value is determined as the cache capacity ratio corresponding to the associated calculation dielet.

[0036] For example, assume that the entire GPU system has compute dielets, and currently it is necessary to divide the capacity for the remote data cache in compute dielet . Assume that the distance between the associated compute dielet ( ) and compute dielet is hops. Then the cache capacity ratio corresponding to the video memory is , where and both represent the numbers of the associated compute dielets, represents that the distance between the associated compute dielet and compute dielet is hops. is the sixth linear scaling factor, which can be the image data access latency between dielets in the case of no resource contention. is the sixth global offset constant, which can be the latency for the compute units in the dielet to access the local video memory.

[0037] For example, assume = 30, = 300. Taking the number of hops between dielets corresponding to compute dielets 0 - 2 shown in Figure 3 as 2, 1, 1 respectively, the unadjusted cache capacity ratios are 1 / 2, 1 / 4, 1 / 4 respectively, and the adjusted cache capacity ratios are 6 / 17, 11 / 34, 11 / 34 respectively. It is worth mentioning that since dielet - to - dielet communication usually includes resource contention, taking 30 is only a special case, and the present application does not limit the values of and .

[0038] It can be seen that in the above embodiments, by introducing a linear scaling factor and a global offset constant to jointly adjust the number of hops between dies, extreme allocation can be suppressed, making the cache capacity division more reasonable. For example, before adjustment, the cache capacity ratio of die 0 in the calculation is 50% (1 / 2), there is a risk of single-die monopoly, while after adjustment, the cache capacity ratio of die 0 in the calculation drops to 35.3% (6 / 17≈0.353), and each of die 1 and die 2 in the calculation accounts for 32.35% (11 / 34≈0.3235). The maximum / minimum capacity ratio drops from 2:1 (before adjustment) to 1.09:1 (after adjustment), avoiding resource allocation imbalance. The smooth attenuation of the hop count difference is achieved, avoiding the polarization of allocation differences.

[0039] In addition, the linear scaling factor and the global offset constant can be dynamically configured, supporting the adjustment of the cache allocation strategy according to factors such as load characteristics during runtime, improving the flexible adjustment ability of the solution.

[0040] S140. According to the cache capacity ratio corresponding to each associated computing die and the capacity of the remote data cache to be divided, determine the capacity corresponding to each associated computing die in the remote data cache to be divided.

[0041] For example Figure 3 As shown, for the remote data cache in die 3 for calculation, assume the total capacity is R, the cache capacity corresponding to video memory 0 is 6 / 17R, and the cache capacities of video memory 1 and video memory 2 are both 11 / 34R. In the specific implementation process, assume the remote data cache is a 256-way cache, then after rounding, 90 ways are allocated to cache the image data of video memory 0, 83 ways are allocated to cache the image data of video memory 1, and 83 ways are allocated to cache the image data of video memory 2.

[0042] S150. Divide the remote data cache to be divided according to the capacity corresponding to each associated computing die.

[0043] Exemplarily, according to the numbering order of the video memories of the associated computing dies, map the video memories of the associated computing dies to the remote data cache to be divided according to the corresponding capacity.

[0044] The multi-die GPU cache division method described in the embodiments of the present application breaks through the limitation that the traditional architecture based on the remote data cache does not divide the remote data cache. By introducing a cache allocation mechanism based on topology awareness, system-level optimization of the performance of accessing image data across dies is achieved. Based on the positive correlation design between the number of hops between dies and the cache capacity ratio in the present application, the associated computing dies with a farther access distance (i.e., a higher number of hops) between the dies for processing image data obtain a larger proportion of the cache space, thereby reducing the cache miss probability of this part of the associated computing dies when processing image data, reducing the access to the remote video memory, and improving the performance of processing large-scale image data.

[0045] In addition, by restricting the number of computing dies included in the multi-die GPU to not exceed a first preset threshold, it is avoided that the cache is divided too finely due to too many computing dies, resulting in fragmentation of the remote data cache for storing image data.

[0046] The method of the above embodiment can be used for a multi-die GPU system with any topological shape. According to the remote data cache partitioning method proposed in this application, after determining the distances between different computing dies, the cache partitioning ratio within each computing die can be calculated.

[0047] For example, Figure 4 the shown topological structure includes 4 computing dies in series, namely computing die 0, computing die 1, computing die 2, and computing die 3. Each computing die corresponds to a video memory. The video memory corresponding to computing die 0 is video memory 0, the video memory corresponding to computing die 1 is video memory 1, the video memory corresponding to computing die 2 is video memory 2, and the video memory corresponding to computing die 3 is video memory 3. Among them, computing die 3 includes multiple computing units, each of which can access the remote data cache and access the last-level cache through the on-chip interconnect network. For computing die 3, the distances from computing dies 0, 1, and 2 are 3 hops, 2 hops, and 1 hop respectively. If = 30, = 300, and the total cache capacity is R, then 13 / 36R, 1 / 3R, and 11 / 36R of the cache capacity are respectively divided to cache the image data of video memory 0, video memory 1, and video memory 2. It can be seen that the capacity of the remote data cache partition corresponding to video memory 0 in the remote data cache of computing die 3 is greater than the capacity of the remote data cache partition corresponding to video memory 1 in the remote data cache of computing die 3. The capacity of the remote data cache partition corresponding to video memory 1 in the remote data cache of computing die 3 is greater than the capacity of the remote data cache partition corresponding to video memory 2 in the remote data cache of computing die 3.

[0048] For example, Figure 5The shown topology includes 4 computing dies that can communicate directly with each other in pairs, namely computing die 0, computing die 1, computing die 2, and computing die 3. Each computing die corresponds to a video memory. The video memory corresponding to computing die 0 is video memory 0, the video memory corresponding to computing die 1 is video memory 1, the video memory corresponding to computing die 2 is video memory 2, and the video memory corresponding to computing die 3 is video memory 3. Among them, computing die 3 includes multiple computing units. Each computing unit can access the remote data cache and access the last-level cache through the on-chip interconnect network. The hop count between every two computing dies is 1. For the remote data cache of each computing die, the capacity of caching the image data of the other 3 video memories is the same, all being 1 / 3R. For example, in computing die 3, it can be seen that the capacities of the remote data cache partitions corresponding to video memory 0, video memory 1, and video memory 2 in the remote data cache of computing die 3 are the same.

[0049] In a feasible design, after dividing the to-be-divided remote data cache according to the capacity corresponding to each associated computing die, the capacity of the to-be-divided remote data cache is periodically re-divided. The steps of re-dividing the capacity of the to-be-divided remote data cache after each division cycle include: Detect the first request quantity for the video memory of each associated computing die accessed by the target computing die in the previous division cycle; Determine the first capacity that needs to be re-divided in the to-be-divided remote data cache; Based on the second linear scaling factor and the second global offset constant, adjust the first request quantity corresponding to each associated computing die; According to the adjusted first request quantity corresponding to each associated computing die, determine the first capacity ratio corresponding to each associated computing die. Among them, the first capacity ratio is the ratio of the capacity that needs to be re-allocated by the associated computing die to the first capacity. The adjusted first request quantity corresponding to each associated computing die has a positive correlation with the corresponding first capacity ratio; According to the first capacity ratio, the original capacity, and the first capacity corresponding to each associated computing die, determine the new capacity corresponding to each associated computing die. Among them, the original capacity is the capacity corresponding to the associated computing die in the previous division cycle; According to the new capacity and the original capacity corresponding to each associated computing die, re-divide part of the cache of the to-be-divided remote data cache.

[0050] Among them, the second linear scaling factor and the second global offset constant are set according to actual requirements, and the present application does not limit this.

[0051] For example, taking the division of the remote data cache in the computing die as an example, assuming that after static division, in the computing die The associated computing die in the remote data cache The number of cache lines corresponding to the video memory on (i.e., the original capacity), where , is the number of computing dies included in the multi-die GPU system. Every clock cycle is a partitioning cycle for re-partitioning the remote data cache (e.g., ). After statically partitioning the remote data cache to be partitioned according to the capacity corresponding to each associated computing die, the part of the cache with the first capacity in the remote data cache to be partitioned is periodically re-partitioned by calculating the new capacity corresponding to each associated computing die. Specifically, each re-partitioning of the remote data cache to be partitioned includes the following steps: (1) Detect the number of times (i.e., the first request quantity) that the computing die (i.e., the target computing die) accesses each associated computing die in the previous partitioning cycle. Among them, the first request quantity of the video memory of the associated computing die is , where represents the number of the associated computing die.

[0052] (2) Determine that the first capacity to be re-partitioned in the remote data cache to be partitioned is 1 / m1 of the total number of cache lines, where m1 is a preset constant.

[0053] (3) Based on the second linear scaling factor and the second global offset constant , adjust the first request quantity corresponding to each associated computing die to obtain the adjusted first request quantity. Among them, the adjusted first request quantity corresponding to the associated computing die is ( ).

[0054] (4) Determine the capacity ratio according to the adjusted first request quantity: sum up the first request quantities corresponding to each associated computing die to obtain . Among them, represents the first request quantity corresponding to the associated computing die . Therefore, the first capacity ratio corresponding to the associated computing die is .

[0055] (5) Determine the new capacity corresponding to each associated computing core particle based on the first capacity ratio, original capacity and first capacity corresponding to each associated computing core particle. Specifically, it can be regarded as first taking out 1 / m1 of the capacity from the original capacity corresponding to each associated computing core particle to participate in the redivision process. Then, the capacity taken out corresponding to each associated computing core particle constitutes the first capacity, and the remaining capacity corresponding to each associated computing core particle is (m1-1) / m1 of the original capacity. At this time, the capacity that the associated computing core particle needs to reallocate refers to the capacity that needs to be allocated to the associated computing core particle in the first capacity during this redivision process. Therefore, for the associated computing core particle , new capacity = original capacity * (m1-1) / m1 + first capacity ratio * first capacity.

[0056] (6) Considering that only part of the cache in the remote data cache is repartitioned, the image data in the cache that has not been repartitioned can be retained to improve the partitioning efficiency. Therefore, by comparing the new capacity and the original capacity corresponding to each associated computing core, the part of the cache to be partitioned in the remote data cache is repartitioned. Among them, if the new capacity corresponding to the associated computing core is the same as the original capacity, there is no need to allocate a new cache line to the associated computing core. If the new capacity corresponding to the associated computing core is less than the original capacity, the image data of the cache line of the corresponding capacity of the associated computing core is cleared to facilitate allocation to other associated computing cores. If the new capacity corresponding to the associated computing core is greater than the original capacity, the cache line that originally belonged to other associated computing cores but had its image data cleared is allocated to the associated computing core.

[0057] The dynamic cache partitioning mechanism of the present application combines periodic detection with progressive adjustment to achieve efficient use of cache resources and dynamic adaptation of loads when a multi-core GPU system processes large-scale image data. Specifically, the number of requests for the target computing core to access the video memory of each associated computing core is periodically detected, and the number of historical access requests is converted into a weight ratio in combination with the second linear scaling factor and the second global offset constant, so that the frequently accessed associated computing cores obtain more cache capacity to cache more image data. In addition, the introduction of linear scaling factors and global offsets enhances the robustness of the algorithm. When the request amount of a certain associated computing core is zero, the offset can prevent its cache capacity from returning to zero to ensure the baseline resources of the zero-request associated computing core.

[0058] Furthermore, by limiting the dynamic reallocation of only part of the cache capacity each time, it is possible to maintain the stability of most of the cache, avoid performance shocks caused by global adjustments, and effectively balance the flexible adjustment performance and stable performance of the system cache.

[0059] In a feasible design, before partitioning the cache according to the request quantities corresponding to the associated computing dies in the previous partitioning cycle, the cache can be further partitioned by combining the access latency factors corresponding to the video memories of the associated computing dies during the operation. The specific steps include: Send a memory access latency perception request to each associated computing die. The memory access latency perception request is used to trigger the round-trip communication between the target computing die and the associated computing die; Determine the access latency value corresponding to each associated computing die according to the sending time and receiving time of each memory access latency perception request; Determine the second capacity that needs to be repartitioned in the remote data cache to be partitioned; Adjust the access latency value corresponding to each associated computing die based on a third linear scaling factor and a third global offset constant; According to the adjusted access latency value corresponding to each associated computing die, determine the second capacity ratio corresponding to each associated computing die. Wherein, the second capacity ratio is the ratio of the capacity that needs to be redistributed for the associated computing die to the second capacity, and the adjusted access latency value corresponding to each associated computing die is positively correlated with the corresponding second capacity ratio; Determine the final capacity corresponding to each associated computing die according to the new capacity, the second capacity ratio, and the second capacity corresponding to each associated computing die; Redistribute a part of the cache of the remote data cache to be partitioned according to the final capacity and the original capacity corresponding to each associated computing die.

[0060] Among them, the third linear scaling factor and the third global offset constant are set according to actual requirements, and the present application does not limit this.

[0061] For example, taking the partitioning of the remote data cache in the target computing die as an example, assume that the new capacity corresponding to the associated computing die determined according to the request quantity is . Further partitioning the remote data cache by combining the access latency factors corresponding to the video memories of the associated computing dies during the operation includes the following steps: (1) Send a memory access latency perception request to each associated computing die. When the associated computing die receives the memory access latency perception request, it immediately returns the request.

[0062] (2) By recording the time when the request is sent and the time when the request is received, calculate the access latency value between the target computing die and the video memory of each associated computing die in the current state (i.e., the time difference between the time when the request is received and the time when the request is sent). Among them, the access latency value corresponding to the associated computing die is denoted as .

[0063] (3) Determine that the second capacity to be repartitioned in the to-be-partitioned remote data cache is 1 / h1 of the total number of cache lines in the cache group, where h1 is a preset constant.

[0064] (4) Based on the third linear scaling factor and the third global offset constant , adjust the access latency value corresponding to each associated computing die to obtain the adjusted access latency value, where the adjusted access latency value corresponding to the associated computing die is ( ).

[0065] (5) Determine the second capacity ratio according to the adjusted access latency value: Sum the access latency values corresponding to each associated computing die to obtain . Among them, represents the access latency value corresponding to the associated computing die . Therefore, the second capacity ratio corresponding to the associated computing die is .

[0066] (6) Determine the final capacity corresponding to each associated computing die according to the new capacity, the second capacity ratio, and the second capacity corresponding to each associated computing die. Specifically, it can be regarded as first taking out 1 / h1 of the capacity from the new capacity corresponding to each associated computing die to participate in the repartitioning process. Then, the capacities taken out from each associated computing die constitute the second capacity, and the remaining capacity corresponding to each associated computing die is (h1 - 1) / h1 of the new capacity. At this time, the capacity that the associated computing die needs to be redistributed refers to the capacity that needs to be allocated to this associated computing die in the second capacity during this repartitioning process. Therefore, for the associated computing die , the final capacity = new capacity * (h1 - 1) / h1 + second capacity ratio * second capacity.

[0067] (7) Considering that only part of the cache in the remote data cache is repartitioned, the image data in the cache that is not repartitioned can be retained to improve the partitioning efficiency. Therefore, by comparing the size of the final capacity and the original capacity corresponding to each associated computing die, part of the cache of the to-be-partitioned remote data cache is repartitioned. For the specific method, refer to the solution in the foregoing embodiment of repartitioning part of the cache of the to-be-partitioned remote data cache by comparing the size of the new capacity and the original capacity corresponding to each associated computing die, which will not be elaborated here.

[0068] In order to further precisely perform dynamic partitioning on the remote data cache and further improve the performance of the multi-die GPU system in processing image data, based on dynamically partitioning the remote data cache according to the request quantity corresponding to each associated computing die, this application also considers the additional latency factor caused by bandwidth contention on the path when the target computing die travels to different associated computing dies. Specifically, it periodically detects the time latency of the target computing die accessing each associated computing die in the current state, combines the third linear scaling factor and the third global offset constant, and converts the access latency value into a weight ratio, so that the associated computing die with a higher access latency obtains more cache capacity to cache more image data. In addition, the introduction of the linear scaling factor and the global offset realizes the smooth attenuation of the hop count difference and avoids the polarization of the allocation difference.

[0069] Furthermore, by limiting that only a part of the cache capacity is dynamically reallocated each time, it is possible to maintain the stability of most of the cache, avoid the performance oscillation caused by global adjustment, and effectively balance the flexible adjustment performance and the stable performance of the system cache.

[0070] It should be understood that it is also possible to periodically perform dynamic partitioning of the cache only based on the access latency values corresponding to the video memories of each associated computing die. The steps include: after the end of each partitioning period, sending a memory access latency perception request to each associated computing die, and the memory access latency perception request is used to trigger the round-trip communication between the target computing die and the associated computing die; According to the sending time and receiving time of each memory access latency perception request, determine the access latency value corresponding to each associated computing die; Determine the second capacity that needs to be re-partitioned in the remote data cache to be partitioned; Based on the third linear scaling factor and the third global offset constant, adjust the access latency value corresponding to each associated computing die; According to the adjusted access latency value corresponding to each associated computing die, determine the second capacity ratio corresponding to each associated computing die, where the adjusted access latency value corresponding to each associated computing die is positively correlated with the corresponding second capacity ratio; According to the original capacity, the second capacity ratio, and the second capacity corresponding to each associated computing die, determine the final capacity corresponding to each associated computing die; According to the final capacity and the original capacity corresponding to each associated computing die, re-partition a part of the remote data cache to be partitioned.

[0071] For specific examples and effects, refer to the foregoing embodiments and will not be elaborated here.

[0072] It should be noted that the above scheme for dynamically partitioning the remote data cache to be partitioned is applicable to multi-die GPU systems with any topological structure, such asFigure 4 and Figure 5 For the topological structure shown, refer to the foregoing examples for the implementation manners, which will not be elaborated herein.

[0073] This application takes into account that the number of cache lines in each cache group for remote data cache to be partitioned is limited. When the number of computing dies in a multi-die GPU system is greater than the number of cache lines, since the video memory of each computing die needs to be allocated at least one cache line in each cache group, the partitioning cannot be performed according to the solution of the foregoing embodiment. Moreover, even when the number of cache lines is greater than the number of computing dies in the multi-die GPU system, when the number of computing dies in the multi-die GPU system is too large, the storage space required for the partitioning identification bits used to record the cache video memory identification in the cache lines is relatively large, wasting the on-chip storage resources and reducing the efficiency of accessing the image data in the remote video memory in large-scale image data processing tasks.

[0074] Based on the above problems, this application realizes the partitioning of remote data cache by partitioning a large-scale multi-die GPU. The method includes: If the number of computing dies included in the multi-die GPU exceeds a first preset threshold, partition the multi-die GPU according to the topological structure to obtain a target region and at least one associated region. The target region is the region to which the target computing die belongs, and each associated region is the region in the multi-die GPU other than the target region, that is, the associated region is the region in the multi-die GPU that does not include the target computing die. The number of computing dies included in each region is the same, and the maximum distance between the computing dies in each region is less than a second preset threshold; Based on the topological structure, determine the number of hops between regions between each region and the target region; Based on a first linear scaling factor and a first global offset constant, adjust the number of hops between regions corresponding to each region; According to the adjusted number of hops between regions corresponding to each region, determine the cache capacity ratio corresponding to each region, where the number of hops between regions corresponding to each region is positively correlated with the corresponding cache capacity ratio; According to the cache capacity ratio corresponding to each region and the capacity of the remote data cache to be partitioned, determine the capacity corresponding to each region in the remote data cache to be partitioned; Partition the remote data cache to be partitioned according to the capacity corresponding to each region, where the associated computing dies in each region share the storage space allocated to the region in the remote data cache to be partitioned.

[0075] Wherein, the first linear scaling factor and the first global offset constant are set according to actual requirements, and this application does not limit this.

[0076] The second preset threshold is set according to actual requirements, and this application does not limit this.

[0077] It should be noted that since the target area includes associated computing dies in addition to the target computing die, it is necessary to allocate a certain amount of cache in the to-be-partitioned remote data cache for the target area to cache the video memory image data of the associated computing dies in the target area.

[0078] The following is an example to illustrate the area partitioning method: As Figure 6 shown, the multi-die GPU includes 16 computing dies, namely computing die 0, computing die 1, computing die 2, computing die 3, computing die 4, computing die 5, computing die 6, computing die 7, computing die 8, computing die 9, computing die 10, computing die 11, computing die 12, computing die 13, computing die 14, and computing die 15. Each computing die corresponds to a video memory, and the video memories corresponding to computing dies 0 - 15 are video memory 0, video memory 1, video memory 2, video memory 3, video memory 4, video memory 5, video memory 6, video memory 7, video memory 8, video memory 9, video memory 10, video memory 11, video memory 12, video memory 13, video memory 14, and video memory 15 respectively. Among them, computing die 15 includes multiple computing units, and each computing unit can access the remote data cache and access the last-level cache through the on-chip interconnect network.

[0079] Figure 6 The shown topology can be divided into four areas, namely area 0, area 1, area 2, and area 3, and the structure inside each area is similar to the structure shown in Figure 3 It can be seen that when partitioning the areas, the computing dies adjacent in the topological space are partitioned into the same area.

[0080] In this case, the capacity of the remote cache in each computing die is partitioned according to the area partitioning, that is, cache lines are allocated for each area in the to-be-partitioned remote data cache corresponding to the target computing die. Taking computing die 15 as the target computing die as an example, it is located in area 3, and the video memory address spaces it needs to cache respectively correspond to area 0 (video memory 0 - 5), area 1 (video memory 2 - 7), area 2 (video memory 8 - 13), and area 3 (video memory 10 - 14).

[0081] The following is an example to illustrate the calculation method of the area corresponding capacity: Suppose the entire system has computing dies, divided into areas, and currently it is necessary to partition the capacity of the remote data caches in area (i.e., the target area). Suppose the distance between area (i.e., the associated area) ( ) and area is hops, then area The corresponding cache capacity ratio is . Among them, , , are all the numbers of regions, represents region and region is hops away.

[0082] For example, Figure 6 as shown, for computing die 15, it belongs to region 3, and the distance from region 0 is 2, the distances from regions 1 and 2 are 1, and the distance from region 3 is 0. Let be 30, be 300. Through the above formula, it can be calculated that 3 / 11 of the capacity in computing die 15 is used to cache the image data of region 0, 1 / 4 of the capacity is used to cache the image data of region 1, 1 / 4 of the capacity is used to cache the image data of region 2, and 5 / 22 of the capacity is used to cache the image data of region 3. Suppose the cache group includes 256-way caches, then the corresponding capacities of regions 0 - 3 are 70 ways, 64 ways, 64 ways, and 58 ways respectively. Among them, the 70-way cache is shared by the requests in the range of video memories 0, 1, 4, 5 in region 1, and so on. It can be seen that the capacity of the remote data cache partition corresponding to region 0 in the remote data cache of computing die 15 is the largest, the capacities of the remote data cache partitions corresponding to regions 1 and 2 in the remote data cache of computing die 15 are the same, and the capacity of the remote data cache partition corresponding to region 3 in the remote data cache of computing die 15 is the smallest.

[0083] In the multi-die GPU cache partitioning method described in the embodiments of the present application, in order to avoid the reduction of utilization rate caused by cache fragmentation due to excessive number of computing dies, when the number of computing dies included in the multi-die GPU exceeds the first preset threshold, the multi-die GPU is partitioned by the region partitioning condition that "the number of computing dies included in each region is the same, and the maximum distance between computing dies in each region is less than the second preset threshold", and multiple adjacent computing dies are combined into the same region, so that the computing dies within the same region share the remote data cache space. By the equal partitioning rule of "each region contains the same number of computing dies", the fairness of cache allocation for each region is ensured; by the constraint of "the maximum distance is less than the second preset threshold", it is ensured that the maximum distance between dies within the region is controllable, and the performance gain of capacity allocation optimization based on the number of hops between regions is avoided from being offset by high-hop access between computing dies within the region. Based on this, the present application allocates cache capacity based on the number of hops between regions, so that regions with a farther distance obtain a larger proportion of cache space, reduces the access frequency of image data in the remote video memory, increases the hit rate of image data in the remote video memory, and thus improves the overall performance.

[0084] In addition, the first linear scaling factor and the first global offset constant are used to adjust the number of hops between regions, thereby enhancing the robustness of the algorithm. When the number of hops between regions is zero (i.e., the number of hops between the target region and the target region is zero), the first global offset constant can prevent the corresponding cache capacity from being reset to zero, thereby preventing the associated computing core particles in the target region from being unallocated capacity in the remote data cache of the target computing core particles.

[0085] In summary, the present application can ensure that the cache capacity of different areas is different through the above method, increase the hit rate of image data in farther video memory, and improve the overall performance; it can also prevent the excessive number of cores from causing the cache to be over-divided and resulting in a decrease in utilization.

[0086] In a feasible design, after the remote data cache to be divided is divided according to the capacity corresponding to each area, the capacity of the remote data cache to be divided is periodically re-divided, and the step of re-dividing the capacity of the remote data cache to be divided after each division cycle includes: Detecting the number of second requests of the target computing core to access the video memory of each region in the last partition cycle, where the video memory of each region includes the video memory of each associated computing core in the region; Determining a third capacity that needs to be re-divided in the remote data cache to be divided; Adjusting the second request quantity corresponding to each region based on a fourth linear scaling factor and a fourth global offset constant; Determine the third capacity ratio corresponding to each region according to the adjusted second request quantity corresponding to each region, wherein the third capacity ratio is the ratio of the capacity required to be reallocated to the region to the third capacity, and the adjusted second request quantity corresponding to each region is positively correlated with the corresponding third capacity ratio; Determine the new capacity corresponding to each area according to the third capacity ratio, the original capacity and the third capacity corresponding to each area, wherein the original capacity is the capacity corresponding to the area in the previous division cycle; According to the new capacity and original capacity corresponding to each area, part of the cache to be divided into remote data caches is re-divided.

[0087] Among them, the fourth linear scaling factor and the fourth global offset constant are set according to actual needs, and this application does not limit this.

[0088] It should be understood that the second request quantity corresponding to each region is the sum of the request quantities of the target computing core particles accessing the region from the associated computing core particles.

[0089] For example, in the region Target computing core As an example, after static partitioning, the calculation core particle In the remote data cache, the region The number of cache lines corresponding to the video memory of (i.e., the original capacity), where ) is the number of regions included in the multi-die GPU system. Every clock cycle is a partitioning cycle for re-partitioning the remote data cache (e.g., ). After statically partitioning the remote data cache to be partitioned according to the capacity corresponding to each region, the part of the cache with the third capacity in the remote data cache to be partitioned is re-partitioned periodically by calculating the new capacity corresponding to each region. Specifically, each re-partitioning of the remote data cache to be partitioned includes the following steps: (1) Detect the number of times (i.e., the second request quantity) that the computing die (i.e., the target computing die) accesses the video memory of each region in the previous partitioning cycle. Among them, the region The corresponding second request quantity is , is the sum of the request quantities corresponding to the associated computing dies within the region , where represents the region number.

[0090] (2) Determine that the third capacity to be re-partitioned in the remote data cache to be partitioned is 1 / m2 of the total number of cache lines in the cache group, where m2 is a preset constant.

[0091] (3) Based on the fourth linear scaling factor and the fourth global offset constant , adjust the second request quantity corresponding to each region to obtain the adjusted second request quantity. Among them, the region The corresponding adjusted second request quantity is ( ).

[0092] (4) Determine the capacity ratio according to the adjusted second request quantity: Sum the second request quantities corresponding to each region to obtain , so the region The corresponding third capacity ratio is .

[0093] (5) Determine the new capacity corresponding to each area based on the third capacity ratio, original capacity and third capacity corresponding to each area. Specifically, it can be considered that 1 / m2 of the capacity corresponding to the original capacity of each area is first taken out to participate in the redivision process. Then the capacity taken out from each area constitutes the third capacity, and the remaining capacity corresponding to each area is (m2-1) / m2 of the original capacity. At this time, the capacity that needs to be reallocated to the area refers to the capacity that needs to be allocated to the area in the third capacity during this redivision process. Therefore, for the area , new capacity = original capacity * (m2-1) / m2 + third capacity ratio * third capacity.

[0094] (6) Considering that only part of the cache in the remote data cache is repartitioned, the image data in the cache that is not repartitioned can be retained to improve the partitioning efficiency. Therefore, by comparing the size of the new capacity and the original capacity corresponding to each area, the part of the cache to be partitioned in the remote data cache is repartitioned. The corresponding new capacity is the same as the original capacity, so there is no need to allocate a new cache line for the region. If the corresponding new capacity is smaller than the original capacity, the image data of the cache line of the corresponding capacity of the area is cleared to facilitate allocation to other areas. If the corresponding new capacity is larger than the original capacity, the cache lines released from other areas will be allocated to this area.

[0095] The dynamic cache partitioning mechanism of the present application realizes efficient use of cache resources and dynamic adaptation of loads in a multi-core GPU system by combining periodic detection with progressive adjustment. Specifically, the target computing core is periodically detected to calculate the number of requests for accessing the video memory of each area, and the number of historical access requests is converted into a weight ratio in combination with the fourth linear scaling factor and the fourth global offset constant, so that high-frequency access areas obtain more cache capacity to cache more image data. In addition, the introduction of linear scaling factors and global offsets enhances the robustness of the algorithm. When the number of requests in a certain area is zero, the offset can prevent its cache capacity from returning to zero to ensure the baseline resources of the computing cores in the zero-request area.

[0096] Furthermore, by limiting the dynamic reallocation of only part of the cache capacity each time, it is possible to maintain the stability of most of the cache, avoid performance shocks caused by global adjustments, and effectively balance the flexible adjustment performance and stable performance of the system cache.

[0097] In a feasible design, before the cache is divided according to the number of requests corresponding to each area in the previous division cycle, the cache can be further divided according to the access delay factor corresponding to the video memory in each area during the operation process. The specific steps include: Send a memory access latency awareness request to each associated computing die, where the memory access latency awareness request is used to trigger round-trip communication between the target computing die and the associated computing die; Determine the access latency value corresponding to each associated computing die according to the sending time and receiving time of each memory access latency awareness request; Determine the sum of the access latency values corresponding to the associated computing dies in each region as the access latency value corresponding to each region; Determine the fourth capacity that needs to be re-partitioned in the to-be-partitioned remote data cache; Adjust the access latency value corresponding to each region based on the fifth linear scaling factor and the fifth global offset constant; Determine the fourth capacity ratio corresponding to each region according to the adjusted access latency value corresponding to each region, where the fourth capacity ratio is the ratio of the capacity that needs to be re-allocated in the region to the fourth capacity, and the adjusted access latency value corresponding to each region has a positive correlation with the corresponding fourth capacity ratio; Determine the final capacity corresponding to each region according to the new capacity, the fourth capacity ratio, and the fourth capacity corresponding to each region; Re-partition the to-be-partitioned remote data cache according to the final capacity and the original capacity corresponding to each region.

[0098] For example, for the remote data cache in the target computing die in the region as an example, assuming the new capacity corresponding to the region determined according to the number of requests is . Further combining the access latency factors corresponding to the video memory of each region during operation, partitioning the remote data cache includes the following steps: (1) Send a memory access latency awareness request to each associated computing die. When the associated computing die receives the memory access latency awareness request, it immediately returns the request.

[0099] (2) By recording the time when the request is sent and received, calculate the die-to-die latency from the target computing die a to the video memory of each associated computing die in the current state.

[0100] (3) Determine the sum of the access latency values corresponding to the associated computing dies in each region as the access latency value corresponding to each region. Among them, the access latency value corresponding to the region is denoted as , is the sum of the access latency values corresponding to the associated computing dies in the region .

[0101] (4) Determine that the fourth capacity to be re-partitioned in the to-be-partitioned remote data cache is 1 / h2 of the total number of cached rows in the cache group, where h2 is a preset constant.

[0102] (5) Based on the fifth linear scaling factor and the fifth global offset constant , adjust the access latency value corresponding to each region to obtain the adjusted access latency value, where the adjusted access latency value corresponding to region is ( ).

[0103] (6) Determine the fourth capacity ratio according to the adjusted access latency value: Sum the access latency values corresponding to each region to obtain . Among them, represents the access latency value corresponding to region . Therefore, the fourth capacity ratio corresponding to region is .

[0104] (7) Determine the final capacity corresponding to each region according to the new capacity, the fourth capacity ratio, and the fourth capacity corresponding to each region. Specifically, it can be regarded as first taking out 1 / h2 of the capacity from the new capacity corresponding to each region to participate in the re-partitioning process. Then, the capacities taken out from each region constitute the fourth capacity, and the remaining capacity corresponding to each region is (h2 - 1) / h2 of the new capacity. At this time, the capacity that needs to be reallocated for a region refers to the capacity that needs to be allocated to this region in the fourth capacity during this re-partitioning process. Therefore, for region , the final capacity = new capacity * (h2 - 1) / h2 + fourth capacity ratio * fourth capacity.

[0105] (8) Considering that since only part of the cache in the remote data cache is re-partitioned, the image data in the cache that is not re-partitioned can be retained to improve the partitioning efficiency. Therefore, by comparing the final capacity and the original capacity corresponding to each region, re-partition part of the cache of the to-be-partitioned remote data cache. For the specific method, refer to the solution of re-partitioning part of the cache of the to-be-partitioned remote data cache by comparing the new capacity and the original capacity corresponding to each region in the foregoing embodiment, which will not be elaborated here.

[0106] To further precisely perform dynamic partitioning on the remote data cache, based on dynamically partitioning the remote data cache according to the request quantity corresponding to each region, the present application also considers the additional latency factor caused by bandwidth contention on the path when the target computing die travels to associated computing dies in different regions. Specifically, it periodically detects the sum of the time latencies for the target computing die to access the associated computing dies within each region in the current state, and combines the fifth linear scaling factor and the fifth global offset constant to convert the access latency value corresponding to each region into a weight ratio, so that the region with a higher access latency obtains more cache capacity to cache more image data. Additionally, the introduction of the linear scaling factor and the global offset realizes the smooth attenuation of the hop count difference, avoiding polarization of the allocation difference.

[0107] Furthermore, by limiting that only a part of the cache capacity is dynamically reallocated each time, it is possible to maintain the stability of most of the cache, avoid the performance oscillation caused by global adjustment, and effectively balance the flexible adjustment performance and the stability performance of the system cache.

[0108] It should be understood that the cache can also be partitioned only based on the access latency values corresponding to the video memories of each region. The steps include: After the end of each partitioning period, send a memory access latency perception request to each associated computing die. The memory access latency perception request is used to trigger the round-trip communication between the target computing die and the associated computing die; According to the sending time and receiving time of each memory access latency perception request, determine the access latency value corresponding to each associated computing die; Determine the sum of the access latency values corresponding to the associated computing dies within each region as the access latency value corresponding to each region; Determine the fourth capacity that needs to be repartitioned in the remote data cache to be partitioned; Based on the fifth linear scaling factor and the fifth global offset constant, adjust the access latency value corresponding to each region; According to the adjusted access latency value corresponding to each region, determine the fourth capacity ratio corresponding to each region, where the adjusted access latency value corresponding to each region is positively correlated with the corresponding fourth capacity ratio; According to the original capacity, the fourth capacity ratio, and the fourth capacity corresponding to each region, determine the final capacity corresponding to each region; According to the final capacity and the original capacity corresponding to each region, repartition the remote data cache to be partitioned.

[0109] For specific examples and effects, refer to the foregoing embodiments, which will not be elaborated here.

[0110] It should be noted that the above solution for dynamically partitioning the remote data cache to be partitioned in the case of multi-die GPU partitioning is applicable to multi-die GPU systems with any topology. For the implementation method, refer to the foregoing examples, and details are not repeated herein.

[0111] As Figure 7 shown, the present application also provides a multi-die GPU cache architecture, including: A topology awareness unit for obtaining the topology of a multi-die GPU. The multi-die GPU includes a target computing die for processing image data and at least one associated computing die. The target computing die is the computing die to be subjected to cache partitioning operations. The associated computing die is connected to the target computing die through an inter-die interconnect network. The remote data cache to be partitioned of the target computing die is used to cache image data from the video memory of at least one associated computing die; A topology awareness unit for, if the number of computing dies included in the multi-die GPU does not exceed a first preset threshold, determining the number of hops between each associated computing die and the target computing die based on the topology; A capacity calculation unit for determining the cache capacity ratio corresponding to each associated computing die according to the number of hops between each associated computing die. Among them, the number of hops between each associated computing die and the corresponding cache capacity ratio are positively correlated. The cache capacity ratio is the ratio of the required capacity of the associated computing die to the capacity of the remote data cache to be partitioned; The capacity calculation unit is further configured to determine the capacity corresponding to each associated computing die in the remote data cache to be partitioned according to the cache capacity ratio corresponding to each associated computing die and the capacity of the remote data cache to be partitioned; A cache control unit for partitioning the remote data cache to be partitioned according to the capacity corresponding to each associated computing die.

[0112] In a feasible design, the remote data cache to be partitioned adopts a set-associative mapping structure, including a plurality of cache groups. Each cache group contains multiple cache lines. Each cache line includes a partitioning identification bit, and the partitioning identification bit is used to store a video memory identifier. The video memory identifier is used to identify the video memory to which the image data cached by the cache line belongs; Among them, the cache control unit is configured to partition the remote data cache to be partitioned according to the capacity corresponding to each associated computing die, including: The cache control unit is configured to determine the video memory identifier stored in the partitioning identification bit of each cache line according to the capacity corresponding to each associated computing die; The cache control unit partitions the remote data cache to be partitioned by writing the corresponding video memory identifier to the partitioning identification bit of each cache line.

[0113] Exemplarily, the remote data cache to be partitioned includes S sets (sets) with a set-associative mapping structure. Each set has E cache lines (E-way), so there are a total of S*E cache lines in the cache. The cache block size in each cache line is B bytes (bytes), and the capacity of this cache is S*E*B bytes. It should be noted that in this application, the capacity of the cache can also be in units of the number of cache lines. Taking Figure 8 as an example, it can be seen that each cache line includes a valid bit, a tag bit, a partitioning identification bit, and a cache block. Among them, the valid bit is used to store a value indicating whether the image data in this cache line is valid; the tag bit is used to store the cache tag (Tag) for tag comparison during cache hit determination; the cache block is used to store the image data of the remote video memory; the partitioning identification bit is used to store the video memory identification or region identification assigned to the cache line.

[0114] In a feasible design, the topology awareness unit is further configured to, if the number of computing dies included in the multi-die GPU exceeds a first preset threshold, perform region partitioning on the multi-die GPU according to the topology structure to obtain a target region and at least one associated region. The target region is the region to which the target computing die belongs, and each associated region is the region other than the target region in the multi-die GPU. The number of computing dies included in each region is the same, and the maximum distance between the computing dies in each region is less than a second preset threshold; The topology awareness unit is further configured to determine the inter-region hop count between each region and the target region based on the topology structure; The capacity calculation unit is further configured to adjust the inter-region hop count corresponding to each region based on a first linear scaling factor and a first global offset constant; The capacity calculation unit is further configured to determine the cache capacity ratio corresponding to each region according to the adjusted inter-region hop count corresponding to each region, where the inter-region hop count corresponding to each region is positively correlated with the corresponding cache capacity ratio; The capacity calculation unit is further configured to determine the capacity corresponding to each region in the remote data cache to be partitioned according to the cache capacity ratio corresponding to each region and the capacity of the remote data cache to be partitioned; The cache control unit is further configured to partition the remote data cache to be partitioned according to the capacity corresponding to each region, where the associated computing dies in each region share the storage space allocated to the region in the remote data cache to be partitioned.

[0115] In a feasible design, the remote data cache to be partitioned adopts a set-associative mapping structure, includes multiple cache groups, each cache group contains multiple cache lines, and each cache line includes a partitioning identification bit. The partitioning identification bit is used to store a region identification, and the region identification is used to identify the region to which the image data cached by the cache line belongs; Among them, the cache control unit is used to divide the to-be-partitioned remote data cache according to the capacity corresponding to each area, including: The cache control unit is used to determine the area identifier stored in the partition identification bit of each cache line according to the capacity corresponding to each area; The cache control unit is used to divide the to-be-partitioned remote data cache by writing the corresponding area identifier to the partition identification bit of each cache line.

[0116] In a feasible design, the architecture further includes a request quantity detection unit. After dividing the to-be-partitioned remote data cache according to the capacity corresponding to each associated computing die, the cache control unit is further used to periodically re-divide the capacity of the to-be-partitioned remote data cache. After each division cycle ends, The request quantity detection unit is used to detect the first request quantity for accessing the video memory of each associated computing die by the target computing die in the previous division cycle; The cache control unit is further used to determine the first capacity that needs to be re-partitioned in the to-be-partitioned remote data cache; The capacity calculation unit is further used to adjust the first request quantity corresponding to each associated computing die based on the second linear scaling factor and the second global offset constant; The capacity calculation unit is further used to determine the first capacity ratio corresponding to each associated computing die according to the adjusted first request quantity corresponding to each associated computing die, where the first capacity ratio is the ratio of the capacity that needs to be re-allocated by the associated computing die to the first capacity, and the adjusted first request quantity corresponding to each associated computing die is positively correlated with the corresponding first capacity ratio; The capacity calculation unit is further used to determine the new capacity corresponding to each associated computing die according to the first capacity ratio, the original capacity, and the first capacity corresponding to each associated computing die, where the original capacity is the capacity corresponding to the associated computing die in the previous division cycle; The cache control unit is further used to re-partition a part of the cache in the to-be-partitioned remote data cache according to the new capacity and the original capacity corresponding to each associated computing die.

[0117] In a feasible design, the architecture further includes a memory access latency awareness unit, where The memory access latency awareness unit is used to send a memory access latency awareness request to each associated computing die, and the memory access latency awareness request is used to trigger the round-trip communication between the target computing die and the associated computing die; The memory access latency awareness unit is further used to determine the access latency value corresponding to each associated computing die according to the sending time and the receiving time of each memory access latency awareness request; The cache control unit is further used to determine the second capacity that needs to be re-partitioned in the to-be-partitioned remote data cache. The capacity calculation unit is further configured to adjust the access latency value corresponding to each associated computing die based on a third linear scaling factor and a third global offset constant; The capacity calculation unit is further configured to determine a second capacity ratio corresponding to each associated computing die according to the adjusted access latency value corresponding to each associated computing die, where the second capacity ratio is the ratio of the capacity that needs to be redistributed for the associated computing die to the second capacity, and the adjusted access latency value corresponding to each associated computing die is positively correlated with the corresponding second capacity ratio; The capacity calculation unit is further configured to determine the final capacity corresponding to each associated computing die according to the new capacity, the second capacity ratio, and the second capacity corresponding to each associated computing die; The cache control unit is further configured to re-partition a partial cache of the remote data cache to be partitioned according to the final capacity and the original capacity corresponding to each associated computing die.

[0118] In a feasible design, after partitioning the remote data cache to be partitioned according to the capacity corresponding to each region, the cache control unit is further configured to periodically re-partition the capacity of the remote data cache to be partitioned. After each partitioning cycle ends, The request quantity detection unit is further configured to detect a second request quantity for accessing the video memory of each region by the target computing die in the previous partitioning cycle. The video memory of each region includes the video memories of the associated computing dies within the region; The cache control unit is further configured to determine a third capacity that needs to be re-partitioned in the remote data cache to be partitioned; The capacity calculation unit is further configured to adjust the second request quantity corresponding to each region based on a fourth linear scaling factor and a fourth global offset constant; The capacity calculation unit is further configured to determine a third capacity ratio corresponding to each region according to the adjusted second request quantity corresponding to each region, where the adjusted second request quantity corresponding to each region is positively correlated with the corresponding third capacity ratio; The capacity calculation unit is further configured to determine the new capacity corresponding to each region according to the third capacity ratio, the original capacity, and the third capacity corresponding to each region, where the third capacity ratio is the ratio of the capacity that needs to be redistributed for the region to the third capacity, and the original capacity is the capacity corresponding to the region in the previous partitioning cycle; The cache control unit is further configured to re-partition a partial cache of the remote data cache to be partitioned according to the new capacity and the original capacity corresponding to each region.

[0119] In a feasible design, the memory access latency-aware unit is further configured to send a memory access latency-aware request to each associated computing die, where the memory access latency-aware request is used to trigger the round-trip communication between the target computing die and the associated computing die; after determining the access latency value corresponding to each associated computing die according to the sending time and receiving time of each memory access latency-aware request, the sum of the access latency values corresponding to the associated computing dies in each region is determined as the access latency value corresponding to each region; The cache control unit is further configured to determine a fourth capacity that needs to be re-partitioned in the to-be-partitioned remote data cache; The capacity calculation unit is further configured to adjust the access latency value corresponding to each region based on a fifth linear scaling factor and a fifth global offset constant; The capacity calculation unit is further configured to determine a fourth capacity ratio corresponding to each region according to the adjusted access latency value corresponding to each region, where the fourth capacity ratio is the ratio of the capacity that needs to be reallocated in the region to the fourth capacity, and the adjusted access latency value corresponding to each region is positively correlated with the corresponding fourth capacity ratio; The capacity calculation unit is further configured to determine the final capacity corresponding to each region according to the new capacity, the fourth capacity ratio, and the fourth capacity corresponding to each region; The cache control unit is further configured to re-partition the to-be-partitioned remote data cache according to the final capacity and the original capacity corresponding to each region.

[0120] In a feasible design, the capacity calculation unit is implemented in the following manner. According to the number of hops between dies corresponding to each associated computing die, the cache capacity ratio corresponding to each associated computing die is determined: Based on a sixth linear scaling factor and a sixth global offset constant, the number of hops between dies corresponding to each associated computing die is adjusted; According to the adjusted number of hops between dies corresponding to each associated computing die, the cache capacity ratio corresponding to each associated computing die is determined.

[0121] For other implementation manners and effects of the above multi-die GPU cache architecture, refer to the description in the embodiment of the cache partitioning method, which will not be elaborated here.

[0122] Exemplarily, based on the above cache architecture, the cache control unit initializes and dynamically manages the partitioning identification bits in the following manner: During initialization, the cache control unit writes the partitioning identification bit of the corresponding cache line as the corresponding video memory number or region number according to the calculation result of the partitioning algorithm.

[0123] During dynamic adjustment, the cache control unit updates the partitioning identification bits of some caches according to the periodic partitioning result to ensure that the cache line stores only the image data from the specified source.

[0124] Based on the above multi-die GPU cache architecture, the present application provides a data access method for a multi-die GPU cache, and the method includes: The computing unit of the target computing die generates a request including a data source identifier, an access address, and an identifier of the computing unit, and the data source identifier is a video memory identifier or a region identifier; If the request misses the L1 cache inside the computing unit and the access address includes the address of the remote video memory, send the request to the cache control unit; The cache control unit parses the request to obtain the data source identifier, the access address, and the identifier of the computing unit of the request; The cache control unit locates the target cache group according to the access address; The cache control unit performs a request hit judgment on each cache line in the target cache group in the storage order according to the access address and the data source identifier; If a hit occurs, the cache control unit extracts the target image data from the target cache line according to the access address, and sends the target image data to the computing unit according to the identifier of the computing unit; If a miss occurs, the cache control unit sends an indication message to the computing unit, and the indication message is used to indicate that the target image data is not stored in the target cache group. After receiving the indication message, the computing unit sends a request to the target associated computing die corresponding to the video memory mapped by the access address through the die-to-die interconnection network.

[0125] In the data access method for the multi-die GPU cache provided in the above embodiment, when generating an access request, by adding data source identifier data to the request, when performing a request hit judgment, in addition to comparing the check bits in the traditional access address, a data source identifier check is added. It can cooperate with the multi-die GPU cache partition structure based on the partition identifier bits to achieve efficient cross-die memory access.

[0126] The following combines Figure 7 and Figure 8 The shown cache architecture is used to give an example of the above data access method: (1) Taking the access address as being composed of the following three parts (as Figure 9 shown) as an example to generate a request: Block offset (b bits): Indicates the byte-level offset position of the target image data within the cache block; Cache group index (s bits): Used to locate the target cache group (set) mapped by the address; Cache tag (t bits): Matches the tag stored in the cache line to verify the data validity.

[0127] Correspondingly, when the computing unit of the target computing die initiates an access request, a request packet including the following information is generated: Data source identifier: identifies the video memory (video memory identifier) ​​or region (region identifier) ​​to which the target image data belongs; Access address: contains the offset within the block, cache group index, and cache tag; Compute unit ID: Identifies the compute unit that initiated the request.

[0128] (2) Cache access process: (2.1) If the request hits the L1 cache inside the computing unit, the target image data is obtained from the L1 cache.

[0129] If the request does not hit the L1 cache inside the computing unit and the access address points to the remote video memory (i.e., non-local video memory address), the request is forwarded to the cache control unit.

[0130] (2.2) The cache control unit parses the request and extracts the following key information: Data source identification (video memory identification / region identification); Access address (block offset, cache group index, cache tag); Compute unit ID.

[0131] (2.3) The cache control unit locates the target cache group according to the cache group index (s-bit set index) in the access address using a preset mapping rule.

[0132] (4.%2) Cache line hit judgment In the target cache group, the cache control unit performs hit judgment on each cache line in sequence according to the storage order. The judgment conditions include: The valid bit is 1: the cache line data is valid; Data source identifier matching: The requested data source identifier is consistent with the identifier stored in the partition identifier of the cache line, wherein the identifier stored in the partition identifier is written by the cache control unit to complete cache initialization when the remote data cache to be partitioned is partitioned, and its value is the video memory number or area number calculated by the partition algorithm, which is used to limit the source of image data that can be stored in the cache line; Tag match: The cache tag of the access address matches the tag stored in the cache line.

[0133] (2.5) If the above conditions are met at the same time, it is determined as a cache hit and executes: (2.5.1) extracting the target image data from the cache block of the hit cache line according to the offset within the block; (2.5.2) Return the target image data to the computing unit that initiated the request according to the computing unit identifier.

[0134] (2.6): If any of the conditions is not met, it is considered a cache miss and the following is executed: (2.6.1) The cache control unit sends an indication message to the computing unit, stating that the target image data is not cached; (2.6.2) The computing unit initiates a data request to the corresponding target associated computing chiplet through the inter-chiplet interconnection network according to the remote memory location mapped by the access address.

[0135] The basic principles of the present application are described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, effects, etc. mentioned in the present application are only examples and not limitations, and it cannot be considered that these advantages, strengths, effects, etc. are required by each embodiment of the present application. In addition, the specific details disclosed above are only for the purpose of illustration and ease of understanding, not for limitation, and the above details do not limit the present application to being implemented by adopting the above specific details.

[0136] It should be understood that, although the steps in the flowchart of the accompanying drawings are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.

[0137] The block diagrams of the devices, apparatuses, equipment, and systems involved in this application are only illustrative examples, and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagram. As will be appreciated by those skilled in the art, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open words, referring to "including but not limited to", and can be used interchangeably with them. The words "or" and "and" used here refer to the words "and / or" and can be used interchangeably with them, unless the context clearly indicates otherwise. The words "such as" used here refer to the phrase "such as but not limited to", and can be used interchangeably with them.

[0138] It should also be noted that in the apparatus, device and method of the present application, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present application.

[0139] The foregoing description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present application. Thus, the present application is not intended to be limited to the aspects shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0140] The foregoing description has been presented for purposes of illustration and description. In addition, this description is not intended to limit the embodiments of the present application to the forms disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.

Claims

1. A method for partitioning a multi-die GPU cache, characterized in that, include: Acquire a topological structure of a multi-core GPU, wherein the multi-core GPU includes a target computing core particle for processing image data and at least one associated computing core particle, wherein the target computing core particle is a computing core particle to be subjected to a cache partitioning operation, wherein the associated computing core particle is connected to the target computing core particle through an inter-core interconnection network, and wherein the remote data cache to be partitioned of the target computing core particle is used to cache image data from a video memory of the at least one associated computing core particle; If the number of computing cores included in the multi-core GPU does not exceed a first preset threshold, determining the number of inter-core hops between each of the associated computing cores and the target computing core based on the topological structure; Determine the cache capacity ratio corresponding to each of the associated computing core particles according to the number of inter-core hops corresponding to each of the associated computing core particles, wherein the number of inter-core hops corresponding to each of the associated computing core particles is positively correlated with the corresponding cache capacity ratio, and the cache capacity ratio is the ratio of the capacity required by the associated computing core particle to the capacity of the remote data cache to be divided; Determine the capacity corresponding to each of the associative computing core particles in the remote data cache to be divided according to the cache capacity ratio corresponding to each of the associative computing core particles and the capacity of the remote data cache to be divided; The remote data cache to be divided is divided according to the capacity corresponding to each of the associated computing core particles.

2. The method according to claim 1, characterized in that, The method comprises: If the number of computing cores included in the multi-core GPU exceeds a first preset threshold, the multi-core GPU is divided into regions according to the topological structure to obtain a target region and at least one associated region, wherein the target region is a region to which the target computing core belongs, and each associated region is a region of the multi-core GPU other than the target region, each region includes the same number of computing cores, and a maximum distance between computing cores in each region is less than a second preset threshold; Based on the topological structure, determining the number of inter-area hops between each area and the target area; Based on the first linear scaling factor and the first global offset constant, adjusting the number of inter-region hops corresponding to each region; Determine the cache capacity ratio corresponding to each area according to the adjusted inter-area hop number corresponding to each area, wherein the inter-area hop number corresponding to each area is positively correlated with the corresponding cache capacity ratio; Determine the capacity corresponding to each area in the remote data cache to be divided according to the cache capacity ratio corresponding to each area and the capacity of the remote data cache to be divided; The remote data cache to be divided is divided according to the capacity corresponding to each area, wherein the associated computing core particles in each area share the storage space allocated to the corresponding area in the remote data cache to be divided.

3. The method according to claim 1, characterized in that, After the remote data cache to be divided is divided according to the capacity corresponding to each of the associated computing core particles, the capacity of the remote data cache to be divided is periodically re-divided, and the step of re-dividing the capacity of the remote data cache to be divided after each division cycle ends includes: Detect the first request quantity for the target computing die to access the video memory of each associated computing die in the previous partitioning period; Determine the first capacity to be repartitioned in the to-be-partitioned remote data cache; Adjust the first request quantity corresponding to each associated computing die based on a second linear scaling factor and a second global offset constant; Determine the first capacity ratio corresponding to each associated computing die according to the adjusted first request quantity corresponding to each associated computing die, where the first capacity ratio is the ratio of the capacity to be redistributed by the associated computing die to the first capacity, and the adjusted first request quantity corresponding to each associated computing die is positively correlated with the corresponding first capacity ratio; Determine the new capacity corresponding to each associated computing die according to the first capacity ratio, the original capacity, and the first capacity corresponding to each associated computing die, where the original capacity is the capacity corresponding to the associated computing die in the previous partitioning period; Repartition a partial cache of the to-be-partitioned remote data cache according to the new capacity and the original capacity corresponding to each associated computing die.

4. The method according to claim 3, characterized in that, The repartitioning a partial cache of the to-be-partitioned remote data cache according to the new capacity and the original capacity corresponding to each associated computing die includes: Send a memory access latency awareness request to each associated computing die, where the memory access latency awareness request is used to trigger the round-trip communication between the target computing die and the associated computing die; Determine the access latency value corresponding to each associated computing die according to the sending time and the receiving time of each memory access latency awareness request; Determine the second capacity to be repartitioned in the to-be-partitioned remote data cache; Adjust the access latency value corresponding to each associated computing die based on a third linear scaling factor and a third global offset constant; Determine the second capacity ratio corresponding to each associated computing die according to the adjusted access latency value corresponding to each associated computing die, where the second capacity ratio is the ratio of the capacity to be redistributed by the associated computing die to the second capacity, and the adjusted access latency value corresponding to each associated computing die is positively correlated with the corresponding second capacity ratio; Determine the final capacity corresponding to each associated computing die according to the new capacity, the second capacity ratio, and the second capacity corresponding to each associated computing die; Repartition a partial cache of the to-be-partitioned remote data cache according to the final capacity and the original capacity corresponding to each associated computing die.

5. The method according to claim 2, wherein After partitioning the to-be-partitioned remote data cache according to the capacity corresponding to each region, periodically repartition the capacity of the to-be-partitioned remote data cache. The steps of repartitioning the capacity of the to-be-partitioned remote data cache after each partitioning period include: Detect the second request quantity for the target computing die to access the video memory of each region in the previous partitioning period, where the video memory of each region includes the video memories of each associated computing die within the region; Determine a third capacity in the remote data cache to be partitioned that needs to be repartitioned; Based on a fourth linear scaling factor and a fourth global offset constant, adjust the second request quantity corresponding to each region; According to the adjusted second request quantity corresponding to each region, determine the third capacity ratio corresponding to each region, where the third capacity ratio is the ratio of the capacity that needs to be reallocated in the region to the third capacity, and the adjusted second request quantity corresponding to each region is positively correlated with the corresponding third capacity ratio; According to the third capacity ratio, the original capacity, and the third capacity corresponding to each region, determine the new capacity corresponding to each region, where the original capacity is the capacity corresponding to the region in the previous partitioning cycle; According to the new capacity and the original capacity corresponding to each region, repartition a part of the remote data cache to be partitioned.

6. The method according to claim 5, wherein The repartitioning a part of the remote data cache to be partitioned according to the new capacity and the original capacity corresponding to each region includes: Send a memory access latency awareness request to each of the associated computing dies, where the memory access latency awareness request is used to trigger the round-trip communication between the target computing die and the associated computing die; According to the sending time and the receiving time of each memory access latency awareness request, determine the access latency value corresponding to each associated computing die; Determine the sum of the access latency values corresponding to the associated computing dies in each region as the access latency value corresponding to each region; Determine a fourth capacity in the remote data cache to be partitioned that needs to be repartitioned; Based on a fifth linear scaling factor and a fifth global offset constant, adjust the access latency value corresponding to each region; According to the adjusted access latency value corresponding to each region, determine the fourth capacity ratio corresponding to each region, where the fourth capacity ratio is the ratio of the capacity that needs to be reallocated in the region to the fourth capacity, and the adjusted access latency value corresponding to each region is positively correlated with the corresponding fourth capacity ratio; According to the new capacity, the fourth capacity ratio, and the fourth capacity corresponding to each region, determine the final capacity corresponding to each region; According to the final capacity and the original capacity corresponding to each region, repartition the remote data cache to be partitioned.

7. The method according to any one of claims 1-6, characterized in that, According to the number of die-to-die hops corresponding to each associated computing die, determine the cache capacity ratio corresponding to each associated computing die, including: Based on a sixth linear scaling factor and a sixth global offset constant, adjust the number of die-to-die hops corresponding to each associated computing die; According to the adjusted number of die-to-die hops corresponding to each associated computing die, determine the cache capacity ratio corresponding to each associated computing die.

8. A multi-die GPU cache architecture, characterized in that, including: A topology awareness unit for obtaining the topology of a multi-die GPU, where the multi-die GPU includes a target computing die for processing image data and at least one associated computing die. The target computing die is the computing die for which cache partitioning operation is to be performed. The associated computing die is connected to the target computing die through an inter-die interconnect network. The to-be-partitioned remote data cache of the target computing die is used to cache image data from the video memory of the at least one associated computing die; The topology awareness unit is configured to, if the number of computing dies included in the multi-die GPU does not exceed a first preset threshold, determine the number of inter-die hops between each associated computing die and the target computing die based on the topology; A capacity calculation unit for determining the cache capacity ratio corresponding to each associated computing die according to the number of inter-die hops corresponding to each associated computing die, where the number of inter-die hops corresponding to each associated computing die is positively correlated with the corresponding cache capacity ratio, and the cache capacity ratio is the ratio of the required capacity of the associated computing die to the capacity of the to-be-partitioned remote data cache; The capacity calculation unit is further configured to determine the capacity corresponding to each associated computing die in the to-be-partitioned remote data cache according to the cache capacity ratio corresponding to each associated computing die and the capacity of the to-be-partitioned remote data cache; A cache control unit is configured to partition the to-be-partitioned remote data cache according to the capacity corresponding to each associated computing die.

9. The multi-die GPU cache architecture according to claim 8, characterized in that, The to-be-partitioned remote data cache adopts a set-associative mapping structure, including a plurality of cache groups, each cache group includes multiple cache lines, and each cache line includes a partitioning identification bit, and the partitioning identification bit is used to store a video memory identifier, and the video memory identifier is used to identify the video memory to which the image data cached by the cache line belongs; Wherein, the cache control unit is configured to partition the to-be-partitioned remote data cache according to the capacity corresponding to each associated computing die, including: The cache control unit is configured to determine the video memory identifier stored in the partitioning identification bit of each cache line according to the capacity corresponding to each associated computing die; The cache control unit is configured to partition the to-be-partitioned remote data cache by writing the corresponding video memory identifier to the partitioning identification bit of each cache line.

10. A data access method for a multi-die GPU cache, characterized in that, Applied to the multi-die GPU cache architecture as claimed in claim 8 or 9, the multi-die GPU cache architecture includes a cache control unit and a remote data cache. The remote data cache adopts a set-associative mapping structure, including a plurality of cache groups, each cache group includes multiple cache lines, and each cache line includes a partitioning identification bit, and the partitioning identification bit is configured to store a video memory identifier or a region identifier. The video memory identifier is used to identify the video memory to which the image data cached by the cache line belongs, and the region identifier is used to identify the region to which the image data cached by the cache line belongs. The method includes: The computing unit of the target compute die generates a request including a data source identifier, an access address, and an identifier of the computing unit, where the data source identifier is a video memory identifier or a region identifier; If the request misses the L1 cache inside the computing unit and the access address includes the address of the remote video memory, send the request to the cache control unit; The cache control unit parses the request to obtain the data source identifier, the access address, and the identifier of the computing unit of the request; The cache control unit locates the target cache group according to the access address; The cache control unit performs a request hit judgment on each cache line in the target cache group in the storage order according to the access address and the data source identifier; If a hit occurs, the cache control unit extracts the target image data from the target cache line according to the access address, and sends the target image data to the computing unit according to the identifier of the computing unit; If a miss occurs, the cache control unit sends indication information to the computing unit, where the indication information is used to indicate that the target image data is not stored in the target cache group. After receiving the indication information, the computing unit sends the request to the target associated compute die corresponding to the video memory mapped by the access address through the die - to - die interconnect network.

Citation Information

Patent Citations

  • Multi-core GPU chip architecture system

    CN118247119A

  • Core particle system memory controller layout optimization method

    CN118332999A

  • Architecture optimization method, device and equipment of multi-chip integration system and storage medium

    CN118519925A

  • Cache architecture optimization method and device of multi-chip integration system and storage medium

    CN119917446A

  • Partitioning Caches for Sub-Entities in Computing Devices

    US20140173211A1

Cited By

  • Performance monitoring method and device of bare chip interconnection system, electronic equipment and computer program product

    CN121029678A

  • Method, device, electronic equipment and computer program product for performance monitoring of a die interconnect system

    CN121029678B

  • Multi-core-particle interconnection system based on RISC-V architecture and packaging chip

    CN122412363A