The invention belongs to the technical field of semiconductors, and discloses a multi-core GPU
interconnection architecture and a self-adaptive cache
allocation method thereof.The multi-core GPU
interconnection architecture comprises a plurality of GPU clusters and a global cluster, the global cluster comprises a global
router, a global scheduler, a
directory, a memory and a plurality of GPU clusters, each
GPU cluster is composed of a plurality of GPU cores, and the GPU cores are arranged in the global
router. Each GPU core grain comprises a plurality of independent primary data caches, secondary data caches, a local
router and a
network interface, and the
network interface judges whether a current request is processed by the GPU core grains in a
GPU cluster or not according to whether a target address in a target node belongs to the
GPU cluster or not, so that a high-level multi-core grain
interconnection architecture supporting cross-cluster communication is formed. According to the method, performance is optimized by adapting to different access
modes, private and shared cache
modes are switched in real time, single
data access delay is reduced, the overall
system efficiency is improved, and the method is used for a multi-core-particle heterogeneous
system with a plurality of GPUs.