Modular integrated circuit unit, distributed storage unit access method and system

By introducing local and virtual remote storage units into modular integrated circuit units, using address interleaving technology and virtual storage mapping, the problem of non-uniform memory access in distributed storage systems is solved, unified memory access and load balancing is achieved, and system performance and ease of use are improved.

CN120256380AActive Publication Date: 2025-07-04BEIJING TSINGMICRO INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510733785.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-04
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

The prior art has problems such as imbalance in latency, high complexity and low resource utilization caused by non-unified memory access in distributed storage systems. Especially in multi-core or multi-chip systems, cross-core or cross-chip storage units cannot be directly accessed, which limits system performance and ease of use.

Method used

Modular integrated circuit units are adopted, by setting up local storage units and virtual remote storage units, using address interleaving technology and virtual storage mapping, unified address access and load balancing of distributed storage units on multiple core particles is achieved, ensuring unified memory access of the computing system.

Benefits of technology

It realizes unified memory access to the computing system, balances the access delay of storage units, reduces complexity, improves resource utilization, and improves the performance and ease of use of distributed storage systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256380A_ABST
    Figure CN120256380A_ABST
Patent Text Reader

Abstract

The invention provides a modularized integrated circuit unit and a distributed storage unit access method and system, and the modularized integrated circuit unit comprises at least one memory controller which sends a storage address access request to a storage unit; the at least one storage unit provides unified memory access for the storage address access request, and the storage unit comprises a local storage unit and a virtual far-end storage unit; the at least one computing unit initiates a storage address access request; and the at least one internet-on-chip is used for interleaving the storage address indicated by the storage address access request, generating interleaved storage address distribution, determining a corresponding target storage unit, and storing the target storage unit by setting a local storage unit and a virtual far-end storage unit. Uniform address access and load balancing of the distributed storage units on the modularized integrated circuit units are achieved through the address interleaving technology and the virtual storage mapping, the resource utilization rate is increased, and the performance and usability of the distributed storage system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer storage, and in particular to a modular integrated circuit unit, a method and a system for accessing distributed storage units. Background Art

[0002] In order to increase the available storage capacity and bandwidth of the entire computing system, multiple independent storage device units are generally placed in the system. At the same time, in order to provide a unified memory access at the computing level, the access to these independent storage units is managed using an address interleaving method. Meanwhile, with the development of Chiplet technology and the emergence of new storage devices, a large number of storage units are placed on multiple different chiplets or chips, resulting in the distributed characteristics of storage devices in the system. Even if the address interleaving of the prior art is applied to distributed storage units, since the information of storage units across chiplets or chips cannot be directly obtained, it leads to non-unified memory access to the computing system. Non-unified memory access has problems such as uneven latency, high complexity, and low resource utilization, which limit the performance and usability of distributed storage systems. Summary of the Invention

[0003] An object of the present invention is to provide a modular integrated circuit unit. By setting up a local storage unit and a virtual remote storage unit, and using address interleaving technology and virtual storage mapping, the distributed storage units on multiple chiplets achieve unified address access and load balancing, ensuring unified memory access to the computing system, balancing the access latency of each storage unit, reducing complexity, improving resource utilization, and enhancing the performance and usability of the distributed storage system. Another object of the present invention is to provide a method for accessing distributed storage units. Another object of the present invention is to provide a system for accessing distributed storage units. Still another object of the present invention is to provide a computer device.

[0004] To achieve the above object, on the one hand, the present invention discloses a modular integrated circuit unit, including:

[0005] At least one computing unit, at least one memory controller, at least one on-chip interconnection network, and at least one storage unit;

[0006] The memory controller is configured to send a storage address access request sent by the on-chip interconnection network to the storage unit;

[0007] The storage unit is configured to provide unified memory access to the storage address access request sent by the memory controller. The storage unit includes a local storage unit and a virtual remote storage unit. The local storage unit is distributed on the current modular integrated circuit unit, and the virtual remote storage unit is distributed on modular integrated circuit units other than the current modular integrated circuit unit;

[0008] The computing unit is used to initiate a storage address access request;

[0009] The on-chip interconnection network is used to interleave the storage addresses indicated by the storage address access requests sent by the computing units, generate an interleaved storage address distribution; determine the target storage unit corresponding to the storage address access request according to the interleaved storage address distribution, and send the storage address access request to the corresponding target storage unit through the memory controller, where the target storage unit is the target local storage unit or the target virtual remote storage unit.

[0010] Preferably, the modular integrated circuit unit is a die in a multi-die package or a chip in a computing system composed of independent chips.

[0011] Preferably, the on-chip interconnection network is specifically used to divide the storage addresses according to a preset interleaving granularity to obtain storage address distributions in multiple intervals; uniformly map the storage address distributions in multiple intervals to different storage units on different modular integrated circuit units to obtain the interleaved storage address distribution.

[0012] Preferably, the on-chip interconnection network is specifically used to send the storage address access requests distributed on the current modular integrated circuit unit where it is located to the target local storage unit through the memory controller according to the interleaved storage address distribution, or send the storage address access requests distributed on modular integrated circuit units other than the current modular integrated circuit unit where it is located to the target virtual remote storage unit through the interconnection interface.

[0013] The present invention also discloses a method for accessing distributed storage units, and the method includes:

[0014] The computing unit initiates a storage address access request to the on-chip interconnection network;

[0015] The on-chip interconnection network interleaves the storage addresses indicated by the storage address access requests to generate an interleaved storage address distribution;

[0016] The on-chip interconnection network determines the target storage unit corresponding to the storage address access request according to the interleaved storage address distribution, and sends the storage address access request to the memory controller in the corresponding modular integrated circuit unit;

[0017] The memory controller sends the storage address access request to the target storage unit, where the target storage unit is the target local storage unit or the target virtual remote storage unit;

[0018] The storage unit includes a local storage unit and a virtual remote storage unit. The local storage units are distributed on the current modular integrated circuit unit, and the virtual remote storage units are distributed on modular integrated circuit units other than the current modular integrated circuit unit.

[0019] Preferably, the on-chip interconnection network interleaves the storage addresses indicated by the storage address access requests to generate an interleaved storage address distribution, including:

[0020] The on-chip interconnection network divides the storage addresses according to a preset interleaving granularity to obtain storage address distributions in multiple intervals; and evenly maps the storage address distributions in multiple intervals to different storage units on different modular integrated circuit units to obtain the interleaved storage address distribution.

[0021] Preferably, the on-chip interconnection network sends the storage address access requests to the memory controllers within the corresponding modular integrated circuit units according to the interleaved storage address distribution, including:

[0022] The on-chip interconnection network sends the storage address access requests distributed on the current modular integrated circuit unit to the corresponding memory controllers, or sends the storage address access requests distributed on modular integrated circuit units other than the current modular integrated circuit unit to the memory controllers within the corresponding modular integrated circuit units through the interconnection interfaces.

[0023] Preferably, the interleaved storage address distribution includes the storage address distributions corresponding to the intervals of each storage unit;

[0024] The on-chip interconnection network sends the storage address access requests distributed on the current modular integrated circuit unit to the corresponding memory controllers according to the interleaved storage address distribution, including:

[0025] The on-chip interconnection network matches the storage addresses indicated by the storage address access requests distributed on the current modular integrated circuit unit with the storage address distributions corresponding to the intervals of each storage unit in the storage address distribution to determine the target local storage unit, and sends the storage address access requests distributed on the current modular integrated circuit unit to the corresponding memory controllers.

[0026] Preferably, the memory controller sends the storage address access requests to the target storage unit, including:

[0027] The memory controller sends the storage address access requests to the target local storage unit.

[0028] Preferably, the on-chip interconnection network stores the interconnection interfaces bound to each storage unit;

[0029] Sending a storage address access request on a modular integrated circuit unit located outside the current modular integrated circuit unit to a memory controller within the corresponding modular integrated circuit unit through an interconnection interface, including:

[0030] The on-chip interconnection network matches the storage address indicated by the storage address access request on the modular integrated circuit unit located outside the current modular integrated circuit unit with the storage address distribution of each storage unit corresponding interval in the storage address distribution, and determines the target interconnection interface bound to the target remote modular integrated circuit unit;

[0031] Through the target interconnection interface, the storage address access request on the modular integrated circuit unit located outside the current modular integrated circuit unit is sent to the target remote modular integrated circuit unit, so that the target remote modular integrated circuit unit can send the storage address access request to the corresponding memory controller.

[0032] Preferably, the memory controller sending the storage address access request to the target storage unit includes:

[0033] The memory controller sends the storage address access request to the target virtual remote storage unit.

[0034] The present invention also discloses a distributed storage unit access system, the system includes: a plurality of modular integrated circuit units as described above, and the plurality of modular integrated circuit units are communicatively connected through an interconnection interface, and the interconnection interface includes an inter-die interconnection interface or an inter-chip interconnection interface;

[0035] If the modular integrated circuit unit is a die in a multi-die package, the dies are communicatively connected through an inter-die interconnection interface;

[0036] If the modular integrated circuit unit is a chip in a computing system composed of independent chips, the chips are communicatively connected through an inter-chip interconnection interface.

[0037] The present invention also discloses a computer device, including the distributed storage unit access system as described above.

[0038] The modular integrated circuit unit of the present invention includes at least one computing unit, at least one memory controller, at least one on-chip interconnection network, and at least one storage unit; the memory controller is used to send the storage address access request sent by the on-chip interconnection network to the storage unit; the storage unit is used to provide unified memory access to the storage address access request sent by the memory controller. The storage unit includes a local storage unit and a virtual remote storage unit. The local storage unit is distributed on the current modular integrated circuit unit, and the virtual remote storage unit is distributed on modular integrated circuit units other than the current modular integrated circuit unit; the computing unit is used to initiate a storage address access request; the on-chip interconnection network is used to interleave the storage addresses indicated by the storage address access requests sent by the computing unit to generate an interleaved storage address distribution; according to the interleaved storage address distribution, determine the target storage unit corresponding to the storage address access request, and send the storage address access request to the corresponding target storage unit through the memory controller. The target storage unit is the target local storage unit or the target virtual remote storage unit. By setting the local storage unit and the virtual remote storage unit, using the address interleaving technology and virtual storage mapping, the distributed storage units on multiple dies can achieve unified address access and load balancing, ensuring unified memory access to the computing system, balancing the access latency of each storage unit, reducing complexity, improving resource utilization, and enhancing the performance and usability of the distributed storage system. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0040] Figure 1 A schematic diagram of storage interleaved access provided by an embodiment of the present invention;

[0041] Figure 2 A schematic diagram of the storage unit organization form composed of one independent chip provided by an embodiment of the present invention;

[0042] Figure 3 A schematic diagram of the storage unit organization form composed of two chips provided by an embodiment of the present invention;

[0043] Figure 4 A schematic diagram of the storage unit organization form composed of two dies provided by an embodiment of the present invention;

[0044] Figure 5 A schematic diagram of the storage unit organization form composed of five dies provided by an embodiment of the present invention;

[0045] Figure 6 Schematic diagram of the structure of a distributed storage unit access system provided by an embodiment of the present invention;

[0046] Figure 7 Schematic diagram of the structure of a modular integrated circuit unit provided by an embodiment of the present invention;

[0047] Figure 8 Flowchart of a method for accessing a distributed storage unit provided by an embodiment of the present invention;

[0048] Figure 9 Flowchart of another method for accessing a distributed storage unit provided by an embodiment of the present invention;

[0049] Figure 10 Schematic diagram of the storage address distribution of multiple intervals provided by an embodiment of the present invention;

[0050] Figure 11 Schematic diagram of the structure of a distributed storage unit access system from the perspective of die 0 of the present invention;

[0051] Figure 12 Schematic diagram of the interleaved storage address distribution from the perspective of die 0 of the present invention;

[0052] Figure 13 Schematic diagram of the structure of a distributed storage unit access system from the perspective of die 1 of the present invention;

[0053] Figure 14 Schematic diagram of the interleaved storage address distribution from the perspective of die 1 of the present invention;

[0054] Figure 15 Schematic diagram of the structure of a distributed storage unit access system from the perspective of die 2 of the present invention;

[0055] Figure 16 Schematic diagram of the interleaved storage address distribution from the perspective of die 2 of the present invention;

[0056] Figure 17 Schematic diagram of the structure of a distributed storage unit access system from the perspective of die 3 of the present invention;

[0057] Figure 18 Schematic diagram of the interleaved storage address distribution from the perspective of die 3 of the present invention;

[0058] Figure 19A schematic structural diagram of a computer device provided by an embodiment of the present invention. Detailed implementation manners

[0059] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0060] To facilitate the understanding of the technical solutions provided in this application, the relevant content of the technical solutions in this application will be described first. A complete computing system includes components such as computing, storage, input / output devices, etc. An ideal storage device should be able to provide infinite storage capacity and bandwidth. However, in fact, for various storage devices that can be implemented, such as dynamic random access memory (DRAM) or static random access memory (SRAM), etc., they can only provide limited capacity and bandwidth. Therefore, multiple independent storage device units are placed in the system to increase the available storage capacity and bandwidth of the entire computing system. The storage access interleaving technology determines the specific storage unit to be accessed based on the accessed address, and the main purpose is to balance the access traffic to the storage unit and provide an abstraction of unified memory access (UMA) for computing. Figure 1 A schematic diagram of storage interleaved access provided by an embodiment of the present invention, as Figure 1 shown, based on the address interleaving algorithm and the preset interleaving granularity, the address range is divided, and access requests with different addresses are sent to different storage units, and then a unified memory access is abstracted, which can improve the concurrency and efficiency of data access. As Figure 1 shown, there are 4 storage units in this system, namely MEM 0, MEM 1, MEM 2, and MEM 3; the addresses from 0x0 to 0xfff are mapped to MEM 0, the addresses from 0x1000 to 0x1fff are mapped to MEM 1, the addresses from 0x2000 to 0x2fff are mapped to MEM 2, and the addresses from 0x3000 to 0x3fff are mapped to MEM 3. This interleaving mode is based on the high several bits of the address to determine the specific storage unit. In Figure 1Among them, the highest bit of the address determines which storage unit it is mapped to. Specifically, if the highest bit of the address is 0, the address is mapped to MEM 0; if the highest bit is 1, it is mapped to MEM 1, and so on. By distributing the address access requests to different storage units, all requests can be prevented from concentrating on one storage unit, thereby balancing the load of each storage unit. When multiple processors or cores access the memory simultaneously, the interleaving technique can reduce the waiting time because different storage units can process different access requests simultaneously.

[0061] The current forms of storage units include those composed of 1 independent chip, those composed of multiple interconnected chips, and those composed of multiple interconnected dielets. Figure 2 The following is a schematic diagram of the organizational form of a storage unit composed of 1 independent chip provided by an embodiment of the present invention, as Figure 2 shown. The computing system is composed of 1 independent chip, and 2 memory controllers (MEM CTRL) are integrated on this chip. It should be noted that the number of memory controllers can be set according to actual requirements; N storage units, namely MEM 0, ……, MEM N - 1, are also integrated on this chip, and generally N is within 8; the addresses of the N storage units are accessed in an interleaved manner within the chip.

[0062] Figure 3 The following is a schematic diagram of the organizational form of a storage unit composed of 2 chips provided by an embodiment of the present invention, as Figure 3 shown. The computing system is composed of 2 chips, namely chip 0 and chip 1; 2 memory controllers (MEM CTRL) are integrated on each chip. It should be noted that the number of memory controllers can be set according to actual requirements; the 2 chips are interconnected through inter-chip connection (Chip to Chip, abbreviated as: C2C), and N storage units, namely MEM 0, ……, MEM N - 1, are integrated on each chip, and generally N is within 8; the N storage units of each chip are interleaved respectively, and the computing system thus formed is called Non-Uniform Memory Access (abbreviated as: NUMA).

[0063] Figure 4 The following is a schematic diagram of the organizational form of a storage unit composed of 2 dielets provided by an embodiment of the present invention, as Figure 4As shown in the figure, the computing system consists of two dies, namely die 0 and die 1; each die integrates two memory controllers (MEM CTRL). It should be noted that the number of memory controllers can be set according to actual requirements; the two dies are interconnected through M die-to-die (D2D) connections. Each die integrates N storage units, namely MEM 0, ……, MEM N-1. Generally, N == M; the N storage units on the two dies are interleaved together, presenting a unified memory access to the computing unit; because N == M, the access addresses to the remote storage targets on different dies can be simulated through the D2D interface, that is: the access addresses to the remote storage targets are sent to the D2D interface. Therefore, the number of storage devices that this system can support is limited.

[0064] Figure 5 FIG. is a schematic diagram of the organization form of storage units composed of five dies provided by an embodiment of the present invention, as Figure 5 shown, the computing system consists of five dies, namely die 0, die 1, die 2, die 3 and die 4; different dies are interconnected through D2D; N storage units are all located on one die, die 0, namely MEM 0, ……, MEM N / 2-1, MEM N / 2, ……, MEM N-1; there are no storage units on the remaining dies; therefore, this organization form is Figure 2 similar, and the storage units can be interleaved on one die; the access to the storage units by the remaining dies is all sent to die 0 through D2D, and die 0 performs specific interleaved access.

[0065] With the development of technology and the demand for high-performance computing, currently, chips are developing towards the direction of multi-die interconnection. Taking a chip with distributed storage units, which is a multi-die packaged chip, as an example, Figure 6 FIG. is a schematic structural diagram of a distributed storage unit access system provided by an embodiment of the present invention, as Figure 6 shown, the system consists of four dies, namely die 0, die1, die 2 and die 3, and different dies are interconnected through D2D; each die integrates two memory controllers (MEMCTRL). It should be noted that the number of memory controllers can be set according to actual requirements; each die integrates two memory controllers (MEM CTRL); each die also integrates multiple computing units, namely computing unit 0, ……, computing unit N; each die integrates an on-chip interconnection network, Figure 6The on-chip interconnect network is shown as black straight lines. The computing units can initiate requests to the on-chip interconnect network, and the on-chip interconnect network can send requests to the memory controller. Each die integrates N memory units, namely MEM 0, ……, MEM N-1. At the same time, due to the demand for storage bandwidth and the like, storage devices such as three-dimensional dynamic random access memory (3DDRAM) have been developed. Due to the physical implementation of 3DDRAM and the like, the intuitive result on the system is that the number N of memory units will be larger than that of traditional memory units. These factors combined result in the distributed characteristics of memory units on the system. In order to abstract a unified memory access model for computing systems and the like, it is necessary to perform unified address interleaving on these numerous distributed memory units within and across dies.

[0066] It is worth noting that the distributed memory unit access system includes multiple modular integrated circuit units. The modular integrated circuit units are dies in a multi-die package or chips in a computing system composed of independent chips. Dies communicate with each other through D2D. Chips communicate with each other through C2C. The technical solution of this application can be applied to a computing system including multiple interconnected distributed modular integrated circuit units, including but not limited to chips using multi-die packaging or a computing system composed of multiple chips. The memory units can use different storage media such as double data rate memory (DDR) or 3DDRAM.

[0067] In response to the current demand for performing unified address interleaving on numerous distributed memory units within and across dies, this application proposes the concepts of using local memory units (LMT) and virtual remote memory units (VRMT). By using virtual remote memory units, the distributed memory units on multiple dies can provide the function of unified address access through address interleaving, thus avoiding the problem of non-uniform memory access to the computing system.

[0068] The modular integrated circuit unit includes at least one computing unit, at least one memory controller (MEM CTRL), at least one on-chip interconnection network, and at least one memory unit (MEM); the memory controller is used to send the memory address access request sent by the on-chip interconnection network to the memory unit, and the memory unit (MEM) is used to provide unified memory access to the memory address access request sent by the memory controller. The computing unit is used to initiate a memory address access request to the on-chip interconnection network; the on-chip interconnection network is used to interleave the memory addresses indicated by the memory address access requests sent by the computing unit to generate an interleaved memory address distribution; according to the interleaved memory address distribution, determine the target memory unit corresponding to the memory address access request, and send the memory address access request to the corresponding target memory unit through the memory controller. The target memory unit is the target LMT or target VRMT. The modular integrated circuit unit is a die in a multi-die package or a chip in a computing system composed of independent chips.

[0069] Figure 7 FIG. is a schematic structural diagram of a modular integrated circuit unit provided by an embodiment of the present invention, as Figure 7 shown. The modular integrated circuit unit includes multiple computing units (Computing Unit 0,..., Computing Unit N), an on-chip interconnection network, multiple memory controllers (MEM CTRL), multiple memory units (MEM 0,..., MEM N), and a D2D interface. The computing unit can initiate a memory address access request to the on-chip interconnection network. The on-chip interconnection network is used to interleave the memory addresses indicated by the memory address access requests sent by the computing unit to generate an interleaved memory address distribution; according to the interleaved memory address distribution, determine the target memory unit corresponding to the memory address access request, and send the memory address access request to the corresponding target memory unit through the memory controller; the memory controller is used to send the memory address access request sent by the on-chip interconnection network to the target memory unit to access the target memory unit. The modular integrated circuit unit can also be communicatively connected to other modular integrated circuit units through the D2D interface; the memory unit is used to provide unified memory access to the memory address access request sent by the memory controller.

[0070] The memory unit includes a local memory unit and a virtual remote memory unit. The local memory unit is distributed on the current modular integrated circuit unit, and the virtual remote memory unit is distributed on modular integrated circuit units other than the current modular integrated circuit unit.

[0071] Specifically, the on-chip interconnection network is used to divide the memory addresses according to a preset interleaving granularity to obtain memory address distributions in multiple intervals; uniformly map the memory address distributions in multiple intervals to different memory units on different modular integrated circuit units to obtain an interleaved memory address distribution.

[0072] The on-chip interconnection network is specifically used to send, according to the interleaved storage address distribution, the storage address access requests distributed on the current modular integrated circuit unit where it is located to the target local storage unit through the memory controller, or to send the storage address access requests distributed on the modular integrated circuit units outside the current modular integrated circuit unit where it is located to the target virtual remote storage unit through the interconnection interface.

[0073] It should be noted that the modular integrated circuit unit is also applied to Figures 8 to 10 Any of the shown distributed storage unit access methods will not be elaborated in the embodiments of the present invention.

[0074] The implementation process of the distributed storage unit access method provided by the embodiments of the present invention will be described below.

[0075] Figure 8 It is a flowchart of a distributed storage unit access method provided by the embodiments of the present invention. As Figure 8 shown, the method includes:

[0076] Step 101, the computing unit sends a storage address access request to the on-chip interconnection network.

[0077] In the embodiments of the present invention, the storage address access request is initiated by the computing unit. The storage address access request refers to that when the processor or the memory controller processes data, it accesses adjacent storage units in sequence according to the memory address. This method usually appears in situations where a large amount of continuous data needs to be processed, such as array operations, file reading and writing, or multimedia data processing, etc.

[0078] Step 102, the on-chip interconnection network interleaves the storage address indicated by the storage address access request to generate an interleaved storage address distribution.

[0079] In the embodiments of the present invention, the storage address access request is sent by the computing unit to the on-chip interconnection network. The storage address access request includes consecutive storage addresses. Based on the address interleaving algorithm, the address space of the storage address is divided into multiple segments according to the preset interleaving granularity. The storage address indicated by the storage address access request is matched with each divided segment address space, and the storage addresses in multiple intervals are evenly mapped to different storage units on different modular integrated circuit units to generate an interleaved storage address distribution. The interleaved storage address distribution includes the number identifiers of multiple storage units and the storage address distribution of each interval corresponding to each storage unit. The on-chip interconnection network stores the interconnection interface bound to each storage unit, and the interconnection interface includes a D2D interface or a C2C interface. The interconnection interface is used to implement the interaction across modular integrated circuit units.

[0080] Step 103: The on-chip interconnection network determines the target storage unit corresponding to the storage address access request according to the interleaved storage address distribution, and sends the storage address access request to the memory controller within the corresponding modular integrated circuit unit.

[0081] Specifically, the on-chip interconnection network sends the storage address access requests distributed on the current modular integrated circuit unit to the corresponding memory controller according to the interleaved storage address distribution, or sends the storage address access requests distributed on the modular integrated circuit units other than the current modular integrated circuit unit to the memory controller within the corresponding modular integrated circuit unit through the interconnection interface.

[0082] Step 104: The memory controller sends the storage address access request sent by the on-chip interconnection network to the target storage unit, and the target storage unit is the target LMT or the target VRMT.

[0083] In the embodiment of the present invention, the storage address access request is sent by the on-chip interconnection network to the memory controller.

[0084] Specifically, if the target storage unit is the target LMT, the memory controller sends the storage address access request to the target LMT.

[0085] Specifically, if the target storage unit is the target VRMT, the memory controller sends the storage address access request to the target VRMT.

[0086] In the embodiment of the present invention, the modular integrated circuit unit includes at least one storage unit, and the storage unit includes an LMT and a VRMT. The LMT is distributed on the current modular integrated circuit unit, and the VRMT is distributed on the modular integrated circuit units other than the current modular integrated circuit unit. The VRMT is invisible to the current modular integrated circuit unit, and the binding relationship between the VRMT and the interconnection interface is visible to the modular integrated circuit unit, so that the system can dynamically adjust or expand the storage resources without affecting the existing architecture. This design enhances the flexibility and scalability of the system and facilitates future upgrades or expansions.

[0087] In the embodiment of the present invention, the modular integrated circuit unit is a die in a multi-die package or a chip in a computing system composed of independent chips. The dies communicate with each other through D2D, and the chips communicate with each other through C2C. It should be noted that the specific interconnection method of D2D can be through the D2D interface, and the specific interconnection method of C2C can be through the C2C interface.

[0088] In the technical solution provided by the embodiment of the present invention, the calculation unit initiates a storage address access request to the on-chip interconnection network; the on-chip interconnection network interleaves the storage address indicated by the storage address access request to generate an interleaved storage address distribution; the on-chip interconnection network determines the target storage unit corresponding to the storage address access request according to the interleaved storage address distribution, and sends the storage address access request to the memory controller in the corresponding modular integrated circuit unit; the memory controller sends the storage address access request to the target storage unit, and the target storage unit is the target local storage unit or the target virtual remote storage unit; the storage unit includes a local storage unit and a virtual remote storage unit, the local storage unit is distributed on the current modular integrated circuit unit, and the virtual remote storage unit is distributed on the modular integrated circuit unit other than the current modular integrated circuit unit. By setting the local storage unit and the virtual remote storage unit, using the address interleaving technology and virtual storage mapping, the distributed storage units on multiple dies can achieve unified address access and load balancing, ensuring unified memory access to the computing system, balancing the access latency of each storage unit, reducing complexity, improving resource utilization, and enhancing the performance and usability of the distributed storage system.

[0089] Figure 9 It is a flowchart of another distributed storage unit access method provided by the embodiment of the present invention. As Figure 9 shown, the method includes:

[0090] Step 201, the on-chip interconnection network divides the storage address according to a preset interleaving granularity to obtain storage address distributions in multiple intervals.

[0091] In the embodiment of the present invention, the continuous storage address is divided according to the interleaving granularity M to generate storage address distributions in multiple intervals. The storage address distributions in multiple intervals include multiple intervals and the storage addresses corresponding to each interval.

[0092] In the embodiment of the present invention, the interleaving granularity M in the address interleaving algorithm can be adjusted according to the type of the storage unit, and the storage unit includes DDR or 3DDRAM.

[0093] Furthermore, according to the storage access load of each modular integrated circuit unit, the allocation strategy in the address interleaving algorithm is adjusted in real time to optimize the access bandwidth and latency.

[0094] The present invention divides the storage address using a preset interleaving granularity and dynamically adjusts the storage address distribution according to requirements, improving the system's recovery ability in the face of failures and the flexibility to adapt to different application scenarios.

[0095] Step 202: The on-chip interconnection network evenly maps the storage addresses of multiple intervals to different storage units on different modular integrated circuit units, obtaining an interleaved storage address distribution.

[0096] In the embodiments of the present invention, the modular integrated circuit unit is a die in a multi-die package or a chip in a computing system composed of independent chips, and multiple storage units are distributed on the modular integrated circuit unit. For each modular integrated circuit unit, the storage units distributed on the current modular integrated circuit unit are LMTs, and the storage units distributed on other modular integrated circuit units are VRMTs. Moreover, the VRMTs are invisible to the current modular integrated circuit unit, and the binding relationship between the VRMTs and the interconnection interfaces is visible to the modular integrated circuit unit.

[0097] Specifically, map the storage addresses of each interval to a storage unit, and then the storage unit processes the access requests for the storage addresses within the interval; map the storage address distributions of multiple intervals to different storage units on different modular integrated circuit units in sequence, obtaining an interleaved storage address distribution. The interleaved storage address distribution includes the number identifiers of multiple storage units and the storage address distribution of each storage unit corresponding to the interval. The on-chip interconnection network stores the interconnection interfaces bound to each storage unit.

[0098] It should be noted that the order of the storage addresses of multiple intervals is a sequence divided according to the interleaving granularity, and the order of the storage units can be sorted according to the number identifiers of the storage units. The embodiments of the present invention do not make any limitations in this regard.

[0099] By interleaving consecutive storage addresses and evenly distributing them to different storage units, the present invention can reduce conflicts during data access, improve the data reading and writing speed, and thus enhance the overall performance; it allows the distribution of storage resources among multiple modular integrated circuit units, supports the addition of more storage units without affecting the system performance, and enhances the expansion ability of the system.

[0100] Step 203: The on-chip interconnection network matches the storage address indicated by the storage address access request distributed on the current modular integrated circuit unit with the storage address distributions of each storage unit corresponding to the intervals in the storage address distribution, determines the target LMT, and sends the storage address access request distributed on the current modular integrated circuit unit to the corresponding memory controller.

[0101] In an embodiment of the present invention, for the current modular integrated circuit unit, based on the numbered identifier of the storage unit, the distribution order of the interleaved storage address distribution is used as the matching order. If the LMT mapped to the storage address corresponding to this interval is matched, the determined LMT is determined as the target LMT; the storage address access requests distributed on the current modular integrated circuit unit are sent to the memory controller corresponding to the target LMT. Using the distribution order of the interleaved storage address distribution as the matching order makes the storage units accessed by the modular integrated circuit units from different perspectives consistent, thereby realizing unified memory access and ensuring the consistency when different modular integrated circuit units access the storage units from various perspectives. This way enables the system to provide a unified memory access interface and simplifies the complexity of memory management.

[0102] Further, the memory controller sends the storage address access request to the target LMT, and then the target LMT processes the storage address access request.

[0103] The present invention provides a unified memory access interface for the computing unit, enabling developers to not need to concern about the complexity of the underlying storage structure, reducing the development difficulty and improving the development efficiency.

[0104] Step 204: The on-chip interconnection network matches the storage address indicated by the storage address access requests distributed on the modular integrated circuit units other than the current modular integrated circuit unit with the storage address distribution of each storage unit corresponding interval in the storage address distribution, and determines the target interconnection interface bound to the target remote modular integrated circuit unit.

[0105] In an embodiment of the present invention, if the VRMT mapped to the storage address corresponding to this interval is matched by querying the interleaved storage address distribution according to the indicated storage address, the matched VRMT is determined as the target VRMT; the interconnection interface corresponding to the target VRMT is determined, and this interconnection interface is the target interconnection interface.

[0106] In an embodiment of the present invention, if the modular integrated circuit unit is a chiplet, the interconnection interface is a D2D interface; if the modular integrated circuit unit is a chip, the interconnection interface is a C2C interface.

[0107] The present invention simplifies the communication management between cross-modular integrated circuit units by clarifying the binding relationship between the interconnection interface and the VRMT. Developers only need to focus on how to send requests through these interfaces without delving into the underlying communication details, thereby reducing the development complexity.

[0108] Step 205: Through the target interconnection interface, send the storage address access requests on the modular integrated circuit units distributed outside the current modular integrated circuit unit to the target remote modular integrated circuit unit, so that the target remote modular integrated circuit unit can send the storage address access requests to the corresponding memory controller.

[0109] In the embodiment of the present invention, the on-chip interconnection network sends the storage address access requests on the modular integrated circuit units distributed outside the current modular integrated circuit unit to the target interconnection interface bound to the target VRMT. The target interconnection interface is used for communication between modular integrated circuit units; communicate with the target remote modular integrated circuit unit through the target interconnection interface and route the storage address access requests to the corresponding memory controller.

[0110] Further, the memory controller of the target remote modular integrated circuit unit sends the storage address access requests to the target VRMT, and then the target VRMT processes the storage address access requests.

[0111] The present invention determines the target VRMT by matching the storage address with the interleaved storage address distribution, and then uses the target interconnection interface bound to the VRMT for communication, ensuring that efficient and accurate data access can be achieved even for storage resources located remotely; the method of determining the target VRMT based on the interleaved storage address distribution helps to more reasonably allocate storage resources and may promote load balancing, avoiding the situation where some nodes are overloaded while other nodes have idle resources.

[0112] By uniformly managing local and remote storage units, the present invention can more efficiently utilize the storage resources on each modular integrated circuit unit, avoiding resource waste; it is applicable to application scenarios that need to process a large amount of data, such as big data analysis, artificial intelligence, etc. Through effective storage management and access strategies, it meets the requirements of these applications for high-performance storage access; through innovative storage access methods, it not only improves the performance of a single modular integrated circuit unit, but also greatly promotes the collaborative work between multiple units, providing strong support for building an efficient and flexible large-scale distributed storage system.

[0113] Taking a computing system composed of 4 die as an example in the present invention, the schematic diagram of the computing system can be referred to Figure 6, that is, the modular integrated circuit unit is a chiplet, and the concepts of the storage units of LMT and VRMT are used to describe the specific implementation process of the distributed storage unit access method. For a computing system composed of 4 chiplets, the purpose is to organize the distributed storage on the 4 chiplets into a unified storage, that is: for continuous storage address access, the system needs to evenly distribute different addresses to the storage units on different chiplets according to the interleaving algorithm to achieve unified memory access. Taking the interleaving granularity as M as an example, Figure 10 is a schematic diagram of the storage address distribution of multiple intervals provided by an embodiment of the present invention, as Figure 10 shown:

[0114] There are N storage units integrated on chiplet die 0, namely MEM 0, MEM 1, ……, MEM N-2, MEM N-1. Map the addresses from 0×M to 1×M-1 to MEM 0, map the addresses from 1×M to 2×M-1 to MEM 1, map the addresses from (N-2)×M to (N-1)×M-1 to MEM N-2, and map the addresses from (N-1)×M to N×M-1 to MEM N-1.

[0115] There are N storage units integrated on chiplet die 1, namely MEM 0, MEM 1, ……, MEM N-2, MEM N-1. Map the addresses from N×M to (N+1)×M-1 to MEM 0, map the addresses from (N+1)×M to (N+2)×M-1 to MEM1, map the addresses from (2N-2)×M to (2N-1)×M-1 to MEM N-2, and map the addresses from (2N-1)×M to 2N×M-1 to MEM N-1.

[0116] There are N storage units integrated on chiplet die 2, namely MEM 0, MEM 1, ……, MEM N-2, MEM N-1. Map the addresses from 2N×M to (2N+1)×M-1 to MEM 0, map the addresses from (2N+1)×M to (2N+2)×M-1 to MEM 1, map the addresses from (3N-2)×M to (3N-1)×M-1 to MEM N-2, and map the addresses from (3N-1)×M to 3N×M-1 to MEM N-1.

[0117] There are N memory cells integrated on die 3, namely MEM 0, MEM 1, ……, MEM N-2, MEM N-1. The addresses from 3N×M to (3N+1)×M-1 are mapped to MEM 0, the addresses from (3N+1)×M to (3N+2)×M-1 are mapped to MEM 1, the addresses from (4N-2)×M to (4N-1)×M-1 are mapped to MEM N-2, and the addresses from (4N-1)×M to 4N×M-1 are mapped to MEM N-1.

[0118] Figure 11 This is a schematic structural diagram of a distributed memory cell access system from the perspective of die 0 provided by an embodiment of the present invention. As Figure 11 shown, from the perspective of die 0, the N memory cells (MEM 0, MEM1, ……, MEM N-2, MEM N-1) in die 0 are local memory cells (LMT), that is: LMT 0, LMT 1, ……, LMT N-2, LMT N-1; the memory cells of die 1, die 2, and die 3 are all virtual remote memory cells (VRMT), that is: VRMT0, ……, VRMT N-1, ……, VRMT N, ……, VRMT 2N-1, ……, VRMT 2N, ……, VRMT 3N-1.

[0119] Specifically, for die 0, to achieve unified address memory access, the access will be interleaved in the order of {LMT 0, LMT1, ……, LMT N-2, LMT N-1, VRMT 0, VRMT 1, ……, VRMT N-2, VRMT N-1, VRMT N, VRMT N+1, ……, VRMT 2N-2, VRMT 2N-1, VRMT 2N, VRMT 2N+1, ……, VRMT 3N-2, VRMT 3N-1}, that is: according to the address of the accessed storage unit and the specific interleaving algorithm, the access to the storage unit is evenly distributed among {LMT 0, LMT 1, ……, LMT N-2, LMT N-1, VRMT 0, VRMT 1, ……, VRMT N-2, VRMT N-1, VRMTN, VRMT N+1, ……, VRMT 2N-2, VRMT 2N-1, VRMT 2N, VRMT 2N+1, ……, VRMT 3N-2, VRMT 3N-1}. For LMT 0 to LMT N-1, the present invention sends the corresponding access requests to the storage units corresponding to this die (die 0); while for VRMT 0 to VRMT 3N-1, the present invention sends the corresponding access requests to D2D and routes them to other dies (die 1, die 2, and die 3) through D2D until the final storage unit.

[0120] It should be noted that the address interleaving of the address access requests is performed on the on-chip interconnection network. For the convenience of clearly expressing the schematic diagram, the on-chip interconnection network and the computing units are not shown in the figure.

[0121] Figure 12 The following is a schematic diagram of the interleaved storage address distribution from the perspective of die 0 provided by the embodiment of the present invention, as Figure 12 shown:

[0122] From the perspective of die 0, the storage units on die 0 are local storage units, which are respectively: LMT 0, LMT 1, ……, LMT N-2, LMT N-1. The addresses from 0×M to 1×M-1 are mapped to LMT 0, the addresses from 1×M to 2×M-1 are mapped to LMT 1, the addresses from (N-2)×M to (N-1)×M-1 are mapped to LMT N-2, and the addresses from (N-1)×M to N×M-1 are mapped to LMT N-1.

[0123] From the perspective of die 0, the storage units on die 1 are virtual remote storage units, which are respectively VRMT 0, VRMT 1, ……, VRMT N-2, VRMT N-1. The addresses from N×M to (N+1)×M-1 are mapped to VRMT 0, the addresses from (N+1)×M to (N+2)×M-1 are mapped to VRMT 1, the addresses from (2N-2)×M to (2N-1)×M-1 are mapped to VRMT N-2, and the addresses from (2N-1)×M to 2N×M-1 are mapped to VRMT N-1.

[0124] From the perspective of die 0, the storage units on die 2 are virtual remote storage units, which are respectively VRMT N, VRMT N+1, ……, VRMT 2N-2, VRMT 2N-1. The addresses from 2N×M to (2N+1)×M-1 are mapped to VRMT N, the addresses from (2N+1)×M to (2N+2)×M-1 are mapped to VRMT N+1, the addresses from (3N-2)×M to (3N-1)×M-1 are mapped to VRMT 2N-2, and the addresses from (3N-1)×M to 3N×M-1 are mapped to VRMT 2N-1.

[0125] From the perspective of die 0, the storage units on die 3 are virtual remote storage units, which are respectively VRMT 2N, VRMT 2N+1, ……, VRMT 3N-2, VRMT 3N-1. The addresses from 3N×M to (3N+1)×M-1 are mapped to VRMT 2N, the addresses from (3N+1)×M to (3N+2)×M-1 are mapped to VRMT 2N+1, the addresses from (4N-2)×M to (4N-1)×M-1 are mapped to VRMT 3N-2, and the addresses from (4N-1)×M to 4N×M-1 are mapped to VRMT 3N-1.

[0126] The storage access of the storage addresses on die 0 corresponding to the core die is sent to the storage unit on die 0 through the interleaving algorithm; the storage access of the storage addresses on other core dies corresponding to the core die is sent to the D2D bound to VRMT through the interleaving algorithm.

[0127] It should be noted that VRMT is not located on die 0 and is invisible to the computing units on die 0. VRMT is shown as a black frame line in the figure.

[0128] Figure 13 The structure diagram of a distributed storage unit access system from the perspective of die 1 provided by the embodiment of the present invention is as Figure 13As shown, from the perspective of die 1, the N memory units (MEM 0, MEM1, ……, MEM N-2, MEM N-1) in die 1 are local memory units (LMT), namely: LMT 0, LMT 1, ……, LMT N-2, LMT N-1; the memory units of die 0, die 2, and die 3 are all virtual remote memory units (VRMT), namely: VRMT0, ……, VRMT N-1, ……, VRMT N, ……, VRMT 2N-1, ……, VRMT 2N, ……, VRMT 3N-1.

[0129] Specifically, for die 1, in order to achieve unified address memory access, it will access in the order of {VRMT 0, VRMT1, ……, VRMT N-2, VRMT N-1, LMT 0, LMT 1, ……, LMT N-2, LMT N-1, VRMT N, VRMT N+1, ……, VRMT 2N-2, VRMT 2N-1, VRMT 2N, VRMT 2N+1, ……, VRMT 3N-2, VRMT 3N-1}, that is: according to the address of the accessed memory unit and the specific interleaving algorithm, the access of the memory unit is evenly distributed on {VRMT 0, VRMT 1, ……, VRMT N-2, VRMT N-1, LMT 0, LMT 1, ……, LMT N-2, LMT N-1, VRMTN, VRMT N+1, ……, VRMT 2N-2, VRMT 2N-1, VRMT 2N, VRMT 2N+1, ……, VRMT 3N-2, VRMT 3N-1}. For LMT 0 to LMT N-1, the present invention sends the corresponding access requests to the corresponding memory units of this die (die 1); for VRMT 0 to VRMT 3N-1, the present invention sends the corresponding access requests to D2D, and routes them to other dies (die 0, die 2, and die 3) through D2D until the final memory unit.

[0130] It should be noted that the address interleaving of the address access requests is performed on the on-chip interconnection network. For the convenience of clearly expressing the schematic diagram, the on-chip interconnection network and the computing unit are not shown in the figure.

[0131] Figure 14 The following is a schematic diagram of the interleaved storage address distribution from the perspective of die 1 provided by the embodiment of the present invention, as Figure 14 shown:

[0132] From the perspective of die 1, the storage units on die 1 are local storage units, which are respectively: LMT 0, LMT 1, ……, LMT N-2, LMT N-1. Map the addresses from N×M to (N+1)×M-1 to LMT 0, map the addresses from (N+1)×M to (N+2)×M-1 to LMT 1, map the addresses from (2N-2)×M to (2N-1)×M-1 to LMT N-2, and map the addresses from (2N-1)×M to 2N×M-1 to LMT N-1.

[0133] From the perspective of die 1, the storage units on die 0 are virtual remote storage units, which are respectively: VRMT 0, VRMT 1, ……, VRMT N-2, VRMT N-1. Map the addresses from 0×M to 1×M-1 to VRMT 0, map the addresses from 1×M to 2×M-1 to VRMT 1, map the addresses from (N-2)×M to (N-1)×M-1 to VRMT N-2, and map the addresses from (N-1)×M to N×M-1 to VRMT N-1.

[0134] From the perspective of die 1, the storage units on die 2 are virtual remote storage units, which are respectively: VRMT N, VRMT N+1, ……, VRMT 2N-2, VRMT 2N-1. Map the addresses from 2N×M to (2N+1)×M-1 to VRMT N, map the addresses from (2N+1)×M to (2N+2)×M-1 to VRMT N+1, map the addresses from (3N-2)×M to (3N-1)×M-1 to VRMT 2N-2, and map the addresses from (3N-1)×M to 3N×M-1 to VRMT 2N-1.

[0135] From the perspective of die 1, the storage units on die 3 are virtual remote storage units, which are respectively: VRMT 2N, VRMT 2N+1, ……, VRMT 3N-2, VRMT 3N-1. Map the addresses from 3N×M to (3N+1)×M-1 to VRMT 2N, map the addresses from (3N+1)×M to (3N+2)×M-1 to VRMT 2N+1, map the addresses from (4N-2)×M to (4N-1)×M-1 to VRMT 3N-2, and map the addresses from (4N-1)×M to 4N×M-1 to VRMT 3N-1.

[0136] The memory access of the memory address corresponding to the die 1 is sent to the memory unit on the die 1 through the interleaving algorithm; the memory access of the memory address corresponding to the other dies is sent to the D2D bound to the VRMT.

[0137] It should be noted that the VRMT is not located on the die 1 and is invisible to the computing unit on the die 1. The VRMT is shown as a black frame line in the figure; from the perspective of the die 1, the memory unit for the same address access is the same as that in the die 0.

[0138] Figure 15 The following is a schematic structural diagram of a distributed memory unit access system from the perspective of the die 2 according to an embodiment of the present invention. As Figure 15 shown, from the perspective of the die 2, the N memory units (MEM 0, MEM1,..., MEM N-2, MEM N-1) in the die 2 are local memory units (LMT), that is: LMT 0, LMT 1,..., LMT N-2, LMT N-1; the memory units of the die 0, die 1, and die 3 are all virtual remote memory units (VRMT), that is: VRMT0,..., VRMT N-1,..., VRMT N,..., VRMT 2N-1,..., VRMT 2N,..., VRMT 3N-1.

[0139] Specifically, for die 2, to achieve unified address memory access, the access will be interleaved in the order of {VRMT 0, VRMT1, ……, VRMT N-2, VRMT N-1, VRMT N, VRMT N+1, ……, VRMT 2N-2, VRMT 2N-1, LMT 0, LMT1, ……, LMT N-2, LMT N-1, VRMT 2N, VRMT 2N+1, ……, VRMT 3N-2, VRMT 3N-1}, that is: according to the address of the accessed storage unit and the specific interleaving algorithm, the access to the storage unit is evenly distributed among {VRMT 0, VRMT 1, ……, VRMT N-2, VRMT N-1, VRMT N, VRMT N+1, ……, VRMT 2N-2, VRMT 2N-1, LMT 0, LMT 1, ……, LMT N-2, LMT N-1, VRMT 2N, VRMT 2N+1, ……, VRMT 3N-2, VRMT 3N-1}. For LMT 0 to LMT N-1, the present invention sends the corresponding access requests to the storage units corresponding to this die (die 2); while for VRMT 0 to VRMT 3N-1, the present invention sends the corresponding access requests to D2D and routes them to other dies (die 0, die 1, and die 3) through D2D until the final storage unit.

[0140] Figure 16 FIG. is a schematic diagram of the interleaved storage address distribution from the perspective of die 2 provided by an embodiment of the present invention, as Figure 16 shown:

[0141] From the perspective of die 2, the storage units on die 2 are local storage units, which are respectively: LMT 0, LMT 1, ……, LMT N-2, LMT N-1. The addresses from 2N×M to (2N+1)×M-1 are mapped to LMT0, the addresses from (2N+1)×M to (2N+2)×M-1 are mapped to LMT 1, the addresses from (3N-2)×M to (3N-1)×M-1 are mapped to LMT N-2, and the addresses from (3N-1)×M to 3N×M-1 are mapped to LMT N-1.

[0142] From the perspective of die 2, the memory cells on die 0 are virtual remote memory cells, which are respectively VRMT 0, VRMT 1, ……, VRMT N-2, VRMT N-1. The addresses from 0×M to 1×M-1 are mapped to VRMT 0, the addresses from 1×M to 2×M-1 are mapped to VRMT 1, the addresses from (N-2)×M to (N-1)×M-1 are mapped to VRMT N-2, and the addresses from (N-1)×M to N×M-1 are mapped to VRMT N-1.

[0143] From the perspective of die 2, the memory cells on die 1 are virtual remote memory cells, which are respectively VRMT N, VRMT N+1, ……, VRMT 2N-2, VRMT 2N-1. The addresses from N×M to (N+1)×M-1 are mapped to VRMT N, the addresses from (N+1)×M to (N+2)×M-1 are mapped to VRMT N+1, the addresses from (2N-2)×M to (2N-1)×M-1 are mapped to VRMT 2N-2, and the addresses from (2N-1)×M to 2N×M-1 are mapped to VRMT 2N-1.

[0144] From the perspective of die 2, the memory cells on die 3 are virtual remote memory cells, which are respectively VRMT 2N, VRMT 2N+1, ……, VRMT 3N-2, VRMT 3N-1. The addresses from 3N×M to (3N+1)×M-1 are mapped to VRMT 2N, the addresses from (3N+1)×M to (3N+2)×M-1 are mapped to VRMT 2N+1, the addresses from (4N-2)×M to (4N-1)×M-1 are mapped to VRMT 3N-2, and the addresses from (4N-1)×M to 4N×M-1 are mapped to VRMT 3N-1.

[0145] The memory access of the corresponding storage addresses on die 2 is sent to the memory cells on die 2 through the interleaving algorithm; the memory access of the corresponding storage addresses on other dies is sent to the D2D bound to VRMT through the interleaving algorithm.

[0146] It should be noted that VRMT is not located on die 2 and is invisible to the computing units on die 2. VRMT is shown as a black box line in the figure; from the perspective of die 2, the memory cells for the same address access are consistent in die 0 and die 1.

[0147] Figure 17This is a schematic structural diagram of a distributed storage unit access system from the perspective of die 3 of the present invention. As Figure 17 shown, from the perspective of die 3, the N storage units (MEM 0, MEM1,..., MEM N-2, MEM N-1) in die 3 are local storage units (LMT), that is: LMT 0, LMT 1,..., LMT N-2, LMT N-1; the storage units of die 0, die 1, and die 2 are all virtual remote storage units (VRMT), that is: VRMT0,..., VRMT N-1,..., VRMT N,..., VRMT 2N-1,..., VRMT 2N,..., VRMT 3N-1.

[0148] Specifically, for die 3, in order to achieve unified address memory access, it will be accessed in an interleaved manner according to the order of {VRMT 0, VRMT1,..., VRMT N-2, VRMT N-1, VRMT N, VRMT N+1,..., VRMT 2N-2, VRMT 2N-1, VRMT 2N, VRMT 2N+1,..., VRMT 3N-2, VRMT 3N-1, LMT 0, LMT 1,..., LMT N-2, LMT N-1}, that is: according to the address of the accessed storage unit and the specific interleaving algorithm, the access of the storage unit is evenly distributed among {VRMT 0, VRMT 1,..., VRMT N-2, VRMT N-1, VRMT N, VRMT N+1,..., VRMT 2N-2, VRMT 2N-1, VRMT 2N, VRMT 2N+1,..., VRMT 3N-2, VRMT 3N-1, LMT 0, LMT 1,..., LMT N-2, LMT N-1}. For LMT 0 to LMT N-1, the present invention sends the corresponding access requests to the corresponding storage units of this die (die 2); for VRMT 0 to VRMT 3N-1, the present invention sends the corresponding access requests to D2D and routes them to other dies (die 0, die 1, and die 2) through D2D until the final storage unit.

[0149] It should be noted that the address interleaving of the address access requests is performed on the on-chip interconnection network. For the convenience of clearly expressing the schematic diagram, the on-chip interconnection network and the computing unit are not shown in the figure.

[0150] Figure 18 This is a schematic diagram of the distributed storage addresses after interleaving from the perspective of die 3 of the present invention. As Figure 18 shown:

[0151] From the perspective of die 3, the storage units on die 3 are local storage units, which are respectively: LMT 0, LMT 1, ……, LMT N-2, LMT N-1. Map the addresses from 3N×M to (3N+1)×M-1 to LMT0, map the addresses from (3N+1)×M to (3N+2)×M-1 to LMT 1, map the addresses from (4N-2)×M to (4N-1)×M-1 to LMT N-2, and map the addresses from (4N-1)×M to 4N×M-1 to LMT N-1.

[0152] From the perspective of die 3, the storage units on die 0 are virtual remote storage units, which are respectively: VRMT 0, VRMT 1, ……, VRMT N-2, VRMT N-1. Map the addresses from 0×M to 1×M-1 to VRMT 0, map the addresses from 1×M to 2×M-1 to VRMT 1, map the addresses from (N-2)×M to (N-1)×M-1 to VRMT N-2, and map the addresses from (N-1)×M to N×M-1 to VRMT N-1.

[0153] From the perspective of die 3, the storage units on die 1 are virtual remote storage units, which are respectively: VRMT N, VRMT N+1, ……, VRMT 2N-2, VRMT 2N-1. Map the addresses from N×M to (N+1)×M-1 to VRMT N, map the addresses from (N+1)×M to (N+2)×M-1 to VRMT N+1, map the addresses from (2N-2)×M to (2N-1)×M-1 to VRMT 2N-2, and map the addresses from (2N-1)×M to 2N×M-1 to VRMT 2N-1.

[0154] From the perspective of die 3, the storage units on die 2 are virtual remote storage units, which are respectively: VRMT 2N, VRMT 2N+1, ……, VRMT 3N-2, VRMT 3N-1. Map the addresses from 2N×M to (2N+1)×M-1 to VRMT 2N, map the addresses from (2N+1)×M to (2N+2)×M-1 to VRMT 2N+1, map the addresses from (3N-2)×M to (3N-1)×M-1 to VRMT 3N-2, and map the addresses from (3N-1)×M to 3N×M-1 to VRMT 3N-1.

[0155] The memory access of the storage address on the corresponding die 3 of the chiplet is sent to the memory cell on die 3 of the chiplet through an interleaving algorithm; the memory access of the storage address on the corresponding other chiplets is sent to the D2D bound to the VRMT.

[0156] It should be noted that the VRMT is not located on die 3 of the chiplet and is invisible to the computing unit on die 3 of the chiplet. The VRMT is shown as a black frame line in the figure; from the perspective of die 3 of the chiplet, the memory cells for the same address access are the same in die 0, die 1, and die 2 of the chiplet.

[0157] In summary, by dividing the memory cells distributed on different chiplets into LMT and VRMT from the perspective of each chiplet, a unified memory cell organization can be formed, thereby realizing unified memory access.

[0158] It should be noted that in the technical solution of this application, the acquisition, storage, use, processing, etc. of data all comply with the relevant regulations of laws and regulations. The user information in the embodiments of this application is obtained through legal and compliant channels, and the acquisition, storage, use, processing, etc. of user information have obtained the authorization and consent of the customer.

[0159] It should be noted that the information collected in this application is information and data authorized by the user or fully authorized by all parties. The processing of relevant data, such as collection, storage, use, processing, transmission, provision, disclosure, and application, all comply with the relevant laws, regulations, and standards of relevant countries and regions, take necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or reject.

[0160] It should be noted that the technical solution provided in this application provides corresponding operation entrances for users to choose to agree or reject the automated decision-making results; if the user chooses to reject, the expert decision-making process will be entered.

[0161] In the technical solution of the distributed storage unit access method provided by the embodiments of the present invention, a computing unit initiates a storage address access request to the on-chip interconnection network; the on-chip interconnection network interleaves the storage address indicated by the storage address access request to generate an interleaved storage address distribution; the on-chip interconnection network determines the target storage unit corresponding to the storage address access request according to the interleaved storage address distribution, and sends the storage address access request to the memory controller within the corresponding modular integrated circuit unit; the memory controller sends the storage address access request to the target storage unit, and the target storage unit is the target local storage unit or the target virtual remote storage unit; the storage unit includes a local storage unit and a virtual remote storage unit, the local storage unit is distributed on the current modular integrated circuit unit, and the virtual remote storage unit is distributed on modular integrated circuit units other than the current modular integrated circuit unit. By setting the local storage unit and the virtual remote storage unit, and using the address interleaving technology and virtual storage mapping, the distributed storage units on multiple dies can achieve unified address access and load balancing, ensuring unified memory access to the computing system, balancing the access latency of each storage unit, reducing complexity, improving resource utilization, and enhancing the performance and usability of the distributed storage system.

[0162] The systems, modules or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer device. Specifically, the computer device can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0163] The embodiments of the present invention provide a computer device, including the distributed storage unit access system as described above. For specific descriptions, reference can be made to the embodiments of the above distributed storage unit access system.

[0164] Next, refer to Figure 19 , which shows a schematic structural diagram of a computer device 600 suitable for implementing the embodiments of the present application.

[0165] As Figure 19As shown, the computer device 600 includes a central processing unit (CPU) 601, which can perform various appropriate operations and processes according to the programs stored in the read-only memory (ROM) 602 or the programs loaded from the storage section 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the computer device 600 are also stored. The CPU 601, ROM 602, and RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0166] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed so that a computer program read from it can be installed in the storage section 608 as needed.

[0167] For the convenience of description, the above devices are described separately as various units according to their functions. Of course, when implementing the present application, the functions of each unit can be implemented in hardware.

[0168] It should also be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, commodity or device including the said element.

[0169] In the technical solution of the present application, the acquisition, storage, use, processing, etc. of data all comply with the relevant provisions of national laws and regulations.

[0170] It should be noted that in the embodiments of the present application, some industry-existing solutions such as certain software, components, models, etc. may be mentioned. They should be regarded as exemplary, and their purpose is only to illustrate the feasibility in the implementation of the technical solution of the present application, but it does not mean that the applicant has already or necessarily used this solution.

[0171] Those skilled in the art should understand that the embodiments of the present application can be provided as methods or systems. Therefore, the present application can take the form of a completely hardware embodiment or an embodiment combining software and hardware aspects.

[0172] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiment.

[0173] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A modular integrated circuit unit, characterized in that, The modular integrated circuit unit includes at least one computing unit, at least one memory controller, at least one on-chip interconnection network, and at least one storage unit; The memory controller is configured to send a storage address access request sent by the on-chip interconnection network to the storage unit; The storage unit is configured to provide unified memory access to the storage address access request sent by the memory controller. The storage unit includes a local storage unit and a virtual remote storage unit. The local storage units are distributed on the current modular integrated circuit unit, and the virtual remote storage units are distributed on modular integrated circuit units other than the current modular integrated circuit unit; The computing unit is configured to initiate a storage address access request; The on-chip interconnection network is configured to interleave the storage addresses indicated by the storage address access requests sent by the computing unit to generate an interleaved storage address distribution; according to the interleaved storage address distribution, determine the target storage unit corresponding to the storage address access request, and send the storage address access request to the corresponding target storage unit through the memory controller. The target storage unit is a target local storage unit or a target virtual remote storage unit.

2. The modular integrated circuit unit according to claim 1, characterized in that The modular integrated circuit unit is a die in a multi-die package or a chip in a computing system composed of independent chips.

3. The modular integrated circuit unit according to claim 1, characterized in that Specifically, the on-chip interconnection network is configured to divide the storage addresses according to a preset interleaving granularity to obtain storage address distributions in multiple intervals; uniformly map the storage address distributions in multiple intervals to different storage units on different modular integrated circuit units to obtain an interleaved storage address distribution.

4. The modular integrated circuit unit according to claim 1, characterized in that, Specifically, the on-chip interconnection network is configured to send, according to the interleaved storage address distribution, the storage address access requests distributed on the current modular integrated circuit unit where it is located to the target local storage unit through the memory controller, or send the storage address access requests distributed on modular integrated circuit units other than the current modular integrated circuit unit where it is located to the target virtual remote storage unit through an interconnection interface.

5. A distributed storage unit access method, characterized in that The method includes: The computing unit sends a storage address access request to the on-chip interconnection network; The on-chip interconnection network interleaves the storage addresses indicated by the storage address access request to generate an interleaved storage address distribution; The on-chip interconnection network determines the target storage unit corresponding to the storage address access request according to the interleaved storage address distribution, and sends the storage address access request to the memory controller within the corresponding modular integrated circuit unit; The memory controller sends the storage address access request to the target storage unit. The target storage unit is a target local storage unit or a target virtual remote storage unit; The storage unit includes a local storage unit and a virtual remote storage unit. The local storage units are distributed on the current modular integrated circuit unit, and the virtual remote storage units are distributed on modular integrated circuit units other than the current modular integrated circuit unit.

6. The distributed storage unit access method according to claim 5, wherein, The on-chip interconnection network interleaves the storage addresses indicated by the storage address access request to generate an interleaved storage address distribution, including: The on-chip interconnection network divides the storage address according to a preset interleaving granularity to obtain the storage address distributions of multiple intervals; and evenly maps the storage address distributions of multiple intervals to different storage units on different modular integrated circuit units to obtain the interleaved storage address distribution.

7. The distributed storage unit access method according to claim 5, wherein, The on-chip interconnection network sends the storage address access request to the memory controller in the corresponding modular integrated circuit unit according to the interleaved storage address distribution, including: The on-chip interconnection network sends the storage address access request distributed on the current modular integrated circuit unit to the corresponding memory controller according to the interleaved storage address distribution, or sends the storage address access request distributed on the modular integrated circuit unit other than the current modular integrated circuit unit to the memory controller in the corresponding modular integrated circuit unit through the interconnection interface.

8. The distributed storage unit access method according to claim 7, wherein The interleaved storage address distribution includes the storage address distribution of the corresponding interval of each storage unit; The on-chip interconnection network sends the storage address access request distributed on the current modular integrated circuit unit to the corresponding memory controller according to the interleaved storage address distribution, including: The on-chip interconnection network matches the storage address indicated by the storage address access request distributed on the current modular integrated circuit unit with the storage address distributions of the corresponding intervals of each storage unit in the storage address distribution to determine the target local storage unit, and sends the storage address access request distributed on the current modular integrated circuit unit to the corresponding memory controller.

9. The distributed storage unit access method according to claim 8, characterized in that, The memory controller sends the storage address access request to the target storage unit, including: The memory controller sends the storage address access request to the target local storage unit.

10. The distributed storage unit access method according to claim 8, wherein The on-chip interconnection network stores the interconnection interface bound to each storage unit; The sending the storage address access request distributed on the modular integrated circuit unit other than the current modular integrated circuit unit to the memory controller in the corresponding modular integrated circuit unit through the interconnection interface includes: The on-chip interconnection network matches the storage address indicated by the storage address access request distributed on the modular integrated circuit unit other than the current modular integrated circuit unit with the storage address distributions of the corresponding intervals of each storage unit in the storage address distribution to determine the target interconnection interface bound to the target remote modular integrated circuit unit; Through the target interconnection interface, the on-chip interconnection network sends the storage address access request distributed on the modular integrated circuit unit other than the current modular integrated circuit unit to the target remote modular integrated circuit unit for the target remote modular integrated circuit unit to send the storage address access request to the corresponding memory controller.

11. The distributed storage unit access method according to claim 10, wherein The memory controller sends the storage address access request to the target storage unit, including: The memory controller sends the storage address access request to the target virtual remote storage unit.

12. A distributed storage unit access system, characterized in that, The system includes: a plurality of modular integrated circuit units as described in any one of claims 1-4, and the plurality of modular integrated circuit units are communicatively connected through an interconnection interface, and the interconnection interface includes an inter-die interconnection interface or an inter-chip interconnection interface; If the modular integrated circuit unit is a die in a multi-die package, the dies are communicatively connected through an inter-die interconnection interface; If the modular integrated circuit unit is a chip in a computing system composed of independent chips, the chips are communicatively connected through an inter-chip interconnection interface.

13. A computer device, characterized in that, It includes the distributed storage unit access system as described in claim 12.

Citation Information

Patent Citations

  • Computing device extension system for system on chip (SOC)

    CN103246623A

  • On-chip cache system for transformation on ultrahigh-definition video frame rates

    CN104268098A

  • System-level space read-write verification method and system, storage medium and equipment

    CN114546890A

  • Apparatus and method for using multiple physical address spaces

    CN115335814A

  • GPU (Graphics Processing Unit) memory access adaptive optimization method and device based on extension page table

    CN115422098A

Cited By

  • Data storage method and device, electronic equipment and storage medium

    CN120994141A

  • Method for accessing memory and on-chip bus interconnection system

    CN121478722A

  • Chip system and channel interleaving method thereof

    CN121858480A

  • Memory access method and device of multi-chip system, electronic equipment and storage medium

    CN122195893A