Near-memory computing module and method, near-memory computing network and construction method
Through three-dimensional design within the same chip, the computing submodule and memory submodule are connected through bonding, and the routing unit is used to exchange data, solving the problem of unbalanced performance between processor and memory, and achieving low latency and high bandwidth computing performance.
Patent Information
- Application Number
- CN202010753117.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-30
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2040-07-30
AI Technical Summary
In the prior art, the performance development of processors and memory is unbalanced, resulting in memory bottlenecks that limit the performance of processors. There are technical difficulties in the existing in-memory computing structure, such as immature process of resistive variable memory RRAM and the reliability of through-silicon TSV technology.
The near-memory computing module adopts a three-dimensional design, which connects the computing submodule and the memory submodule through bonding in different layers, and exchanges data between the routing unit and the computing unit and the memory unit through an exchange interface to achieve low latency and high bandwidth computing performance.
Implementing large-scale computing within the same chip reduces data access latency, improves computing performance, meets computing needs of different scales, and ensures electrical performance through mature bonding methods.
Smart Images

Figure CN113688065B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of memories. Specifically, the present invention relates to a near-memory computing module and method, a near-memory computing network, and a construction method thereof. Background Art
[0002] In the prior art, processor manufacturers and memory manufacturers are separated from each other, resulting in the asynchronous development of memory technology and processor technology. The performance of processors has been rapidly improved, while the performance of memories has been relatively slow to improve. The unbalanced development of processor performance and memory performance has caused the access speed of memories to lag far behind the computing speed of processors. The memory bottleneck makes it difficult for high-performance processors to exert their due effects, which greatly restricts the growing high-performance computing. This phenomenon that the memory performance severely limits the performance of processors is called the "Memory Wall".
[0003] With the continuous rapid growth of the computing power of the central processing unit (CPU) and the scale of application computing, the problem of the "Memory Wall" has become increasingly prominent.
[0004] To solve the "Memory Wall" problem, the concept of "near-memory computing" or "bringing the memory as close as possible to the computing" has emerged.
[0005] Traditional computing units and memory units are separated, that is, not in the same chip. Therefore, in traditional computing processes, the computing unit needs to extract data from the memory unit, and then write it back to the memory unit after processing. While "in-memory computing" combines the memory unit and the computing unit. By bringing the memory unit as close as possible to the computing unit, the data transmission path is shortened, thereby reducing the data access latency. At the same time, "in-memory computing" tries to increase the access bandwidth, thereby effectively improving the computing performance.
[0006] A known "in-memory computing" structure in the prior art is as Figure 1 shown. In this "in-memory computing" structure, the memory unit uses a resistive random access memory (RRAM), and the computing unit is a neural-network processing unit (NPU). Since the resistive random access memory (RRAM) and the neural-network processing unit (NPU) use an integrated structure, the access latency of the neural-network processing unit (NPU) is low. However, the process technology of the resistive random access memory (RRAM) is not yet mature, and the structural design is not scalable, making it difficult to meet the requirements of higher-performance computing.
[0007] In the prior art, there is also a 3D stacking technology using Through-Silicon Via (TSV) technology to implement an "in-memory computing" structure. This 3D stacking technology stacks multiple wafers together and uses TSV technology to interconnect different wafers. This is a three-dimensional multi-layer stacking that enables communication between multiple wafers in the vertical direction through TSV. However, there are many technical difficulties in this 3D stacking technology. For example, the filling technology of deep TSV holes, because the filling effect of deep TSV holes is directly related to the reliability and yield of the 3D stacking technology, which is crucial for the integration and practical application of the 3D stacking technology. Another example is that TSV technology needs to maintain good integrity during the substrate thinning process to avoid crack propagation.
[0008] Therefore, it is urgent to solve the above-mentioned technical problems in the prior art. Summary of the Invention
[0009] The present invention proposes a near-memory solution, which relates to a near-memory computing module and method, a near-memory computing network, and a construction method. The near-memory computing module of the present invention adopts a three-dimensional design, with the computing sub-module and the memory sub-module arranged in different layers. Preferably, the layers are connected by bonding, and the total data bit width of the connection is a positive integer multiple of the data bit width of a single computing unit. In this way, the problems of storage latency and bandwidth are solved. The memory sub-module has multiple memory cells, enabling a large memory capacity to be achieved in a single memory sub-module. The computing units in the computing sub-module exchange data through the switching interfaces of the router, and each computing sub-module accesses data through the routing interfaces, further improving the computing performance. This near-memory computing network utilizes this near-memory computing module to meet computing requirements of different scales.
[0010] According to the first aspect of the present invention, there is provided a near-memory computing module, which includes:
[0011] A computing sub-module, which includes multiple computing units;
[0012] At least one memory sub-module, which is arranged on at least one side of the computing sub-module, where each memory sub-module includes multiple memory cells and each memory sub-module is connected to the computing sub-module;
[0013] Wherein the computing sub-module and the at least one memory sub-module are located within the same chip.
[0014] Thus, the near-memory computing module including multiple memory sub-modules can achieve large-scale computing within the same chip, and the computing units can achieve low latency when accessing the memory units, improving the computing performance.
[0015] According to a preferred embodiment of the near-memory computing module of the present invention, the computing sub-module further includes: a routing unit;
[0016] Wherein the routing unit is connected to each computing unit, and the routing unit is connected to each memory unit of each memory sub-module, and the routing unit is connected to the routing unit of at least another near-memory computing module;
[0017] Wherein the routing unit is configured to perform an access by a first computing unit of the near-memory computing module to a second computing unit of the near-memory computing module, or an access to a first memory unit of the near-memory computing module, or an access to a third computing unit or a second memory unit of at least another near-memory computing module.
[0018] According to a preferred embodiment of the near-memory computing module of the present invention, the routing unit includes:
[0019] A plurality of switching interfaces, the plurality of switching interfaces connecting the routing unit to each computing unit;
[0020] A routing interface, the routing interface connecting the routing unit to the routing unit of at least another near-memory computing module;
[0021] A memory control interface, the memory control interface connecting the routing unit to each memory unit in each memory sub-module.
[0022] According to a preferred embodiment of the near-memory computing module of the present invention, the routing unit further includes:
[0023] A crossbar unit;
[0024] A switching and routing calculation unit, the switching and routing calculation unit being connected to the plurality of switching interfaces, the routing interface, and the crossbar unit, and the switching and routing calculation unit storing at least routing information about the near-memory computing module and the plurality of computing units, the switching and routing calculation unit parsing the received data access request and controlling the switching of the crossbar unit based on the parsed data access request information;
[0025] A memory control unit, the memory control unit being connected to the crossbar unit and the memory control interface, and the memory control unit storing at least routing information about the plurality of memory units, the memory control unit, in response to the crossbar unit switching to the memory control unit, performing a secondary parsing of the received and parsed data access request from the crossbar unit to determine the target memory unit, and accessing the target memory unit via the memory control interface.
[0026] In a preferred embodiment of the near-memory computing module according to the present invention, each computing unit directly accesses at least one memory unit via a routing unit. That is, the routing unit parses the data access request issued by the computing unit, directly obtains the access data from the at least one memory unit, and forwards the access data to the computing unit that issued the data access request, where the data access request includes at least the address of the at least one memory unit.
[0027] In a preferred embodiment of the near-memory computing module according to the present invention, each computing unit indirectly accesses at least one other memory unit via a routing unit. That is, the routing unit parses the data access request issued by the computing unit, forwards the parsed data access request to another computing unit, indirectly obtains the access data from at least one other memory unit via the other computing unit, and forwards the access data to the computing unit that issued the data access request, where the data access request includes at least the address of the at least one other memory unit, and the other computing unit can directly access the at least one other memory unit via the routing unit.
[0028] Thus, the routing unit realizes the unified distribution of the access requests of the computing units to the memory units, and realizes the memory control function, which can achieve further low latency when the computing units access the memory units.
[0029] In a preferred embodiment of the near-memory computing module according to the present invention, the routing unit is connected to each memory unit of each memory sub-module by bonding.
[0030] Thus, the memory sub-module and the computing sub-module can be connected by a mature bonding method to achieve the required electrical performance.
[0031] In a preferred embodiment of the near-memory computing module according to the present invention, the total data bit width of the connection between the routing unit and each memory unit of each memory sub-module is n times the data bit width of a single computing unit, where n is a positive integer.
[0032] Thus, by setting the relationship between the data bit width of the connection between the memory unit and the routing unit and the data bit width of a single computing unit, a higher data bandwidth can be achieved.
[0033] In a preferred embodiment of the near-memory computing module according to the present invention, in the computing sub-module, the routing unit is located at the center, and the multiple computing units are distributed around the routing unit.
[0034] According to a preferred embodiment of the near-memory computing module of the present invention, the computing sub-module further includes: at least two routing units, each routing unit is connected to at least one computing unit, and each routing unit is connected to at least one memory unit of each memory sub-module;
[0035] The at least two routing units are connected to each other to form an overall routing unit, the overall routing unit is connected to each computing unit, and the overall routing unit is connected to each memory unit of each memory sub-module, and the overall routing unit is connected to at least another routing unit of at least another near-memory computing module;
[0036] Wherein the overall routing unit is configured to perform access by a first computing unit of the near-memory computing module to a second computing unit of the near-memory computing module, or access to a first memory unit of the near-memory computing module, or access to a third computing unit or a second memory unit of at least another near-memory computing module.
[0037] According to a preferred embodiment of the near-memory computing module of the present invention, each routing unit of the at least two routing units includes:
[0038] A plurality of switching interfaces, the plurality of switching interfaces connect the routing unit to at least one computing unit;
[0039] A routing interface, the routing interface connects the routing unit to at least another routing unit of the near-memory computing module, and / or to at least another routing unit of at least another near-memory computing module;
[0040] A memory control interface, the memory control interface connects the routing unit to at least one memory unit.
[0041] According to a preferred embodiment of the near-memory computing module of the present invention, the at least two routing units are connected to each other through the routing interface.
[0042] According to a preferred embodiment of the near-memory computing module of the present invention, each routing unit of the at least two routing units further includes:
[0043] A crossbar unit;
[0044] A switching routing calculation unit, the switching routing calculation unit is connected to the plurality of switching interfaces, the routing interface, and the crossbar unit, and the switching routing calculation unit stores at least routing information about the near-memory computing module and the plurality of computing units, and the switching routing calculation unit parses the received data access request and controls the switching of the crossbar unit based on the parsed data access request information;
[0045] A memory control unit is connected to the crossbar switch unit and the memory control interface. The memory control unit stores at least routing information about the multiple memory units. In response to the crossbar switch unit switching to the memory control unit, the memory control unit performs secondary parsing on the parsed data access request received from the crossbar switch unit to determine the memory unit to be accessed, and accesses the memory unit to be accessed via the memory control interface.
[0046] According to a preferred embodiment of the in-memory computing module of the present invention, each computing unit directly accesses at least one memory unit via the global routing unit. That is, the global routing unit parses the data access request issued by the computing unit, directly obtains the access data from the at least one memory unit, and forwards the access data to the computing unit that issued the data access request, where the data access request includes at least the address of the at least one memory unit.
[0047] According to a preferred embodiment of the in-memory computing module of the present invention, each computing unit indirectly accesses at least one other memory unit via the global routing unit. That is, the global routing unit parses the data access request issued by the computing unit, forwards the parsed data access request to another computing unit, indirectly obtains the access data from at least one other memory unit via the other computing unit, and forwards the access data to the computing unit that issued the data access request, where the data access request includes at least the address of the at least one other memory unit, and the other computing unit can directly access the at least one other memory unit via the global routing unit.
[0048] According to a preferred embodiment of the in-memory computing module of the present invention, the global routing unit is connected to each memory unit of each memory sub-module by bonding.
[0049] According to a preferred embodiment of the in-memory computing module of the present invention, the total data bit width between each memory unit and the global routing unit is n times the data bit width of a single computing unit, where n is a positive integer.
[0050] According to a preferred embodiment of the in-memory computing module of the present invention, in the computing sub-module, the at least two routing units are located at the center, and the multiple computing units are distributed around the at least two routing units.
[0051] According to a preferred embodiment of the in-memory computing module of the present invention, in the computing sub-module, the multiple computing units are located at the center, and the at least two routing units are distributed around the multiple computing units.
[0052] According to a preferred embodiment of the near-memory computing module of the present invention, the memory unit includes a dynamic random access memory, and the computing unit includes a central processing unit.
[0053] Since the process of dynamic random access memory is relatively mature, such a memory is preferably used in the present invention.
[0054] According to a preferred embodiment of the near-memory computing module of the present invention, the number of memory units in the memory sub-module is determined at least according to the total data bit width of the multiple computing units in the computing sub-module and the data bit width of a single memory unit.
[0055] Since the number of memory units can be selected according to requirements, the design is made more flexible.
[0056] According to a preferred embodiment of the near-memory computing module of the present invention, the storage capacity of the memory unit is customizable.
[0057] Since the storage capacity of the memory unit is customizable, the flexibility of the design is improved.
[0058] According to a second aspect of the present invention, there is provided a near-memory computing method, which is used for the above-mentioned near-memory computing module (in the near-memory computing module, there is a routing unit in a computing sub-module), and the near-memory computing method includes:
[0059] The routing unit receives a data access request, which is issued by a first computing unit and includes at least the address of the target memory unit; and,
[0060] The routing unit parses the data access request, obtains the access data from the target memory unit, and forwards the access data to the first computing unit.
[0061] In this technical solution, if the target memory unit and the first computing unit are in the same near-memory computing module, the "routing unit" refers to the routing unit in the near-memory computing module; if the target memory unit and the first computing unit are not in the same near-memory computing module, the "routing unit" refers to all the routing units required for communication between the first computing unit and the target memory unit.
[0062] According to a preferred embodiment of the near-memory computing method of the present invention, the near-memory computing method further includes:
[0063] After the routing unit parses the data access request and before obtaining the access data from the target memory unit, the routing unit connected to the first computing unit determines whether the first computing unit can directly access the target memory unit via the routing unit connected to the first computing unit;
[0064] When the first computing unit can directly access the target memory unit via the routing unit connected to the first computing unit, the routing unit connected to the first computing unit directly obtains access data from the target memory unit and forwards the access data to the first computing unit; and
[0065] When the first computing unit cannot directly access the target memory unit via the routing unit connected to the first computing unit, the routing unit connected to the first computing unit forwards the resolved data access request to the second computing unit, and obtains access data from the target memory unit via the second computing unit and forwards the access data to the first computing unit.
[0066] According to a preferred embodiment of the in-memory computing method of the present invention, the in-memory computing method further includes:
[0067] When the first computing unit cannot directly access the target memory unit via the routing unit connected to the first computing unit and before the routing unit connected to the first computing unit forwards the resolved data access request to the second computing unit, the routing unit connected to the first computing unit determines whether the target memory unit and the first computing unit are located in the same in-memory computing module;
[0068] When the target memory unit and the first computing unit are located in the same in-memory computing module, the routing unit connected to the first computing unit directly forwards the resolved data access request to the second computing unit, and obtains access data from the target memory unit via the second computing unit and forwards the access data to the first computing unit, wherein the second computing unit can directly access the target memory unit via the routing unit connected to the first computing unit;
[0069] When the target memory unit and the first computing unit are not located in the same in-memory computing module, the routing unit connected to the first computing unit forwards the resolved data access request to the routing unit of at least another in-memory computing module connected to the routing unit connected to the first computing unit and forwards it to the second computing unit connected to the routing unit of the at least another in-memory computing module, and obtains access data from the target memory unit via the second computing unit and forwards the access data to the first computing unit, wherein the second computing unit can directly access the target memory unit via the routing unit of the at least another in-memory computing module.
[0070] According to a third aspect of the present invention, there is provided an in-memory computing method, which is used for the above-mentioned in-memory computing module (in the in-memory computing module, there is one routing unit in one computing sub-module), and the in-memory computing method includes:
[0071] The routing unit receives a data access request, which is sent by the first computing unit and includes at least the address of the target computing unit; and,
[0072] The routing unit parses the data access request, obtains the access data from the target computing unit, and forwards the access data to the first computing unit.
[0073] In this technical solution, if the target computing unit and the first computing unit are located within the same near-memory computing module, the "routing unit" refers to the routing unit within this near-memory computing module; if the target computing unit and the first computing unit are not located within the same near-memory computing module, the "routing unit" refers to all the routing units required for communication between the first computing unit and the target computing unit.
[0074] According to a preferred implementation of the near-memory computing method of the present invention, the near-memory computing method further includes:
[0075] After the routing unit parses the data access request and before obtaining the access data from the target computing unit, the routing unit connected to the first computing unit determines whether the target computing unit and the first computing unit are located within the same near-memory computing module;
[0076] When the target computing unit and the first computing unit are located within the same near-memory computing module, the routing unit connected to the first computing unit directly obtains the access data from the target computing unit and forwards the access data to the first computing unit;
[0077] When the target computing unit and the first computing unit are not located within the same near-memory computing module, the routing unit connected to the first computing unit forwards the parsed data access request to the routing unit of at least another near-memory computing module connected to this routing unit, and obtains the access data from the target computing unit via the routing unit of this at least another near-memory computing module and forwards the access data to the first computing unit.
[0078] According to a fourth aspect of the present invention, a near-memory computing method is provided. This near-memory computing method is used for the above-mentioned near-memory computing module (in this near-memory computing module, there are at least two routing units in one computing sub-module), and this near-memory computing method includes:
[0079] The overall routing unit receives a data access request, which is sent by the first computing unit and includes at least the address of the target memory unit; and,
[0080] The overall routing unit parses the data access request, obtains the access data from the target memory unit, and forwards the access data to the first computing unit.
[0081] In this technical solution, if the target memory unit and the first computing unit are within the same near-memory computing module, the "overall routing unit" refers to the overall routing unit within this near-memory computing module; if the target memory unit and the first computing unit are not within the same near-memory computing module, the "overall routing unit" refers to all the overall routing units required for communication between the first computing unit and the target memory unit.
[0082] According to a preferred embodiment of the near-memory computing method of the present invention, the near-memory computing method further includes:
[0083] After the overall routing unit parses the data access request and before obtaining the access data from the target memory unit, the overall routing unit connected to the first computing unit determines whether the first computing unit can directly access the target memory unit via the overall routing unit connected to the first computing unit;
[0084] When the first computing unit can directly access the target memory unit via the overall routing unit connected to the first computing unit, the overall routing unit connected to the first computing unit directly obtains the access data from the target memory unit and forwards the access data to the first computing unit; and
[0085] When the first computing unit cannot directly access the target memory unit via the overall routing unit connected to the first computing unit, the overall routing unit connected to the first computing unit forwards the parsed data access request to the second computing unit, and obtains the access data from the target memory unit via the second computing unit and forwards the access data to the first computing unit.
[0086] According to a preferred embodiment of the near-memory computing method of the present invention, the near-memory computing method further includes:
[0087] When the first computing unit cannot directly access the target memory unit via the overall routing unit connected to the first computing unit and before the overall routing unit connected to the first computing unit forwards the parsed data access request to the second computing unit, the overall routing unit connected to the first computing unit determines whether the target memory unit and the first computing unit are within the same near-memory computing module;
[0088] When the target memory unit and the first computing unit are within the same near-memory computing module, the overall routing unit connected to the first computing unit directly forwards the parsed data access request to the second computing unit, and obtains the access data from the target memory unit via the second computing unit and forwards the access data to the first computing unit, where the second computing unit can directly access the target memory unit via the overall routing unit connected to the first computing unit;
[0089] When the target memory unit and the first computing unit are not in the same near-memory computing module, the overall routing unit connected to the first computing unit forwards the parsed data access request to the overall routing unit of at least another near-memory computing module connected to the overall routing unit connected to the first computing unit, and forwards it to the second computing unit connected to the overall routing unit of the at least another near-memory computing module, and obtains the access data from the target memory unit via the second computing unit and forwards the access data to the first computing unit, where the second computing unit can directly access the target memory unit via the overall routing unit of the at least another near-memory computing module.
[0090] According to the fifth aspect of the present invention, there is provided a near-memory computing method, which is used for the above-mentioned near-memory computing module (in the near-memory computing module, there are at least two routing units in one computing sub-module), and the near-memory computing method includes:
[0091] The overall routing unit receives a data access request, which is issued by the first computing unit and includes at least the address of the target computing unit; and,
[0092] The overall routing unit parses the data access request, obtains the access data from the target computing unit and forwards the access data to the first computing unit.
[0093] In this technical solution, if the target computing unit and the first computing unit are in the same near-memory computing module, the "overall routing unit" refers to the overall routing unit in the near-memory computing module. If the target computing unit and the first computing unit are not in the same near-memory computing module, the "overall routing unit" refers to all the overall routing units required for communication between the first computing unit and the target computing unit.
[0094] According to a preferred embodiment of the near-memory computing method of the present invention, the near-memory computing method further includes:
[0095] After the overall routing unit parses the data access request and before obtaining the access data from the target computing unit, the overall routing unit connected to the first computing unit determines whether the target computing unit and the first computing unit are in the same near-memory computing module;
[0096] When the target computing unit and the first computing unit are in the same near-memory computing module, the overall routing unit connected to the first computing unit directly obtains the access data from the target computing unit and forwards the access data to the first computing unit;
[0097] When the target computing unit and the first computing unit are not in the same near-memory computing module, the overall routing unit connected to the first computing unit forwards the parsed data access request to the overall routing unit of at least another near-memory computing module connected to the overall routing unit, and obtains access data from the target computing unit via the overall routing unit of the at least another near-memory computing module and forwards the access data to the first computing unit.
[0098] According to a sixth aspect of the present invention, there is provided a near-memory computing network, the near-memory computing network comprising:
[0099] A plurality of near-memory computing modules, the plurality of near-memory computing modules being the plurality of above-mentioned near-memory computing modules, and the plurality of near-memory computing modules are connected by routing units.
[0100] According to a preferred embodiment of the near-memory computing network of the present invention, the plurality of near-memory computing modules are connected into a bus-type, star-type, ring-type, tree-type, mesh-type and hybrid topology structure.
[0101] According to a preferred embodiment of the near-memory computing network of the present invention, the plurality of near-memory computing modules are connected via metal wires by routing units.
[0102] According to a seventh aspect of the present invention, there is provided a method for constructing a near-memory computing module, the construction method comprising:
[0103] Arranging at least one memory sub-module on at least one side of the computing sub-module, wherein each memory sub-module includes a plurality of memory cells, and the computing sub-module includes a plurality of computing units;
[0104] Connecting each memory sub-module to the computing sub-module;
[0105] Wherein the computing sub-module and the at least one memory sub-module are arranged in the same chip.
[0106] According to a preferred embodiment of the construction method of the present invention, the computing sub-module further comprises: a routing unit;
[0107] The construction method further comprises:
[0108] Connecting the routing unit to each computing unit, and connecting the routing unit to each memory cell of each memory sub-module, and connecting the routing unit to the routing unit of at least another near-memory computing module, and configuring the routing unit to perform access of the first computing unit of the near-memory computing module to the second computing unit of the near-memory computing module, or access to the first memory cell of the near-memory computing module, or access to the third computing unit or the second memory cell of at least another near-memory computing module.
[0109] According to a preferred embodiment of the construction method of the present invention, the construction method further includes:
[0110] Connecting a plurality of switching interfaces of the routing unit to each computing unit;
[0111] Connecting the routing interface of the routing unit to the routing interface of the routing unit of at least another near-memory computing module;
[0112] Connecting the memory control interface of the routing unit to each memory unit in each memory sub-module.
[0113] According to a preferred embodiment of the construction method of the present invention, the construction method further includes:
[0114] Connecting the switching routing calculation unit of the routing unit to the plurality of switching interfaces, the routing interface, and a crossbar unit, and storing at least the routing information about the near-memory computing module and the plurality of computing units in the switching routing calculation unit, and configuring the switching routing calculation unit to parse the received data access request and control the switching of the crossbar unit based on the parsed data access request information;
[0115] Connecting the memory control unit of the routing unit to the crossbar unit and the memory control interface, and storing at least the routing information of the plurality of memory units in the memory control unit, and configuring the memory control unit to perform secondary parsing on the parsed data access request received from the crossbar unit in response to the crossbar unit switching to the memory control unit to determine the target memory unit, and accessing the target memory unit via the memory control interface.
[0116] According to a preferred embodiment of the construction method of the present invention, the construction method further includes: configuring each computing unit to directly access at least one memory unit via the routing unit. That is, the routing unit parses the data access request issued by the computing unit, directly obtains the access data from the at least one memory unit, and forwards the access data to the computing unit that issued the data access request, where the data access request at least includes the address of the at least one memory unit.
[0117] According to a preferred embodiment of the construction method of the present invention, the construction method further includes: configuring each computing unit to indirectly access at least one other memory unit via a routing unit. That is, the routing unit parses the data access request issued by the computing unit, forwards the parsed data access request to another computing unit, indirectly obtains access data from at least one other memory unit via the other computing unit, and forwards the access data to the computing unit that issued the data access request, where the data access request at least includes the address of the at least one other memory unit, and the other computing unit can directly access the at least one other memory unit via the routing unit.
[0118] According to a preferred embodiment of the construction method of the present invention, the construction method further includes: connecting the routing unit to each memory unit of each memory sub-module by bonding.
[0119] According to a preferred embodiment of the construction method of the present invention, the construction method further includes: setting the total data bit width of the connection between the routing unit and each memory unit of each memory sub-module to n times the data bit width of a single computing unit, where n is a positive integer.
[0120] According to a preferred embodiment of the construction method of the present invention, the construction method further includes: in the computing sub-module, arranging the routing unit at the center and arranging the multiple computing units to be distributed around the routing unit.
[0121] According to a preferred embodiment of the construction method of the present invention, the computing sub-module further includes: at least two routing units, where each routing unit is connected to at least one computing unit, and each routing unit is connected to at least one memory unit of each memory sub-module;
[0122] The construction method further includes:
[0123] Connecting the at least two routing units to each other to form an overall routing unit, connecting the overall routing unit to each computing unit, connecting the overall routing unit to each memory unit of each memory sub-module, and connecting the overall routing unit to at least another routing unit of at least another near-memory computing module;
[0124] Configuring the overall routing unit to perform access by the first computing unit of the near-memory computing module to the second computing unit of the near-memory computing module, or to the first memory unit of the near-memory computing module, or to the third computing unit or the second memory unit of another near-memory computing module.
[0125] According to a preferred embodiment of the construction method of the present invention, the construction method further includes:
[0126] Connect the multiple switching interfaces of each routing unit among the at least two routing units to at least one computing unit;
[0127] Connect the routing interface of each routing unit among the at least two routing units to the routing interface of at least another routing unit of the near-memory computing module, and / or connect the routing interface of each routing unit among the at least two routing units to the routing interface of at least another routing unit of at least another near-memory computing module;
[0128] Connect the memory control interface of each routing unit among the at least two routing units to at least one memory unit.
[0129] According to a preferred embodiment of the construction method of the present invention, the construction method further includes:
[0130] Connect the at least two routing units to each other through routing interfaces.
[0131] According to a preferred embodiment of the construction method of the present invention, the construction method further includes:
[0132] Connect the switching routing calculation unit of each routing unit among the at least two routing units to the multiple switching interfaces, the routing interface, and a crossbar unit, and store at least the routing information about the near-memory computing module and the multiple computing units in the switching routing calculation unit, and configure the switching routing calculation unit to parse the received data access request and control the switching of the crossbar unit based on the parsed data access request information;
[0133] Connect the memory control unit of each routing unit among the at least two routing units to the crossbar unit and the memory control interface, and store at least the routing information of the multiple memory units in the memory control unit, and configure the memory control unit to perform secondary parsing on the parsed data access request received from the crossbar unit in response to the crossbar unit switching to the memory control unit to determine the memory unit to be accessed, and access the memory unit to be accessed via the memory control interface.
[0134] According to a preferred embodiment of the construction method of the present invention, the construction method further includes:
[0135] Configure each computing unit to directly access at least one memory unit via the overall routing unit. That is, the overall routing unit parses the data access request issued by the computing unit, directly obtains the access data from the at least one memory unit, and forwards the access data to the computing unit that issued the data access request, where the data access request includes at least the address of the at least one memory unit.
[0136] According to a preferred embodiment of the construction method of the present invention, the construction method further includes:
[0137] Configuring each computing unit to indirectly access at least one other memory unit via the global routing unit. That is, the global routing unit parses the data access requests issued by the computing units, forwards the parsed data access requests to another computing unit, indirectly obtains access data from at least one other memory unit via the other computing unit, and forwards the access data to the computing unit that issued the data access request, where the data access request includes at least the address of the at least one other memory unit, and where the other computing unit can directly access the at least one other memory unit via the global routing unit.
[0138] According to a preferred embodiment of the construction method of the present invention, the construction method further includes:
[0139] Connecting the global routing unit to each memory unit of each memory sub-module by bonding.
[0140] According to a preferred embodiment of the construction method of the present invention, the construction method further includes:
[0141] Setting the total data bit width between each memory unit and the global routing unit to n times the data bit width of a single computing unit, where n is a positive integer.
[0142] According to a preferred embodiment of the construction method of the present invention, the construction method further includes:
[0143] In the computing sub-module, arranging the at least two routing units at the center and arranging the multiple computing units to be distributed around the at least two routing units.
[0144] According to a preferred embodiment of the construction method of the present invention, the construction method further includes:
[0145] In the computing sub-module, arranging the multiple computing units at the center and arranging the at least two routing units to be distributed around the multiple computing units.
[0146] According to a preferred embodiment of the construction method of the present invention, the construction method further includes:
[0147] Setting the memory unit to include a dynamic random access memory and setting the computing unit to include a central processing unit.
[0148] According to a preferred embodiment of the construction method of the present invention, the construction method further includes:
[0149] Determine the number of memory units in the memory sub-module based at least on the total data bit width of the multiple computing units in the computing sub-module and the data bit width of a single memory unit.
[0150] According to a preferred embodiment of the construction method of the present invention, the storage capacity of the memory unit is set to be customizable.
[0151] According to an eighth aspect of the present invention, there is provided a method for constructing a near-memory computing network, the construction method comprising:
[0152] Connect multiple near-memory computing modules through a routing unit, wherein the multiple near-memory computing modules are the multiple above-mentioned near-memory computing modules.
[0153] According to a preferred embodiment of the construction method of the present invention, the construction method further comprises:
[0154] Connect the multiple near-memory computing modules into bus, star, ring, tree, mesh, and hybrid topological structures.
[0155] According to a preferred embodiment of the construction method of the present invention, the construction method further comprises:
[0156] Connect the multiple near-memory computing modules through a routing unit via metal wires. BRIEF DESCRIPTION OF THE DRAWINGS
[0157] The present invention will be more easily understood through the following description in conjunction with the drawings, wherein:
[0158] Figure 1 is a schematic diagram of a near-memory computing structure in the prior art.
[0159] Figure 2 is a schematic diagram of a near-memory computing module according to an embodiment of the present invention.
[0160] Figure 3 is a schematic diagram of a routing unit according to an embodiment of the present invention.
[0161] Figure 4 is a flowchart of a near-memory computing method according to an embodiment of the present invention.
[0162] Figure 5 is a flowchart of a near-memory computing method according to another embodiment of the present invention.
[0163] Figure 6 is a flowchart of a method for constructing a near-memory computing module according to an embodiment of the present invention.
[0164] Figure 7 is a schematic diagram of a near-memory computing network according to an embodiment of the present invention. Detailed Embodiments
[0165] The embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0166] Figure 2 It is a schematic diagram of the near-memory computing module 20 according to an embodiment of the present invention.
[0167] Figure 2 The near-memory computing module 20 shown in [reference] includes a computing sub-module (the "layer" or sub-module where the computing unit 201 and the routing unit 202 are located) and a memory sub-module (the "layer" or sub-module where the memory unit 203 is located), and the memory sub-module is arranged on one side of the computing sub-module (as Figure 2 shown in [reference], the memory sub-module is arranged below the computing sub-module).
[0168] However, the present invention is not limited to one memory sub-module and may also include more than one memory sub-module. In the case of including more than one memory sub-module, these memory sub-modules can be arranged on both sides of the computing sub-module (for example, referring to Figure 2 the case where the memory sub-module is arranged below the computing sub-module in [reference], in the case where the memory sub-modules are arranged on both sides of the computing sub-module, these memory sub-modules can also be arranged Figure 2 above the computing sub-module in [reference]).
[0169] Figure 2 The computing sub-module and the memory sub-module shown in [reference] are very close, located within the same chip and form a complete system. This complete system is also called the near-memory computing module, also known as the near-memory computing node.
[0170] As Figure 2 shown in [reference], the computing sub-module includes four computing units 201 and one routing unit 202, and the memory sub-module includes twelve memory units 203. Since the computing sub-module and the memory sub-module are located within the same chip, the computing unit 201 and the memory unit 203 are also integrated together, so that the latency of the computing unit 201 when accessing the memory unit 203 is very small.
[0171] The memory unit 203 is a unit for storing the operation data in the computing unit 201 and the data exchanged with external memories such as hard disks. Since the process of dynamic random access memory is relatively mature, in the present invention, the memory unit 203 is preferably a dynamic random access memory.
[0172] The number of memory units 203 is determined at least according to the total data bit width of the computing units 201 in the computing sub-module and the data bit width of a single memory unit 203.
[0173] For example, if there are four computing units 201 in the computing sub-module, the total data bit width of the four computing units 201 is 96 bits, and the data bit width of a single memory unit 203 is 8 bits, then the number of memory units 203 required is twelve (as Figure 2 shown).
[0174] In addition, the storage capacity of the memory unit 203 can also be customized according to requirements.
[0175] The computing unit 201 is the final execution unit for information processing and program operation, and preferably a central processing unit.
[0176] The routing unit 202 is connected to each computing unit 201 and each memory unit 203. In addition, the routing unit 202 will be connected to the routing unit 202 of at least another near-memory computing module. The main function of the routing unit 202 is to execute the access of a computing unit 201 of a near-memory computing module to another computing unit 202 of the near-memory computing module, or to a memory unit 203 of the near-memory computing module, or to the computing unit 202 or memory unit 203 of at least another near-memory computing module.
[0177] Each computing unit 201 can "directly" access the memory unit 203 corresponding to the computing unit via the routing unit 202.
[0178] For example, referring to Figure 2 , assume that the computing unit 201 in the upper left corner can "directly" access the three memory units 203 in the upper left corner via the routing unit 202, the computing unit 201 in the lower left corner can "directly" access the three memory units 203 in the lower left corner via the routing unit 202, the computing unit 201 in the upper right corner can "directly" access the three memory units 203 in the upper right corner via the routing unit 202, and the computing unit 201 in the lower right corner can "directly" access the three memory units 203 in the lower right corner via the routing unit 202. Here, "upper left corner", "lower left corner", "upper right corner", and "lower right corner" are the orientations when the reader observes from Figure 2 the horizontal direction of Figure 2 the horizontal direction of Figure 2 is longer than Figure 2 the vertical direction) of
[0179] and are the orientations relative to the routing unit 202, and this is an assumption made for illustrative purposes.
[0180] In addition, each computing unit 201 can "indirectly" access some other memory units 203 via the routing unit 202.
[0181] Referring again to Figure 2 and continuing with the previous assumption, if the computing unit 201 in the upper left corner issues a data access request for any one of the three memory units 203 in the upper right corner, the routing unit 202 can resolve the data access request and send the resolved data access request to the computing unit 201 in the upper right corner, and "indirectly" obtain the access data from any one of the three memory units 203 in the upper right corner via the computing unit 201 in the upper right corner and return the access data to the computing unit 201 in the upper left corner.
[0182] That is to say, the computing unit 201 can "directly" or "indirectly" access all the memory units 203 in a near-memory computing module via the routing unit 202. The specific structure of the routing unit 202 will be further described in detail later with respect to Figure 3 An embodiment of the distribution of multiple computing units 201 and the routing unit 202 in the computing sub-module is shown in
[0183] Figure 2 wherein multiple computing units 201 are arranged on both sides of the routing unit 202. However, this distribution is illustrative rather than restrictive.
[0184] The multiple computing units 201 surrounding the routing unit 202, or the multiple computing units 201 distributed on one side of the routing unit 202, etc. also fall within the scope of the present invention.
[0185] In Figure 2 the computing sub-module and the memory sub-module need to be connected through a data interface by means of the routing unit 202, and preferably a three-dimensional connection 206 process is used for the connection.
[0186] Common three-dimensional connection 206 processes include bonding methods, through-silicon vias (TSV), flip chips, and wafer-level packaging. In the present invention, the three-dimensional connection 206 process preferably uses a bonding method.
[0187] Bonding is a common three-dimensional connection process and is a wafer stacking process within a chip. Specifically, bonding is to connect wafers together by a certain process using metal wires to achieve the required electrical characteristics.
[0188] In addition, in order to relieve the data transmission pressure, the total data bit width of the connection between each memory cell 203 in the memory sub-module and the routing unit 202 in the computing sub-module should be n times the data bit width of a single computing unit 201 in the computing sub-module, where n is a positive integer.
[0189] In Figure 2 the access between the computing unit 201 in the computing sub-module and the memory cell 203 in the memory sub-module needs to pass through the connection between the routing units 202. If the total data bit width of the connection is the same as that of a single computing unit 201, when the data access operation between the computing unit 201 in the computing sub-module and the memory cell 203 in the memory sub-module is frequent (i.e., the data throughput rate is high), the data throughput rate between the computing unit 201 in the computing sub-module and the memory cell 203 in the memory sub-module will also increase accordingly, and congestion may occur in the connection. Therefore, the total data bit width of the connection is set to a positive integer multiple of the data bit width of a single computing unit 201.
[0190] It should be understood that the specific value of the positive integer n is set according to the service requirements. For example, in general system design, the bandwidth requirements for data transmission between different computing sub-modules in the chip can be obtained through service simulation, and the required data bit width can be obtained based on the bandwidth requirements.
[0191] Suppose the data bandwidth requirement between the computing unit 201 in the computing sub-module and the memory cell 203 in the memory sub-module is 144 Gb / s, and the existing total data bit width of the connection is 72 b and the clock is 1 GHz, then the total data bandwidth of the connection will be 72 Gb / s. At this time, it is necessary to consider increasing the total data bit width of the connection to 144 b to adapt to the data bandwidth requirement.
[0192] Figure 3 is a schematic diagram of the routing unit 202 according to an embodiment of the present invention.
[0193] Figure 3 The external interfaces of the routing unit 202 are shown, and these external interfaces mainly include a switching interface, a routing interface, and a memory control interface. It should be understood that in order not to obscure the main idea of the present invention, Figure 3 only the external interfaces involved in the present invention are shown in
[0194] In Figure 3 the switching interface is shown as a Memory Front Bus (MFB) interface 301, and these memory front bus interfaces 301 are respectively connected to the computing units 201 ( Figure 2 shown in
[0195] The routing interface is shown as a Memory Front Routing (MFR) interface 302, and these Memory Front Routing interfaces 302 are connected to the routing interfaces of the routing units of at least another near-memory computing module.
[0196] The memory control interface is shown as a DDR operation interface (DDRIO-bonding), and this DDR operation interface is connected to the memory unit 203.
[0197] Figure 3 The internal structure of the routing unit 202 is additionally shown, and these internal structures mainly include a switching routing calculation unit 304, a crossbar unit 305, and a memory control unit 306. It should be understood that in order not to obscure the gist of the present invention, Figure 3 only the internal structures involved in the present invention are shown herein, and these internal structures are exemplary rather than restrictive. It should be understood that the routing unit 202 should also include some buffer circuits, digital-analog circuits, etc. For example, the buffer circuit buffers and prioritizes the data access requests of multiple computing units 201, and the digital-analog circuit can cooperate with the memory control unit 306 for operation, etc.
[0198] In the present invention, the switching routing calculation unit 304 stores routing information about the computing unit 201. This routing information can be stored in the form of a routing table, for example. Thus, the switching routing calculation unit 304 can determine information such as whether the computing unit 201 in the data access request can "directly" access a memory unit via the router.
[0199] In addition, the switching routing calculation unit 304 also stores routing information about the near-memory computing module. This routing information can also be stored in the form of a routing table, for example. Thus, the switching routing calculation unit 304 can determine in which near-memory computing module the memory address or the target computing unit address is (whether it is in the present near-memory computing module or in another near-memory computing module), etc. For example, based on a certain specific position information (such as the first bit) in the target memory address or the target computing unit address and this routing information, the switching routing calculation unit 304 can judge in which near-memory computing module the target memory address or the target computing unit address is. For example, if the first bit of the target memory address or the target computing unit address is 1, it means that the target memory or the target computing unit is in the first near-memory computing module; if the first bit of the target memory address or the target computing unit address is 3, it means that the target memory or the target computing unit is in the third near-memory computing module.
[0200] In addition, the switching routing calculation unit 304 can receive data access requests from the switching interface and can also receive data access requests from the routing interface. In the case where the computing sub-module 201 of the near-memory computing module includes a single routing unit (such asFigure 2 In the case shown in FIG. (not shown), the switching routing calculation unit 304 receives, via the switching interface, a data access request issued by the calculation unit of the calculation sub-module 201 of the present near-memory calculation module, and receives, via the routing interface, a data access request issued by the calculation unit of another near-memory calculation module.
[0201] The memory control unit 306 stores routing information regarding the memory unit 203. Thus, the memory control unit 306 can determine port information corresponding to the memory unit in the data access request, and so on.
[0202] Figure 4 is a flowchart of a near-memory calculation method according to an embodiment of the present invention.
[0203] The near-memory calculation method includes the following steps:
[0204] Step S401: A routing unit in the calculation sub-module receives a data access request, which is issued by a first calculation unit in the calculation sub-module and includes at least the address of a target memory unit.
[0205] Step S402: The routing unit receives the data access request through the switching interface, parses the data access request, and the routing unit determines whether the target memory unit and the first calculation unit are located in the same near-memory calculation module.
[0206] If the target memory unit and the first calculation unit are not located in the same near-memory calculation module, then step S403 is executed: The routing unit forwards the parsed data access request to the routing unit of another near-memory calculation module connected to the routing unit and to a second calculation unit connected to the routing unit of the other near-memory calculation module, and obtains access data from the target memory unit via the second calculation unit and forwards the access data to the first calculation unit.
[0207] If the target memory unit and the first calculation unit are located in the same near-memory calculation module, then step S404 is executed: The routing unit determines whether the first calculation unit can directly access the target memory unit.
[0208] If the first calculation unit can directly access the target memory unit, then step S405 is executed: The routing unit directly obtains access data from the target memory unit and forwards the access data to the first calculation unit.
[0209] If the first computing unit cannot directly access the target memory unit, step S406 is executed: the routing unit directly forwards the parsed data access request to the second computing unit, and obtains the access data from the target memory unit via the second computing unit and forwards the access data to the first computing unit.
[0210] Figure 4 The method flowchart in
[0211] Figure 5 is only schematic and does not have to be executed in this order. For example, it can first be determined whether the first computing unit can directly access the target memory unit, and then it can be determined whether the target memory unit and the first computing unit are in the same near-memory computing module.
[0212] The near-memory computing method includes the following steps:
[0213] Step S501: The routing unit in the computing sub-module receives a data access request, which is issued by the first computing unit in the computing sub-module and includes at least the address of the target computing unit.
[0214] Step S502: The routing unit receives the data access request through the switching interface, parses the data access request, and determines whether the target computing unit and the first computing unit are in the same near-memory computing module.
[0215] If the target computing unit and the first computing unit are in the same near-memory computing module, step S503 is executed: the routing unit directly obtains the access data from the target computing unit and forwards the access data to the first computing unit.
[0216] If the target computing unit and the first computing unit are not in the same near-memory computing module, step S504 is executed: the routing unit forwards the parsed data access request to the routing unit of another near-memory computing module connected to the routing unit, and obtains the access data from the target computing unit via the routing unit of the other near-memory computing module and forwards the access data to the first computing unit.
[0217] The following combines Figure 2 and Figure 3 , and further understands the external interface and internal structure of the routing unit 202 and Figure 4 and Figure 5 in the near-memory computing method for the following five cases of data flow processing in the near-memory computing module.
[0218] Refer again to Figure 2, it is still assumed that the computing unit 201 in the upper left corner can "directly" access the three memory units 203 in the upper left corner via the routing unit 202, the computing unit 201 in the lower left corner can "directly" access the three memory units 203 in the lower left corner via the routing unit 202, the computing unit 201 in the upper right corner can "directly" access the three memory units 203 in the upper right corner via the routing unit 202, and the computing unit 201 in the lower right corner can "directly" access the three memory units 203 in the lower right corner via the routing unit 202.
[0219] In addition, it is assumed that the computing unit 201 in the upper left corner is connected to the routing unit 202 through the switching interface MFB0, and the computing unit 201 in the lower left corner is connected to the routing unit 202 through the switching interface MFB1.
[0220] Case (i): The computing unit 201 in the upper left corner accesses any one of the three memory units 203 in the upper left corner
[0221] The computing unit 201 in the upper left corner sends a data access request to the switching routing computing unit 304 through the switching interface MFB0. The switching routing computing unit 304 parses the data access request to obtain the target memory address, and determines whether the target memory address is within the same near-memory computing module (but does not parse which specific memory unit it is) (in this case, it is determined to be within the same near-memory computing module), and determines whether the computing unit 201 in the upper left corner can "directly" access the three memory units 203 in the upper left corner (in this case, it is determined that it can "directly" access).
[0222] After that, the switching routing computing unit 304 queries the target memory address in the routing information it stores about the computing unit and the near-memory computing module, determines the port information corresponding to the target memory address, and then controls the opening and closing of the crossbar unit 305, so that the parsed data access request is sent to the memory control unit 306 for secondary parsing to obtain which specific memory unit needs to be accessed (any one of the three memory units 203 in the upper left corner), and then accesses any one of the three memory units 203 in the upper left corner through the memory control (DDRIO-bonding) interface.
[0223] Case (ii): The computing unit 201 in the upper left corner accesses any one of the three memory units 203 in the lower left corner
[0224] The computing unit 201 in the upper left corner sends a data access request to the switching routing computing unit 304 through the switching interface MFB0. The switching routing computing unit 304 parses the data access request to obtain the target memory address, and determines whether the target memory address is within the same near-memory computing module (but does not parse out which specific memory unit it is) (in this case, it is determined that it is within the same near-memory computing module), and determines whether the computing unit 201 in the upper left corner can "directly" access the three memory units 203 in the upper left corner (in this case, it is determined that it cannot "directly" access).
[0225] After that, the switching routing computing unit 304 queries the target memory address in the routing information it stores about the computing unit and the near-memory computing module, determines the port information corresponding to the target memory address, and then controls the opening and closing of the crossbar switch unit 305, so that the parsed data access request is sent to the computing unit 201 in the lower left corner through the switching interface MFB1. The computing unit 201 in the lower left corner performs the operations in case (i), and then accesses any one of the three memory units 203 in the upper left corner through the memory control (DDRIO-bonding) interface.
[0226] Case (iii): The computing unit 201 in the upper left corner accesses the memory unit 203 of another near-memory computing module
[0227] The computing unit 201 in the upper left corner sends a data access request to the switching routing computing unit 304 through the switching interface MFB0. The switching routing computing unit 304 parses the data access request to obtain the target memory address, and determines whether the target memory address is within the same near-memory computing module (but does not parse out which specific memory unit it is) (in this case, it is determined that it is not within the same near-memory computing module).
[0228] After that, the switching routing computing unit 304 controls the opening and closing of the crossbar switch unit 305, and sends the parsed data access request to another near-memory computing module through the routing interface MFR.
[0229] Another near-memory computing module performs the operations in cases (i) and (ii) above, and then accesses the memory unit 203 of another near-memory computing module through the memory control (DDRIO-bonding) interface of the routing unit of another near-memory computing module.
[0230] Case (iv): The computing unit 201 in the upper left corner accesses the computing unit 201 in the lower left corner
[0231] The computing unit 201 in the upper left corner sends a data access request to the switching routing computing unit 304 through the switching interface MFB0. The switching routing computing unit 304 parses the data access request to obtain the target computing unit address, and determines whether the target computing unit address is within the same near-memory computing module (in this case, it is determined that it is within the same near-memory computing module).
[0232] After that, the switching routing computing unit 304 queries the target computing unit address in the routing information it stores about the computing unit and the near-memory computing module, determines the port information corresponding to the target computing unit address, and then controls the opening and closing of the crossbar switch unit 305 to access the computing unit 201 in the lower left corner through the switching interface MFB1.
[0233] Case (5): The computing unit 201 in the upper left corner accesses the computing unit 201 of another near-memory computing module
[0234] The computing unit 201 in the upper left corner sends a data access request to the switching routing computing unit 304 through the switching interface MFB0. The switching routing computing unit 304 parses the data access request to obtain the target computing unit address, and determines whether the target computing unit address is within the same near-memory computing module (in this case, it is determined that it is not within the same near-memory computing module).
[0235] After that, the switching routing computing unit 304 controls the opening and closing of the crossbar switch unit 305, and sends the parsed data access request to another near-memory computing module through the routing interface MFR.
[0236] Another near-memory computing module performs the operations as in Case (4) above, and then realizes the access to the computing unit 201 of another near-memory computing module through the switching interface 301 of the routing unit of another near-memory computing module.
[0237] Figure 6 It is a flowchart of a method for constructing a near-memory computing module according to an embodiment of the present invention.
[0238] The method for constructing the near-memory computing module includes the following steps:
[0239] Step S601: Arrange at least one memory sub-module on one side or both sides of the computing sub-module, where each memory sub-module includes a plurality of memory cells 203, and the computing sub-module includes a plurality of computing units 201.
[0240] Step S602: Connect each memory sub-module to the computing sub-module.
[0241] Step S603: Arrange the computing sub-module and the at least one memory sub-module within the same chip.
[0242] Figure 7 It is a schematic diagram of a near-memory computing network according to an embodiment of the present invention.
[0243] Figure 7 The near-memory computing network shown in [the figure] includes: a plurality of near-memory computing modules 70, which are the plurality of near-memory computing modules as described above, and the plurality of near-memory computing modules 70 are connected through a routing unit.
[0244] By interconnecting and topologizing the near-memory computing modules, high data bandwidth and high-performance computing requirements can be formed.
[0245] Figure 7 What is shown in [the figure] is a typical mesh topology. It should be understood that the plurality of near-memory computing modules can also be connected into a bus type, star type, ring type, tree type, and hybrid topology.
[0246] In the present invention, the plurality of near-memory computing modules are connected through a routing unit via a metal wire connection 701. The metal wire connection here is the metal wire connection conventionally used for two-dimensional connections.
[0247] In the present invention, the embodiment is that the computing sub-module includes a single routing unit 202. However, the present invention is not limited to one routing unit and may also include more than one routing unit. In the case of including more than one routing unit, these routing units can be connected through a routing interface MFR (similar to the operation between the routing interface MFR of one near-memory computing module and the routing interface MFR of another near-memory computing module), constituting an overall routing unit.
[0248] The overall routing unit presents the same function to the outside as Figure 2 the single routing unit 202 shown in [the figure]. Different from Figure 2 the single routing unit 202 in [the figure], each routing unit in the overall routing unit is not connected to each memory unit in the memory sub-module. In addition, since the computing sub-module includes more than one routing unit, the switching routing calculation unit 304 will receive the data access requests sent by the computing units of the computing sub-modules in this near-memory computing module via a switching interface or a routing interface, and receive the data access requests sent by the computing units of another near-memory computing module via a routing interface.
[0249] For example, if there are two routing units in the computing sub-module, these two routing units constitute an overall routing unit. For example, assume Figure 2There are two routing units in [the system]: routing unit A and routing unit B. Routing unit A is connected to the computing unit 201 in the upper left corner, the computing unit 201 in the lower left corner, as well as three memory units 203 in the upper left corner and three memory units 203 in the lower left corner. Routing unit B is connected to the computing unit 201 in the upper right corner, the computing unit 201 in the lower right corner, as well as three memory units 203 in the upper right corner and three memory units 203 in the lower right corner.
[0250] If any one of the three memory units 203 in the lower right corner needs to be accessed by the computing unit 201 in the upper left corner connected to routing unit A, it needs to be connected through the routing interface MFR between routing unit A and routing unit B for access, and the operation is similar to the above case (iii). Specifically as follows:
[0251] In router A, the computing unit 201 in the upper left corner sends a data access request to the switching routing computing unit 304 through the switching interface MFB0. The switching routing computing unit 304 parses the data access request to obtain the target memory address and determines whether the target memory address is within the addressing range of router A (in this case, it is determined that it is not within the addressing range of router A). After that, the switching routing computing unit 304 controls the opening and closing of the crossbar switch unit 305, and sends the parsed data access request to router B through the routing interface MFR to access any one of the three memory units 203 in the lower right corner.
[0252] It should be noted that the above-described embodiments illustrate rather than limit the present invention, and those skilled in the art will be able to design many alternative embodiments without departing from the scope of the appended claims. It should be understood that the scope of the present invention is defined by the claims.
Claims
1. A near-memory computing module, characterized in that the near-memory computing module includes: a computing sub-module, the computing sub-module includes a plurality of computing units; at least one memory sub-module, the at least one memory sub-module is located on a different layer from the computing sub-module and is arranged on at least one side of the computing sub-module, wherein each memory sub-module includes a plurality of memory units and each memory sub-module is connected to the computing sub-module; wherein the computing sub-module and the at least one memory sub-module are located within the same chip; wherein the computing sub-module further includes: a routing unit; wherein the routing unit is connected to each computing unit, and the routing unit is connected to each memory unit of each memory sub-module, and the routing unit is connected to the routing unit of at least another near-memory computing module; wherein the routing unit is configured to perform access by a first computing unit of the near-memory computing module to a second computing unit of the near-memory computing module, or access to a first memory unit of the near-memory computing module, or access to a third computing unit or a second memory unit of at least another near-memory computing module.
2. The near-memory computing module according to claim 1, wherein The routing unit includes: a plurality of switching interfaces, the plurality of switching interfaces connect the routing unit to each computing unit; a routing interface, the routing interface connects the routing unit to the routing interface of the routing unit of at least another near-memory computing module; a memory control interface, the memory control interface connects the routing unit to each memory unit in each memory sub-module.
3. The near-memory computing module according to claim 2, wherein The routing unit further includes: a crossbar switch unit; a switching routing calculation unit, the switching routing calculation unit is connected to the plurality of switching interfaces, the routing interface, and the crossbar switch unit, and the switching routing calculation unit stores at least routing information about the near-memory computing module and the plurality of computing units, and the switching routing calculation unit parses the received data access request and controls the switching of the crossbar switch unit based on the parsed data access request information; a memory control unit, the memory control unit is connected to the crossbar switch unit and the memory control interface, and the memory control unit stores at least routing information about the plurality of memory units, and the memory control unit, in response to the crossbar switch unit switching to the memory control unit, performs secondary parsing on the parsed data access request received from the crossbar switch unit to determine the target memory unit, and accesses the target memory unit via the memory control interface.
4. The near-memory computing module according to any one of claims 1-3, characterized in that each computing unit directly accesses at least one memory unit via the routing unit.
5. The near-memory computing module according to claim 4, characterized in that each computing unit indirectly accesses at least one other memory unit via the routing unit.
6. The near-memory computing module according to any one of claims 1-3, characterized in that the routing unit is connected to each memory unit of each memory sub-module by a bonding method.
7. The near-memory computing module according to any one of claims 1-3, characterized in that The total data bit width of the connections between the routing unit and each memory cell of each memory sub-module is n times the data bit width of a single computing unit, where n is a positive integer.
8. The in-memory computing module according to any one of claims 1-3, characterized in that In the computing sub-module, the routing unit is located at the center, and the multiple computing units are distributed around the routing unit.
9. The in-memory computing module according to any one of claims 1-3, characterized in that The memory cell includes a dynamic random access memory, and the computing unit includes a central processing unit.
10. The in-memory computing module according to any one of claims 1-3, characterized in that The number of memory cells in the memory sub-module is determined at least according to the total data bit width of the multiple computing units in the computing sub-module and the data bit width of a single memory cell.
11. The in-memory computing module according to any one of claims 1-3, characterized in that The storage capacity of the memory cell is customizable.
12. An in-memory computing module, characterized in that The in-memory computing module includes: A computing sub-module, which includes multiple computing units; At least one memory sub-module, which is located on a different layer from the computing sub-module and is arranged on at least one side of the computing sub-module, where each memory sub-module includes multiple memory cells and each memory sub-module is connected to the computing sub-module; Wherein the computing sub-module and the at least one memory sub-module are located within the same chip; Wherein the computing sub-module further includes: at least two routing units, each routing unit is connected to at least one computing unit, and each routing unit is connected to at least one memory cell of each memory sub-module; The at least two routing units are connected to each other to form an overall routing unit, the overall routing unit is connected to each computing unit, and the overall routing unit is connected to each memory cell of each memory sub-module, and the overall routing unit is connected to at least another routing unit of at least another in-memory computing module; Wherein the overall routing unit is configured to perform access by a first computing unit of the in-memory computing module to a second computing unit of the in-memory computing module, or to a first memory unit of the in-memory computing module, or to a third computing unit or a second memory unit of at least another in-memory computing module.
13. The near-memory computing module according to claim 12, wherein Each routing unit of the at least two routing units includes: Multiple switching interfaces, which connect the routing unit to at least one computing unit; A routing interface, which connects the routing unit to at least another routing unit of the in-memory computing module, and / or to at least another routing unit of at least another in-memory computing module; A memory control interface, which connects the routing unit to at least one memory unit.
14. The near-memory computing module according to claim 13, wherein The at least two routing units are connected to each other through the routing interface.
15. The near-memory computing module according to claim 13, wherein Each routing unit of the at least two routing units further includes: A crossbar switch unit; A switching and routing calculation unit, which is connected to the multiple switching interfaces, the routing interface, and the crossbar unit, and the switching and routing calculation unit stores at least routing information about the near-memory calculation module and the multiple calculation units. The switching and routing calculation unit parses the received data access requests and controls the switching of the crossbar unit based on the parsed data access request information; A memory control unit, which is connected to the crossbar unit and the memory control interface, and the memory control unit stores at least routing information about the multiple memory units. In response to the crossbar unit switching to the memory control unit, the memory control unit performs secondary parsing on the parsed data access requests received from the crossbar unit to determine the memory unit to be accessed, and accesses the memory unit to be accessed via the memory control interface.
16. The near-memory calculation module according to any one of claims 12-15, characterized in that Each calculation unit directly accesses at least one memory unit via the global routing unit.
17. The near-memory calculation module according to claim 16, characterized in that Each calculation unit indirectly accesses at least one other memory unit via the global routing unit.
18. The near-memory calculation module according to any one of claims 12-15, characterized in that The global routing unit is connected to each memory unit of each memory sub-module by bonding.
19. The near-memory calculation module according to any one of claims 12-15, characterized in that The total data bit width between each memory unit and the global routing unit is n times the data bit width of a single calculation unit, where n is a positive integer.
20. The near-memory calculation module according to any one of claims 12-15, characterized in that In the calculation sub-module, the at least two routing units are located at the center, and the multiple calculation units are distributed around the at least two routing units.
21. The near-memory calculation module according to any one of claims 12-15, characterized in that In the calculation sub-module, the multiple calculation units are located at the center, and the at least two routing units are distributed around the multiple calculation units.
22. The near-memory calculation module according to any one of claims 12-15, characterized in that The memory unit includes a dynamic random access memory, and the calculation unit includes a central processing unit.
23. The near-memory calculation module according to any one of claims 12-15, characterized in that The number of memory units in the memory sub-module is determined at least according to the total data bit width of the multiple calculation units in the calculation sub-module and the data bit width of a single memory unit.
24. The near-memory calculation module according to any one of claims 12-15, characterized in that The storage capacity of the memory unit is customizable.
25. A near-memory calculation method, which is used for the near-memory calculation module according to claim 5, characterized in that The near-memory calculation method includes: The routing unit receives a data access request issued by a first computing unit and including at least the address of a target memory unit; and, The routing unit parses the data access request, obtains access data from the target memory unit and forwards the access data to the first computing unit; Wherein the near-memory computing method further includes: After the routing unit parses the data access request and before obtaining access data from the target memory unit, the routing unit connected to the first computing unit determines whether the first computing unit can directly access the target memory unit via the routing unit connected to the first computing unit; When the first computing unit can directly access the target memory unit via the routing unit connected to the first computing unit, the routing unit connected to the first computing unit directly obtains access data from the target memory unit and forwards the access data to the first computing unit; and When the first computing unit cannot directly access the target memory unit via the routing unit connected to the first computing unit, the routing unit connected to the first computing unit forwards the parsed data access request to a second computing unit, and obtains access data from the target memory unit via the second computing unit and forwards the access data to the first computing unit.
26. The near-memory computing method according to claim 25, wherein, The near-memory computing method further includes: When the first computing unit cannot directly access the target memory unit via the routing unit connected to the first computing unit and before the routing unit connected to the first computing unit forwards the parsed data access request to the second computing unit, the routing unit connected to the first computing unit determines whether the target memory unit and the first computing unit are located in the same near-memory computing module; When the target memory unit and the first computing unit are located in the same near-memory computing module, the routing unit connected to the first computing unit directly forwards the parsed data access request to the second computing unit, and obtains access data from the target memory unit via the second computing unit and forwards the access data to the first computing unit, wherein the second computing unit can directly access the target memory unit via the routing unit connected to the first computing unit; When the target memory unit and the first computing unit are not located in the same near-memory computing module, the routing unit connected to the first computing unit forwards the parsed data access request to the routing unit of at least another near-memory computing module connected to the routing unit connected to the first computing unit and forwards it to the second computing unit connected to the routing unit of at least another near-memory computing module, and obtains access data from the target memory unit via the second computing unit and forwards the access data to the first computing unit, wherein the second computing unit can directly access the target memory unit via the routing unit of at least another near-memory computing module.
27. A near-memory computing method, which is used for the near-memory computing module described in any one of claims 1-11, wherein, The near-memory computing method includes: A routing unit receives a data access request, which is sent by a first computing unit and includes at least the address of a target computing unit; and, The routing unit parses the data access request, obtains access data from the target computing unit, and forwards the access data to the first computing unit; Wherein the near-memory computing method further includes: After the routing unit parses the data access request and before obtaining access data from the target computing unit, the routing unit connected to the first computing unit determines whether the target computing unit and the first computing unit are located in the same near-memory computing module; When the target computing unit and the first computing unit are located in the same near-memory computing module, the routing unit connected to the first computing unit directly obtains access data from the target computing unit and forwards the access data to the first computing unit; When the target computing unit and the first computing unit are not located in the same near-memory computing module, the routing unit connected to the first computing unit forwards the parsed data access request to the routing unit of at least another near-memory computing module connected to the routing unit, and obtains access data from the target computing unit via the routing unit of the at least another near-memory computing module and forwards the access data to the first computing unit.
28. A near-memory computing method, which is used for the near-memory computing module described in claim 17, and is characterized in that The near-memory computing method includes: An overall routing unit receives a data access request, which is sent by a first computing unit and includes at least the address of a target memory unit; and, The overall routing unit parses the data access request, obtains access data from the target memory unit, and forwards the access data to the first computing unit; Wherein the near-memory computing method further includes: After the overall routing unit parses the data access request and before obtaining access data from the target memory unit, the overall routing unit connected to the first computing unit determines whether the first computing unit can directly access the target memory unit via the overall routing unit connected to the first computing unit; When the first computing unit can directly access the target memory unit via the overall routing unit connected to the first computing unit, the overall routing unit connected to the first computing unit directly obtains access data from the target memory unit and forwards the access data to the first computing unit; and When the first computing unit cannot directly access the target memory unit via the overall routing unit connected to the first computing unit, the overall routing unit connected to the first computing unit forwards the parsed data access request to a second computing unit, and obtains access data from the target memory unit via the second computing unit and forwards the access data to the first computing unit.
29. The near-memory computing method according to claim 28, wherein The near-memory computing method further includes: When the first computing unit cannot directly access the target memory unit via the global routing unit connected to the first computing unit and before the global routing unit connected to the first computing unit forwards the resolved data access request to the second computing unit, the global routing unit connected to the first computing unit determines whether the target memory unit and the first computing unit are located in the same near-memory computing module; When the target memory unit and the first computing unit are located in the same near-memory computing module, the global routing unit connected to the first computing unit directly forwards the resolved data access request to the second computing unit, and obtains the access data from the target memory unit via the second computing unit and forwards the access data to the first computing unit, where the second computing unit can directly access the target memory unit via the global routing unit connected to the first computing unit; When the target memory unit and the first computing unit are not located in the same near-memory computing module, the global routing unit connected to the first computing unit forwards the resolved data access request to the global routing unit of at least another near-memory computing module connected to the global routing unit connected to the first computing unit and forwards it to the second computing unit connected to the global routing unit of the at least another near-memory computing module, and obtains the access data from the target memory unit via the second computing unit and forwards the access data to the first computing unit, where the second computing unit can directly access the target memory unit via the global routing unit of the at least another near-memory computing module.
30. A near-memory computing method, which is used for the near-memory computing module described in any one of claims 12-24, and is characterized in that, The near-memory computing method includes: The global routing unit receives a data access request, which is sent by a first computing unit and at least includes the address of a target computing unit; and, The global routing unit resolves the data access request, obtains the access data from the target computing unit and forwards the access data to the first computing unit; Wherein the near-memory computing method further includes: After the global routing unit resolves the data access request and before obtaining the access data from the target computing unit, the global routing unit connected to the first computing unit determines whether the target computing unit and the first computing unit are located in the same near-memory computing module; When the target computing unit and the first computing unit are located in the same near-memory computing module, the global routing unit connected to the first computing unit directly obtains the access data from the target computing unit and forwards the access data to the first computing unit; When the target computing unit and the first computing unit are not located in the same near-memory computing module, the global routing unit connected to the first computing unit forwards the resolved data access request to the global routing unit of at least another near-memory computing module connected to the global routing unit, and obtains the access data from the target computing unit via the global routing unit of the at least another near-memory computing module and forwards the access data to the first computing unit.
31. A near-memory computing network, characterized in that, The near-memory computing network includes: Multiple near-memory computing modules, where the multiple near-memory computing modules are multiple near-memory computing modules according to any one of claims 1-24, and the multiple near-memory computing modules are connected through a routing unit.
32. The near-memory computing network according to claim 31, wherein The multiple near-memory computing modules are connected into a bus topology, star topology, ring topology, tree topology, mesh topology, and hybrid topology.
33. The near-memory computing network according to claim 31 or 32, wherein The multiple near-memory computing modules are connected through a routing unit via metal wires.
34. A method for constructing a near-memory computing module, wherein The construction method includes: Arranging at least one memory sub-module in a different layer from the computing sub-module and on at least one side of the computing sub-module, where each memory sub-module includes multiple memory cells, and the computing sub-module includes multiple computing cells; Connecting each memory sub-module to the computing sub-module; Wherein the computing sub-module and the at least one memory sub-module are arranged in the same chip; Wherein the computing sub-module further includes: a routing unit; The construction method further includes: Connecting the routing unit to each computing cell, connecting the routing unit to each memory cell of each memory sub-module, connecting the routing unit to the routing unit of at least another near-memory computing module, and configuring the routing unit to execute the access of the first computing cell of the near-memory computing module to the second computing cell of the near-memory computing module, or to the first memory cell of the near-memory computing module, or to the third computing cell or the second memory cell of at least another near-memory computing module.
35. The construction method according to claim 34, wherein The construction method further includes: Connecting multiple switching interfaces of the routing unit to each computing cell; Connecting the routing interface of the routing unit to the routing interface of the routing unit of at least another near-memory computing module; Connecting the memory control interface of the routing unit to each memory cell in each memory sub-module.
36. The construction method according to claim 35, wherein The construction method further includes: Connecting the switching routing calculation unit of the routing unit to the multiple switching interfaces, the routing interface, and a crossbar unit, storing at least the routing information about the near-memory computing module and the multiple computing cells in the switching routing calculation unit, and configuring the switching routing calculation unit to parse the received data access request and control the switching of the crossbar unit based on the parsed data access request information. Connect the memory control unit of the routing unit to the crossbar unit and the memory control interface, store at least the routing information of the multiple memory units in the memory control unit, and configure the memory control unit to, in response to the crossbar unit switching to the memory control unit, perform secondary parsing on the parsed data access request received from the crossbar unit to determine the target memory unit, and access the target memory unit via the memory control interface.
37. The construction method according to any one of claims 34 - 36, characterized in that The construction method further includes: Configure each computing unit to directly access at least one memory unit via the routing unit.
38. The construction method according to claim 37, characterized in that The construction method further includes: Configure each computing unit to indirectly access at least one other memory unit via the routing unit.
39. The construction method according to any one of claims 34 - 36, characterized in that The construction method further includes: Connect the routing unit to each memory unit of each memory sub-module by bonding.
40. The construction method according to any one of claims 34 - 36, characterized in that The construction method further includes: Set the total data bit width of the connection between the routing unit and each memory unit of each memory sub-module to n times the data bit width of a single computing unit, where n is a positive integer.
41. The construction method according to any one of claims 34 - 36, characterized in that The construction method further includes: In the computing sub-module, arrange the routing unit at the center and arrange the multiple computing units to be distributed around the routing unit.
42. The construction method according to any one of claims 34 - 36, characterized in that The construction method further includes: Set the memory unit to include a dynamic random access memory, and set the computing unit to include a central processing unit.
43. The construction method according to any one of claims 34 - 36, characterized in that The construction method further includes: Determine at least the number of memory units in the memory sub-module according to the total data bit width of the multiple computing units in the computing sub-module and the data bit width of a single memory unit.
44. The construction method according to any one of claims 34 - 36, characterized in that Set the storage capacity of the memory unit to be customizable.
45. A construction method of a near-memory computing module, characterized in that The construction method includes: Arrange at least one memory sub-module in a layer different from the computing sub-module and on at least one side of the computing sub-module, where each memory sub-module includes multiple memory units, and the computing sub-module includes multiple computing units; Connect each memory sub-module to the computing sub-module; Wherein arrange the computing sub-module and the at least one memory sub-module in the same chip; The computing sub-module further includes: at least two routing units, each of which is connected to at least one computing unit, and each of which is connected to at least one memory unit of each memory sub-module; The construction method further includes: Connecting the at least two routing units to each other to form an overall routing unit, connecting the overall routing unit to each computing unit, connecting the overall routing unit to each memory unit of each memory sub-module, and connecting the overall routing unit to at least another routing unit of at least another near-memory computing module; Configuring the overall routing unit to perform access by a first computing unit of the near-memory computing module to a second computing unit of the near-memory computing module, or access to a first memory unit of the near-memory computing module, or access to a third computing unit or a second memory unit of another near-memory computing module.
46. The construction method according to claim 45, wherein: The construction method further includes: Connecting a plurality of switching interfaces of each of the at least two routing units to at least one computing unit; Connecting the routing interface of each of the at least two routing units to the routing interface of at least another routing unit of the near-memory computing module, and / or connecting the routing interface of each of the at least two routing units to the routing interface of at least another routing unit of at least another near-memory computing module; Connecting the memory control interface of each of the at least two routing units to at least one memory unit.
47. The construction method according to claim 46, wherein: The construction method further includes: Connecting the at least two routing units to each other through routing interfaces.
48. The construction method according to claim 47, wherein: The construction method further includes: Connecting the switching routing calculation unit of each of the at least two routing units to the plurality of switching interfaces, the routing interface, and a crossbar unit, storing at least the routing information about the near-memory computing module and the plurality of computing units in the switching routing calculation unit, and configuring the switching routing calculation unit to parse the received data access request and control the switching of the crossbar unit based on the parsed data access request information; Connecting the memory control unit of each of the at least two routing units to the crossbar unit and the memory control interface, storing at least the routing information of the plurality of memory units in the memory control unit, and configuring the memory control unit to, in response to the crossbar unit switching to the memory control unit, perform secondary parsing on the parsed data access request received from the crossbar unit to determine the memory unit to be accessed, and access the memory unit to be accessed through the memory control interface.
49. The construction method according to any one of claims 45-48, wherein: The construction method further includes: Configuring each computing unit to directly access at least one memory unit via the overall routing unit.
50. The construction method according to claim 49, wherein: The construction method further includes: Configuring each computing unit to indirectly access at least one other memory unit via an overall routing unit.
51. The construction method according to any one of claims 45 - 48, wherein: The construction method further includes: Connecting the overall routing unit to each memory unit of each memory sub-module by bonding.
52. The construction method according to any one of claims 45 - 48, wherein: The construction method further includes: Setting the total data bit width between each memory unit and the overall routing unit to n times the data bit width of a single computing unit, where n is a positive integer.
53. The construction method according to any one of claims 45 - 48, wherein: The construction method further includes: In the computing sub-module, arranging the at least two routing units at the center and arranging the multiple computing units to be distributed around the at least two routing units.
54. The construction method according to any one of claims 45 - 48, wherein: The construction method further includes: In the computing sub-module, arranging the multiple computing units at the center and arranging the at least two routing units to be distributed around the multiple computing units.
55. The construction method according to any one of claims 45 - 48, wherein: The construction method further includes: Setting the memory unit to include a dynamic random access memory and setting the computing unit to include a central processing unit.
56. The construction method according to any one of claims 45 - 48, wherein: The construction method further includes: Determining the number of memory units in the memory sub-module at least according to the total data bit width of the multiple computing units in the computing sub-module and the data bit width of a single memory unit.
57. The construction method according to any one of claims 45 - 48, wherein: Setting the storage capacity of the memory unit to be customizable.
58. A construction method of a near-memory computing network, wherein: The construction method includes: Connecting multiple near-memory computing modules through a routing unit, wherein the multiple near-memory computing modules are multiple near-memory computing modules according to any one of claims 1 - 24.
59. The construction method according to claim 58, wherein: The construction method further includes: Connecting the multiple near-memory computing modules into bus, star, ring, tree, mesh, and hybrid topologies.
60. The construction method according to claim 58 or 59, wherein: The construction method further includes: Connecting the multiple near-memory computing modules via metal wires through the routing unit.
Citation Information
Patent Citations
Processor chip, layout method, and method of accessing data
WO2017020193A1
Chip having extensible memory
WO2018058430A1