In-memory computing module and method, in-memory computing network and construction method

By using a layer-symmetric design and routing unit connections within the same chip, the problem of memory performance lag was solved, enabling low-latency in-memory computing, improving the processor's computing performance and data bandwidth, and meeting computing needs of different scales.

CN113704137BActive Publication Date: 2025-10-21XI AN UNIIC SEMICON CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202010754206.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-30
Publication Date
2025-10-21
Estimated Expiration
2040-07-30

AI Technical Summary

Technical Problem

In existing technologies, the slow improvement of memory performance leads to a lag in processor performance improvement, forming a memory bottleneck and limiting the development of high-performance computing. The structural design of resistive random access memory (RRAM) and neural network processing unit (NPU) lacks scalability, and through-silicon via (TSV) technology faces challenges such as deep hole filling and crack propagation.

Method used

The in-memory computing module, which adopts a layer-symmetric design, includes multiple sub-computing modules. The computing unit and the memory unit are located on the same chip. Low-latency access is achieved through a routing unit. The multiple sub-computing modules are connected by a bonding method. The data bit width can be a positive integer multiple of the computing unit. Dynamic random access memory is used to improve flexibility and data bandwidth.

Benefits of technology

It enables large-scale computing within the same chip, reduces the latency of computing units accessing memory units, improves computing performance, and connects sub-modules through a mature bonding method to meet computing needs of different scales.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113704137B_ABST
    Figure CN113704137B_ABST
Patent Text Reader

Abstract

The present application relates to in-memory computing module and method, in-memory computing network and its construction method. The in-memory computing module includes at least two sub-computing modules, and the computing unit in each sub-computing module can realize low delay when accessing the memory unit. The multiple sub-computing modules present a layer symmetry design, and the layer symmetry structure is convenient for constructing a topological network to realize large-scale or super-large-scale computing. The storage capacity of the memory unit in each sub-computing module can be customized, and the design is more flexible. These computing sub-computing modules are connected through bonding, and the data bit width of the bonding connection can be an integer multiple of the data bit width of the computing unit, realizing higher data bandwidth. The in-memory computing network utilizes the in-memory computing module, and can meet different scale computing requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of memory, and more particularly to an in-memory computing module and method, an in-memory computing network, and a construction method thereof. Background Art

[0002] In existing technologies, processor and memory manufacturers are separate, resulting in out-of-sync development of memory and processor technologies. Processor performance is rapidly improving, while memory performance is advancing at a relatively slow pace. This imbalance in processor and memory performance has caused memory access speeds to lag significantly behind processor computing speeds. Memory bottlenecks hinder the full performance of high-performance processors, significantly constraining the growing demand for high-performance computing. This phenomenon of memory performance severely limiting processor performance is known as the "memory wall."

[0003] As the computing power of central processing units (CPUs) and the scale of application computing continue to grow rapidly, the problem of "memory wall" has become increasingly prominent.

[0004] In order to solve the "memory wall" problem, the concept of "in-memory computing" or "storage and computing in one" emerged.

[0005] Traditionally, computing units and memory units are separate, meaning they aren't located on the same chip. Therefore, during traditional computing, the computing unit needs to extract data from the memory unit, process it, and then write it back to the memory unit. In-memory computing, on the other hand, combines the memory and computing units. By placing the memory unit as close as possible to the computing unit, the data transmission path is shortened, thereby reducing data access latency. Furthermore, in-memory computing increases access bandwidth, effectively improving computing performance.

[0006] A known "in-memory computing" structure in the prior art is as follows Figure 1 As shown in the figure. In this "in-memory computing" structure, the memory unit uses resistive random access memory (RRAM), and the computing unit is a neural network processing unit (NPU). Because the resistive random access memory (RRAM) and the neural network processing unit (NPU) use an integrated structure, the access latency of the neural network processing unit (NPU) is low. However, the process technology of resistive random access memory (RRAM) is not yet mature, and the structural design is not scalable, making it difficult to meet higher performance computing requirements.

[0007] In the existing technology, there is also a 3D stacking technology that uses through silicon via (TSV) technology to realize the "in-memory computing" structure. This 3D stacking technology stacks multiple chips together and uses through silicon via (TSV) technology to interconnect different chips. This is a three-dimensional multi-layer stack that enables communication between multiple chips in the vertical direction through silicon via (TSV). However, there are many technical difficulties in this 3D stacking technology. For example, the filling technology of the through silicon via (TSV) deep hole is directly related to the reliability and yield of the 3D stacking technology, which is crucial for the integration and practical application of the 3D stacking technology. For another example, the through silicon via (TSV) technology needs to maintain good integrity during the substrate thinning process to avoid crack propagation.

[0008] Therefore, it is urgent to solve the above technical problems in the prior art. Summary of the Invention

[0009] The present invention relates to an in-memory computing module and method, an in-memory computing network and a construction method thereof. The in-memory computing module includes multiple sub-computing modules, and the computing units in each sub-computing module can achieve low latency when accessing the memory unit. The multiple sub-computing modules present a layer-symmetrical design, and this layer-symmetrical structure is convenient for constructing a topological network to achieve large-scale or ultra-large-scale computing. The storage capacity of the memory unit in each sub-computing module can be customized, and the design is more flexible. The multiple sub-computing modules are connected by bonding, and the data bit width of the bonded connection can be a positive integer multiple of the data bit width of the computing unit, achieving higher data bandwidth. The in-memory computing network utilizes the in-memory computing module to meet computing needs of different scales.

[0010] According to a first aspect of the present invention, there is provided an in-memory computing module, the in-memory computing module comprising:

[0011] At least two computing submodules, the at least two computing submodules are stacked in sequence in one direction, each computing submodule is connected to an adjacent computing submodule, and each computing submodule includes at least one computing unit and a plurality of memory units;

[0012] The at least two computing sub-modules are located in the same chip.

[0013] As a result, the in-memory computing module including multiple computing modules can realize large-scale computing within the same chip, and the computing unit can achieve low latency when accessing the memory unit, thereby improving computing performance.

[0014] According to a preferred embodiment of the in-memory computing module of the present invention, each computing submodule includes:

[0015] a computing unit;

[0016] Multiple memory units;

[0017] a routing unit, the routing unit being connected to the computing unit, the routing unit being connected to each of the plurality of memory units, the routing unit being connected to a routing unit of at least another computing submodule of the in-memory computing module in which the routing unit is located, and the routing unit being connected to a routing unit of at least another computing submodule of at least another in-memory computing module;

[0018] The routing unit is configured to execute an access by a first computing unit of the computing sub-module where the routing unit is located to a first memory unit of the computing sub-module where the routing unit is located, an access by a second computing unit or a second memory unit of at least another computing sub-module of the in-memory computing module where the routing unit is located, or an access by a third computing unit or a third memory unit of at least another computing sub-module of at least another in-memory computing module.

[0019] According to a preferred embodiment of the in-memory computing module of the present invention, the routing unit includes:

[0020] a routing interface connecting the routing unit to a routing unit of at least another computing submodule of the in-memory computing module in which the routing unit is located, and / or connecting the routing unit to a routing unit of at least another computing submodule of at least another in-memory computing module;

[0021] A memory control interface connects the routing unit to each of the plurality of memory units.

[0022] According to a preferred embodiment of the in-memory computing module of the present invention, the routing unit further comprises:

[0023] a crossbar switch unit;

[0024] a switching routing calculation unit, the switching routing calculation unit being connected to the routing interface and the crossbar switch unit, and the switching routing calculation unit storing at least routing information about the in-memory computing module where the routing unit is located and the computing unit of the computing submodule where the routing unit is located, the switching routing calculation unit parsing received data access requests and controlling switching of the crossbar switch unit based on the parsed data access request information;

[0025] A memory control unit is connected to the crossbar switch unit and the memory control interface, and the memory control unit at least stores routing information about the multiple memory units. In response to the crossbar switch unit switching to the memory control unit, the memory control unit performs a secondary analysis on the parsed data access request received from the crossbar switch unit to determine the target memory unit, and accesses the target memory unit via the memory control interface.

[0026] According to a preferred embodiment of the in-memory computing module of the present invention, the computing unit directly accesses at least one memory unit via the routing unit. Specifically, the routing unit parses a data access request issued by the computing unit, directly obtains access data from the at least one memory unit, and forwards the access data to the computing unit that issued the data access request, wherein the data access request includes at least the address of the at least one memory unit.

[0027] According to a preferred embodiment of the in-memory computing module of the present invention, the computing unit indirectly accesses at least one other memory unit via the routing unit. That is, the routing unit parses a data access request issued by the computing unit, forwards the parsed data access request to the routing unit of at least another computing sub-module and to another computing unit to which the routing unit of the at least another computing sub-module is connected, indirectly obtains access data from the at least one other memory unit via the other computing unit, and forwards the access data to the computing unit that issued the data access request, wherein the data access request at least includes the address of the at least one other memory unit, and wherein the other computing unit can directly access the at least one other memory unit via the routing unit.

[0028] Thus, the routing unit enables unified distribution of computing unit access requests to memory units, implements memory control functionality, and further reduces latency when computing units access memory units. Furthermore, the routing unit enables intercommunication between computing submodules.

[0029] According to a preferred embodiment of the in-memory computing module of the present invention, the routing unit is connected to the routing unit of at least another computing submodule of the in-memory computing module where the routing unit is located by bonding.

[0030] As a result, multiple sub-computing modules can be connected through mature bonding methods to achieve the required electrical performance.

[0031] According to a preferred embodiment of the in-memory computing module of the present invention, the total data bit width of the connection between the routing unit and the routing unit of at least another computing sub-module of the in-memory computing module where the routing unit is located is n times the data bit width of the computing unit, where n is a positive integer.

[0032] Therefore, by setting the relationship between the data bit width of the connection between the routing units and the data bit width of a single computing unit, a higher data bandwidth can be achieved.

[0033] According to a preferred embodiment of the in-memory computing module of the present invention, the number of the plurality of memory units is determined at least according to the data bit width of the computing unit and the data bit width of a single memory unit.

[0034] Since the number of memory cells can be selected according to requirements, the design is more flexible.

[0035] According to a preferred embodiment of the in-memory computing module of the present invention, in each computing sub-module, the computing unit, the plurality of memory units and the routing unit are located in the same position in the corresponding computing sub-module.

[0036] According to a preferred embodiment of the in-memory computing module of the present invention, in each computing sub-module, the computing unit and the routing unit are located at the center of the corresponding computing sub-module, and the multiple memory units are distributed around the computing unit and the routing unit in the corresponding computing sub-module.

[0037] According to a preferred embodiment of the in-memory computing module of the present invention, each computing submodule of the at least two computing submodules is identical to each other.

[0038] As a result, multiple sub-computing modules present a layer-symmetrical design, and this layer-symmetrical structure is convenient for constructing a topological network to achieve large-scale or ultra-large-scale computing.

[0039] According to a preferred embodiment of the in-memory computing module of the present invention, each computing submodule includes:

[0040] at least two computing units;

[0041] Multiple memory units;

[0042] at least two routing units, each routing unit connected to at least one computing unit, and each routing unit connected to at least one memory unit;

[0043] The at least two routing units are connected to each other to form an overall routing unit, the overall routing unit is connected to each of the plurality of memory units, the overall routing unit is connected to the overall routing unit of at least another computing sub-module of the in-memory computing module where the overall routing unit is located, and the overall routing unit is connected to the overall routing unit of at least another computing sub-module of at least another in-memory computing module;

[0044] The overall routing unit is configured to execute an access by the first computing unit of the computing sub-module where the overall routing unit is located to the first memory unit or the second computing unit of the computing sub-module where the overall routing unit is located, an access by the second memory unit or the third computing unit of at least another computing sub-module of the in-memory computing module where the overall routing unit is located, or an access by the third memory unit or the fourth computing unit of at least another computing sub-module of at least another in-memory computing module.

[0045] According to a preferred embodiment of the in-memory computing module of the present invention, each of the at least two routing units comprises:

[0046] a routing interface, the routing interface connecting the routing unit to at least another routing unit of the computing submodule in which the routing unit is located, and / or connecting the routing unit to a routing unit of at least another computing submodule of the in-memory computing module in which the routing unit is located, and / or connecting the routing unit to a routing unit of at least another computing submodule of at least another in-memory computing module;

[0047] A memory control interface connects the routing unit to at least one memory unit of the plurality of memory units.

[0048] According to a preferred embodiment of the in-memory computing module of the present invention, each of the at least two routing units further comprises:

[0049] a crossbar switch unit;

[0050] a switching routing calculation unit, the switching routing calculation unit being connected to the routing interface and the crossbar switch unit, and the switching routing calculation unit storing at least routing information about an in-memory computing module in which the routing unit is located and at least one computing unit of a computing submodule in which the routing unit is located, the switching routing calculation unit parsing a received data access request and controlling switching of the crossbar switch unit based on the parsed data access request information;

[0051] A memory control unit is connected to the crossbar switch unit and the memory control interface, and the memory control unit at least stores routing information about at least one memory unit among the multiple memory units. In response to the crossbar switch unit switching to the memory control unit, the memory control unit performs a secondary analysis on the parsed data access request received from the crossbar switch unit to determine the target memory unit, and accesses the target memory unit via the memory control interface.

[0052] According to a preferred embodiment of the in-memory computing module of the present invention, each computing unit directly accesses at least one memory unit via the overall routing unit. Specifically, the overall routing unit parses a data access request issued by the computing unit, directly retrieves access data from the at least one memory unit, and forwards the access data to the computing unit that issued the data access request, wherein the data access request includes at least the address of the at least one memory unit.

[0053] According to a preferred embodiment of the in-memory computing module of the present invention, each computing unit indirectly accesses at least one other memory unit via the overall routing unit. That is, the overall routing unit parses a data access request issued by the computing unit, forwards the parsed data access request to the overall routing unit of at least another computing submodule and to another computing unit connected to the overall routing unit of the at least another computing submodule, indirectly obtains access data from the at least one other memory unit via the other computing unit, and forwards the access data to the computing unit that issued the data access request, wherein the data access request at least includes the address of the at least one other memory unit, and wherein the other computing unit can directly access the at least one other memory unit via the overall routing unit.

[0054] According to a preferred embodiment of the in-memory computing module of the present invention, the overall routing unit is connected to the overall routing unit of at least another computing submodule of the in-memory computing module where the overall routing unit is located by bonding.

[0055] According to a preferred embodiment of the in-memory computing module of the present invention, the total data bit width of the connection between the overall routing unit and the overall routing unit of at least another computing sub-module of the in-memory computing module where the overall routing unit is located is n times the data bit width of the computing unit, where n is a positive integer.

[0056] According to a preferred embodiment of the in-memory computing module of the present invention, the number of the plurality of memory units is determined at least according to the data bit widths of the at least two computing units and the data bit width of a single memory unit.

[0057] According to a preferred embodiment of the in-memory computing module of the present invention, in each computing submodule, the positions of the at least two computing units, the plurality of memory units and the overall routing unit in the corresponding computing submodule are the same.

[0058] According to a preferred embodiment of the in-memory computing module of the present invention, in each computing sub-module, the at least two computing units and the overall routing unit are located at the center of the corresponding computing sub-module, and the multiple memory units are distributed around the at least two computing units and the overall routing unit in the corresponding computing sub-module.

[0059] According to a preferred embodiment of the in-memory computing module of the present invention, each computing submodule of the at least two computing submodules is identical to each other.

[0060] According to a preferred embodiment of the in-memory computing module of the present invention, the memory unit includes a dynamic random access memory, and the computing unit includes a central processing unit.

[0061] Since the technology of dynamic random access memory is relatively mature, this type of memory is preferably used in the present invention.

[0062] According to a preferred embodiment of the in-memory computing module of the present invention, the at least two computing sub-modules are two computing sub-modules.

[0063] According to a preferred embodiment of the in-memory computing module of the present invention, the storage capacity of the memory unit is customizable.

[0064] Since the storage capacity of the memory cell can be customized, the design flexibility is further improved.

[0065] According to a second aspect of the present invention, an in-memory computing method is provided. The in-memory computing method is used in the in-memory computing module (in the in-memory computing module, a computing submodule has a routing unit). The in-memory computing method includes:

[0066] The routing unit receives a data access request, which is issued by the first computing unit and includes at least an address of a target memory unit; and

[0067] The routing unit parses the data access request, obtains access data from the target memory unit, and forwards the access data to the first computing unit.

[0068] In this technical solution, if the target memory unit and the first computing unit are located in the same in-memory computing module, then the "routing unit" refers to the routing unit in the in-memory computing module; if the target memory unit and the first computing unit are not located in the same in-memory computing module, then the "routing unit" refers to all routing units required for communication between the first computing unit and the target memory unit.

[0069] According to one embodiment of the in-memory computing method of the present invention, the in-memory computing method further comprises:

[0070] After the routing unit parses the data access request and before obtaining access data from the target memory unit, the routing unit connected to the first computing unit determines whether the first computing unit can directly access the target memory unit via the routing unit connected to the first computing unit;

[0071] When the first computing unit can directly access the target memory unit via the routing unit connected to the first computing unit, the routing unit connected to the first computing unit directly obtains access data from the target memory unit and forwards the access data to the first computing unit; and

[0072] When the first computing unit cannot directly access the target memory unit via the routing unit connected to the first computing unit, the routing unit connected to the first computing unit forwards the parsed data access request to the second computing unit, and obtains access data from the target memory unit via the second computing unit and forwards the access data to the first computing unit.

[0073] According to one embodiment of the in-memory computing method of the present invention, the in-memory computing method further comprises:

[0074] When the first computing unit cannot directly access the target memory unit via the routing unit connected to the first computing unit and before the routing unit connected to the first computing unit forwards the parsed data access request to the second computing unit, the routing unit connected to the first computing unit determines whether the target memory unit and the first computing unit are located in the same in-memory computing module;

[0075] When the target memory unit and the first computing unit are located in the same in-memory computing module, the routing unit connected to the first computing unit forwards the parsed data access request to a routing unit of another computing sub-module connected to the routing unit connected to the first computing unit and to a second computing unit connected to the routing unit of the another computing sub-module, and obtains access data from the target memory unit via the second computing unit and forwards the access data to the first computing unit, wherein the second computing unit can directly access the target memory unit via the routing unit of the another computing sub-module;

[0076] When the target memory unit and the first computing unit are not located in the same in-memory computing module, the routing unit connected to the first computing unit forwards the parsed data access request to the routing unit of another computing sub-module of another in-memory computing module connected to the routing unit connected to the first computing unit and forwards it to the second computing unit connected to the routing unit of another computing sub-module of the other in-memory computing module, and obtains access data from the target memory unit via the second computing unit and forwards the access data to the first computing unit, wherein the second computing unit can directly access the target memory unit via the routing unit of another computing sub-module of the other in-memory computing module.

[0077] According to a third aspect of the present invention, an in-memory computing method is provided. The in-memory computing method is used in the in-memory computing module (in the in-memory computing module, a computing submodule has a routing unit). The in-memory computing method includes:

[0078] The routing unit receives a data access request, which is issued by the first computing unit and includes at least an address of a target computing unit; and

[0079] The routing unit parses the data access request, obtains access data from the target computing unit, and forwards the access data to the first computing unit.

[0080] In this technical solution, if the target computing unit and the first computing unit are located in the same in-memory computing module, then the "routing unit" refers to the routing unit in the in-memory computing module; if the target computing unit and the first computing unit are not located in the same in-memory computing module, then the "routing unit" refers to all routing units required for communication between the first computing unit and the target computing unit.

[0081] According to one embodiment of the in-memory computing method of the present invention, the in-memory computing method further comprises:

[0082] After the routing unit parses the data access request and before obtaining access data from the target computing unit, the routing unit connected to the first computing unit determines whether the target computing unit and the first computing unit are located in the same in-memory computing module;

[0083] When the target computing unit and the first computing unit are located in the same in-memory computing module, the routing unit connected to the first computing unit forwards the parsed data access request to a routing unit of another computing submodule connected to the routing unit connected to the first computing unit, and obtains access data from the target computing unit via the routing unit of the other computing submodule and forwards the access data to the first computing unit.

[0084] When the target computing unit and the first computing unit are not located in the same in-memory computing module, the routing unit connected to the first computing unit forwards the parsed data access request to the routing unit of another computing sub-module of another in-memory computing module connected to the routing unit to which the first computing unit is connected, and obtains access data from the target computing unit via the routing unit of another computing sub-module of the other in-memory computing module and forwards the access data to the first computing unit.

[0085] According to a fourth aspect of the present invention, an in-memory computing method is provided. The in-memory computing method is used in the in-memory computing module (in the in-memory computing module, a computing submodule has at least two routing units). The in-memory computing method includes:

[0086] The overall routing unit receives a data access request, where the data access request is issued by the first computing unit and includes at least an address of a target memory unit; and

[0087] The overall routing unit parses the data access request, obtains access data from the target memory unit, and forwards the access data to the first computing unit.

[0088] In this technical solution, if the target memory unit and the first computing unit are located in the same in-memory computing module, then the "overall routing unit" refers to the overall routing unit in the in-memory computing module; if the target memory unit and the first computing unit are not located in the same in-memory computing module, then the "overall routing unit" refers to all overall routing units required for communication between the first computing unit and the target memory unit.

[0089] According to one embodiment of the in-memory computing method of the present invention, the in-memory computing method further comprises:

[0090] After the overall routing unit parses the data access request and before obtaining access data from the target memory unit, the overall routing unit connected to the first computing unit determines whether the first computing unit can directly access the target memory unit via the overall routing unit connected to the first computing unit;

[0091] When the first computing unit can directly access the target memory unit via the overall routing unit connected to the first computing unit, the overall routing unit connected to the first computing unit directly obtains access data from the target memory unit and forwards the access data to the first computing unit; and

[0092] When the first computing unit cannot directly access the target memory unit via the overall routing unit connected to the first computing unit, the overall routing unit connected to the first computing unit forwards the parsed data access request to the second computing unit, and obtains access data from the target memory unit via the second computing unit and forwards the access data to the first computing unit.

[0093] According to one embodiment of the in-memory computing method of the present invention, the in-memory computing method further comprises:

[0094] When the first computing unit cannot directly access the target memory unit via the overall routing unit connected to the first computing unit and before the overall routing unit connected to the first computing unit forwards the parsed data access request to the second computing unit, the overall routing unit connected to the first computing unit determines whether the target memory unit and the first computing unit are located in the same in-memory computing module;

[0095] When the target memory unit and the first computing unit are located in the same in-memory computing module, the overall routing unit connected to the first computing unit forwards the parsed data access request to the overall routing unit of another computing submodule connected to the overall routing unit connected to the first computing unit and to a second computing unit connected to the overall routing unit of the another computing submodule, and obtains access data from the target memory unit via the second computing unit and forwards the access data to the first computing unit, wherein the second computing unit can directly access the target memory unit via the overall routing unit of the another computing submodule;

[0096] When the target memory unit and the first computing unit are not located in the same in-memory computing module, the overall routing unit connected to the first computing unit forwards the parsed data access request to the overall routing unit of another computing sub-module of another in-memory computing module connected to the overall routing unit connected to the first computing unit and forwards it to the second computing unit connected to the overall routing unit of another computing sub-module of the other in-memory computing module, and obtains access data from the target memory unit via the second computing unit and forwards the access data to the first computing unit, wherein the second computing unit can directly access the target memory unit via the overall routing unit of another computing sub-module of the other in-memory computing module.

[0097] According to a fifth aspect of the present invention, an in-memory computing method is provided. The in-memory computing method is used in the in-memory computing module (in the in-memory computing module, a computing submodule has at least two routing units). The in-memory computing method includes:

[0098] The overall routing unit receives a data access request, where the data access request is issued by the first computing unit and includes at least an address of a target computing unit; and

[0099] The overall routing unit parses the data access request, obtains access data from the target computing unit, and forwards the access data to the first computing unit.

[0100] In this technical solution, if the target computing unit and the first computing unit are located in the same in-memory computing module, then the "overall routing unit" refers to the overall routing unit in the in-memory computing module; if the target computing unit and the first computing unit are not located in the same in-memory computing module, then the "overall routing unit" refers to all overall routing units required for communication between the first computing unit and the target computing unit.

[0101] According to one embodiment of the in-memory computing method of the present invention, the in-memory computing method further comprises:

[0102] After the overall routing unit parses the data access request and before obtaining access data from the target computing unit, the overall routing unit connected to the first computing unit determines whether the target computing unit and the first computing unit are located in the same in-memory computing module;

[0103] When the target computing unit and the first computing unit are located in the same in-memory computing module, the overall routing unit connected to the first computing unit forwards the parsed data access request to a routing unit of another computing submodule connected to the overall routing unit to which the first computing unit is connected, and obtains access data from the target computing unit via the overall routing unit of the other computing submodule and forwards the access data to the first computing unit.

[0104] When the target computing unit and the first computing unit are not located in the same in-memory computing module, the overall routing unit connected to the first computing unit forwards the parsed data access request to the overall routing unit of another in-memory computing module connected to the overall routing unit to which the first computing unit is connected, and obtains access data from the target computing unit via the overall routing unit of the other in-memory computing module and forwards the access data to the first computing unit.

[0105] According to a sixth aspect of the present invention, there is provided an in-memory computing network, the in-memory computing network comprising:

[0106] Multiple in-memory computing modules, which are multiple in-memory computing modules mentioned above, are connected through the routing units of the multiple in-memory computing modules.

[0107] According to a preferred embodiment of the in-memory computing network of the present invention, the plurality of in-memory computing modules are connected into a bus, star, ring, tree, mesh and hybrid topology.

[0108] According to a preferred embodiment of the in-memory computing network of the present invention, the multiple in-memory computing modules are connected via metal wires through a routing unit.

[0109] According to a seventh aspect of the present invention, a method for constructing an in-memory computing module is provided, the method comprising:

[0110] Arrange at least two computing submodules to be stacked sequentially in one direction;

[0111] Arranging each computing submodule to be connected to its adjacent computing submodule, wherein each computing submodule includes at least one computing unit and a plurality of memory units;

[0112] The at least two computing sub-modules are arranged in the same chip.

[0113] According to a preferred embodiment of the construction method of the present invention, each computing submodule includes: a computing unit; a plurality of memory units; a routing unit;

[0114] The construction method also includes:

[0115] The routing unit is connected to the computing unit, and the routing unit is connected to each of the multiple memory units, and the routing unit is connected to the routing unit of at least another computing sub-module of the in-memory computing module where the routing unit is located, and the routing unit is connected to the routing unit of at least another computing sub-module of at least another in-memory computing module; and the routing unit is configured to execute access by the first computing unit of the computing sub-module where the routing unit is located to the first memory unit of the computing sub-module where the routing unit is located, access to the second computing unit or the second memory unit of at least another computing sub-module of the in-memory computing module where the routing unit is located, or access to the third computing unit or the third memory unit of at least another computing sub-module of at least another in-memory computing module.

[0116] According to a preferred embodiment of the construction method of the present invention, the construction method further comprises:

[0117] connecting the routing interface of the routing unit to the routing unit of at least another computing submodule of the in-memory computing module where the routing unit is located, and / or connecting the routing unit to the routing unit of at least another computing submodule of at least another in-memory computing module;

[0118] The memory control interface of the routing unit is connected to each memory unit of the plurality of memory units.

[0119] According to a preferred embodiment of the construction method of the present invention, the construction method further comprises:

[0120] connecting a switching routing calculation unit of the routing unit to the routing interface and a crossbar switch unit, storing at least routing information about the in-memory computing module where the routing unit is located and the computing unit of the computing submodule where the routing unit is located in the switching routing calculation unit, and configuring the switching routing calculation unit to parse a received data access request and control switching of the crossbar switch unit based on the parsed data access request information;

[0121] The memory control unit of the routing unit is connected to the crossbar switch unit and the memory control interface, and at least routing information about the multiple memory units is stored in the memory control unit. The memory control unit is also configured to perform a secondary analysis on the parsed data access request received from the crossbar switch unit in response to the crossbar switch unit switching to the memory control unit to determine the target memory unit, and access the target memory unit via the memory control interface.

[0122] According to a preferred embodiment of the construction method of the present invention, the construction method further comprises:

[0123] The computing unit is configured to directly access the at least one memory unit via the routing unit. That is, the routing unit parses a data access request issued by the computing unit, directly obtains access data from the at least one memory unit, and forwards the access data to the computing unit that issued the data access request, wherein the data access request includes at least the address of the at least one memory unit.

[0124] According to a preferred embodiment of the construction method of the present invention, the construction method further comprises:

[0125] The computing unit is configured to indirectly access at least one other memory unit via the routing unit. That is, the routing unit parses a data access request issued by the computing unit, forwards the parsed data access request to the routing unit of at least another computing sub-module and to another computing unit to which the routing unit of the at least another computing sub-module is connected, indirectly obtains access data from the at least one other memory unit via the other computing unit, and forwards the access data to the computing unit that issued the data access request, wherein the data access request at least includes the address of the at least one other memory unit, and wherein the other computing unit can directly access the at least one other memory unit via the routing unit.

[0126] According to a preferred embodiment of the construction method of the present invention, the construction method further comprises:

[0127] The routing unit is connected to the routing unit of at least another computing submodule of the in-memory computing module where the routing unit is located by bonding.

[0128] According to a preferred embodiment of the construction method of the present invention, the construction method further comprises:

[0129] The total data bit width of the connection between the routing unit and the routing unit of at least another computing submodule of the in-memory computing module where the routing unit is located is set to n times the data bit width of the computing unit, where n is a positive integer.

[0130] According to a preferred embodiment of the construction method of the present invention, the construction method further comprises:

[0131] The number of the plurality of memory units is determined at least according to the data bit width of the calculation unit and the data bit width of a single memory unit.

[0132] According to a preferred embodiment of the construction method of the present invention, the construction method further comprises:

[0133] In each computing sub-module, the positions of the computing unit, the multiple memory units, and the routing unit in the corresponding computing sub-module are set to be the same.

[0134] According to a preferred embodiment of the construction method of the present invention, the construction method further comprises:

[0135] In each computing submodule, the computing unit and the routing unit are arranged at the center of the corresponding computing submodule, and the multiple memory units are arranged to be distributed around the computing unit and the routing unit in the corresponding computing submodule.

[0136] According to a preferred embodiment of the construction method of the present invention, the construction method further comprises:

[0137] The respective computing submodules in the at least two computing submodules are configured to be identical to each other.

[0138] According to a preferred embodiment of the construction method of the present invention, each computing submodule includes: at least two computing units; a plurality of memory units; at least two routing units, wherein each routing unit is connected to at least one computing unit, and each routing unit is connected to at least one memory unit;

[0139] The construction method also includes:

[0140] The at least two routing units are connected to each other to form an overall routing unit, and the overall routing unit is connected to each of the plurality of memory units, the overall routing unit is connected to the overall routing unit of at least another computing sub-module of the in-memory computing module where the overall routing unit is located, and the overall routing unit is connected to the overall routing unit of at least another computing sub-module of at least another in-memory computing module; and the overall routing unit is configured to execute access by a first computing unit of the computing sub-module where the overall routing unit is located to a first memory unit or a second computing unit of the computing sub-module where the overall routing unit is located, access by a second memory unit or a third computing unit of at least another computing sub-module of the in-memory computing module where the overall routing unit is located, or access by a third memory unit or a fourth computing unit of at least another computing sub-module of at least another in-memory computing module.

[0141] According to a preferred embodiment of the construction method of the present invention, the construction method further comprises:

[0142] connecting a routing interface of each routing unit of the at least two routing units to at least another routing unit of the computing submodule in which the routing unit is located, and / or to a routing unit of at least another computing submodule of the in-memory computing module in which the routing unit is located, and / or to a routing unit of at least another computing submodule of at least another in-memory computing module;

[0143] The memory control interface of each routing unit of the at least two routing units is connected to at least one memory unit of the plurality of memory units.

[0144] According to a preferred embodiment of the construction method of the present invention, the construction method further comprises:

[0145] connecting a switching routing calculation unit of each of the at least two routing units to the routing interface and a crossbar switch unit, storing routing information of at least one computing unit of an in-memory computing module and a computing submodule of the routing unit in the switching routing calculation unit, and configuring the switching routing calculation unit to parse a received data access request and control switching of the crossbar switch unit based on the parsed data access request information;

[0146] The memory control unit of each routing unit of the at least two routing units is connected to the crossbar switch unit and the memory control interface, and at least routing information about at least one memory unit among the multiple memory units is stored in the memory control unit. The memory control unit is also configured to, in response to the crossbar switch unit switching to the memory control unit, perform a secondary analysis on the parsed data access request received from the crossbar switch unit to determine a target memory unit, and access the target memory unit through the memory control interface.

[0147] According to a preferred embodiment of the construction method of the present invention, the construction method further comprises:

[0148] Each computing unit is configured to directly access at least one memory unit via the overall routing unit. That is, the overall routing unit parses a data access request issued by the computing unit, directly obtains access data from the at least one memory unit, and forwards the access data to the computing unit that issued the data access request, wherein the data access request includes at least the address of the at least one memory unit.

[0149] According to a preferred embodiment of the construction method of the present invention, the construction method further comprises:

[0150] Each computing unit is configured to indirectly access at least one other memory unit via the overall routing unit. That is, the overall routing unit parses a data access request issued by the computing unit, forwards the parsed data access request to the overall routing unit of at least another computing submodule and to another computing unit connected to the overall routing unit of the at least another computing submodule, indirectly obtains access data from the at least one other memory unit via the other computing unit, and forwards the access data to the computing unit that issued the data access request, wherein the data access request includes at least the address of the at least one other memory unit, and wherein the other computing unit can directly access the at least one other memory unit via the overall routing unit.

[0151] According to a preferred embodiment of the construction method of the present invention, the construction method further comprises:

[0152] The overall routing unit is connected to the overall routing unit of at least another computing submodule of the in-memory computing module where the overall routing unit is located by bonding.

[0153] According to a preferred embodiment of the construction method of the present invention, the construction method further comprises:

[0154] The total data bit width of the connection between the overall routing unit and the overall routing unit of at least another computing submodule of the in-memory computing module where the overall routing unit is located is set to n times the data bit width of the computing unit, where n is a positive integer.

[0155] According to a preferred embodiment of the construction method of the present invention, the construction method further comprises:

[0156] The number of the plurality of memory units is determined at least according to the data bit width of the calculation unit and the data bit width of a single memory unit.

[0157] According to a preferred embodiment of the construction method of the present invention, the construction method further comprises:

[0158] In each computing submodule, positions of the at least two computing units, the plurality of memory units, and the overall routing unit in the corresponding computing submodule are set to be the same.

[0159] According to a preferred embodiment of the construction method of the present invention, the construction method further comprises:

[0160] In each computing submodule, the at least two computing units and the overall routing unit are arranged at the center of the corresponding computing submodule, and the multiple memory units are arranged in the corresponding computing submodule to be distributed around the at least two computing units and the overall routing unit.

[0161] According to a preferred embodiment of the construction method of the present invention, the construction method further comprises:

[0162] The respective computing submodules in the at least two computing submodules are configured to be identical to each other.

[0163] According to a preferred embodiment of the construction method of the present invention, the construction method further comprises:

[0164] The memory unit is configured to include a dynamic random access memory, and the computing unit is configured to include a central processing unit.

[0165] According to a preferred embodiment of the construction method of the present invention, the construction method further comprises:

[0166] The at least two computing submodules are configured as two computing submodules.

[0167] According to a preferred embodiment of the construction method of the present invention, the construction method further comprises:

[0168] The storage capacity of the memory unit is set to be customizable.

[0169] According to an eighth aspect of the present invention, a method for constructing an in-memory computing network is provided, the method comprising:

[0170] The multiple in-memory computing modules are connected via the routing units of the multiple in-memory computing modules, wherein the multiple in-memory computing modules are the multiple in-memory computing modules mentioned above.

[0171] According to a preferred embodiment of the construction method of the present invention, the construction method further comprises:

[0172] The multiple in-memory computing modules are connected into bus, star, ring, tree, mesh and hybrid topologies.

[0173] According to a preferred embodiment of the construction method of the present invention, the construction method further comprises:

[0174] The plurality of in-memory computing modules are connected via metal lines through a routing unit. BRIEF DESCRIPTION OF THE DRAWINGS

[0175] The present invention will be more easily understood through the following description in conjunction with the accompanying drawings, in which:

[0176] Figure 1 It is a schematic diagram of the in-memory computing structure in the prior art.

[0177] Figure 2 is a schematic diagram of an in-memory computing module according to one embodiment of the present invention.

[0178] Figure 3 Schematic diagram of the relative positions of the routing unit and the computing unit and the interface of the routing unit according to one embodiment of the present invention.

[0179] Figure 4 is a schematic diagram of a routing unit according to one embodiment of the present invention.

[0180] Figure 5 is a flow chart of an in-memory computing method according to one embodiment of the present invention.

[0181] Figure 6 is a flow chart of an in-memory computing method according to another embodiment of the present invention.

[0182] Figure 7 It is a flowchart of a method for constructing an in-memory computing module according to one embodiment of the present invention.

[0183] Figure 8 is a schematic diagram of an in-memory computing network according to one embodiment of the present invention. DETAILED DESCRIPTION

[0184] The embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0185] Figure 2is a schematic diagram of an in-memory computing module 20 according to one embodiment of the present invention.

[0186] Figure 2 The in-memory computing module 20 shown in FIG includes two computing submodules (or referred to as two in-memory computing "layers"), with the upper computing submodule and the lower computing submodule stacked in sequence. Figure 2 As shown in , each computing submodule includes a computing unit 210 , a memory unit 212 and a routing unit 211 .

[0187] However, the present invention is not limited to two computing submodules, and may also include more than two computing submodules. In the case of more than two computing submodules, these computing submodules may be stacked in sequence in one direction (for example, referring to FIG. Figure 2 In the case of two computing submodules stacked in sequence, in the case of more than two computing submodules, these computing submodules can be stacked in sequence. Figure 2 ).

[0188] Figure 2 The two computing submodules shown in FIG are located very close to each other, are located on the same chip, and form a complete system. This complete system is also called an in-memory computing module, or an in-memory computing node.

[0189] like Figure 2 As shown in , each computing submodule includes a computing unit 210, a routing unit 211, and eight memory units 212. Because the two computing submodules 21 and 22 are located in the same chip, all computing units 210 and memory units 212 are also integrated together, so that the latency when the computing unit 210 accesses the memory unit 212 is very small.

[0190] Figure 2 The two computing submodules 21 and 22 shown in FIG are identical to each other, and the following description will focus on the computing submodule 21. It should be understood that all features of the computing submodule 21 are also applicable to the computing submodule 22.

[0191] The memory unit 212 is a unit for storing the calculation data in the computing unit 210 and the data exchanged with the external memory such as the hard disk. Since the technology of dynamic random access memory is relatively mature, in the present invention, the memory unit 212 preferably adopts dynamic random access memory.

[0192] The number of the memory units 212 is determined at least according to the data bit width of the calculation unit 210 and the data bit width of a single memory unit 212 .

[0193] For example, if there is one computing unit 210 in the computing submodule, the data bit width of the computing unit 210 is 64 bits, and the data bit width of a single memory unit 203 is 8 bits, then the number of memory units 212 required is eight (e.g., Figure 2 ).

[0194] In addition, the storage capacity of the memory unit 212 can also be customized according to needs.

[0195] The computing unit 210 is the final execution unit for information processing and program execution, and is preferably a central processing unit.

[0196] In computing submodule 21, routing unit 211 is connected to computing unit 210 and each memory unit 212. Furthermore, routing unit 211 in computing submodule 21 is connected to routing unit 211 in computing submodule 22. Furthermore, routing unit 211 in computing submodule 21 is also connected to routing unit 211 in at least one other computing submodule in at least one other in-memory computing module. The primary function of routing unit 211 in computing submodule 21 is to enable a computing unit 210 in computing submodule 21 to access a memory unit 212 in computing submodule 21, access a computing unit 210 or a memory unit 212 in at least one other computing submodule 22, or access a computing unit 210 or a memory unit 212 in at least one other in-memory computing module.

[0197] The computing unit 210 in the computing submodule 21 can “directly” access the memory unit 212 corresponding to the computing unit via the routing unit 211 .

[0198] For example, reference Figure 2 , the computing unit 210 in the computing submodule 21 can “directly” access all memory units 212 in the computing submodule 21 via the routing unit 211 .

[0199] That is to say, if the computing unit 210 in the computing sub-module 21 issues a data access request for any memory unit 212 in the computing sub-module 21, the routing unit 211 in the computing sub-module 21 can parse the data access request, "directly" obtain the access data from any memory unit 212 in the computing sub-module 21 and return the access data to the computing unit 210.

[0200] In addition, the computing unit 210 in the computing submodule 21 can “indirectly” access other memory units 212 via the routing unit 211 .

[0201] Reference again Figure 2If the computing unit 210 in the computing sub-module 21 issues a data access request for any memory unit 212 in the computing sub-module 22, the routing unit 211 in the computing sub-module 21 can parse the data access request and forward the parsed data access request to the routing unit 211 of the computing sub-module 22 and the computing unit 210 connected to the routing unit 211, "indirectly" obtain access data from any memory unit 212 via the computing unit 210 connected to the routing unit 211, and return the access data to the computing unit 210 that issued the access request.

[0202] Figure 2 An embodiment of the distribution of the computing unit 210, routing unit 211 and multiple memory units 212 in the computing sub-modules 21 and 22 is shown, wherein the computing unit 210 and the routing unit 211 are located at the center of the corresponding computing sub-module, and the multiple memory units 212 are distributed around the computing unit 210 and the routing unit 211 in the corresponding computing sub-module.

[0203] However, this distribution is illustrative and not restrictive. It is also within the scope of the present invention that the multiple memory units 212 are distributed on one side of the computing unit 210 and the routing unit 211, or on both sides of the computing unit 210 and the routing unit 211.

[0204] In addition, if Figure 2 As shown in FIG, in the computing submodule 21 and the computing submodule 22, the positions of the computing unit 210, the routing unit 211 and the plurality of memory units 212 are the same in the corresponding computing submodules ( Figure 2 The calculation submodule 21 and the calculation submodule 22 shown in FIG are identical to each other).

[0205] That is to say, Figure 2 The computing submodules 21 and 22 of the in-memory computing module 20 shown in FIG have a layer-symmetrical structure. This layer-symmetrical structure is convenient for constructing a topological network to implement large-scale or ultra-large-scale computing.

[0206] exist Figure 2 In the embodiment, the data interfaces of the computing submodules 21 and 22 need to be connected with the aid of a routing unit 210 , preferably using a three-dimensional connection 230 process.

[0207] Commonly used three-dimensional connection 230 processes include bonding, through silicon vias (TSV), flip chip, and wafer level packaging. In the present invention, the three-dimensional connection 230 process preferably adopts bonding.

[0208] Bonding is a commonly used three-dimensional connection process and a wafer stacking process within a chip. Specifically, bonding involves connecting wafers together using metal wires through a specific process to achieve the required electrical characteristics.

[0209] In addition, to alleviate data transmission pressure, the total data bit width of the connection between the routing unit 211 of the computing submodule 21 and the routing unit 211 of the computing submodule 22 is n times the data bit width of the computing unit 210, where n is a positive integer.

[0210] exist Figure 2 In the example, access between different computing submodules 21 and 22 needs to go through the connection between the routing units 210. If the total data bit width of the connection is the same as that of the computing unit 210, when the data access operations between the computing submodules 21 and 22 are frequent (i.e., the data throughput is high), the data throughput between the computing submodules 21 and 22 will also increase, and the connection may be congested. Therefore, the total data bit width of the connection is set to a positive integer multiple of the data bit width of the computing unit 210. Figure 2 In the specific case shown in , the total data bit width of the connection is at most 8 times the data bit width of the computing unit 210.

[0211] It should be understood that the specific value of the positive integer n is set according to business requirements. For example, a general system design can derive the bandwidth requirements for data transmission between different computing submodules within a chip based on business simulation, and the required data bit width can be derived based on the bandwidth requirements.

[0212] Assume that the data bandwidth requirement between the two computing submodules 21 and 22 is 144 Gb / s, and the total data bit width of the existing connection is 72 bits and the clock is 1 GHz. In this case, it is necessary to consider increasing the data bit width of the connection to 144 bits to adapt to the data bandwidth requirement.

[0213] Figure 3 2 is a schematic diagram illustrating the relative positions of the routing unit 211 and the computing unit 210 and the interface of the routing unit 211 according to an embodiment of the present invention.

[0214] like Figure 3 As shown in FIG, the routing unit 211 is arranged around the computing unit 210. However, this arrangement is exemplary and not restrictive. The routing unit 211 can also be arranged on one side or both sides of the computing unit 210, which also falls within the scope of the present invention.

[0215] Figure 3 The memory control interface 213 and the routing interface 214 of the routing unit 211 are also schematically shown. These interfaces will refer to Figure 4 A more detailed explanation is given.

[0216] Figure 4 is a schematic diagram of a routing unit according to one embodiment of the present invention.

[0217] Figure 4 4 shows the external interfaces of the routing unit 211, which mainly include a routing interface 401 and a memory control interface 405. It should be understood that in order not to obscure the main purpose of the present invention, Figure 3 Only the external interfaces involved in the present invention are shown, and these external interfaces are exemplary rather than restrictive.

[0218] exist Figure 4 , the routing interface is shown as a memory front-end routing (MFR) interface 401, which connects the routing unit to the routing unit of at least another computing sub-module of the in-memory computing module where the routing unit is located, and / or connects the routing unit to the routing unit of at least another computing sub-module of at least another in-memory computing module.

[0219] For example, return reference Figure 2 , the memory front-end routing interface 401 of the routing unit 211 in the computing sub-module 21 is connected to the memory front-end routing interface 401 of the routing unit 211 in the computing sub-module 22, and / or is connected to the memory front-end routing interface 401 of the routing unit of at least another computing sub-module of at least another in-memory computing module.

[0220] The memory control interface is shown as a DDR operation interface (DDRIO-bonding), which is connected to each memory unit 212.

[0221] Figure 4 4. In addition, the internal structure of the routing unit 211 is shown, which mainly includes a switching routing calculation unit 402, a crossbar switch unit 403 and a memory control unit 404. It should be understood that in order not to obscure the main purpose of the present invention, Figure 4 Only the internal structures involved in the present invention are shown, which are exemplary and non-limiting. It should be understood that the routing unit 211 should also include some buffer circuits, digital and analog circuits, etc. For example, the buffer circuits can buffer and prioritize data access requests from multiple computing units 210, and the digital and analog circuits can cooperate with the memory control unit 404, etc.

[0222] In the present invention, the switch routing calculation unit 402 stores routing information about the computing unit 210. This routing information can be stored in the form of a routing table, for example. Thus, the switch routing calculation unit 402 can determine information such as whether the computing unit 210 in a data access request can "directly" access a memory unit via a router.

[0223] In addition, the switching routing calculation unit 402 also stores routing information about the in-memory computing modules. The routing information can also be stored in the form of a routing table. Thus, the switching routing calculation unit 402 can determine in which in-memory computing module the memory address or the target computing unit address is located (whether it is in the current in-memory computing module or in another in-memory computing module). For example, based on a certain specific position information (such as the first bit) in the target memory address or the target computing unit address and the routing information, the switching routing calculation unit 402 can determine in which in-memory computing module the target memory address or the target computing unit address is located. For example, if the first bit of the target memory address or the target computing unit address is 1, it means that the target memory or the target computing unit is in the first in-memory computing module; if the first bit of the target memory address or the target computing unit address is 3, it means that the target memory or the target computing unit is in the third in-memory computing module.

[0224] The memory control unit 404 stores routing information about the memory unit 212. Thus, the memory control unit 404 can determine port information corresponding to the memory unit in the data access request, and so on.

[0225] Figure 5 is a flow chart of an in-memory computing method according to one embodiment of the present invention.

[0226] The in-memory calculation method includes the following steps:

[0227] Step S501: the routing unit in the computing submodule receives a data access request.

[0228] The data access request in this embodiment is issued by the first computing unit in the computing submodule and includes at least the address of the target memory unit.

[0229] Step S502: The routing unit in the computing submodule parses the data access request, and the routing unit determines whether the target memory unit and the first computing unit are located in the same in-memory computing module.

[0230] If the target memory unit and the first computing unit are not located in the same in-memory computing module, step S503 is executed: the routing unit forwards the parsed data access request to the routing unit of another computing sub-module of another in-memory computing module connected to the routing unit and forwards it to the second computing unit connected to the routing unit of another computing sub-module of the other in-memory computing module, and obtains access data from the target memory unit via the second computing unit and forwards the access data to the first computing unit.

[0231] If the target memory unit and the first computing unit are located in the same in-memory computing module, step S504 is executed: the routing unit determines whether the first computing unit can directly access the target memory unit.

[0232] If the first computing unit can directly access the target memory unit via the routing unit, step S505 is executed: the routing unit directly obtains access data from the target memory unit and forwards the access data to the first computing unit.

[0233] If the first computing unit cannot directly access the target memory unit via the routing unit, step S506 is executed: the routing unit forwards the parsed data access request to the second computing unit, and obtains access data from the target memory unit via the second computing unit and forwards the access data to the first computing unit, wherein the second computing unit can directly access the target memory unit via the routing unit.

[0234] Figure 5 The method flow chart in the embodiment is only illustrative and does not necessarily have to be executed in this order. For example, it is possible to first determine whether the first computing unit can directly access the target memory unit, and then determine whether the target memory unit and the first computing unit are located in the same in-memory computing module.

[0235] Figure 6 is a flow chart of an in-memory computing method according to another embodiment of the present invention.

[0236] The in-memory calculation method includes the following steps:

[0237] Step S601: the routing unit in the computing submodule receives a data access request.

[0238] The data access request in this embodiment is issued by the first computing unit in the computing submodule and includes at least the address of the target computing unit.

[0239] Step S602: The routing unit in the computing submodule parses the data access request, and the routing unit determines whether the target computing unit and the first computing unit are located in the same in-memory computing module.

[0240] If the target computing unit and the first computing unit are located in the same in-memory computing module, step S603 is executed: the routing unit forwards the parsed data access request to the routing unit of another computing sub-module connected to the routing unit, and obtains access data from the target computing unit via the routing unit of the other computing sub-module and forwards the access data to the first computing unit.

[0241] If the target computing unit and the first computing unit are not located in the same in-memory computing module, step S604 is executed: the routing unit forwards the parsed data access request to the routing unit of another computing sub-module of another in-memory computing module connected to the routing unit, and obtains access data from the target computing unit via the routing unit of another computing sub-module of the other in-memory computing module and forwards the access data to the first computing unit.

[0242] The following combination Figure 2-Figure 4 , further understand the external interface and internal structure of the routing unit 211 and the following five data flow processing scenarios in the in-memory computing module Figure 5 and Figure 6 In-memory calculation method in .

[0243] Reference again Figure 2 , assuming that the computing unit 210 in the computing submodule 21 can “directly” access all memory units 212 via the routing unit 211 , and the computing unit 210 in the computing submodule 22 can “directly” access all memory units 212 via the routing unit 211 .

[0244] Scenario (1): The computing unit 210 in the computing submodule 21 accesses any memory unit 212 in the computing submodule 21

[0245] The computing unit 210 in the computing submodule 21 sends a data access request to the switching routing computing unit 402 through the routing interface MFR. The switching routing computing unit 402 parses the data access request to obtain a target memory address, and determines whether the target memory address is within the same in-memory computing module (but does not parse out which specific memory unit it is) (in this case, it is determined to be within the same in-memory computing module), and determines whether the computing unit 210 in the computing submodule 21 can "directly" access any memory unit 212 in the computing submodule 21 (in this case, it is determined to be "directly" accessible).

[0246] Afterwards, the switching routing calculation unit 402 queries the target memory address in the routing information stored therein about the computing unit and the in-memory computing module, determines the port information corresponding to the target memory address, and then controls the opening and closing of the cross switch unit 403, so that the parsed data access request is sent to the memory control unit 404 for secondary parsing to determine which specific memory unit needs to be accessed, and then accesses any memory unit 212 in the computing sub-module 21 through the memory control (DDRIO-bonding) interface.

[0247] Scenario (2): The computing unit 210 in the computing submodule 21 accesses any memory unit 212 in the computing submodule 22

[0248] The computing unit 210 in the computing submodule 21 sends a data access request to the switching routing computing unit 402 through the routing interface MFR. The switching routing computing unit 402 parses the data access request to obtain a target memory address, and determines whether the target memory address is within the same in-memory computing module (but does not parse out which specific memory unit it is) (in this case, it is determined to be within the same in-memory computing module), and determines whether the computing unit 210 in the computing submodule 21 can "directly" access any memory unit 212 in the computing submodule 21 (in this case, it is determined that "direct" access is not possible).

[0249] Afterwards, the switching routing calculation unit 402 queries the target memory address in the routing information stored therein about the computing unit and the in-memory computing module, determines the port information corresponding to the target memory address, and then controls the opening and closing of the crossbar switch unit 403, so that the parsed data access request is sent to the routing unit 211 of the computing sub-module 22 through the routing interface MFR and to the computing unit 210 connected to the routing unit 211 of the computing sub-module 22. The computing unit 210 connected to the routing unit 211 of the computing sub-module 22 performs the operation as in scenario (1), and then accesses any memory unit 212 in the computing sub-module 22 through the memory control (DDRIO-bonding) interface.

[0250] Scenario (3): The computing unit 210 in the computing submodule 21 accesses the memory unit 212 of another in-memory computing module

[0251] The computing unit 210 in the computing submodule 21 sends a data access request to the switching routing computing unit 402 through the routing interface MFR. The switching routing computing unit 402 parses the data access request to obtain a target memory address and determines whether the target memory address is within the same in-memory computing module (but does not parse out the specific memory unit). In this case, it is determined that the target memory address is not within the same in-memory computing module.

[0252] Afterwards, the switching routing calculation unit 402 controls the opening and closing of the crossbar switch unit 403 and sends the parsed data access request to another in-memory calculation module through the routing interface MFR.

[0253] The other in-memory computing module performs the operations in the above scenario (1) and scenario (2), and then accesses the memory unit 212 of the other in-memory computing module through the memory control (DDRIO-bonding) interface of the routing unit of the other in-memory computing module.

[0254] Scenario (4): The computing unit 210 in the computing submodule 21 accesses the computing unit 210 in the computing submodule 22

[0255] The computing unit 210 in the computing submodule 21 sends a data access request to the switching routing computing unit 402 through the routing interface MFR. The switching routing computing unit 402 parses the data access request to obtain the target computing unit address and determines whether the target computing unit address is in the same in-memory computing module (in this case, it is determined to be in the same in-memory computing module).

[0256] Afterwards, the switching routing calculation unit 402 queries the target computing unit address in the routing information stored therein about the computing units and the in-memory computing modules, determines the port information corresponding to the target computing unit address, and then controls the opening and closing of the cross switch unit 403, accessing the routing unit 211 of the computing sub-module 22 and the computing unit 210 connected to the routing unit 211 of the computing sub-module 22 through the routing interface MFR.

[0257] Scenario 5: The computing unit 210 in the computing submodule 21 accesses the computing unit 210 in another in-memory computing module

[0258] The computing unit 210 in the computing submodule 21 sends a data access request to the switching routing computing unit 402 through the routing interface MFR. The switching routing computing unit 402 parses the data access request to obtain the target computing unit address and determines whether the target computing unit address is in the same in-memory computing module (in this case, it is determined that the target computing unit address is not in the same in-memory computing module).

[0259] Afterwards, the switching routing calculation unit 402 controls the opening and closing of the crossbar switch unit 403 and sends the parsed data access request to another in-memory calculation module through the routing interface MFR.

[0260] Another in-memory computing module performs the operations in the above scenario (IV), and then accesses the routing unit 211 of the other in-memory computing module and the computing unit 210 connected to the routing unit 211 of the other in-memory computing module through the routing interface MFR.

[0261] Figure 7 It is a flowchart of a method for constructing an in-memory computing module according to one embodiment of the present invention.

[0262] The method for constructing the in-memory computing module includes the following steps:

[0263] Step S701: Arrange at least two computing sub-modules to be stacked in sequence in one direction.

[0264] Step S702 : Arrange each computing sub-module to be connected to its adjacent computing sub-module, wherein each computing sub-module includes at least one computing unit 210 and a plurality of memory units 212 .

[0265] Step S703: arranging the at least two computing sub-modules in the same chip.

[0266] Figure 8 is a schematic diagram of an in-memory computing network according to one embodiment of the present invention.

[0267] Figure 8 The in-memory computing network shown in the figure includes: a plurality of in-memory computing modules 80, which are a plurality of in-memory computing modules as described above, and the plurality of in-memory computing modules 80 are connected through the routing units of the plurality of in-memory computing modules.

[0268] By interconnecting and topologically combining in-memory computing modules, high data bandwidth and high-performance computing requirements can be achieved.

[0269] Figure 8 It should be understood that the multiple in-memory computing modules 80 can also be connected into a bus, star, ring, tree or mixed topology.

[0270] In the present invention, the multiple in-memory computing modules are connected via a routing unit via metal wires 801. The metal wire connections here are conventional metal wire connections used in two-dimensional connections.

[0271] The present invention is based on an embodiment in which a computing submodule includes a single routing unit 211. However, the present invention is not limited to a single routing unit and may include more than one routing unit. In the case of including more than one routing unit, these routing units can be connected via a routing interface MFR (similar to the operation between the routing interface MFR of one in-memory computing module and the routing interface MFR of another in-memory computing module) to form an overall routing unit.

[0272] The overall routing unit presents the same Figure 2 The same functionality as the single routing unit 210 shown. Figure 2 The difference between the single routing unit 210 in the overall routing unit is that each routing unit in the overall routing unit is not connected to each memory unit in the memory sub-module.

[0273] For example, if there are three routing units in the computing submodule, these two routing units constitute the overall routing unit. Figure 2 There are three routing units in the computing submodule 21: routing unit A, routing unit B and routing unit C. Routing unit A is connected to the three memory units 212 on the left and the related computing units 210, routing unit B is connected to the two middle memory units 212 and the related computing units 210, and routing unit C is connected to the three memory units 212 on the right and the related computing units 210.

[0274] If the computing unit 210 connected to the routing unit A needs to access any of the two memory units in the middle, it needs to connect through the routing interface MFR between the routing unit A and the routing unit B to access it, and the operation is similar to the above situation (three). The details are as follows:

[0275] In router A, computing unit 210, connected to routing unit A, sends a data access request to switching routing computing unit 402 via routing unit A's routing interface MFR. Switching routing computing unit 402 parses the data access request to obtain the target memory address and determines whether the target memory address is within router A's addressing range (in this case, it is determined to be not within router A's addressing range). Switching routing computing unit 402 then controls the opening and closing of crossbar switch unit 403 and sends the parsed data access request to router B via routing interface MFR, allowing router B to access either of the two intermediate memory units.

[0276] It should be noted that the above-mentioned embodiments illustrate rather than limit the invention, and that those skilled in the art will be able to design many alternative embodiments without departing from the scope of the appended claims.It should be understood that the scope of the invention is defined by the claims.

Claims

1. An in-memory computing module, characterized in that: The in-memory computing module includes: At least two computing sub-modules, the at least two computing sub-modules are stacked sequentially in one direction, each computing sub-module is connected to an adjacent computing sub-module, and each computing sub-module includes at least one computing unit, multiple memory units, and at least one routing unit; The at least two computing submodules are located in the same chip; and Wherein, in each computing submodule, the at least one computing unit and the at least one routing unit are located at the center of the corresponding computing submodule, and the multiple memory units are distributed around the at least one computing unit and the at least one routing unit in the corresponding computing submodule; wherein the at least one routing unit is arranged on one side of the at least one computing unit or around the at least one computing unit; In each computing sub-module, the at least one computing unit accesses the plurality of memory units via the at least one routing unit.

2. The in-memory computing module according to claim 1, wherein: Each computing submodule includes: a computing unit; Multiple memory units; a routing unit, the routing unit being connected to the computing unit, the routing unit being connected to each of the plurality of memory units, the routing unit being connected to a routing unit of at least another computing submodule of the in-memory computing module in which the routing unit is located, and the routing unit being connected to a routing unit of at least another computing submodule of at least another in-memory computing module; The routing unit is configured to execute an access by a first computing unit of the computing sub-module where the routing unit is located to a first memory unit of the computing sub-module where the routing unit is located, an access by a second computing unit or a second memory unit of at least another computing sub-module of the in-memory computing module where the routing unit is located, or an access by a third computing unit or a third memory unit of at least another computing sub-module of at least another in-memory computing module.

3. The in-memory computing module according to claim 2, wherein: The routing unit includes: a routing interface connecting the routing unit to a routing unit of at least another computing submodule of the in-memory computing module in which the routing unit is located, and / or connecting the routing unit to a routing unit of at least another computing submodule of at least another in-memory computing module; A memory control interface connects the routing unit to each of the plurality of memory units.

4. The in-memory computing module according to claim 3, characterized in that: The routing unit also includes: a crossbar switch unit; a switching routing calculation unit, the switching routing calculation unit being connected to the routing interface and the crossbar switch unit, and the switching routing calculation unit storing at least routing information about the in-memory computing module where the routing unit is located and the computing unit of the computing submodule where the routing unit is located, the switching routing calculation unit parsing received data access requests and controlling switching of the crossbar switch unit based on the parsed data access request information; A memory control unit is connected to the crossbar switch unit and the memory control interface, and the memory control unit at least stores routing information about the multiple memory units. In response to the crossbar switch unit switching to the memory control unit, the memory control unit performs a secondary analysis on the parsed data access request received from the crossbar switch unit to determine the target memory unit, and accesses the target memory unit via the memory control interface.

5. The in-memory computing module according to any one of claims 2 to 4, characterized in that: The computing unit directly accesses at least one memory unit via the routing unit.

6. The in-memory computing module according to claim 5, characterized in that: The computing unit accesses at least one further memory unit indirectly via the routing unit.

7. The in-memory computing module according to any one of claims 2 to 4, characterized in that: The routing unit is connected to the routing unit of at least another computing submodule of the in-memory computing module where the routing unit is located by bonding.

8. The in-memory computing module according to claim 7, characterized in that: The total data bit width of the connection between the routing unit and the routing unit of at least another computing submodule of the in-memory computing module where the routing unit is located is n times the data bit width of the computing unit, where n is a positive integer.

9. The in-memory computing module according to any one of claims 2 to 4, characterized in that: The number of the plurality of memory units is determined at least according to the data bit width of the calculation unit and the data bit width of a single memory unit.

10. The in-memory computing module according to any one of claims 2 to 4, characterized in that: In each computing sub-module, the computing unit, the plurality of memory units, and the routing unit are located in the same position in the corresponding computing sub-module.

11. The in-memory computing module according to claim 10, wherein: The computing submodules in the at least two computing submodules are identical to each other.

12. The in-memory computing module according to claim 1, wherein: Each computing submodule includes: at least two computing units; Multiple memory units; at least two routing units, each routing unit connected to at least one computing unit, and each routing unit connected to at least one memory unit; The at least two routing units are connected to each other to form an overall routing unit, the overall routing unit is connected to each of the plurality of memory units, the overall routing unit is connected to the overall routing unit of at least another computing sub-module of the in-memory computing module where the overall routing unit is located, and the overall routing unit is connected to the overall routing unit of at least another computing sub-module of at least another in-memory computing module; The overall routing unit is configured to execute an access by the first computing unit of the computing sub-module where the overall routing unit is located to the first memory unit or the second computing unit of the computing sub-module where the overall routing unit is located, an access by the second memory unit or the third computing unit of at least another computing sub-module of the in-memory computing module where the overall routing unit is located, or an access by the third memory unit or the fourth computing unit of at least another computing sub-module of at least another in-memory computing module.

13. The in-memory computing module according to claim 12, wherein: Each of the at least two routing units includes: a routing interface, the routing interface connecting the routing unit to at least another routing unit of the computing submodule in which the routing unit is located, and / or connecting the routing unit to a routing unit of at least another computing submodule of the in-memory computing module in which the routing unit is located, and / or connecting the routing unit to a routing unit of at least another computing submodule of at least another in-memory computing module; A memory control interface connects the routing unit to at least one memory unit of the plurality of memory units.

14. The in-memory computing module according to claim 13, wherein: Each of the at least two routing units further includes: a crossbar switch unit; a switching routing calculation unit, the switching routing calculation unit being connected to the routing interface and the crossbar switch unit, and the switching routing calculation unit storing at least routing information about an in-memory computing module in which the routing unit is located and at least one computing unit of a computing submodule in which the routing unit is located, the switching routing calculation unit parsing a received data access request and controlling switching of the crossbar switch unit based on the parsed data access request information; A memory control unit is connected to the crossbar switch unit and the memory control interface, and the memory control unit at least stores routing information about at least one memory unit among the multiple memory units. In response to the crossbar switch unit switching to the memory control unit, the memory control unit performs a secondary analysis on the parsed data access request received from the crossbar switch unit to determine the target memory unit, and accesses the target memory unit via the memory control interface.

15. The in-memory computing module according to any one of claims 12 to 14, characterized in that: Each computing unit directly accesses at least one memory unit via the global routing unit.

16. The in-memory computing module according to claim 13, wherein: Each computing unit indirectly accesses at least one other memory unit via the global routing unit.

17. The in-memory computing module according to any one of claims 12 to 14, characterized in that: The overall routing unit is connected to the overall routing unit of at least another computing submodule of the in-memory computing module where the overall routing unit is located by bonding.

18. The in-memory computing module according to claim 17, wherein: The total data bit width of the connection between the overall routing unit and the overall routing unit of at least another computing submodule of the in-memory computing module where the overall routing unit is located is n times the data bit width of the computing unit, where n is a positive integer.

19. The in-memory computing module according to any one of claims 12 to 14, characterized in that: The number of the plurality of memory units is determined at least according to the data bit widths of the at least two computing units and the data bit width of a single memory unit.

20. The in-memory computing module according to any one of claims 12 to 14, characterized in that: In each computing sub-module, the positions of the at least two computing units, the plurality of memory units, and the overall routing unit in the corresponding computing sub-module are the same.

21. The in-memory computing module according to claim 20, wherein: The computing submodules in the at least two computing submodules are identical to each other.

22. The in-memory computing module according to any one of claims 2-4 and 12-14, characterized in that: The memory unit includes a dynamic random access memory, and the computing unit includes a central processing unit.

23. The in-memory computing module according to any one of claims 2-4 and 12-14, characterized in that: The at least two computing submodules are two computing submodules.

24. The in-memory computing module according to any one of claims 2-4 and 12-14, characterized in that: The storage capacity of this memory unit is customizable.

25. An in-memory computing method, used in the in-memory computing module according to claim 6, characterized in that: The in-memory calculation method includes: The routing unit receives a data access request, which is issued by the first computing unit and includes at least an address of a target memory unit; and The routing unit parses the data access request, obtains access data from the target memory unit and forwards the access data to the first computing unit; The target memory unit includes a memory unit of a computing submodule where the first computing unit is located, a memory unit of another computing submodule, and a memory unit of another computing submodule of another in-memory computing module.

26. The in-memory computing method according to claim 25, wherein: The in-memory calculation method further includes: After the routing unit parses the data access request and before obtaining access data from the target memory unit, the routing unit connected to the first computing unit determines whether the first computing unit can directly access the target memory unit via the routing unit connected to the first computing unit; When the first computing unit can directly access the target memory unit via the routing unit connected to the first computing unit, the routing unit connected to the first computing unit directly obtains access data from the target memory unit and forwards the access data to the first computing unit; and When the first computing unit cannot directly access the target memory unit via the routing unit connected to the first computing unit, the routing unit connected to the first computing unit forwards the parsed data access request to the second computing unit, and obtains access data from the target memory unit via the second computing unit and forwards the access data to the first computing unit.

27. The in-memory computing method according to claim 26, wherein: The in-memory calculation method further includes: When the first computing unit cannot directly access the target memory unit via the routing unit connected to the first computing unit and before the routing unit connected to the first computing unit forwards the parsed data access request to the second computing unit, the routing unit connected to the first computing unit determines whether the target memory unit and the first computing unit are located in the same in-memory computing module; When the target memory unit and the first computing unit are located in the same in-memory computing module, the routing unit connected to the first computing unit forwards the parsed data access request to a routing unit of another computing sub-module connected to the routing unit connected to the first computing unit and to a second computing unit connected to the routing unit of the another computing sub-module, and obtains access data from the target memory unit via the second computing unit and forwards the access data to the first computing unit, wherein the second computing unit can directly access the target memory unit via the routing unit of the another computing sub-module; When the target memory unit and the first computing unit are not located in the same in-memory computing module, the routing unit connected to the first computing unit forwards the parsed data access request to the routing unit of another computing sub-module of another in-memory computing module connected to the routing unit connected to the first computing unit and forwards it to the second computing unit connected to the routing unit of another computing sub-module of the other in-memory computing module, and obtains access data from the target memory unit via the second computing unit and forwards the access data to the first computing unit, wherein the second computing unit can directly access the target memory unit via the routing unit of another computing sub-module of the other in-memory computing module.

28. An in-memory computing method, the in-memory computing method being used in the in-memory computing module according to any one of claims 2-11 and 22-24, characterized in that: The in-memory calculation method includes: The routing unit receives a data access request, which is issued by the first computing unit and includes at least an address of a target computing unit; and The routing unit parses the data access request, obtains access data from the target computing unit and forwards the access data to the first computing unit; The target computing unit includes a computing unit of another computing submodule of the in-memory computing module and a computing unit of another computing submodule of another in-memory computing module.

29. The in-memory computing method according to claim 28, wherein: The in-memory calculation method further includes: After the routing unit parses the data access request and before obtaining access data from the target computing unit, the routing unit connected to the first computing unit determines whether the target computing unit and the first computing unit are located in the same in-memory computing module; When the target computing unit and the first computing unit are located in the same in-memory computing module, the routing unit connected to the first computing unit forwards the parsed data access request to a routing unit of another computing submodule connected to the routing unit connected to the first computing unit, and obtains access data from the target computing unit via the routing unit of the other computing submodule and forwards the access data to the first computing unit. When the target computing unit and the first computing unit are not located in the same in-memory computing module, the routing unit connected to the first computing unit forwards the parsed data access request to the routing unit of another computing sub-module of another in-memory computing module connected to the routing unit to which the first computing unit is connected, and obtains access data from the target computing unit via the routing unit of another computing sub-module of the other in-memory computing module and forwards the access data to the first computing unit.

30. An in-memory computing method, the in-memory computing method being used in the in-memory computing module according to claim 16, characterized in that: The in-memory calculation method includes: The overall routing unit receives a data access request, where the data access request is issued by the first computing unit and includes at least an address of a target memory unit; and The overall routing unit parses the data access request, obtains access data from the target memory unit and forwards the access data to the first computing unit; The target memory unit includes a memory unit of a computing submodule where the first computing unit is located, a memory unit of another computing submodule, and a memory unit of another computing submodule of another in-memory computing module.

31. The in-memory computing method according to claim 30, wherein: The in-memory calculation method further includes: After the overall routing unit parses the data access request and before obtaining access data from the target memory unit, the overall routing unit connected to the first computing unit determines whether the first computing unit can directly access the target memory unit via the overall routing unit connected to the first computing unit; When the first computing unit can directly access the target memory unit via the overall routing unit connected to the first computing unit, the overall routing unit connected to the first computing unit directly obtains access data from the target memory unit and forwards the access data to the first computing unit; and When the first computing unit cannot directly access the target memory unit via the overall routing unit connected to the first computing unit, the overall routing unit connected to the first computing unit forwards the parsed data access request to the second computing unit, and obtains access data from the target memory unit via the second computing unit and forwards the access data to the first computing unit.

32. The in-memory computing method according to claim 31, wherein: The in-memory calculation method further includes: When the first computing unit cannot directly access the target memory unit via the overall routing unit connected to the first computing unit and before the overall routing unit connected to the first computing unit forwards the parsed data access request to the second computing unit, the overall routing unit connected to the first computing unit determines whether the target memory unit and the first computing unit are located in the same in-memory computing module; When the target memory unit and the first computing unit are located in the same in-memory computing module, the overall routing unit connected to the first computing unit forwards the parsed data access request to the overall routing unit of another computing submodule connected to the overall routing unit connected to the first computing unit and to a second computing unit connected to the overall routing unit of the another computing submodule, and obtains access data from the target memory unit via the second computing unit and forwards the access data to the first computing unit, wherein the second computing unit can directly access the target memory unit via the overall routing unit of the another computing submodule; When the target memory unit and the first computing unit are not located in the same in-memory computing module, the overall routing unit connected to the first computing unit forwards the parsed data access request to the overall routing unit of another computing sub-module of another in-memory computing module connected to the overall routing unit connected to the first computing unit and forwards it to the second computing unit connected to the overall routing unit of another computing sub-module of the other in-memory computing module, and obtains access data from the target memory unit via the second computing unit and forwards the access data to the first computing unit, wherein the second computing unit can directly access the target memory unit via the overall routing unit of another computing sub-module of the other in-memory computing module.

33. An in-memory computing method, the in-memory computing method being used in the in-memory computing module according to any one of claims 12 to 24, characterized in that: The in-memory calculation method includes: The overall routing unit receives a data access request, where the data access request is issued by the first computing unit and includes at least an address of a target computing unit; and The overall routing unit parses the data access request, obtains access data from the target computing unit and forwards the access data to the first computing unit; The target computing unit includes a computing unit of another computing submodule of the in-memory computing module and a computing unit of another computing submodule of another in-memory computing module.

34. The in-memory computing method according to claim 33, wherein: The in-memory calculation method further includes: After the overall routing unit parses the data access request and before obtaining access data from the target computing unit, the overall routing unit connected to the first computing unit determines whether the target computing unit and the first computing unit are located in the same in-memory computing module; When the target computing unit and the first computing unit are located in the same in-memory computing module, the overall routing unit connected to the first computing unit forwards the parsed data access request to a routing unit of another computing submodule connected to the overall routing unit to which the first computing unit is connected, and obtains access data from the target computing unit via the overall routing unit of the other computing submodule and forwards the access data to the first computing unit. When the target computing unit and the first computing unit are not located in the same in-memory computing module, the overall routing unit connected to the first computing unit forwards the parsed data access request to the overall routing unit of another in-memory computing module connected to the overall routing unit to which the first computing unit is connected, and obtains access data from the target computing unit via the overall routing unit of the other in-memory computing module and forwards the access data to the first computing unit.

35. An in-memory computing network, characterized in that The in-memory computing network includes: Multiple in-memory computing modules, wherein the multiple in-memory computing modules are multiple in-memory computing modules according to any one of claims 1-24, and the multiple in-memory computing modules are connected through the routing units of the multiple in-memory computing modules.

36. The in-memory computing network according to claim 35, wherein: The multiple in-memory computing modules are connected into bus, star, ring, tree, mesh and hybrid topologies.

37. The in-memory computing network according to claim 35 or 36, characterized in that: The multiple in-memory computing modules are connected via metal lines through a routing unit.

38. A method for constructing an in-memory computing module, characterized in that: The construction method includes: Arrange at least two computing submodules to be stacked sequentially in one direction; Arranging each computing submodule to be connected to its adjacent computing submodule, wherein each computing submodule includes at least one computing unit, a plurality of memory units, and at least one routing unit; wherein the at least two computing submodules are arranged in the same chip; The construction method further comprises: In each computing submodule, the at least one computing unit and the at least one routing unit are arranged at the center of the corresponding computing submodule, and the plurality of memory units are arranged to be distributed around the at least one computing unit and the at least one routing unit in the corresponding computing submodule; Arranging the at least one routing unit on one side of the at least one computing unit, or arranging the at least one routing unit around the at least one computing unit; Wherein, in each computing sub-module, the at least one computing unit accesses the plurality of memory units via the at least one routing unit.

39. The construction method according to claim 38, characterized in that Each computing submodule includes: a computing unit; multiple memory units; a routing unit; The construction method also includes: The routing unit is connected to the computing unit, and the routing unit is connected to each of the multiple memory units, and the routing unit is connected to the routing unit of at least another computing sub-module of the in-memory computing module where the routing unit is located, and the routing unit is connected to the routing unit of at least another computing sub-module of at least another in-memory computing module; and the routing unit is configured to execute access by the first computing unit of the computing sub-module where the routing unit is located to the first memory unit of the computing sub-module where the routing unit is located, access to the second computing unit or the second memory unit of at least another computing sub-module of the in-memory computing module where the routing unit is located, or access to the third computing unit or the third memory unit of at least another computing sub-module of at least another in-memory computing module.

40. The construction method according to claim 39, characterized in that The construction method also includes: connecting the routing interface of the routing unit to the routing unit of at least another computing submodule of the in-memory computing module where the routing unit is located, and / or connecting the routing unit to the routing unit of at least another computing submodule of at least another in-memory computing module; The memory control interface of the routing unit is connected to each memory unit of the plurality of memory units.

41. The construction method according to claim 40, characterized in that The construction method also includes: connecting a switching routing calculation unit of the routing unit to the routing interface and a crossbar switch unit, storing at least routing information about the in-memory computing module where the routing unit is located and the computing unit of the computing submodule where the routing unit is located in the switching routing calculation unit, and configuring the switching routing calculation unit to parse a received data access request and control switching of the crossbar switch unit based on the parsed data access request information; The memory control unit of the routing unit is connected to the crossbar switch unit and the memory control interface, and at least routing information about the multiple memory units is stored in the memory control unit. The memory control unit is also configured to perform a secondary analysis on the parsed data access request received from the crossbar switch unit in response to the crossbar switch unit switching to the memory control unit to determine the target memory unit, and access the target memory unit via the memory control interface.

42. The construction method according to any one of claims 39 to 41, characterized in that: The construction method also includes: The computing unit is configured to access at least one memory unit directly via the routing unit.

43. The construction method according to claim 42, characterized in that The construction method also includes: The computing unit is configured to access at least one other memory unit indirectly via the routing unit.

44. The construction method according to any one of claims 39 to 41, characterized in that: The construction method also includes: The routing unit is connected to the routing unit of at least another computing submodule of the in-memory computing module where the routing unit is located by bonding.

45. The construction method according to claim 44, characterized in that The construction method also includes: The total data bit width of the connection between the routing unit and the routing unit of at least another computing submodule of the in-memory computing module where the routing unit is located is set to n times the data bit width of the computing unit, where n is a positive integer.

46. ​​The construction method according to any one of claims 39 to 41, characterized in that: The construction method also includes: The number of the plurality of memory units is determined at least according to the data bit width of the calculation unit and the data bit width of a single memory unit.

47. The construction method according to any one of claims 39 to 41, characterized in that The construction method also includes: In each computing sub-module, the positions of the computing unit, the multiple memory units, and the routing unit in the corresponding computing sub-module are set to be the same.

48. The construction method according to claim 47, characterized in that The construction method also includes: The respective computing submodules in the at least two computing submodules are configured to be identical to each other.

49. The construction method according to claim 38, characterized in that Each computing submodule includes: at least two computing units; a plurality of memory units; at least two routing units, wherein each routing unit is connected to at least one computing unit, and each routing unit is connected to at least one memory unit; The construction method also includes: The at least two routing units are connected to each other to form an overall routing unit, and the overall routing unit is connected to each of the plurality of memory units, the overall routing unit is connected to the overall routing unit of at least another computing sub-module of the in-memory computing module where the overall routing unit is located, and the overall routing unit is connected to the overall routing unit of at least another computing sub-module of at least another in-memory computing module; and the overall routing unit is configured to execute access by a first computing unit of the computing sub-module where the overall routing unit is located to a first memory unit or a second computing unit of the computing sub-module where the overall routing unit is located, access by a second memory unit or a third computing unit of at least another computing sub-module of the in-memory computing module where the overall routing unit is located, or access by a third memory unit or a fourth computing unit of at least another computing sub-module of at least another in-memory computing module.

50. The construction method according to claim 49, characterized in that The construction method also includes: connecting a routing interface of each routing unit of the at least two routing units to at least another routing unit of the computing submodule in which the routing unit is located, and / or to a routing unit of at least another computing submodule of the in-memory computing module in which the routing unit is located, and / or to a routing unit of at least another computing submodule of at least another in-memory computing module; The memory control interface of each routing unit of the at least two routing units is connected to at least one memory unit of the plurality of memory units.

51. The construction method according to claim 50, characterized in that The construction method also includes: connecting a switching routing calculation unit of each of the at least two routing units to the routing interface and a crossbar switch unit, storing routing information of at least one computing unit of an in-memory computing module and a computing submodule of the routing unit in the switching routing calculation unit, and configuring the switching routing calculation unit to parse a received data access request and control switching of the crossbar switch unit based on the parsed data access request information; The memory control unit of each routing unit of the at least two routing units is connected to the crossbar switch unit and the memory control interface, and at least routing information about at least one memory unit among the multiple memory units is stored in the memory control unit. The memory control unit is also configured to, in response to the crossbar switch unit switching to the memory control unit, perform a secondary analysis on the parsed data access request received from the crossbar switch unit to determine a target memory unit, and access the target memory unit through the memory control interface.

52. The construction method according to any one of claims 49 to 51, characterized in that: The construction method also includes: Each computing unit is configured to directly access at least one memory unit via the global routing unit.

53. The construction method according to claim 50, characterized in that The construction method also includes: Each computing unit is configured to access at least one other memory unit indirectly via the global routing unit.

54. The construction method according to any one of claims 49 to 51, characterized in that: The construction method also includes: The overall routing unit is connected to the overall routing unit of at least another computing submodule of the in-memory computing module where the overall routing unit is located by bonding.

55. The construction method according to claim 54, characterized in that The construction method also includes: The total data bit width of the connection between the overall routing unit and the overall routing unit of at least another computing submodule of the in-memory computing module where the overall routing unit is located is set to n times the data bit width of the computing unit, where n is a positive integer.

56. The construction method according to any one of claims 49 to 51, characterized in that: The construction method also includes: The number of the plurality of memory units is determined at least according to the data bit width of the calculation unit and the data bit width of a single memory unit.

57. The construction method according to any one of claims 49 to 51, characterized in that: The construction method also includes: In each computing submodule, positions of the at least two computing units, the plurality of memory units, and the overall routing unit in the corresponding computing submodule are set to be the same.

58. The construction method according to claim 57, characterized in that The construction method also includes: The respective computing submodules in the at least two computing submodules are configured to be identical to each other.

59. The construction method according to any one of claims 39-41 and 49-51, characterized in that The construction method also includes: The memory unit is configured to include a dynamic random access memory, and the computing unit is configured to include a central processing unit.

60. The construction method according to any one of claims 39-41 and 49-51, characterized in that The construction method also includes: The at least two computing submodules are configured as two computing submodules.

61. The construction method according to any one of claims 39 to 41 and 49 to 51, wherein: The construction method also includes: The storage capacity of the memory unit is set to be customizable.

62. A method for constructing an in-memory computing network, characterized in that: The construction method includes: Multiple in-memory computing modules are connected through routing units of the multiple in-memory computing modules, wherein the multiple in-memory computing modules are multiple in-memory computing modules according to any one of claims 1-24.

63. The construction method according to claim 62, characterized in that The construction method also includes: The multiple in-memory computing modules are connected into bus, star, ring, tree, mesh and hybrid topologies.

64. The construction method according to claim 62 or 63, characterized in that: The construction method also includes: The plurality of in-memory computing modules are connected via metal lines through a routing unit.

Citation Information

Patent Citations

  • Three-dimensional multiprocessor system chip

    CN101145147A

  • Near data stream computing acceleration array based on RISC-V

    CN111159094A

  • Chip having extensible memory

    WO2018058430A1