Near memory computing chip and end side equipment

By using a three-dimensional packaging process to stack interface control chips, memory chips and logic chips in the end-side equipment, and separating memory interfaces and control circuits through efficient interconnection technology, the high computing power and high bandwidth requirements of large models on the end-side equipment in the prior art are solved, and performance and energy efficiency are improved, while reducing costs.

CN120029968APending Publication Date: 2025-05-23HANG ZHOU NANO CORE CHIP ELECTRONIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510126380.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-23

Smart Images

  • Figure CN120029968A_ABST
    Figure CN120029968A_ABST
Patent Text Reader

Abstract

The invention relates to a near-memory computing chip and end side equipment. The near-memory computing chip comprises an interface control chip, a memory chip and a logic chip which are sequentially stacked by a three-dimensional packaging process, the interface control chip is configured to be used for being connected with an SoC system, a storage chip and a logic chip so that the SoC system can access the storage chip and control the logic chip. According to the near memory computing chip provided by the invention, the memory interface circuit and the control circuit are integrated into an independent chip, so that the size of the memory chip is reduced under a fixed memory capacity, meanwhile, the cost of a logic chip for iteration by using an advanced process is also reduced, and the cost of the near memory computing chip is reduced on the whole.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence chip technology, and in particular to a near-memory computing chip and a terminal device. Background Art

[0002] In the existing technology, with the rapid development of artificial intelligence (AI) technology, the application scenarios based on large-scale pre-trained models are constantly expanding, such as natural language processing, image recognition, and multimodal interaction. Large models have been widely used in both cloud and end-side devices due to their excellent performance. However, as the scale of model parameters continues to grow, migrating large models from the cloud to end-side devices (such as AI phones and AI PCs) faces many challenges.

[0003] like Figure 1 As shown in the figure, due to the limitations of power consumption, volume, and real-time response requirements, the computing and storage resources of the end-side devices are relatively limited compared to cloud servers. In particular, during the large model inference process, frequent memory accesses will result in a large amount of data being transferred between the processor and the memory. The current mainstream storage medium is DRAM, which is limited by data transmission bandwidth and access latency, becoming an important bottleneck restricting the performance of large model inference. In addition, frequent data transmission will significantly increase system energy consumption and have a negative impact on the battery life of mobile devices.

[0004] To solve the above problems, a common technical approach is to integrate the computing unit directly into the memory system, such as embedding a processor unit or computing engine inside the DRAM. This "near-memory computing" architecture effectively reduces the distance and frequency of data transmission and improves the efficiency of data processing. The core concept of storage and computing integration is to execute computing tasks on-site, that is, to complete the calculation directly at the location where the data is stored, thereby breaking through the bottleneck of the traditional "storage-computing" separation architecture to a certain extent, and significantly improving the performance and energy efficiency of large models running on end-side devices. The design and optimization of storage and computing integration architecture for end-side acceleration of large models has important research value and application prospects.

[0005] The existing near-storage architecture is as follows Figure 2As shown in the figure: In the end-side device 1, the computing logic is integrated on the DRAM core particle, and part of the computing process is completed inside the DRAM core particle. The SoC accesses through the memory interface and uses the computing unit in the SoC to complete the remaining calculations. Since some operations are implemented on the DRAM core particle, the data transmission bandwidth requirement of the SoC is reduced. However, due to factors such as DRAM storage density, refresh time, and leakage, the process used by DRAM is relatively backward compared to the computing chip. The integration of large-scale computing logic on the DRAM process seriously affects the original DRAM storage density, which limits the number of computing units that can be integrated and the wiring space for DRAM memory access, and cannot meet the high computing power and high bandwidth requirements of large models. In terms of energy efficiency, the computing logic is tightly coupled with the DRAM storage unit in the process, so that the computing logic cannot play its computing energy efficiency advantage. At the same time, the DRAM process evolution speed lags behind the CMOS process, making this technology unsustainable.

[0006] like Figure 3 As shown, in the end-side device 2, DRAM core particles and computing logic are interconnected with high bandwidth through three-dimensional stacking. Compared with device 1, the process platform of computing logic and the DRAM storage core particle platform are decoupled. Therefore, the computing logic can use more advanced CMOS logic technology to achieve large-scale computing power integration. Under this device, the near-memory computing architecture can integrate large-scale computing power, so that the computing pressure of SoC is greatly reduced or the end-to-end application of large models is completed by the near-memory computing architecture. In this case, the bottleneck of large model inference performance caused by the low data transmission bandwidth of the existing memory interface is solved.

[0007] However, in the end-side device 2, the CMOS process is used in the computing logic, enjoying the influence of Moore's Law, and the computing density is greatly reduced with the evolution of the process. Under the advanced CMOS process, the area required for computing logic is much smaller than the area required for DRAM storage particles. However, due to the Wafer to Wafer packaging format, the area of ​​DRAM storage particles must be equal to the computing logic, resulting in a large area waste of computing logic, resulting in cost loss. As the process used becomes more advanced, the loss cost will also increase.

[0008] Therefore, it is necessary to improve the existing end-side devices and near-storage computing chips. Summary of the invention

[0009] In order to solve the above technical problems, the present invention provides a near-memory computing chip, including an interface control chip, a memory chip and a logic chip stacked in sequence using a three-dimensional packaging process; the interface control chip is configured to be connected to a SoC system, and to be connected to the memory chip and the logic chip, so that the SoC system can access the memory chip and control the logic chip.

[0010] The near-memory computing chip provided by the present invention integrates the memory interface circuit and the control circuit into an independent chip, so that the size of the memory chip can be reduced under a fixed storage capacity, while also reducing the cost of iterating the logic chip using advanced processes, thereby reducing the cost of the near-memory computing chip as a whole.

[0011] Optionally, the interface control chip includes a high-speed interface circuit, a memory control circuit and a logic chip control circuit; the high-speed interface circuit is configured to be connected to the SoC system to exchange information with the SoC system; the memory chip control circuit is connected to the memory chip and the high-speed interface circuit respectively, and can send control instructions to the memory chip based on information sent by the SoC system transmitted via the high-speed interface circuit; the logic chip control circuit is connected to the logic chip and the high-speed interface circuit respectively, and can send control instructions to the logic chip based on information sent by the SoC system transmitted via the high-speed interface circuit.

[0012] Optionally, the memory chip includes a plurality of memory groups, each of the memory groups is respectively connected to the memory chip control circuit and can interact with the memory chip control circuit for data in response to instructions sent by the memory chip control circuit.

[0013] Optionally, the logic chip includes a controller, a computing engine and a memory group controller; the controller is connected to the logic chip control circuit, and can receive a driving signal sent by the logic chip control circuit to drive the computing engine; the computing engine and the memory group controller are arranged in a one-to-one correspondence and are connected to the corresponding memory group controller, the memory group controller is connected to the corresponding memory group, and the computing engine can drive the memory group controller to control the corresponding memory group.

[0014] Optionally, the memory chip is designed as a memory bank arrangement structure of a dynamic random access memory with high density characteristics.

[0015] Optionally, the interface control chip is connected to the memory chip and / or the logic chip using RDL, TSV or hybrid bonding technology.

[0016] Optionally, the memory chip and the logic chip are connected by RDL, TSV or hybrid bonding process.

[0017] Optionally, the interface control chip, the memory chip and the logic chip are packaged with the substrate by means of metal balls. The interface control chip, the memory chip and the logic chip all include a substrate, a circuit layer and a metal layer stacked in sequence, and the metal layer is arranged toward a side close to the substrate. The metal layer of the logic chip is electrically connected to the metal layer of the interface control chip and the memory chip respectively through RDL, TSV or hybrid bonding processes. The metal layers of the interface control chip and the memory chip are electrically connected through RDL, TSV or hybrid bonding processes, and the interface control chip is electrically connected to the substrate through metal balls.

[0018] Optionally, the interface control chip, the memory chip and the logic chip are packaged with the substrate by Wire Bonding. The interface control chip, the memory chip and the logic chip all include a substrate, a circuit layer and a metal layer stacked in sequence. The metal layer of the memory chip and the logic chip is arranged toward a side close to the substrate, and the metal layer of the interface control chip is arranged toward a side away from the substrate. The logic chip is electrically connected to the memory chip and the interface control chip by RDL, TSV or hybrid bonding process. The interface control chip and the memory chip are electrically connected by RDL, TSV or hybrid bonding process, and the logic chip and the interface control chip are electrically connected to the substrate by Wire Bonding.

[0019] In order to achieve the above-mentioned purpose of the invention, the present application provides an end-side device, including the near memory computing chip and SoC system mentioned above, wherein the SoC system is connected to the near memory computing chip and can control the operation of the near memory computing chip through the interface control chip.

[0020] The near-memory computing chip provided by this application is a three-dimensional stacking package with at least three layers of chips. This architecture separates the memory interface from the control circuit, and uses multi-layer stacking technology to achieve high-bandwidth data transmission, while reducing the waste of logic chip area. By adopting TSV, hybrid bonding or RDL technology, efficient interconnection is achieved between the layers. This architecture effectively reduces the energy efficiency problems caused by the coupling of memory chips and logic chips, and reduces the waste of logic chip area under advanced processes, significantly improving the performance and energy efficiency of large models running on end-side devices, reducing chip manufacturing costs, and is suitable for large-model acceleration applications on end-side devices such as AI mobile phones and AIPCs. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1It is a structural diagram of the terminal side device of the storage and computing separation architecture in the prior art.

[0022] Figure 2 It is a structural diagram of another end-side device of the storage-computing separation architecture in the prior art.

[0023] Figure 3 It is a structural diagram of another end-side device of the storage-computing separation architecture in the prior art.

[0024] Figure 4 It is a schematic diagram of the structure of a near-memory computing chip provided in an embodiment of the present invention.

[0025] Figure 5 It is a schematic diagram of the packaging structure of a near-memory computing chip provided in an embodiment of the present invention.

[0026] Figure 6 It is a schematic diagram of the packaging structure of a near-memory computing chip provided in an embodiment of the present invention.

[0027] Figure 7 It is a schematic diagram of the structure of the terminal side device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0028] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0029] like Figure 4 As shown, this embodiment provides a near-memory computing chip, including an interface control chip 100, a memory chip 200 and a logic chip 300 stacked in sequence using a three-dimensional packaging process; the interface control chip 100 is configured to be connected to a SoC system, and to be connected to the memory chip 200 and the logic chip 300, so that the SoC system can access the memory chip 200 and control the logic chip 300.

[0030] The near-memory computing chip provided in this embodiment integrates the memory interface circuit and the control circuit into an independent chip, so that the size of the memory chip can be reduced under a fixed storage capacity, while also reducing the cost of iterating the logic chip using advanced processes, thereby reducing the cost of the near-memory computing chip as a whole.

[0031] Optionally, the interface control chip 100 includes a high-speed interface circuit 110, a memory control circuit 120 and a logic chip control circuit 130; the high-speed interface circuit 110 is configured to be connected to the SoC system to exchange information with the SoC system; the memory chip control circuit 120 is connected to the memory chip 200 and the high-speed interface circuit 110 respectively, and can send control instructions to the memory chip 200 based on the information sent by the SoC system transmitted via the high-speed interface circuit 110; the logic chip control circuit 130 is connected to the logic chip 300 and the high-speed interface circuit 110 respectively, and can send control instructions to the logic chip 300 based on the information sent by the SoC system transmitted via the high-speed interface circuit 110.

[0032] Optionally, the memory chip 200 includes a plurality of memory groups 210, each memory group 210 is connected to the memory chip control circuit 120, and can interact with the memory chip control circuit 120 in response to instructions sent by the memory chip control circuit 120. More specifically, each memory group 210 also includes a plurality of memory blocks 211.

[0033] Optionally, the logic chip 300 includes a controller 310, a computing engine 320 and a memory group controller 330; the controller 310 is connected to the logic chip control circuit 130, and can receive a driving signal sent by the logic chip control circuit 130 to drive the computing engine 320; the computing engine 320 and the memory group controller 330 are arranged in a one-to-one correspondence, the computing engine 320 is connected to the corresponding memory group controller 330, the memory group controller 330 is connected to the corresponding memory group 210, and the computing engine 320 can drive the memory group controller 330 to control the corresponding memory group 210.

[0034] In this embodiment, a separate computing engine 320 and a memory group controller 330 are provided for each memory group 210, so that each memory group can execute instructions of the logic chip 300 simultaneously or independently, thereby achieving high-bandwidth data transmission.

[0035] Optionally, the memory chip 200 is designed as a memory bank arrangement structure of a dynamic random access memory (DRAM) having a high density characteristic.

[0036] Optionally, the logic chip 300 is designed using a CMOS process.

[0037] Optionally, the interface control chip 100 is connected to the memory chip 200 and / or the logic chip 300 by using RDL, TSV or hybrid bonding technology.

[0038] Optionally, the memory chip 200 and the logic chip 300 are connected by RDL, TSV or hybrid bonding process.

[0039] Optionally, refer to Figure 5 The interface control chip 100, the memory chip 200 and the logic chip 300 are packaged with the substrate by means of metal balls. In this case, the interface control chip 100, the memory chip 200 and the logic chip 300 all include a substrate, a circuit layer and a metal layer stacked in sequence, the metal layer is arranged toward the side close to the substrate, the metal layer of the logic chip 300 is electrically connected to the metal layers of the interface control chip 100 and the memory chip 200 respectively through RDL, TSV or hybrid bonding processes, the metal layers of the interface control chip 100 and the memory chip 200 are electrically connected through RDL, TSV or hybrid bonding processes, and the interface control chip 100 is electrically connected to the substrate through metal balls.

[0040] Optionally, refer to Figure 6 The interface control chip 100, the memory chip 200 and the logic chip 300 are packaged with the substrate by Wire Bonding. In this case, the interface control chip 100, the memory chip 200 and the logic chip 300 all include a substrate, a circuit layer and a metal layer stacked in sequence. The metal layers of the memory chip 200 and the logic chip 300 are arranged toward the side close to the substrate, and the metal layer of the interface control chip 100 is arranged toward the side away from the substrate. The logic chip 300 is electrically connected to the memory chip 200 and the interface control chip 100 by RDL, TSV or hybrid bonding process. The interface control chip 100 and the memory chip 200 are electrically connected by RDL, TSV or hybrid bonding process. The logic chip 300 and the interface control chip 100 are electrically connected to the substrate by Wire Bonding.

[0041] Optionally, refer to Figure 7 This embodiment also provides a terminal side device, including a SoC system and a near memory computing chip provided in this embodiment. The SoC system is connected to the near memory computing chip and can control the operation of the near memory computing chip through an interface control chip 100.

[0042] So far, the technical solution of the present invention has been described in conjunction with the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to the above-mentioned specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.

Claims

1. A near memory computing chip, characterized in that: It includes an interface control chip, a memory chip and a logic chip stacked in sequence using a three-dimensional packaging process; the interface control chip is configured to be connected to a SoC system, and to be connected to the memory chip and the logic chip, so that the SoC system can access the memory chip and control the logic chip.

2. The near memory computing chip according to claim 1, characterized in that: The interface control chip includes a high-speed interface circuit, a memory control circuit and a logic chip control circuit; the high-speed interface circuit is configured to be connected to the SoC system to exchange information with the SoC system; the memory chip control circuit is connected to the memory chip and the high-speed interface circuit respectively, and can send control instructions to the memory chip based on the information sent by the SoC system transmitted via the high-speed interface circuit; the logic chip control circuit is connected to the logic chip and the high-speed interface circuit respectively, and can send control instructions to the logic chip based on the information sent by the SoC system transmitted via the high-speed interface circuit.

3. The near memory computing chip according to claim 2, characterized in that: The memory chip includes a plurality of memory groups, each of which is connected to the memory chip control circuit and can perform data interaction with the memory chip control circuit in response to instructions sent by the memory chip control circuit.

4. The near memory computing chip according to claim 3, characterized in that: The logic chip includes a controller, a computing engine and a memory group controller; the controller is connected to the logic chip control circuit, and can receive a driving signal sent by the logic chip control circuit to drive the computing engine; the computing engine and the memory group controller are arranged in a one-to-one correspondence and are connected to the corresponding memory group controller, the memory group controller is connected to the corresponding memory group, and the computing engine can drive the memory group controller to control the corresponding memory group.

5. The near memory computing chip according to any one of claims 1 to 4, characterized in that: The memory chip is designed as a memory bank arrangement structure of a dynamic random access memory with high density characteristics.

6. The near memory computing chip according to any one of claims 1 to 4, characterized in that: The interface control chip is connected to the memory chip and / or the logic chip by using RDL, TSV or hybrid bonding technology.

7. The near memory computing chip according to any one of claims 1 to 4, characterized in that: The memory chip and the logic chip are connected by RDL, TSV or hybrid bonding process.

8. The near memory computing chip according to any one of claims 1 to 4, characterized in that: The interface control chip, the memory chip and the logic chip are packaged with the substrate by means of metal balls. The interface control chip, the memory chip and the logic chip all include a substrate, a circuit layer and a metal layer stacked in sequence, and the metal layer is arranged toward a side close to the substrate. The metal layer of the logic chip is electrically connected to the metal layer of the interface control chip and the memory chip respectively through RDL, TSV or hybrid bonding processes. The metal layers of the interface control chip and the memory chip are electrically connected through RDL, TSV or hybrid bonding processes, and the interface control chip is electrically connected to the substrate through metal balls.

9. The near memory computing chip according to any one of claims 1 to 4, characterized in that: The interface control chip, the memory chip and the logic chip are packaged with the substrate by Wire Bonding. The interface control chip, the memory chip and the logic chip all include a substrate, a circuit layer and a metal layer stacked in sequence. The metal layer of the memory chip and the logic chip is arranged toward a side close to the substrate, and the metal layer of the interface control chip is arranged toward a side away from the substrate. The logic chip is electrically connected to the memory chip and the interface control chip by RDL, TSV or hybrid bonding process. The interface control chip and the memory chip are electrically connected by RDL, TSV or hybrid bonding process. The logic chip and the interface control chip are electrically connected to the substrate by Wire Bonding.

10. A terminal side device, characterized in that: It comprises the near memory computing chip and SoC system according to any one of claims 1 to 9, wherein the SoC system is connected to the near memory computing chip and can control the operation of the near memory computing chip through the interface control chip.