Edge end large language model inference system

The edge-end large language model inference system addresses the bandwidth limitations of NAND Flash storage by using vertically stacked NAND flash and DRAM chips with mixed bonding interconnects, enhancing read and transfer speeds to support real-time inference of large-scale models.

CN120316038APending Publication Date: 2025-07-15TSINGHUA UNIVERSITY
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510371653.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Currently, edge devices cannot effectively support real-time inference of large-scale language models, mainly due to the bottleneck of NAND Flash SSD's read performance, including insufficient external transmission bandwidth and internal reading bandwidth, resulting in the inability to meet the requirements of computing speed and accuracy.

Method used

Using hybrid bonded three-dimensional interconnection technology, NAND flash memory chips are separated and stacked with the CMOS chips, and connected to the main computing chip through the LPDDR interface. Combined with small array design and double buffer strategy, the power supply and reading methods of the flash memory array are optimized, internal reading bandwidth is improved, and speculative decoding algorithms are supported to improve computing efficiency.

Benefits of technology

It realizes that edge-end devices can run large-scale large-language model inference in real time, improves computing speed and accuracy, meets the needs of delay-sensitive application scenarios, and reduces hardware and system integration costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120316038A_ABST
    Figure CN120316038A_ABST
Patent Text Reader

Abstract

The invention discloses an edge end large language model inference system. The system comprises a main computing chip, an LPDDR interconnection resource and a vertically stacked and packaged storage chip stack, the storage chip stack comprises an NAND flash memory chip and a DRAM (Dynamic Random Access Memory) chip, and the NAND flash memory chip and the DRAM chip are paired and share an LPDDR interconnection resource; the NAND flash memory chip comprises a flash memory chip bare chip and a CMOS (Complementary Metal-Oxide-Semiconductor Transistor) chip bare chip which are three-dimensionally interconnected based on hybrid bonding; a flash memory array is arranged on the flash memory chip bare chip; a logic block is arranged on the CMOS chip bare chip; the main calculation chip reads a parameter matrix of a model full connection layer from the NAND flash memory chip and reads other parameters from the DRAM chip for calculation in a pre-filling stage of large language model reasoning; and the NAND flash memory chip performs full connection layer calculation in a decoding stage. The system provided by the invention can improve the external transmission and internal reading bandwidth.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of hardware design and computer data processing, and particularly to an edge-side large language model inference system. Background Art

[0002] This section aims to provide background or context for the embodiments of the present invention stated in the claims. The descriptions herein are not admitted to be prior art merely because they are included in this section.

[0003] With the leapfrog development of model quality, the applications of large language models (LLMs, Large Language Model, hereinafter referred to as large models) have become increasingly widespread. However, due to the high requirements of large language models for computing platforms, current large language model services are mainly deployed in the cloud. This brings many problems: (1) For users, cloud computing increases the risk of privacy leakage, and the long queuing time during peak service request periods and service failure problems caused by congestion will also greatly affect the user experience; (2) For service providers, the high acquisition and maintenance costs of inference infrastructure bring great pressure to capital turnover; (3) Looking to the future, unmanned systems and embodied intelligent systems are important potential application scenarios for large language models, and these systems are very likely to need to be deployed in extreme task conditions (such as exploration in extreme environments, disaster relief, etc.) and cannot communicate with the cloud.

[0004] Due to the problems listed above, the demand for edge-side large language model inference is increasing. However, even with relatively advanced algorithms such as model quantization, current edge-side devices (such as mobile phones, laptops, etc.) can still only support real-time inference of relatively small-scale large language models (<13B (Billion)), while larger-scale large language models can only be used in application scenarios where latency is not sensitive. This results in the current edge-side large language model inference system being difficult to provide a satisfactory user experience. On the one hand, most application scenarios of large language models are latency-sensitive; on the other hand, both the model parameter scale and inference accuracy have a crucial impact on model inference quality.

[0005] Currently, at the edge side, the bottleneck of large-scale large language model inference systems mainly lies in the read performance bottleneck of NAND Flash SSDs (NAND flash solid-state drives). Specifically, the size limitations of edge devices prevent them from accommodating enough traditional high-speed memory (i.e., DRAM, Dynamic Random Access Memory), and they cannot store the large number of parameters of large-scale large language models in traditional high-speed memory. Therefore, the parameters can only be stored in slow SSDs based on 3D NAND Flash, which completely fails to meet the read speed requirements of large-scale large language model inference. There are two main factors restricting the read speed of NAND Flash SSDs. First, the traditional data transfer interface between NAND Flash SSDs and computing chips (such as PCIe Gen4 x4) often has low bandwidth. This is because the main off-chip interconnect resources of computing chips are used to implement the interconnection with high-speed memory (DRAM), and more hardware interconnect resources cannot be allocated to NAND Flash SSDs, which are traditionally in the secondary storage position. Second, the internal read bandwidth of NAND Flash chips is also low. Traditionally, the design of NAND Flash chips is density-driven. This leads to their tendency to adopt larger storage array designs, resulting in higher read latency and lower read parallelism, and thus lower internal read bandwidth.

[0006] In summary, there is a need for a high-precision (FP16) edge-side real-time inference system for large language models (especially for large-scale large language models) to solve the problems of the external transfer bandwidth and internal read bandwidth bottlenecks of NAND Flash storage systems. Summary of the Invention

[0007] Embodiments of the present invention provide an edge-side large language model inference system to solve the problems of the external transfer bandwidth and internal read bandwidth bottlenecks of traditional NAND Flash storage systems during large language model inference. The system includes: a main computing chip, LPDDR interconnect resources connected to the main computing chip, and at least one vertically stacked package of storage chip stacks;

[0008] Each storage chip stack includes multiple NAND flash chips and multiple DRAM chips. One NAND flash chip and one DRAM chip exist in pairs and share the LPDDR memory channel of the LPDDR interconnect resources;

[0009] The NAND flash chip includes a flash chip die and a CMOS chip die based on hybrid bonding three-dimensional interconnection;

[0010] The flash memory chip bare die is provided with a plurality of flash memory arrays; the flash memory arrays are used to store parameter matrices required for fully connected layer calculations of a large language model;

[0011] The DRAM chip is used to store model parameters and intermediate calculation results other than the parameter matrix of the fully connected layer required for the fully connected layer calculation of the large language model;

[0012] The CMOS chip bare die is provided with a plurality of logic blocks, and one logic block corresponds to one flash memory array;

[0013] The main computing chip is used to read the parameter matrix from the NAND flash chip through the LPDDR memory channel and read the model parameters and intermediate calculation results from the DRAM chip during the pre-filling stage of the large language model inference, and perform calculations on the large language model. During the decoding stage of the large language model inference, the input data of the fully connected layer is placed on the NAND flash chip to perform calculations on the parts of the large language model other than the fully connected layer.

[0014] NAND flash memory chips are used to perform calculations on the fully connected layers of large language models through logic blocks during the decoding phase of large language model inference.

[0015] In the embodiment of the present invention, the flash memory chip bare die (including the flash memory array) and the CMOS chip bare die of three-dimensional interconnection by hybrid bonding can not only eliminate or reduce the plane area overhead occupied by the peripheral circuit of the existing flash memory array; more importantly, considering that the size of the existing flash memory array design is limited by area efficiency considerations and cannot be excessively reduced, the combined use of the hybrid bonding method provides a new opportunity to greatly reduce the area efficiency limitation, so that the size of the flash memory array can be further reduced, thereby further reducing the access delay on the basis of the existing small array design, and increasing the number of flash memory arrays that can theoretically work in parallel in the flash memory chip bare die, thereby solving the external transmission bandwidth and internal read bandwidth bottleneck problems of the traditional NAND Flash storage system. In addition, since the traditional SSD system is abandoned, the flash memory chip bare die designed based on this system can break through the pin limit specified by the traditional SSD, and then a large number of new power pins are added, thereby significantly increasing the power supply capacity of the external circuit to the NAND flash memory chip. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings:

[0017] Figure 1 Schematic diagram of the structure of the edge - side large language model inference system in the embodiments of the present invention;

[0018] Figure 2 Schematic diagram of the structure of the NAND flash chip in the embodiments of the present invention;

[0019] Figure 3 Schematic diagram of the physical architecture of the edge - side large language model inference system in the embodiments of the present invention;

[0020] Figure 4 Schematic diagram of the logical architecture of the edge - side large language model inference system in the embodiments of the present invention;

[0021] Figure 5 Schematic diagram of the execution flow of the edge - side large language model inference system for large language model inference at the edge in the embodiments of the present invention;

[0022] Figure 6 Schematic diagram of the layout and processing principle of the parameter matrix of the fully - connected layer in the embodiments of the present invention;

[0023] Figure 7 Schematic diagram of the control principle in the edge - side large language model inference system in the embodiments of the present invention. Detailed implementation manners

[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer and more understandable, the following further elaborates on the embodiments of the present invention with reference to the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but not to limit the present invention.

[0025] First, the concepts, terms, and variables involved in the embodiments of the present invention are explained.

[0026] SSD, Solid State Drive, solid - state drive;

[0027] FP16, 16 - bit Floating Point, 16 - bit floating - point;

[0028] GBps, Giga Bytes per second, gigabytes per second (in the computer field, "giga" often takes "2 to the 30th power");

[0029] DIMM, Dual Inline Memory Module, dual - in - line memory module (abbreviation: "memory stick");

[0030] PCIe, Peripheral Component Interconnect express, a widely adopted standard for interconnecting computer peripherals, namely, High-Speed Peripheral Interconnect (Standard);

[0031] Gen 4, Generation 4, the 4th generation;

[0032] 3D, 3 Dimension, 3D;

[0033] TSV, Through Silicon Via, Through-Silicon Via;

[0034] QKV, Query / Key / Value, Query / Key / Value;

[0035] GEMM, General Matrix Multiplication, General Matrix Multiplication;

[0036] GEMV, General Matrix-Vector Multiplication, General Matrix-Vector Multiplication;

[0037] ECC, Error Correction Code, Error Correction Code;

[0038] BCH ECC, Bose-Chaudhuri-Hocquenghem Error Correction Code, Bose-Chaudhuri-Hocquenghem Error Correction Code;

[0039] LPDDR, Low Power Double Data Rate, Low Power Double Data Rate (protocol);

[0040] SRAM, Static Random Access Memory, Static Random Access Memory;

[0041] b, bit, bit.

[0042] The system proposed in the embodiments of the present invention relates to the design technology of near-memory computing systems. Existing near-memory computing technologies mainly make trade-off designs for performance, cost, and capacity / storage density, which can be summarized into three categories: (1) high-performance, high-cost, low-capacity near-memory computing systems; (2) medium-performance, medium-low-cost, medium-capacity near-memory computing systems; (3) low-performance, medium-low-cost, high-density high-capacity near-storage computing systems. The core features of (1) and (2) are implemented based on DRAM (or emerging memory devices with a density higher than DRAM but much lower than 3D NAND Flash), and there is a large interconnection bandwidth (~100GBps) between them and the main computing chip; the core feature of (3) is based on 3D NAND Flash and is implemented based on the traditional NAND Flash SSD architecture.

[0043] (1) High performance, high cost, low capacity: Such as near-memory computing technology based on 3D DRAM stacking

[0044] This technology includes three categories: (a) three-dimensionally stacking the main computing logic chip and DRAM; (b) directly placing some lightweight computing logic on the DRAM layer of 3D DRAM (such as HBM); (c) a single-chip accelerator system formed by combining the two.

[0045] Cost: Due to the use of expensive advanced packaging and even directly implementing computing logic using DRAM technology, the production cost is relatively high.

[0046] Capacity: Due to the limitations of a single chip and the storage density limitations of DRAM devices themselves, the capacity is small and completely unable to meet the capacity requirements for large-scale large language model inference at the edge.

[0047] Performance: Using the ultra-high bandwidth of 3D stacking, the parallelism between DRAM banks, and the ultra-high interconnection bandwidth within the chip, such a system can achieve an extremely high near-memory access bandwidth. At the same time, since the main computing chip and DRAM are interconnected using an advanced packaging process with a relatively high bandwidth (such as 3D / 2.5D stacking interconnection), there is a relatively high interconnection bandwidth between them.

[0048] (2) Medium performance, medium-low cost, medium capacity: Such as near-memory computing technology based on DIMM buffer chips

[0049] In traditional computing systems, memory DRAM is organized in the form of DIMMs based on the DDR interface, which can densely integrate a relatively large number of DRAM chips to achieve medium or high capacities. However, generally speaking, the computing unit can only access part of the DRAM chipset each time, resulting in some DRAM chipsets being idle, and their bandwidth performance cannot be fully utilized. The DIMM near-memory system places some lightweight computing logic on the buffer chips of traditional DIMMs, enabling all DRAM chipsets to be accessed in parallel, thereby achieving a larger memory access bandwidth than the main computing chip through near-memory computing logic. In addition to the near-memory computing technology based on DIMM buffer chips, there are also near-memory computing technologies based on the CXL interface, etc., which can also be classified as near-memory computing technologies with medium performance, medium low cost, and medium capacity, and will not be elaborated here.

[0050] Cost: There is no need to modify DRAM. Only by placing the computing logic on the buffer chips can the implementation cost be greatly reduced. At the same time, since it can be compatible with traditional DDR systems, the system integration cost can be relatively low.

[0051] Capacity: Since it can be compatible with traditional large-scale DDR DIMM systems, a large number of DIMMs and DRAM chips can be integrated on a single server on a large scale, thus achieving medium or large capacities. However, this technology is actually still limited by the storage density of the DRAM device itself. In edge scenarios, even based on the DDR system, the number of DRAM chips that can be placed within the system is still significantly limited and cannot meet the capacity requirements for large-scale large language model edge inference.

[0052] Performance: Since the DIMM near-memory system essentially exploits the working parallelism of traditional DRAM chipsets, and (1) further fully utilizes the bank parallelism inside DRAM by changing the DRAM architecture, the potential performance of the DIMM near-memory system itself is lower than that of (1). However, due to the high-speed characteristics of the DRAM chips themselves, it still has good near-memory computing performance. At the same time, since these DIMMs are still connected to the main computing chip through a high-speed DDR interface, there is still a considerable data transmission bandwidth between the two.

[0053] (3) Low performance, medium low cost, high density and high capacity: Near-memory computing technology based on traditional SSD

[0054] Traditional systems utilize NAND Flash-based SSD systems to achieve high-density and high-capacity storage. In a near-storage computing system based on a traditional SSD, some lightweight computing units are integrated into the SSD controller chip to make full use of the internal channel bandwidth of the SSD; or some lightweight computing logic is integrated into the logic layer of the Flash chip to make full use of the internal bandwidth of the Flash chip. The SSD system is still interconnected with the main computing chip through a traditional interconnect interface with a relatively low bandwidth.

[0055] Cost: If the computing logic is integrated into the control chip, there is no need to modify the Flash chip, and the implementation cost is relatively low; even if it is added to the logic layer of the Flash chip. Since the current Flash chip itself already has an additional logic layer for placing peripheral circuits and the area is not fully utilized. Therefore, this design does not require additional packaging processes or additional chip area, so the cost remains relatively low. In addition, since this design is compatible with traditional SSD systems, the system integration cost is also relatively low.

[0056] Capacity: Since the Flash storage density itself is much higher than that of DRAM (in addition, Flash generally adopts multi-die packaging), the density and capacity are much higher than those of (1)(2). In the edge scenario, it can meet the capacity requirements for large-scale large language model inference.

[0057] Performance: Since the Flash chips of traditional SSDs generally adopt a large array design, their internal read bandwidth is relatively low, and the near-storage computing performance is also significantly limited; at the same time, since this near-storage computing technology still connects the near-storage computing system and the main computing chip through a traditional low-performance interconnect, the interconnection bandwidth between the two is also relatively low.

[0058] In an embodiment of the present invention, the large language model includes multiple layers of decoder modules. Each layer of decoder module includes an element generation module, a self-attention module, a projection module, a feed-forward module, and a feed-forward module. The element generation module is the QKV generation module, and the elements include Query / Key / Vector. Among them, the feed-forward module includes at least one feed-forward network and an activation module. The element generation sub-module, the projection module, and the feed-forward module all contain fully connected layers. In terms of edge-side inference, its computing and memory access costs are mainly concentrated in the Query / Key / Vector (QKV) element generation sub-module, the projection module, and the feed-forward module of the attention module, and their main components are all fully connected interconnections. One round of the inference process of the large language model includes a pre-fill stage and a decoding stage. The pre-fill stage is used to preprocess the user input. By inputting all input tokens into the model simultaneously (1 token is similar to 1 Chinese character or word, called a word unit), only one model calculation is required to complete. In this one model calculation, the calculation of the fully connected layer is mainly matrix multiplication (GEMM), which is a computationally intensive process and requires a relatively powerful computing unit to undertake. The decoding stage is used to generate the output of the large language model and requires a considerable number of model calculations. Without speculative decoding, each model calculation can only output one token, and the input for each calculation is the output of the previous time. This calculation will be repeated many times until the generation limit is reached or a special token indicating the end of the output is obtained.

[0059] An embodiment of the present invention proposes an inference system for an edge - side large - language model (especially for real - time large - scale large - language models) based on an LPDDR interface circuit and near - memory computing. Different from the near - memory computing technology based on DRAM, the system proposed in the embodiment of the present invention is based on NAND Flash chips (NAND flash memory chips); different from the existing near - memory computing systems based on NAND Flash SSDs, this system abandons the system organization form of SSDs and instead adopts an organization form similar to that of an LPDDR DRAM system to provide sufficient transmission bandwidth for data transmission between NAND flash memory chips and external computing chips. At the same time, a dual - buffer strategy and Flash - side ECC error - correction calculation are used to reduce the long read latency of Flash and the negative performance impact of coarse - grained ECC on the LPDDR read process. Aiming at the problem of low internal read bandwidth of traditional NAND flash memory chips, this system combines a small flash memory array design with a three - dimensional stacking design based on hybrid bonding technology and abandons the power supply design of traditional NAND flash memory chips, so that sufficient internal data read bandwidth can be generated inside the NAND flash memory chips, while ensuring the area efficiency of the flash memory array and solving the problem of power supply capacity. In addition, although this near - memory computing technology adopts a logic placement strategy similar to that of traditional near - memory computing technologies, the innovation of the system proposed in the embodiment of the present invention lies in the simultaneous combination of near - memory computing methods and support for speculative decoding algorithms, thereby further improving the inference speed of large - language models. Finally, in order to enable the parameter matrix arrangement of the large - language model in the NAND flash memory chip to simultaneously meet the efficient access requirements of the main computing chip and the near - memory computing logic, and thus avoid the storage waste caused by storing the same parameters in two different arrangement forms, the system proposed in the embodiment of the present invention designs a special parameter matrix arrangement format and access method for the large - language model. Ultimately, the system proposed in the embodiment of the present invention enables edge - side devices to run large - scale large - language model inference in real time (such as large - language models with more than 50 billion to 100 billion parameters).

[0060] Figure 1 FIG. 4 is a schematic structural diagram of the edge - side large - language model inference system in the embodiment of the present invention. The system includes: a main computing chip, LPDDR interconnection resources connected to the main computing chip, and at least one vertically stacked package of memory chip stacks;

[0061] Each memory chip stack includes a plurality of NAND flash memory chips and a plurality of DRAM chips. One NAND flash memory chip and one DRAM chip exist in pairs and share the LPDDR memory channel of the LPDDR interconnection resources;

[0062] The NAND flash memory chip includes a flash memory chip die and a CMOS chip die with three - dimensional interconnection based on hybrid bonding;

[0063] Multiple flash arrays are provided on the flash chip die; the flash arrays are used to store the parameter matrices required for the calculation of the fully connected layer of the large language model;

[0064] The DRAM chip is used to store the model parameters and intermediate calculation results required for the calculation of all modules of the large language model except the parameter matrices of the fully connected layer;

[0065] Multiple logic blocks are provided on the CMOS chip die, and one logic block corresponds to one flash array;

[0066] The main computing chip is used to read the parameter matrix from the NAND flash chip through the LPDDR memory channel and read the model parameters and intermediate calculation results from the DRAM chip during the prefill stage of the large language model inference for the calculation of the large language model; during the decoding stage of the large language model inference, the input data of the fully connected layer is placed on the NAND flash chip for the calculation of the parts of the large language model except the fully connected layer;

[0067] The NAND flash chip is used to perform the calculation of the fully connected layer of the large language model through the logic block during the decoding stage of the large language model inference.

[0068] In the embodiment of the present invention, the flash array adopts a small array design. The main purpose of the small array design is to reduce the read latency of the array and increase the number of arrays on the chip. For example, an array with a read latency less than 8 microseconds and more than 8 arrays on the chip is a small array.

[0069] The existing small-array NAND Flash array design technology mainly reduces the length and width of the NAND Flash array, that is, the lengths of the bitline and wordline, to achieve a lower array read latency, and at the same time enables more Flash arrays to be placed on the same Flash die (flash memory chip die). The current typical defects of this existing technology are as follows: (a) The peripheral circuit and the flash memory array (Flash array) are placed on the same die (chip die). The increase in the number of Flash arrays will significantly increase the area occupied by the peripheral circuit, thus greatly reducing the area ratio of the flash memory storage unit, further reducing the area efficiency of the Flash storage, and reducing the storage density of the Flash chip. Due to the consideration of area efficiency, it is difficult to make the array of the current small-array Flash chip too small, so more arrays cannot be placed within a limited area. (2) In order to be adapted to the traditional SSD system, its power supply system adopts the power supply scheme of the traditional SSD system. This scheme cannot support all Flash arrays to perform parallel access with the lowest latency (i.e., the read latency when only accessing one Flash array). Either only part of the Flash arrays can be accessed with the lowest latency, or all Flash arrays need to be read at a time significantly higher than the lowest latency.

[0070] The inter-wafer hybrid bonding technology can stack two chips "face to face" and obtain a high-density interconnection bandwidth by connecting the metal layers between the two chips. At the same time, since it does not require Through-Silicon Vias (TSVs), its cost is also lower than that of the 3D stacking based on TSVs.

[0071] For a 3D NAND Flash chip, if the flash memory array and the peripheral CMOS circuit are placed on the same chip die, the peripheral circuit generally occupies a large planar area, thus reducing the storage density of the Flash chip. To solve this problem, in the embodiments of the present invention, the flash memory array and the peripheral CMOS circuit are respectively placed on different chip dies, and then 3D stacking is performed through the inter-wafer hybrid bonding technology, thereby completely eliminating the influence of the peripheral circuit on the chip planar area. In addition, since this method allows the flash memory chip die and the CMOS chip die to be manufactured separately, this method can also eliminate the mutual influence between the 3D NAND Flash manufacturing process and the CMOS manufacturing process, reduce the manufacturing difficulty of both, and improve the manufacturing quality.

[0072] In one embodiment, an LPDDR interface circuit is further provided on the CMOS chip die;

[0073] The main computing chip is used to: access the LPDDR interface through the LPDDR memory channel and read the parameter matrix from the NAND flash memory chip.

[0074] In one embodiment, multiple voltage domains are partitioned inside the NAND flash memory chip. Each voltage domain corresponds to at least one flash memory array within the chip and is powered by different pins.

[0075] Through the above embodiments, one of the beneficial effects of the system proposed by the embodiments of the present invention can be obtained: by combining the design method of the flash memory array with the 3D NAND flash memory design method based on hybrid bonding, and at the same time by breaking the power supply design limitations imposed by the traditional SSD standard, the internal read bandwidth of the NAND flash memory chip is significantly improved on the premise of ensuring the area efficiency of the flash memory array, specifically reflected in:

[0076] (1) The combined design of three-dimensional interconnection of the flash memory array and hybrid bonding is adopted, which can not only eliminate or reduce the planar area overhead occupied by the peripheral circuits of the existing flash memory array; more importantly, considering that the size of the existing flash memory array design is limited by the area efficiency consideration and cannot be overly reduced, the combined use of the hybrid bonding method provides a new opportunity to greatly reduce the area efficiency limitation, enabling the size of the flash memory array to be further reduced, thereby further reducing the access latency on the basis of the existing small array design and increasing the number of flash memory arrays that can theoretically work in parallel within the flash memory chip die.

[0077] (2) The method of breaking the traditional SSD power supply design limitations includes two aspects.

[0078] (a) Since the traditional SSD system is abandoned, the flash memory chip die designed based on this system can break through the pin limitations specified by the traditional SSD, and thus a large number of new power pins can be added, significantly increasing the power supply capacity of the external circuit to the NAND flash memory chip.

[0079] (b) Inside the NAND flash memory chip, multiple voltage domains are partitioned. Each voltage domain corresponds to a small number of flash memory arrays within the chip and is powered by different pins.

[0080] In one embodiment, the flash memory chip die further includes a row decoder corresponding to each flash memory array, and the row decoder is used to select and activate a specific row in the flash memory array.

[0081] Specifically, the row decoder usually receives address signals from the control module, decodes these signals to determine the specific row in the NAND flash memory chip to be accessed. It isolates the target row from the other parts of the flash memory array by controlling the corresponding row select signal, enabling the data of that row to be read out.

[0082] In one embodiment, each logic block includes a peripheral circuit and multiple near-memory computing units. The multiple near-memory computing units include a multiplier-accumulator, an error correction unit, and a register;

[0083] The multiplier - adder and register are used to perform calculations for the fully - connected layer during the decoding stage;

[0084] The error - correction unit is used to perform error - correction calculations on the parameter matrix read from the flash memory array during the pre - filling stage and the decoding stage.

[0085] In the embodiments of the present invention, the peripheral circuit is used for power management, signal input / output, clock and reset, data storage and caching, interface and communication, etc., including page caches, etc. There can be multiple such multiplier - adders. The register is a local register bank.

[0086] Through the above - mentioned embodiments, lightweight calculation logic required for the fully - connected layer calculation in the decoding stage is implemented on the CMOS chip die where the peripheral CMOS circuit of the NAND flash memory chip is located. Since the flash memory array design generally adopts a single - bit flash memory cell design (Single - Level Cell), a relatively simple BCH error - correction code unit can be used, which has a small area and can be conveniently placed on the CMOS chip die.

[0087] The embodiments of the present invention propose to use a small - scale large - language model (or other lightweight prediction models) that is much smaller than the target large - language model as a draft model to predict the output of the target large - language model. In this way, the calculation in the decoding stage becomes multi - round calculation. In each round of calculation, first use the draft model to generate multiple tokens according to the original method, and then use the target large - language model to verify these generated tokens. The first few tokens that are continuously verified as correct will be used as the output of this round of calculation, and then continue the next round of calculation. Here, the verification calculation based on the target large - language model is similar to the calculation in the pre - filling stage and only requires one - round calculation. In this way, the speculative decoding technology is almost equivalent to generating multiple tokens through one calculation of the target large - language model, thus greatly improving the execution efficiency of the decoding stage.

[0088] At the hardware level, compared with the NAND flash memory chip that only supports the calculation in the decoding stage, the system proposed in the embodiments of the present invention can support the verification calculation of the word unit tokens speculated by the draft model by the target large - language model on the premise of making full use of the internal read bandwidth of the NAND flash memory chip by increasing the number of multiplier - adders and the size of the register in the NAND flash memory chip. Since the amount of calculation for the verification calculation is generally much smaller than that of the computationally intensive pre - filling stage, the increased hardware overhead is acceptable for the NAND flash memory chip. At the software level, the entire decoding stage is completed by making full use of the hardware support of the NAND flash memory chip.

[0089] Based on the foregoing embodiments, the second beneficial effect of the system described in the embodiments of the present invention can be obtained:

[0090] Combining the near-memory computing design method with the speculative decoding algorithm in the decoding stage of the large language model, bypassing the external interconnect bandwidth limitation between the main computing chip and the NAND flash memory chip, making full use of the large amount of internal read bandwidth of the NAND flash memory chip, and further improving the overall decoding performance by enabling the near-memory computing unit to support the speculative decoding algorithm, ultimately achieving a decoding latency that meets the requirements of real-time inference.

[0091] In one embodiment, the main computing chip is connected to the package substrate, and the flash memory chip die and the CMOS chip die are connected to the package substrate through wire bonding.

[0092] In an embodiment of the present invention, a NAND flash memory chip and a DRAM chip exist in pairs and share the LPDDR memory channel of the LPDDR interconnection resource. Specifically, an LPDDR interface circuit is provided on the CMOS chip die, and this circuit is finally connected to the package substrate through the wire bonding method. The NAND flash memory chip serves as a memory rank in the system and shares the same LPDDR memory channel interface with the DRAM chip as another rank.

[0093] In one embodiment, a global cache is provided on the CMOS chip die;

[0094] The main computing chip is used for: in the pre-filling stage, reading the parameter matrix from the global cache; in the decoding stage, placing the input data of the fully connected layer in the global cache; while the NAND flash memory chip performs the calculation of the fully connected layer, performing the calculation of the part other than the fully connected layer in the large language model;

[0095] The NAND flash memory chip is used for: in the pre-filling stage, while the main computing chip reads a matrix block of data from the global cache, reading each row of the next matrix block to be processed from the flash memory array and temporarily storing it in the global cache; in the decoding stage, reading the parameter matrix from the flash memory array, reading the input data of the fully connected layer from the global cache, and performing the calculation of the fully connected layer through the logic block.

[0096] Combining the foregoing embodiments, the third beneficial effect of the system according to the embodiments of the present invention can be obtained:

[0097] Using the LPDDR interface for DRAM and the package similar to the LPDDR standard to 3D stack the DRAM chip and the NAND flash memory chip in the same package, and at the same time sharing the LPDDR interconnection resource connected to the main computing chip, so that the NAND flash memory chip can utilize the large amount of external bandwidth provided by the LPDDR interconnection while minimizing the area overhead, and enabling the main computing chip to efficiently read the parameter matrix of the fully connected layer of the large language model from the NAND flash memory chip in the pre-filling stage of the large language model. Specifically, it is reflected in:

[0098] (1) The NAND flash memory chip shares an LPDDR memory channel of the LPDDR interconnection resource with the memory DRAM chip, which can make full use of the bandwidth of the LPDDR channel. For example, when the NAND flash memory chip does not need to be accessed, the channel bandwidth can be used to access the DRAM chip.

[0099] (2) In order to mask the flash array read latency that is much higher than the DRAM latency, a global cache (which can be implemented by SRAM and has a much lower latency than DRAM) placed on the CMOS chip die of the NAND flash memory chip can cache the low-latency on-chip buffer storage of at least two data blocks (i.e., the double-buffer strategy). When the main computing chip reads a matrix block of a parameter matrix from the global cache, the NAND flash memory chip can simultaneously read the next matrix block from the memory array, so that the two can cover each other, and finally completely mask the read latency of the memory chip. Using an error correction unit, when the memory chip reads a matrix block, it can simultaneously complete the error correction calculation, and the latency of the error correction calculation can also be fully masked.

[0100] In one embodiment, each parameter matrix is divided into multiple matrix blocks; the rows or columns of each matrix block are allocated to each NAND flash memory chip in a round-robin manner, and the matrix data allocated to each NAND flash memory chip is allocated to each flash memory array in a row-wise round-robin manner.

[0101] In one embodiment, the main computing chip is used for: in the pre-filling stage, reading matrix block data from the global cache of each NAND flash memory chip;

[0102] The NAND flash memory chip is used for: in the pre-filling stage, while the main computing chip reads a matrix block of data from the global cache, reading each row of the next matrix block to be processed from each flash memory array and temporarily storing it in the global cache; in the decoding stage, reading each row of each matrix block from each flash memory array.

[0103] Through the above embodiments, the fourth beneficial effect of the system described in the embodiments of the present invention can be obtained:

[0104] Storing the parameter matrix of the fully connected layer of a large language model in the system in a specific arrangement manner, enabling both the main computing chip and the NAND flash memory chip to achieve efficient data reading. That is, the process of the NAND flash memory chip reading a matrix block from the flash memory array to the global cache on the NAND flash memory die coincides with the process of the main computing chip reading the matrix block from the global cache and performing calculations; the NAND flash memory chip can simultaneously read data from each flash memory array.

[0105] In one embodiment, the main computing chip includes a neural network processing unit and an LPDDR memory controller. A control module, an instruction address cache, and a status register are also provided on the CMOS chip die;

[0106] The neural network processing unit is configured to: during the calculation of the fully connected layer in the pre-filling stage, send instructions to the LPDDR memory controller; read matrix block data from the global caches of each NAND flash chip; perform calculations of the large language model based on the parameter matrix, the model parameters read from the DRAM chip, and the intermediate calculation results;

[0107] The LPDDR memory controller is configured to: process the instructions received from the neural network processing unit, where the instructions include at least one of a write request to the instruction address cache, a status read request to the status register, a read / write request to the global cache, and a read / write request to the DRAM;

[0108] The control module is configured to: after reading an instruction from the instruction address cache, while the main computing chip reads matrix block data from the global cache, read each row of the matrix block from each flash array according to the instruction and temporarily store it in the global cache, and store the execution status of the instruction in the status register;

[0109] The neural network processing unit is further configured to: in the decoding stage, put the input data of the fully connected layer into the global cache;

[0110] The control module is configured to: broadcast the input data in the global cache to each near-memory computing unit one by one according to the calculation timing of each near-memory computing unit within the logic block; in each flash array, after reading each row of the matrix block, perform error correction calculation through the error correction unit, and sequentially take out each element from each row after the error correction calculation through the multiplier-accumulator and the register, and perform multiplier-accumulation calculation with the elements in the global cache.

[0111] It should be noted that the instruction data sent by the neural network processing unit to the flash instruction address buffer of the LPDDR memory controller can include one flash control instruction or multiple flash control instructions at the same time. The size of the instruction address cache is limited. After it is full of instructions, the LPDDR memory controller will no longer send instructions to the instruction address cache and can continue to send instructions after the instructions in the instruction address cache are consumed. There is no limit here.

[0112] In one embodiment, the software running on the main computing chip includes a user large model program and an operating system;

[0113] Some modules of the operating system are used to manage the NAND flash chips when the large language model is idle for inference;

[0114] The access permission management of the user large model program and the operating system to the NAND flash memory chip is realized through the mapping from virtual address to physical address, and the memory row address is located at the highest bit of the physical address. At the start and end of the large language model inference, the access permission of the NAND flash memory chip switches between the user large model program and the operating system.

[0115] Through the above embodiments, the fifth beneficial effect of the system described in the embodiments of the present invention can be obtained:

[0116] Different from the traditional SSD-based NAND flash memory chip control system implemented by an independent controller chip, the embodiments of the present invention can be implemented based on the main computing chip and the NAND flash memory chip itself. When the inference is idle, the operating system on the main computing chip is responsible for managing the NAND flash memory chip; when the inference system is operating, the management authority is transferred to the user large model program, and the user large model program is responsible for both computing and NAND flash memory chip control, thus eliminating the access overhead of the operating system. To further improve efficiency, the hardware module on the main computing chip can help the software to perform some control tasks, including the implementation of data reading and the implementation of the flash array reading method.

[0117] The following gives several specific embodiments to illustrate the specific application of the system proposed in the embodiments of the present invention.

[0118] Embodiment 1

[0119] Figure 2 It is a schematic structural diagram of the NAND flash memory chip in the embodiments of the present invention. This NAND flash memory chip combines the NAND flash memory design based on a small array and the NAND flash memory chip design technology based on hybrid bonding. The flash array is placed on the flash memory chip die, and the logic blocks formed by peripheral circuits and the like are placed on another CMOS chip die, and then face-to-face stacking is performed through hybrid bonding technology. There are multiple logic blocks on the CMOS chip die, each logic block corresponds to a flash array, and the peripheral circuits inside the logic block are connected to the corresponding flash array through hybrid bonding technology. This design greatly reduces the limitation of the area efficiency problem on the reduction of the array size. The number of flash arrays in this embodiment doubles to 32, and at the same time, the word line length of the array is almost halved. In order to support the parallel operation of the arrays, the number of voltage domains originally used to power the Flash array is increased from 1 to 16 voltage domains, each corresponding to 2 arrays (because the prior art only supports the parallel reading of 2 arrays), and at the same time, the number of chip pins and off-chip capacitors also increase accordingly.

[0120] As Figure 2As shown in the figure, in this embodiment, the multiplier - accumulators for supporting the calculation of the fully - connected layer and the error - correction units for verifying the parameter matrices read out from the flash memory array are both placed in each logic block. The multiplier - accumulators in each logic block are only used to process the matrix blocks in the corresponding flash memory array.

[0121] The NAND flash memory chip contains an LPDDR interface circuit for interacting with the external LPDDR memory channel. On the CMOS chip die, a global cache is placed near the LPDDR interface circuit. This is used to overlap the execution of the process of reading the parameter matrix blocks from the flash memory array inside the NAND flash memory chip and the process of the main computing chip reading the parameter matrix blocks from the NAND flash memory chip. That is, the matrix blocks of the parameter matrix read from the flash memory array can be temporarily stored in this global cache, and at the same time, the main computing chip can read the matrix blocks of the parameter matrix read from the flash memory array before from this global cache through the LPDDR interface circuit. At the same time, when the flash memory array reads out data, it will simultaneously use Figure 2 the error - correction unit module in the corresponding logic block to correct the matrix blocks of the parameter matrix.

[0122] Embodiment 2

[0123] Figure 3 It is a schematic diagram of the physical architecture for the inference of the large - language model at the edge in the embodiment of the present invention. Figure 4 It is a schematic diagram of the logical architecture for the inference of the large - language model at the edge in the embodiment of the present invention. Inside a package similar to an LPDDR DRAM, 2 3D - stacked memory chip stacks (stacks) are placed at the same time. Each memory chip stack contains 4 NAND flash memory chips (i.e., Figure 3 the NAND flash memory chips in

[0124] Embodiment 3

[0125] Figure 5 It is a schematic diagram of the execution flow for the large - language model inference system at the edge to perform large - language model inference in the embodiment of the present invention. Figure 6 It is a schematic diagram of the layout and processing principle of the fully - connected layer parameter matrix. Whether it is the calculation of the target large - language model or the draft model, it follows Figure 5The execution process shown. The NAND flash chip is only related to the calculations of the fully connected layer and only stores the parameter matrices of the fully connected layer. Figure 5 In Figure 5 , the feed-forward module of the large language model includes feed-forward network FC1, feed-forward network FC2, and activation module. The parameter matrices include QKV generation parameter matrix, projection parameter matrix, feed-forward network FC1 parameter matrix, and feed-forward network FC2 parameter matrix.

[0126] In the pre-fill stage, the main computing chip is responsible for all calculations including the fully connected layer, namely QKV generation, self-attention calculation, projection, two feed-forward networks and activation. The main computing chip reads the parameter matrices from the NAND flash chip through the LPDDR interface and reads other model parameters and intermediate calculation results from the DRAM chip. The model parameters include input / output embedding vectors, QKV parameters, bias parameters, etc.

[0127] In the decoding stage, the NAND flash chip is responsible for all calculations of the fully connected layer, which are reflected in the calculations of the fully connected layer in the generation module, projection module, feed-forward network FC1, and feed-forward network FC2. Other calculations of the large language model, such as self-attention calculation and activation calculation, are still performed by the main computing chip.

[0128] See Figure 6 , when arranging the parameter matrices, the entire parameter matrix is first divided into multiple matrix blocks. The matrix rows of a matrix block are distributed to each NAND flash chip in a round-robin manner and then distributed to each flash array in a round-robin manner.

[0129] When the main computing chip accesses data, the main computing chip reads data from each NAND flash chip in parallel. The read data can be composed into a matrix block of appropriate size through splicing or splitting, etc., to facilitate calculation.

[0130] Embodiment 4

[0131] Figure 7This is the control schematic diagram in the edge large language model inference system in the embodiments of the present invention. In the edge large language model inference system, the main computing chip includes a neural network processing unit and an LPDDR memory controller, which serves as the control module of the main computing chip as a whole. The neural network processing unit includes a memory access unit. In addition, the software operating system and the user large model program run on the main computing chip, and the control module is placed on the NAND flash chip. The main computing chip can directly read / write the instruction address cache, registers, and global cache of the control module through the LPDDR interface circuit, but cannot directly access the cache array. The main computing chip can only send the instructions to be executed to the instruction address cache through the LPDDR, and the control module of the NAND flash chip reads the instructions from the instruction address cache and then executes them. The data for NAND flash chip management (such as the NAND flash address mapping table, etc.) is stored in the DRAM memory space. The access permission management of the NAND flash chip by the user large model program and the operating system is achieved through the mapping from virtual address to physical address. At the beginning and end of the large language model inference, the access permission of the NAND flash chip will be switched between the user large model program and the operating system. In addition, this embodiment changes the traditional LPDDR physical address mapping, placing the memory row address (rank address) at the highest bit of the physical address, so that all DRAM chips and NAND flash chips in the whole system respectively form two separate continuous address spaces, facilitating system management. The LPDDR memory controller includes a NAND flash reading management unit, which is used to achieve efficient reading of the data of the NAND flash chip, such as detecting whether the NAND flash array reading is completed and efficiently managing the process of reading the parameter matrix of the fully connected layer from the global cache.

[0132] In the specific embodiments described above, the purpose, technical solution, and beneficial effects of the present invention have been further detailed. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. An edge-side large language model inference system, characterized in that, Comprising: A main computing chip, LPDDR interconnection resources connected to the main computing chip, and at least one vertically stacked package of memory chip stacks; Each memory chip stack includes a plurality of NAND flash memory chips and a plurality of DRAM chips. One NAND flash memory chip and one DRAM chip exist in pairs and share the LPDDR memory channel of the LPDDR interconnection resources; The NAND flash memory chip includes a flash memory chip die and a CMOS chip die based on three-dimensional interconnection by hybrid bonding; A plurality of flash memory arrays are provided on the flash memory chip die; the flash memory arrays are used to store the parameter matrices required for the calculation of the fully connected layer of the large language model; The DRAM chip is used to store the model parameters and intermediate calculation results required for the calculation of all modules of the large language model except the parameter matrices of the fully connected layer; A plurality of logic blocks are provided on the CMOS chip die, and one logic block corresponds to one flash memory array; The main computing chip is used to, in the prefill stage of large language model inference, read the parameter matrices from the NAND flash memory chips through the LPDDR memory channel, and read the model parameters and intermediate calculation results from the DRAM chips to perform the calculation of the large language model; in the decoding stage of large language model inference, place the input data of the fully connected layer on the NAND flash memory chips to perform the calculation of the part of the large language model except the fully connected layer; The NAND flash memory chip is used to, in the decoding stage of large language model inference, perform the calculation of the fully connected layer of the large language model through the logic block.

2. The system according to claim 1, wherein An LPDDR interface circuit is also provided on the CMOS chip die; The main computing chip is used to: access the LPDDR interface through the LPDDR memory channel and read the parameter matrices from the NAND flash memory chips.

3. The system according to claim 1, characterized in that, Multiple voltage domains are divided inside the NAND flash memory chip, each voltage domain corresponds to at least one flash memory array within the chip, and is powered by different pins.

4. The system according to claim 1, characterized in that The main computing chip is connected to the package substrate, and the flash memory chip die and the CMOS chip die are connected to the package substrate through wire bonding.

5. The system according to claim 1, wherein Each logic block includes peripheral circuits and a variety of near-memory computing units. The variety of near-memory computing units include multipliers, error correction units, and registers; The multipliers and registers are used to perform the calculation of the fully connected layer in the decoding stage; The error correction unit is used to perform error correction calculation on the parameter matrices read from the flash memory arrays in the prefill stage and the decoding stage.

6. The system according to claim 5, wherein A global cache is provided on the CMOS chip die; The main computing chip is used to: read the parameter matrices from the global cache in the prefill stage; place the input data of the fully connected layer in the global cache in the decoding stage; perform the calculation of the part of the large language model except the fully connected layer while the NAND flash memory chip is performing the calculation of the fully connected layer; The NAND flash memory chip is used to: while the main computing chip reads the parameter matrices from the global cache in the prefill stage, read the parameter matrices from the flash memory arrays and temporarily store them in the global cache; In the decoding stage, read the parameter matrices from the flash memory arrays, read the input data of the fully connected layer from the global cache, and perform the calculation of the fully connected layer through the logic block.

7. The system according to claim 6, wherein Each parameter matrix is divided into multiple matrix blocks; each row or column of each matrix block is allocated to each NAND flash chip in a rotating manner, and the matrix data allocated to each NAND flash chip is allocated to each flash array in a rotating manner by row.

8. The system according to claim 7, wherein The main computing chip is used for: in the pre-filling stage, reading matrix block data from the global cache of each NAND flash chip; The NAND flash chip is used for: in the pre-filling stage, while the main computing chip reads a matrix block of data from the global cache, reading each row of the next matrix block to be processed from each flash array and temporarily storing it in the global cache; In the decoding stage, reading each row of each matrix block from each flash array.

9. The system according to claim 8, wherein The main computing chip includes a neural network processing unit and an LPDDR memory controller. A control module, an instruction address cache, and a status register are also provided on the CMOS chip die; The neural network processing unit is used for: during the calculation of the fully connected layer in the pre-filling stage, sending instructions to the LPDDR memory controller; reading matrix block data from the global cache of each NAND flash chip; performing calculations of the large language model based on the parameter matrix and the model parameters and intermediate calculation results read from the DRAM chip; The LPDDR memory controller is used for: processing the instructions received from the neural network processing unit, where the instructions include at least one of a write request to the instruction address cache, a read request to the status register, a read / write request to the global cache, and a read / write request to the DRAM; The control module is used for: after reading an instruction from the instruction address cache, while the main computing chip reads matrix block data from the global cache, reading each row of the matrix block from each flash array according to the instruction and temporarily storing it in the global cache, and storing the execution status of the instruction in the status register; The neural network processing unit is also used for: in the decoding stage, putting the input data of the fully connected layer into the global cache; The control module is used for: broadcasting the input data in the global cache to each near-memory computing unit one by one according to the calculation timing of each near-memory computing unit within the logic block; in each flash array, after reading each row of the matrix block, performing error correction calculation through the error correction unit, and taking out each element from each row after error correction calculation one by one through the multiplier-accumulator and register, and performing multiplier-accumulation calculation with the elements in the global cache.

10. The system according to claim 9, characterized in that, The software running on the main computing chip includes a user large model program and an operating system; The operating system is used for managing the NAND flash chip when the large language model is idle for inference; The access permission management of the user large model program and the operating system for the NAND flash chip is realized through the mapping from virtual address to physical address, and the memory row address is located at the highest bit of the physical address, and at the start and end moments of the large language model inference, the access permission of the NAND flash chip switches between the user large model program and the operating system.

Citation Information

Cited By

  • Memory device, computer system and memory die

    CN121483322A

  • Data read-write method and system based on hybrid bonding large model reasoning, controller and storage medium

    CN121958147A

  • Data read-write method and system based on hybrid bonding large model inference, controller and storage medium

    CN121958147B

  • Near memory computing chip architecture and method

    CN122086839A