AI device based on near-memory computing and electronic device
Patent Information
- Application Number
- CN202610810726.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-05
- Publication Date
- 2026-09-29
AI Technical Summary
高功耗不仅缩短了设备的续航时间,还带来了严重的散热问题
[0029]本发明提供的AI装置,通过含有神经网络处理功能的计算模块结合第一存储模块和第二存储模块,以实现神经网络计算,进而实现AI功能。其中,第一存储模块的读出电路内设置有近存储计算电路,至少用于在解码阶段根据从存储电路读取的大模型的权重数据进行前馈神经网络的计算,而其它的神经网络计算可以在计算模块的第一处理器里实现。即第一处理器根据输入数据、第一存储模块及第二存储模块的数据进行神经网络计算,进而可以将不同阶段不同类型的计算分工到不同的地方,且在第一存储模块的内部的近存储计算电路里实现解码阶段里与权重的读取紧耦合的前馈神经网络计算。所以,其可以实现在第一存储模块内的大模型的权重数据尚未离开第一存储模块时即完成解码中的前馈神经网络计算,仅将精简后的中间结果传输至计算模块。其可以大幅减少计算模块与第一存储模块之间的数据搬移量、搬移次数、搬移距离、搬移时间、搬移功耗、搬移产生的热量、缓解带宽压力和瓶颈,还可以减少计算模块等待数据搬移的时间、减少数据搬移和计算的延迟、降低功耗、减少推理延迟、提高数据的吞吐能力。而且,两存储模块分工协作存储不同的数据及进行不同的处理,二者协作可以优化存储资源的配置。使得AI装置既可以存储更多更大的参数以解决存储容量不足的问题,也可以缓解带宽压力,还可以减少计算模块与各存储模块之间的数据的传输路径和距离,以减少功耗和提高能效。计算模块包括的具备神经网络处理能力的第一处理器,能够充分利用输入数据、来自第一存储模块和第二存储模块的数据,结合近存储计算电路的中间处理结果,完成剩余部分的神经网络计算任务,确保整体推理过程的高效执行,其可以进一步提高神经网络处理的能效并减少对应的功耗和热量的产生,以从源头上解决散热的问题。所以,本发明提供的AI装置,其可以解决功耗、散热、吞吐能力、推理延迟、存储容量及带宽压力的问题,提高推理效能并提升用户体验。
Smart Images

Figure CN122840137A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, specifically to AI devices and electronic devices based on near-memory computing. Background Technology
[0002] With the rapid development of artificial intelligence technology, large models have achieved great success in fields such as natural language processing. However, the number of parameters in large models is increasing rapidly, with some models containing billions or even hundreds of billions of parameters, placing extremely high demands on AI devices. Currently, however, AI devices based on large models face numerous technical bottlenecks affecting inference performance, such as: 1. High power consumption: The larger the number of parameters in a large model, the higher the power consumption required for the AI device to operate. Edge devices typically rely on battery power, which imposes strict requirements on power consumption and battery life. High power consumption not only shortens the device's battery life but also leads to serious heat dissipation problems.
[0003] 2. The low number of tokens processed per second results in low speed and throughput of the AI device.
[0004] 3. High latency: Due to limitations in storage capacity, bandwidth, energy efficiency, and other aspects, AI devices often experience high latency, especially AI devices used for inference, which makes it difficult to meet the needs of real-time interaction.
[0005] 4. Bandwidth bottleneck: Large models require frequent access to their parameters during training and inference, but current AI devices, especially edge-side inference chips, have insufficient storage bandwidth to meet these data transfer demands. This leads to a significant decrease in the inference speed of large models, impacting user experience and limiting the widespread application of AI.
[0006] In summary, AI devices based on large models face multiple technical challenges, including high power consumption, poor heat dissipation, low number of tokens processed per second, high latency, and bandwidth bottlenecks. Summary of the Invention
[0007] The purpose of this invention is to provide AI devices and electronic devices based on near-memory computing to solve some or all of the above-mentioned technical problems.
[0008] To achieve the above objectives, the present invention provides an AI device based on near-memory computation, comprising: a computing module, a first storage module electrically connected to the computing module, and a second storage module electrically connected to the computing module; the computing module includes a first processor for performing neural network processing, the first processor being used to perform neural network computation based on input data and data from the first and second storage modules; the first storage module is used to store at least the parameters of a large model, and the second storage module is used to store at least the intermediate data generated by the computing module in performing neural network computation; the first storage module includes a storage circuit for storage and a readout circuit for reading data from the storage circuit, the readout circuit having a near-memory computation circuit internally disposed therein; the near-memory computation circuit within the first storage module is used at least during the decoding stage to perform feedforward neural network computation based on the weight data of the large model read from the storage circuit.
[0009] In one implementation, the large model is a large model based on the MOE architecture, and the near-memory computing circuit within the readout circuit is used to perform calculations tightly coupled with the reading of expert weight data.
[0010] In one implementation, the large model includes a large model based on the MOE architecture, wherein the near-memory computing circuit within the readout circuit is used to perform feedforward neural network computation based on expert weight data read from the memory circuit during the decoding phase.
[0011] In one implementation, the near-memory computing circuitry within the readout circuitry is used to perform lightweight gated network computations required for expert routing logic, and / or multiply-accumulate computations to activate expert weights, and / or computations of preliminary scalar combinations of expert outputs.
[0012] In one implementation, the computation module is used for processing the pre-filling stage of large model inference, the near-memory computation circuit is used for the first layer linear transformation of the MOE expert network layer in the decoding stage of inference, and the computation module is also used to perform the remaining computations of the decoding stage in combination with the result of the first layer linear transformation.
[0013] In one embodiment, the weight data of the large model in the storage circuit is divided into several blocks. When the readout circuit reads the (N+1)th block of weight data in the storage circuit, the near-memory computing circuit in the readout circuit performs feedforward neural network calculation on the weight data of the Nth block that has been read, so as to realize simultaneous reading and calculation in the first storage module.
[0014] In one embodiment, the near-memory computing circuit is also used to perform attention calculation based on the weight data of the large model read from the storage circuit during the decoding stage of large model inference.
[0015] In one embodiment, the computing module includes: an I / O system for inputting and outputting data; a second processor electrically connected to the I / O system for preprocessing the input data so that the preprocessed input data conforms to the format of a large model; and a plurality of first processors electrically connected to the second processor for performing neural network calculations on the preprocessed input data.
[0016] In one embodiment, the first processor is a neural network accelerator (NPU), the NPU comprising: a preprocessing module for converting the preprocessed input data into matrix data conforming to the processing format of the in-memory computing module; a plurality of neural network sub-accelerators (SNPUs) for performing neural network calculations on the data processed by the preprocessing module, the SNPU including a plurality of in-memory computing modules (CIMs) arranged in a matrix to form an in-memory computing matrix; and a vector processing module for performing vector processing on the data processed by the SNPU.
[0017] In one embodiment, the in-memory computation matrix includes: N columns of SRAM-based in-memory computation submodules, each column of the in-memory computation submodule including M rows of in-memory computation units for storing and performing computations on the stored data and input feature data, and capacitors or adders for accumulating the results of the computations performed by the in-memory computation units.
[0018] In one embodiment, the computing module is further configured to predict the neuron parameters to be activated in the parameters of the large model in the first storage module based on the input data and the sparsity prediction model. The computing module reads the predicted neuron parameters to be activated from the first storage module based on the prediction results, and performs neural network calculations based on the read neuron parameters, the input data, and the data from the second storage module.
[0019] In one embodiment, the system further includes a storage controller disposed within the computing module, the storage controller being used to control the scheduling of data in the first storage module; the read circuit in the first storage module reads data from the storage circuit according to instructions from the storage controller of the computing module.
[0020] In one embodiment, the first storage module is stacked with the computing module via hybrid bonding. The computing module directly reads data from the first storage module to perform calculations based on the read data. The computing module then transmits the intermediate data generated by the calculations to the second storage module for storage.
[0021] In one embodiment, the second storage module is stacked on top of the first storage module, and the computing module, the first storage module, and the second storage module are stacked from bottom to top through a layered encapsulation method.
[0022] In one embodiment, the second storage module and the computing module are disposed on the first substrate in a flat manner, and the first storage module is stacked on the computing module in a hybrid bonding manner.
[0023] In one embodiment, the first storage module includes at least two layers of Nand dies, which are stacked together by a hybrid bonding method to form the storage circuit. The readout circuit of the first storage module is stacked with and electrically connected to the storage circuit, and the readout circuit is located between the storage circuit and the computing module.
[0024] In one embodiment, the second storage module is a DRAM memory; the AI device further includes: a first substrate disposed below the computing module, and a second substrate disposed between the first storage module and the second storage module, wherein the first substrate, the computing module, the first storage module, the second substrate, and the second storage module are stacked from bottom to top in a layered packaging manner; the second substrate is provided with a through hole, and the DRAM memory is electrically connected to the computing module through a wire passing through the through hole of the second substrate.
[0025] In one embodiment, the first storage module is a Nand Flash type memory, and / or the second storage module is a DRAM type memory.
[0026] In one embodiment, the first storage module is a Nand Flash type memory, and the second storage module is a DRAM type memory; the Nand Flash type memory is also used to store a sparsity prediction model, the sparsity prediction model is a low-rank prediction model, the calculation module predicts the parameters of the neuron to be activated based on the low-rank prediction model read from the first storage module and the current input data, and the calculation module selectively reads the parameters of the neuron to be activated from the first storage module and performs neural network calculation based on the prediction results, and stores the intermediate data generated during the calculation process in the second storage module.
[0027] In one implementation, the large model includes a large model based on the MOE architecture, and the second storage module is also used to store parameters of the large model; the first storage module stores at least the parameters of the sparse part of the large model, and the second storage module stores at least the parameters of the dense part of the large model and the intermediate data generated by the neural network calculation; or the first storage module stores the sparse large model, and the second storage module stores the dense large model and the intermediate data generated by the neural network calculation.
[0028] The present invention also provides an electronic device including the AI device described in any of the preceding claims.
[0029] The AI device provided by this invention achieves AI functionality by combining a computing module with neural network processing capabilities with a first storage module and a second storage module to perform neural network calculations. Specifically, the readout circuit of the first storage module contains a near-memory computing circuit, used at least during the decoding stage to perform feedforward neural network calculations based on the weight data of the large model read from the storage circuit. Other neural network calculations can be performed within the first processor of the computing module. That is, the first processor performs neural network calculations based on the input data, the data from the first storage module, and the data from the second storage module, thus dividing different types of calculations at different stages into different parts. Furthermore, the feedforward neural network calculations, tightly coupled with the weight reading during the decoding stage, are implemented within the near-memory computing circuit inside the first storage module. Therefore, it can complete the feedforward neural network calculations during decoding even before the weight data of the large model within the first storage module leaves the first storage module, only transmitting the simplified intermediate results to the computing module. This design significantly reduces the amount, number of data transfers, transfer distance, transfer time, transfer power consumption, and heat generated during data transfer between the computing module and the first storage module, alleviating bandwidth pressure and bottlenecks. It also reduces the time the computing module spends waiting for data transfers, reduces latency in data transfer and computation, lowers power consumption, reduces inference latency, and improves data throughput. Furthermore, the two storage modules collaborate to store different data and perform different processing tasks, optimizing storage resource allocation. This allows the AI device to store more and larger parameters to address insufficient storage capacity, alleviate bandwidth pressure, and reduce the data transmission paths and distances between the computing module and each storage module, thereby reducing power consumption and improving energy efficiency. The computing module includes a first processor with neural network processing capabilities. This processor fully utilizes input data, data from the first and second storage modules, and combines the intermediate processing results from the near-storage computing circuit to complete the remaining neural network computation tasks, ensuring efficient execution of the overall inference process. This further improves the energy efficiency of neural network processing and reduces corresponding power consumption and heat generation, addressing the heat dissipation problem at its source. Therefore, the AI device provided by this invention can solve the problems of power consumption, heat dissipation, throughput, inference latency, storage capacity and bandwidth pressure, improve inference performance and enhance user experience. Attached Figure Description
[0030] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0031] Figure 1This is a schematic diagram of the structure of an AI device provided in an embodiment of the present invention; Figure 2 This is a diagram illustrating the overhead of accessing and calculating weighted data. Figure 3 This is a schematic diagram of the structure of an AI device provided in another embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of an AI device provided in an embodiment of the present invention; Figure 5 This is a block diagram of a CIM-based computing module provided in an embodiment of the present invention; Figure 6 yes Figure 5 A magnified view of the NPU in the middle; Figure 7 This is a structural block diagram of a neural network accelerator in one embodiment of the present invention.
[0032] Figure 8 This is a schematic diagram of a column of stored operator modules in an in-memory computation matrix according to an embodiment of the present invention; Figure 9 This is a schematic diagram of the vertical structure of a column storage operator module in one embodiment of the present invention; Figure 10 This is a schematic diagram illustrating the principle of in-memory computation matrix calculation in one embodiment of the present invention. Detailed Implementation
[0033] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. In the absence of conflict, the following embodiments and their technical features can be combined with each other.
[0034] To make the objectives, technical solutions, and advantages of the present invention clearer, various exemplary embodiments described below will be referenced to the accompanying drawings, which form part of the exemplary embodiments, illustrating various exemplary embodiments that may be used to implement the present invention. Unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. It should be understood that they are merely examples of processes, methods, and apparatuses consistent with some aspects of the present invention disclosed as detailed in the appended claims, and other embodiments may be used, or structural and functional modifications may be made to the embodiments listed herein without departing from the scope and spirit of the present invention.
[0035] It should be understood that in the description of this invention, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. The term "a plurality of" means two or more. The terms "connected" and "linked" should be interpreted broadly, for example, they can refer to fixed connections, detachable connections, integral connections, mechanical connections, electrical connections, communication connections, direct connections, indirect connections via an intermediate medium, or connections within two elements or interactions between two elements. The term "and / or" includes any and all combinations of one or more of the related listed items. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0036] AI devices can be used to implement various AI functions, especially through training and / or inference. Currently, the use of AI devices for inference is increasing. The inference process of AI devices based on large models mainly includes a prefill stage and a decode stage. The inventors of this invention have discovered that the prefill stage has a large computational load, with computing power requirements exceeding bandwidth requirements; in the decode stage, computing power requirements are significantly reduced, but bandwidth requirements are significantly increased. Furthermore, since most of the time in large model inference is spent in the decode stage, it is memory-intensive, and its inference efficiency is severely limited by bandwidth. These two stages differ significantly in computational paradigm, data access pattern, and system bottlenecks. For example, the prefill stage processes the entire input sequence in a context-parallel manner, exhibiting prominent computational intensity. This stage requires loading and computing all or most of the network in a batch at once, placing extremely high demands on the peak computing power of the computing core and the continuous throughput capacity of memory access bandwidth. The performance of the prefill stage is mainly limited by the computational capacity (Compute-Bound). The decode stage generates tokens one by one in an autoregressive manner, which is a memory-intensive stage. This process triggers massive, random, and fine-grained access requests for weight parameters. In traditional compute-storage separation architectures, this access processing mode leads to extremely low memory bandwidth utilization. For example, the processor's massively parallel computing units become idle while waiting for tiny data packets, and the processor's computing units frequently become idle while waiting for the transfer of these tiny data packets. This access processing mode also suffers from severe "data migration walls," where a significant amount of energy and time is consumed in data migration from storage media to the processor, rather than actual computation. Therefore, inference performance, especially the number of tokens processed per second, is significantly affected. Furthermore, AI devices need to incorporate increasingly large model parameters, resulting in increasing power consumption and heat generation. In existing AI devices based on large models, the computational components are located far from the heat sink, making it difficult to effectively dissipate the heat generated by these computations, thus posing a serious challenge to heat dissipation. Current technologies address this by adding heat sinks and improving their heat dissipation capabilities, but the improvement is still unsatisfactory.
[0037] See Figure 1As shown, the AI device based on near-memory computing provided by the present invention includes: a computing module, a first storage module electrically connected to the computing module, and a second storage module electrically connected to the computing module. The computing module includes a first processor for performing neural network processing. The first processor performs neural network calculations based on input data and data from the first and second storage modules to achieve AI functions, such as training and / or inference. The first storage module is electrically connected to the computing module, enabling data interaction. The first storage module is used to store at least the parameters of a large model. The second storage module is used to store at least the intermediate data generated by the neural network calculations performed by the computing module. The second storage module is electrically connected to the computing module to enable data interaction. The first storage module can be stacked with the computing module or electrically connected to the computing module in a tiled manner. The second storage module can also be stacked with the computing module or electrically connected to the computing module in a tiled manner. The first storage module includes a storage circuit for storage and a readout circuit for reading data from the storage circuit. Data output by the readout circuit can be transmitted to the computing module. The readout circuit internally contains a near-memory computing circuit. The near-memory computing circuit within the readout circuit is used at least during the decoding stage to perform feedforward neural network calculations based on the weight data of the large model read from the memory circuit.
[0038] The computation module can directly read the parameters of the large model stored in the storage circuit of the first storage module through the readout circuit of the first storage module, and the computation module can directly read the data in the second storage module to perform neural network calculations. Moreover, the near-memory computation circuit in the first storage module has already performed feedforward neural network calculations based on the weight data from the storage circuit, so the weight data in this part does not need to be read from the first storage module and moved to the computation module; only the simplified intermediate results (i.e., after feedforward neural network calculations) are transmitted to the computation module. This invention offloads the feedforward neural network computation task in the decoding stage to a near-memory computing circuit within a first storage module electrically connected to the computing module. Specifically, it decouples the feedforward neural network computation, which involves reading weight data from the storage circuit of the first storage module, from the storage circuit of the first storage module. This allows the near-memory computing circuit within the readout circuit of the first storage module to perform the feedforward neural network computation. This significantly reduces the amount, number, distance, time, power consumption, and heat generated during data movement between the computing module and the first storage module (e.g., significantly reducing memory access and data movement between the computing chip and the storage chip, instead performing feedforward neural network computation-related operations directly within the storage chip or storage die). This alleviates bandwidth pressure, reduces power consumption, increases the number of tokens processed per second, and reduces inference latency. Furthermore, the feedforward neural network computation in the decoding stage is a bandwidth-intensive computation, involving massive, random, and fine-grained access requests for weight parameters. Since this invention implements the feedforward neural network computation in the decoding stage within the first storage module, it avoids the massive, random, and fine-grained access and relocation of weight parameters from the computing module to the memory, as is common in existing technologies. This significantly alleviates bandwidth pressure and bottlenecks. Additionally, the near-memory computing circuit in the read circuit of the first storage module implements the feedforward neural network computation in the decoding stage, enabling computation to be completed before the data leaves the first storage module. This allows for read-while-computing within the first storage module, greatly reducing the waiting time, latency, and idle time of the computing module during data relocation. This alleviates the "memory wall" bottleneck, significantly improves bandwidth utilization and the effective use of computing resources, resulting in a substantial increase in the number of tokens processed per second by the AI device, significantly reducing power consumption and latency. It also reduces heat generation and solves the heat dissipation problem at its source. Moreover, the two storage modules collaborate to store different data and perform different processing tasks, optimizing the allocation of storage resources. This allows AI devices to store more and larger parameters to solve the problem of insufficient storage capacity, alleviate bandwidth pressure, and reduce the data transmission path and distance between computing modules and storage modules, thereby reducing power consumption and improving energy efficiency.
[0039] Preferably, the computing module and the first storage module are electrically connected via stacking. This allows the parameters of the large model within the first storage module to be accessed by the computing module with extremely high bandwidth and extremely low latency. It also effectively alleviates the bandwidth bottleneck and power consumption caused by long-distance data movement in traditional architectures. This enables the entire AI device to support greater bandwidth and bandwidth utilization, significantly improving the data throughput of the AI device and greatly reducing data transmission distance and energy consumption, data transmission latency, heat generated during data processing, area occupation, and inference latency. Furthermore, the first storage module can handle the aforementioned feedforward neural network calculations by setting / integrating lightweight near-memory computing circuitry. This requires no additional hardware, does not occupy additional area or volume, and is cost-effective.
[0040] The second storage module can store intermediate data generated by neural network calculations, or it can store both intermediate data and parameters of the large model, while the first storage module stores the parameters of the large model. This allows different types of data to be stored separately in the two storage modules, even if the two modules have different functions. For example, parameters that are frequently accessed can be stored in one memory, while parameters / data that are not frequently accessed can be stored in another. For instance, parameters with a large number of parameters in the large model that are not frequently accessed can be stored in the first storage module, reducing the bandwidth requirements of the computing module. Meanwhile, parameters / data that are frequently accessed in the large model and intermediate data with a relatively small amount of data can be stored in the second storage module, without affecting the bandwidth requirements of the computing module for accessing the second storage module. This further solves the bottlenecks of storage capacity and bandwidth simultaneously. Moreover, the second storage module used to store the small amount of intermediate data generated by neural network calculations and the parameters of the large model can use low-bandwidth, low-capacity, and low-cost storage devices, or a smaller number of memory modules, thereby reducing the cost and size of the AI device.
[0041] The computing module includes a first processor that can employ a neural network accelerator (NPU), possessing highly efficient neural network processing capabilities. It can fully and quickly utilize input data, data from the first and second storage modules, and combine this with intermediate processing results from the near-storage computing circuitry to complete the remaining computations in the decoding stage and the pre-filling stage, ensuring the efficient operation of the AI device. This further improves the energy efficiency of neural network processing and reduces corresponding power consumption and heat generation, addressing the heat dissipation problem at its source and enhancing the performance of the AI device. Furthermore, by assigning different types of computations at different stages to the computing module and the first storage module stacked adjacent to it, some bandwidth-intensive, lightweight computational tasks can be offloaded to the storage side. This significantly reduces data transfer overhead between the computing module and the first storage module, alleviates bandwidth pressure, reduces waiting time for data transfer in the computing module, reduces data transfer and computation latency, lowers power consumption, and reduces inference latency. Furthermore, the second storage module collaborates with the first storage module: the first storage module focuses on storing large model parameters and lightweight near-memory computation, while the second storage module is responsible for storing frequently read / written intermediate data, or frequently read / written but smaller amounts of intermediate data and a small number of parameters. This collaboration optimizes the efficiency of storage resource allocation. This division of labor, storing different parameters in two storage modules, allows the AI device to store more and larger parameters to address insufficient storage capacity, while also reducing the data transmission paths and distances between the computation module and each storage module, thereby improving energy efficiency and reducing power consumption. Therefore, the AI device provided by this invention can solve problems related to storage capacity, bandwidth, computing power, energy consumption, throughput, energy efficiency, heat dissipation, and inference latency.
[0042] In a preferred embodiment of the present invention, the large model includes a large model based on an MOE (Mixture of Experts) architecture. The near-memory computing circuit within the readout circuit of the first storage module is used to perform feedforward neural network calculations based on the expert weight data read from the storage circuit during the decoding stage. This efficiently supports the MOE large model in the decoding stage of inference, performing calculations tightly coupled with the reading of expert weight data, such as a small number of computational tasks related to expert routing at the MOE layer. Therefore, the near-memory computing circuit can only calculate "partial weight data" in the first storage module, correspondingly completing the calculation of these frequently read weight data with small parameter amounts within the first storage module during the decoding stage. Furthermore, it is evident that the near-memory computing circuit does not need to read all the weight data in the storage circuit, but only reads a portion of the expert weight data, and directly performs feedforward neural network calculations within the first storage module based on the read portion of expert weight data. It only needs to transmit the intermediate calculation results calculated by the feedforward neural network to the computing module, which then performs the remaining calculations during the decoding stage. Therefore, it can greatly alleviate the problems of bandwidth, latency, power consumption, and heat dissipation caused by moving massive, random, and fine-grained data from the storage module to the main computing module for computation in the decoding stage of existing technologies, thereby greatly improving the inference efficiency and performance of the inference device.
[0043] Furthermore, the near-memory computation circuit within the readout circuit is used to perform feedforward neural network computation tightly coupled with the reading of activated expert weight data in the MOE during the autoregressive decoding process, before the weight data of the large model in the first storage module leaves the first storage module. This type of computation typically only involves a portion of the weights (such as activated experts in the MOE, accounting for 5%~10% of the total parameters), and the computational load is relatively small. That is, the AI device provided by this invention can start the computation of the decoding stage before the weight data in the first storage module leaves the first storage module. It can start the above computation immediately while the data is read from the storage circuit inside the first storage module, realizing "read-and-calculate". The computation and reading are highly overlapped in time to optimize the latency problem of inference. In the decoding stage, each input token needs to dynamically activate and compute a few expert networks (such as Top-1 or Top-2 experts). For example, a gating network selects activated experts (such as Top-2) for the current token. Then, the weight matrix of the selected expert needs to be read and matrix multiplied with the input vector of the token to obtain the expert output. This invention avoids moving the entire expert weight matrix to the computation module, significantly reducing data movement overhead and the time the computation module spends waiting for the massive amounts of fine-grained expert weight data to be moved. This reduces inference latency and improves inference performance. Therefore, this invention can fully utilize the limited computing resources within the storage device, reducing both the computational requirements and costs of the storage device and leveraging the advantages of proximity to storage. It reduces bottlenecks in data movement and computation, mitigating the resulting drawbacks and achieving more refined utilization and optimization of hardware. In particular, it can efficiently and cost-effectively accelerate the "decoding stage, which has lower computational intensity but highly random and fine-grained data access," thereby optimizing inference performance.
[0044] In a preferred embodiment, the near-memory computing circuit within the readout circuit is used to perform lightweight gated network computations required for expert routing logic, and / or multiply-accumulate computations to activate expert weights, and / or computations of preliminary scalar combinations of expert outputs. Correspondingly, the near-memory computing circuit in the readout circuit of the first storage module can be a lightweight computing circuit (such as a small number of multiply-accumulators or an addition tree) specifically designed to handle such tasks. Its small circuit size and low power consumption can meet the needs of "small computations," alleviating the memory wall bottleneck, improving bandwidth utilization, and enhancing the energy efficiency and latency of AI devices.
[0045] As can be seen from the above, by having the feedforward neural network computation task in the decoding stage executed internally by the first storage module, this invention significantly reduces the system-level data movement overhead caused by frequent reading of expert weights, fine-grained reading, and long-distance relocation to the computing module, effectively alleviating the memory bandwidth bottleneck. This results in an overall reduction of inference latency and system energy consumption in the decoding stage of the MOE large model, and an increase in the number of tokens processed per second by the MOE large model. It achieves deep collaborative optimization of storage and computing resources, which can significantly improve the inference performance of AI devices.
[0046] It is worth noting that the AI device provided by this invention can also support other types of large models, and the AI device provided by this invention can support one, two or more large models.
[0047] In one embodiment of the present invention, the working method of the AI device based on the MOE large model for inference is briefly described as follows: During the pre-filling stage, the computation module performs neural network computation for inference based on the input data and data from the first and second storage modules, and stores the intermediate data generated by the computation (such as KV cache and intermediate computation results such as feature map input and output) in the second storage module.
[0048] During the pre-filling phase, the computation module computes the weight parameters of a large number of expert networks, which are read at high speed directly from the first storage module stacked with the computation module (e.g., bandwidth greater than 100 GB / s). The computation module processes the context input sequence of the MOE large model. This phase is computationally intensive, requiring the loading and computation of all or most of the expert networks in a single, batch manner. The computation flow includes, for example: layer input, layer normalization, multi-head self-attention layer (MLA), feedforward network layer, layer output, etc.
[0049] During the decoding phase, the computing module sends instructions to the readout circuit, which reads the corresponding expert weight data from the storage circuit according to the instructions. The near-memory computing circuit in the readout circuit performs feedforward neural network calculations based on the read expert weight data and transmits the results of the feedforward neural network calculations to the computing module. The computing module then performs the remaining calculations in the decoding phase, thereby enabling the AI device to complete the inference function.
[0050] For example, after the sparse gating network computes the expert routing result (such as the Top-K expert index) for the current token in the computing module, the computing module sends this index information as part of an instruction to the read circuit of the first storage module. The near-memory computing circuit in the read circuit of the first storage module can parse this instruction and directly locate and read the corresponding expert weight data within the storage circuit based on the instruction. The near-memory computing circuit inside the read circuit can combine the activated expert weight data read from the storage circuit to perform feedforward neural network computation. Therefore, the feedforward neural network computation in the decoding stage is completed before the expert weight data leaves the first storage module, realizing "read-and-compute".
[0051] In one embodiment, the computation module is used for pre-filling processing in the large model inference stage, and the near-memory computation circuit is used for performing the first linear transformation of the MOE expert network layer in the decoding stage and transmitting the result of the first linear transformation to the computation module. The computation module is also used to combine the result of the first linear transformation to perform the remaining computations in the decoding stage. Thus, in the decoding stage of MOE large model inference, the first linear transformation of the MOE expert network layer (usually the layer most directly dependent on expert weights) can be completed entirely at the storage end. Only the intermediate computation results (the amount of data in the intermediate computation results is much smaller than the amount of data in the original expert weights) are transmitted back to the computation module for subsequent activation function calculations, expert output weighted summation, and other decoding operations. This fundamentally restructures the computational pipeline of the decoding stage, avoiding the long-latency process of multiple steps in the prior art: "reading all expert weights → transmitting all expert weights to the main computation layer → the main computation layer performing decoding computations." This invention can adjust the decoding stage to an efficient process of "memory-side computation → transmitting simplified intermediate results to the computation module → the computation module performing subsequent decoding computations." This invention, through the aforementioned deep integration, shifts the paradigm from the "computation-centric" approach of existing technologies to the "data-centric" approach of this invention, bringing multiple system-level advantages. For example, this invention can greatly alleviate bandwidth bottlenecks: the computation of the first-layer linear transformation of the most densely packed MOE expert network layer in the MOE large model decoding stage is digested within the first storage module. This eliminates most of the cross-layer, long-distance weight data transport and significantly reduces the movement of large-volume weight data. Because the large-volume weight data from the first storage module does not need to be moved to the computation module, this large-volume weight data is directly computed within the first storage module using a tightly coupled feedforward neural network. Therefore, only the intermediate results (which are smaller in volume) generated by the tightly coupled computation need to be moved to the computation module. This directly removes the main performance constraints of the decoding stage. It can significantly (e.g., by 1-2 orders of magnitude or more, x 10-100) increase the number of tokens generated per second by the MOE large model while maintaining the bandwidth provided by the first storage module. This invention can also significantly reduce end-to-end latency and power consumption: by performing the above-mentioned tightly coupled reading calculations in the near-memory computing circuit of the read circuit of the first storage module, the sharp reduction in the movement of expert weight data is directly translated into a reduction in the latency and energy consumption of token generation. This is especially significant for long text generation or interactive applications, resulting in a very significant improvement in user experience.
[0052] The near-memory computing circuit provided by this invention can also be used in the decoding stage of large model inference to calculate attention based on the weight data of the large model read from the storage circuit. This further reduces the impact of data movement from the storage circuit to the computing module, such as the amount of data moved, the number of moves, the time of movement, and the power consumption, and shortens the inference time. Correspondingly, it improves the bandwidth utilization of the first storage module, increases the number of tokens processed per second by the AI device, improves the efficiency of inference, and improves the user's inference experience.
[0053] In a preferred embodiment of the present invention, the weight data of the large model in the storage circuit is divided into several blocks. When the readout circuit reads the (N+1)th block of weight data in the storage circuit (N is a positive integer), the near-memory computation circuit within the readout circuit performs feedforward neural network computation on the read Nth block of weight data, thereby achieving simultaneous reading and computation within the first storage module. See also Figure 2 As shown, the yellow blocks represent memory access and loading overhead, and the light green blocks represent matrix multiplication computation overhead. Assuming reading a block of weighted data takes 1 tick, and assuming calculating the weighted data takes 1 tick, if we follow... Figure 2 Option 1 involves reading the weight data block by block. After all blocks (e.g., the 16 blocks in the diagram) have been read from the storage circuit, the calculation is performed within the read circuit of the storage chip. This completes the memory access and calculation of the 16 blocks of weight data, requiring at least 32 ticks (because memory access and calculation each require 16 ticks). If the weight data is read block by block from the storage chip to the calculation chip for calculation, the weight data also needs to traverse chip-level distances and time, requiring even more time. If we follow... Figure 2 In Scheme 2 (i.e., an embodiment of the present invention), when the read circuit in the first storage module reads the second block of weight data, the near-memory computing circuit in the first storage module can simultaneously calculate the weight data of the first block that has already been read. Thus, completing the memory access and calculation of the 16 blocks of weight data requires only 17 ticks (most of the computational overhead is hidden). Therefore, this implementation can significantly increase the number of tokens processed per unit time by the AI device and greatly reduce the latency of the AI device's operation.
[0054] In a preferred embodiment, the first storage module and the computing module are stacked. More preferably, the first storage module and the computing module are stacked using a hybrid bonding method. See also Figure 3The NPU-based computing module and the first storage module are vertically stacked via hybrid bonding. The computing module directly reads data from the first storage module to perform calculations based on that data. The computing module then transmits the intermediate data generated during the calculations to the second storage module for storage. Hybrid bonding is an advanced stacking technology that enables direct contact bonding of the metal layers (usually copper) and dielectric layers of two chips. For example, hybrid bonding bonds copper to copper and dielectric layers to dielectric layers, replacing traditional bump or solder ball interconnects with direct copper-to-copper connections. This allows for ultra-fine pitch stacking and packaging within a very small space, achieving three-dimensional integration. Hybrid bonding allows the metal layers of two or more chips to be precisely aligned and directly pressed together, forming direct electrical contacts. Compared to traditional bonding technologies, hybrid bonding can achieve sub-micron or even nanometer-level interconnect pitches, allowing for more connection points to be placed in a smaller area, significantly increasing the data communication bandwidth between chips. Its contact density can reach 10K-1MM / mm2. Moreover, since hybrid bonding eliminates intermediate materials such as solder, the direct copper-to-copper connection has lower resistance, reducing energy loss during signal transmission and also reducing signal propagation time delay. The compact structure and direct conductive path of hybrid bonding help improve thermal management and reduce heat generation, which is especially important for high-performance computing such as artificial intelligence.
[0055] Thanks to the high-density interconnects of hybrid bonding, its bandwidth can be increased exponentially. For example, as shown in Table 1, hybrid bonding can achieve bandwidths of tens of TB. Compared to the 70-150GB bandwidth of existing LPDDR6 and the advanced GDDR6 (150-300GB bandwidth), the bandwidth and energy requirements for transmitting a single bit of data are greatly improved, significantly reducing the power consumption of data transfer. The bandwidth supported by the advanced LPDDR6 connection method in LPDDR is usually only tens of GB / s, which is difficult to support the bandwidth required for decoding large models, which often reaches hundreds of GB or even TB per second. Correspondingly, the bandwidth of the AI device provided in this application can be increased exponentially. By hybrid bonding the first storage module and the computing module, and having the near-memory computing circuit in the first storage module perform feedforward neural network calculations during the decoding stage, this invention can better alleviate bandwidth pressure, better support the ultra-high bandwidth required by large models with increasingly large parameters, and achieve lower power consumption. Furthermore, preferably, the first processor in the computing module is a neural network accelerator (NPU). The NPU-based computing module enhances the computing power of the module, enabling it to perform neural network calculations faster and more efficiently. Figure 3The AI device shown can complete training and / or inference faster and with lower power consumption, and can support large models with a larger number of parameters, thereby realizing more powerful AI functions. In particular, it can better solve the problems of bandwidth bottlenecks, power consumption, heat dissipation and battery life encountered in the deployment of edge AI large model (LLM) chips. Table 1 In one embodiment of the present invention, the first storage module is preferably a non-volatile memory, and / or the second storage module is a volatile memory. This allows the two storage modules to store parameters / data with different characteristics, thereby supporting the storage needs of the AI device on the edge while balancing storage capacity, cost, read / write speed, and reliability. The bandwidth of the first storage module is greater than that of the second storage module. Correspondingly, the first storage module, stacked close to the computing module, can store parameters of large models with a large number of parameters and can support the high bandwidth required for transmission between the parameters with a large amount of data in the first storage module and the computing module. The second storage module can use a memory with a slightly smaller storage capacity to store parameters with a smaller amount of data and intermediate results with a smaller amount of data in the large model.
[0056] Figure 4 This is a schematic diagram of the structure of an AI device provided in a preferred embodiment of the present invention. The first storage module is stacked with the computing module using a hybrid bonding method. Since the second storage module is typically used to store smaller amounts of data, such as intermediate calculation results and a small number of large model parameters, it can employ a memory with slightly lower bandwidth, such as DRAM. Figure 4 In this design, the first storage module is a NAND-based storage device, and the second storage module is a DRAM memory. Furthermore, the second storage module is stacked on top of the first storage module. This design can further reduce circuit wiring, shorten the distance the computing module travels to read data from the DRAM, reduce power consumption during data reading, reduce the area occupied by the AI device, and increase bandwidth.
[0057] It is worth noting that the first storage module is stacked with the computing module containing the CIM-based NPU through hybrid bonding, which makes the design of the first and second storage modules simpler and more efficient, and the data transmission method simpler and more efficient.
[0058] In one embodiment of the present invention, the computing module, the first storage module, and the second storage module are stacked from bottom to top through a layered packaging method, which further enables the AI device to support higher bandwidth, more compact structure, smaller area, smaller size, and lower power consumption, making it more suitable for edge AI devices.
[0059] In one embodiment of the present invention, the second storage module is electrically connected to the computing module in a flat manner (i.e., the two are not stacked). The second storage module and the computing module are disposed on the first substrate in a flat manner, and the first storage module is stacked on the computing module in a hybrid bonding manner. This allows the AI device to ensure that the first storage module supports high bandwidth and low power consumption through hybrid bonding, while the flat arrangement of the second storage module simplifies the packaging process and reduces costs.
[0060] Preferably, the first storage module is a Nand Flash memory, and / or the second storage module is a DRAM memory. This allows the AI device to simultaneously meet the requirements of storage capacity and bandwidth, while maintaining low cost. In a preferred embodiment, the first storage module includes at least two layers of Nand dies, which are stacked together using a hybrid bonding method to form the storage circuit. The readout circuit of the first storage module is stacked and electrically connected to the storage circuit, and the readout circuit is located between the storage circuit and the computing module. This further reduces the distance and power consumption for reading weight data of large models from the first storage module to the computing module, and the entire AI device is more compact, smaller, lower in cost, and more suitable for edge electronic devices. This makes the AI device particularly suitable for small and medium-sized edge devices, such as mobile phones, tablets, and computers.
[0061] exist Figure 4 In this design, the storage circuit is formed by hybrid bonding of more than two hundred layers of NAND-Die, with the readout circuit located below it. NAND Flash memory features high storage density (e.g., 20-36 Gb / mm²) and low storage cost, making it well-suited for edge-deployed tasks involving storing large model parameters. The second storage module can selectively store smaller parameters and smaller amounts of intermediate data from the large model, while the first storage module, hybrid-bonded with the computation module, stores complex and large-volume parameters (e.g., numerous expert weight parameters). Therefore, AI devices can further reduce their memory requirements, for example, by using fewer, smaller, and lower-cost memory modules. The small amount of parameters stored in the DRAM memory can include, for example, parameters of a linear attention model and / or partially varying weight parameters from online reinforcement learning. This further reduces the size, power consumption, and cost of the AI device, allowing it to simultaneously possess the advantages of high storage capacity, high bandwidth, small size, and low cost. It is worth noting that the first and second storage modules can also be other types of memory.
[0062] In a preferred embodiment, the storage density of the first storage module is greater than that of the second storage module. This allows the first storage module to store larger models with a greater number of parameters, while the first and second storage modules as a whole still maintain good storage capacity and bandwidth. More preferably, the ratio of the number of parameters stored in the first storage module to its bandwidth is greater than the ratio of the number of parameters stored in the second storage module to its bandwidth. For example, the ratio of the number of parameters stored in the Flash storage module to its bandwidth is A, and the ratio of the number of parameters of the large model stored in the DRAM storage module to its bandwidth is B, where A is greater than B. This allows the first and second storage modules to be two different types and functions of storage modules, enabling the selective transfer of different parameters between storage modules and computing modules. For example, sparse parameters in a large model can be stored in the Flash storage module, while dense parameters and intermediate data in the large model can be stored in the DRAM storage module. This reduces the amount of data transferred, the power consumption during data transfer, and the heat generated during data transfer, while also reducing memory costs and effectively utilizing bandwidth and alleviating bandwidth bottlenecks.
[0063] See Figure 4 In the preferred embodiment shown, to better package the AI device, making its structure more compact, smaller, more stable, and applicable to more application scenarios, the AI device further includes: a first substrate disposed below the computing module, and a second substrate disposed between the first storage module and the second storage module. The first substrate, computing module, first storage module, second substrate, and second storage module are stacked from bottom to top using a layered packaging method. The first storage module is a Nand Flash memory, and the second storage module is a DRAM memory. The second substrate has through-holes, and the DRAM memory is electrically connected to the computing module through wires passing through the through-holes of the second substrate. This reduces the size of the AI device while significantly lowering its power consumption, and simplifies the packaging process and reduces the cost of the AI device, making it more suitable for edge AI devices and electronic devices requiring low power consumption and small size.
[0064] Figure 4In this embodiment, the NPU in the computing module is a neural network accelerator based on in-memory computing (CIM). The computing module can utilize a chip with medium to high computing power (e.g., less than tens of TOPS) to achieve edge-side AI functions based on large models. This structure in this embodiment solves the problems of insufficient parameter storage capacity and high cost faced when deploying large models on the edge. The first storage module provides, for example, 50GB-200GB of large model parameter storage space within a chip area of 20-70 square millimeters, which is acceptable for small to medium-sized edge devices, corresponding to support edge-side deployment of large models with 50 billion to 150 billion parameters. At the same time, the lower cost of 3D-NAND-FLASH memory (e.g., $1.3 / 128Gb) significantly reduces the cost of deploying large models on the edge. For example, the cost of storing GPT-OSS 120B MOE large model parameters (120 billion) can be reduced to less than $10 (traditional solutions require at least $50); the cost of storing Qwen3-30B MOE large model parameters (30 billion) can be reduced to less than $3.
[0065] See Figure 5 The computing module provided in this embodiment includes: an I / O system for inputting and outputting data; a second processor (i.e., a RISC-V-CPU in the figure) electrically connected to the I / O system for preprocessing the input data to make the preprocessed data conform to the format of the large model, such as normalizing, resizing, and encoding the input data (e.g., text, images, or sensor data) to conform to the input format of the large model; and several first processors (in the figure, the first processor is a neural network accelerator NPU) electrically connected to the second processor for performing neural network calculations on the preprocessed input data. This implementation can make data processing more accurate and efficient, improving the computing power and energy efficiency of AI devices.
[0066] In AI devices, such as AI large-model inference chips, computation involves a large number of matrix multiplication calculations (multiply-accumulate MAC), which account for 30-70% of the overall chip power consumption. As Moore's Law gradually approaches its limits, the "memory wall" and "power wall" problems of AI chips based on the von Neumann architecture are becoming increasingly prominent, and the growth rate of chip computing power is slowing down. In one embodiment of the present invention, the first processor in the computing module is an NPU, and the NPU is based on in-memory computing (CIM). That is, the computing module includes a CIM-based neural network accelerator NPU, which includes an in-memory computing matrix. The in-memory computing matrix is used to perform neural network calculations based on input data and data from the first and second storage modules. Therefore, the neural network accelerator NPU included in the computing module achieves in-memory computing through the in-memory computing matrix. In-memory computing is an innovative architecture designed to overcome the "memory wall" problem by integrating storage and computing functions on the same chip. This technology embeds computing power into memory and employs a new computing architecture to perform matrix-vector multiplication and accumulation operations. This reduces the frequent data transfer between memory and computing chips, significantly decreasing transfer time, power consumption, and heat generation, and improving parallel data processing efficiency. It enables in-situ computation, eliminating bandwidth limitations and data movement costs. This fundamentally eliminates unnecessary data transfer latency and power consumption, improving AI computing efficiency by hundreds or thousands of times, reducing costs, and breaking down the "memory wall" and "power wall." Compared to existing devices performing computations under the same conditions, the NPU and AI devices containing it can have greater computing power, consume less power, and generate less heat. This avoids the limitations of power consumption and area in existing edge devices that prevent the integration of large-scale high-performance computing units, leading to insufficient computing power and excessively long inference latency when processing large models. It also addresses existing heat dissipation problems at their source.
[0067] See Figure 6In the illustrated embodiment, each neural network accelerator (NPU) includes: a pre-processing module (i.e., Pre-Processing in the figure), used to convert the pre-processed input data into matrix data conforming to the processing format of the in-memory computing module; several neural network sub-accelerators (SNPUs), used to perform neural network calculations on the data processed by the pre-processing module; and a vector processing module (VPU), used to perform vector processing on the data processed by the SNPUs. Each SNPU includes several in-memory computing modules (CIMs) (CIMD in the figure is illustrated as digital CIMs), arranged in a matrix to form the in-memory computing matrix. The CIM-based NPU provided in this embodiment, combining a pre-processing module, a matrix-like in-memory computing module, and a vector processing module, can significantly improve the computing power and energy efficiency of the NPU and AI device, while greatly reducing power consumption and heat generation. Furthermore, it has a simple structure and low cost. The figure illustrates that the four neural network sub-accelerators, SNPU0 to SNPU3, are connected to the bus in parallel. The NPU also includes on-chip memory (such as...). Figure 5 The on-chip SRAM can be used to store intermediate data generated during in-memory computation matrix calculations. The on-chip SRAM cache and DMA (Direct Memory Access) system in the computing module handle memory management and data transfer optimization during large language model inference. On-chip memory can store large model parameters and intermediate data, reducing access to off-chip memory (such as DRAM) and lowering memory bandwidth bottlenecks. The computing module provided in this embodiment has advantages such as high computing power, high energy efficiency, low cost, low power consumption, and low heat generation.
[0068] See Figure 7The embodiment of the present invention shown illustrates a schematic diagram of a neural network accelerator (NPU). The NPU includes: a preprocessing module, an in-memory computation matrix, and a vector processor, all electrically connected in sequence; and a shared memory electrically connected to the preprocessing module, the in-memory computation matrix, and the vector processor, respectively. It also includes a chip controller electrically connected to the shared memory. The NPU is also electrically connected to an external memory, and data stored in the external memory is transferred to the NPU's shared memory in batches. For example, weights are read from the external memory in batches and then transferred to the shared memory. One possible operation is briefly described as follows: the preprocessing module obtains input feature data from the shared memory, processes the input feature data into matrix data in the format of an in-memory computation matrix, and inputs the matrix data into the in-memory computation matrix; the in-memory computation matrix performs convolution calculations based on the matrix data and the weight data to obtain a calculation result, which is then input into the vector processor; the vector processor performs vector processing on the calculation result to obtain output feature map data, which is then written into the shared memory. The NPU in the diagram has a fast neural network processing speed, low power consumption, and high accuracy, which can improve the working efficiency and accuracy of AI devices. It also has low power consumption, low heat generation, and long battery life.
[0069] See Figure 8 In this embodiment, the in-memory computation matrix includes: N columns of SRAM-based in-memory computation submodules ( Figure 8The diagram illustrates one column. Each column of the in-memory operator module includes: M rows of in-memory units for storing and calculating the stored data with the input feature data, and capacitors electrically connected to the in-memory units. The charges output from the capacitors in the M rows are pooled together to achieve charge accumulation. Therefore, the in-memory unit can store and perform multiplication calculations on the weights and input feature data. The corresponding product is accumulated through the charges of the capacitors connected to the in-memory unit to realize the multiply-accumulate (MAC) calculation in neural network computation. That is, it can perform matrix calculations within the in-memory computation matrix, which can avoid the frequent data transfer between two different and geographically distant memories and computing chips in existing technologies. Adders / addition trees can also be used instead of capacitors to accumulate the results calculated by the in-memory units. Multiply-accumulate calculations typically account for a large portion of the total power consumption of the computation, often reaching 50-70%. This embodiment, based on an in-memory computing architecture, incorporates a neural network accelerator with an in-memory computing matrix within the computing module. This new computing architecture enables matrix multiplication / addition operations, significantly reducing the power consumption of neural network computations. Furthermore, it allows for a smaller NPU area and, within the same area and power consumption, supports greater computing power while generating less heat, thus addressing the heat dissipation problem at its root. The computing module, including the in-memory computing matrix, is stacked with the first storage module using a hybrid bonding method, further alleviating the bandwidth issue of this AI device. It also boasts high computing power, low power consumption, small size, and low cost.
[0070] See Figure 9 The diagram shown is a vertical structural diagram of a column of stored-operator submodules in an in-memory computation matrix, according to an embodiment of the present invention. Each stored-operator unit includes an SRAM memory and a logic unit (i.e., digital logic in the diagram) adjacent to the SRAM memory. Multiple stored-operator units are stacked to form a column, and several columns of stored-operator units form a stored-operator matrix. The aforementioned logic unit can be used to perform multiplication and addition calculations on the weights and input feature data in the SRAM memory. Figure 9 The diagram also showcases a block diagram of Integrated In-Memory Computing (CIMD). CIMD significantly reduces the power consumption of data transfer while greatly increasing the speed of parallel multiply-accumulate (MAC) calculations, effectively alleviating the two major challenges of the "memory wall" and the "power wall." SRAM memory and the logic units used for computation are stacked adjacent to each other, which allows for a more compact structure of the in-memory computation matrix, resulting in greater computing power, lower power consumption, and, combined with digital logic units, faster and more accurate computations.
[0071] Figure 10This is a schematic diagram illustrating the principle of in-memory computation matrix calculation in one embodiment of the present invention. The in-memory computation unit performs multiplication calculations on the input feature data and weights (e.g., parameters / data from the first and second storage modules), and then accumulates the product through a multi-level addition tree to obtain the result of multiplication and accumulation.
[0072] As shown in Table 2, at the 12nm process node, the matrix multiplication efficiency of the digital CIM (i.e., CIMD) in-memory computing unit can reach 24.7 TOPS / W, which is an order of magnitude improvement compared to traditional digital architectures (i.e., general-purpose computing units, such as GPUs and TPUs). In-memory computing significantly reduces the power consumption of data transfer while greatly improving the speed of parallel multiply-accumulate (MAC) computation. In-memory computing (CIMD) technology can significantly improve the energy efficiency of AI large model (LLM) inference chips, that is, while keeping the computing power constant, it significantly reduces its power consumption requirements.
[0073] Table 2 It is worth noting that existing computing chips typically have high power consumption, resulting in significant heat generation from these chips and AI devices containing them, making it easy for ambient temperatures to exceed 85 degrees Celsius. However, traditional memory devices generally cannot withstand ambient temperatures above 85 degrees Celsius. Under these high temperatures, leakage current in memory devices increases, requiring higher refresh rates, which in turn leads to even higher power consumption and heat generation. This further increases ambient temperature, further exacerbating leakage current and creating a vicious cycle that significantly impacts AI devices. The AI device provided in this application combines a computing module including a CIM-based NPU with a first storage module including near-memory computing circuitry stacked with the computing module. This improves computing power and efficiency while reducing power consumption and heat generation, thus fundamentally solving the power consumption and heat dissipation problems, and consequently addressing the aforementioned heat / temperature-related issues.
[0074] In another preferred embodiment of the present invention, the computing module further predicts the neuron parameters to be activated in the large model parameters of the first storage module based on the input data and the sparsity prediction model. The computing module reads the predicted neuron parameters to be activated from the first storage module based on the prediction results, and performs neural network calculations based on the read neuron parameters, the input data, and the data from the second storage module. Therefore, the present invention can fundamentally reduce the computing module's access to and read operations on the first storage module, fundamentally reduce the amount of data transmission between the first storage module and the computing module, and reduce the amount of data required for calculation, thereby reducing the bandwidth requirements of the first and second storage modules. The computing module only needs to read the predicted neuron parameters to be activated from the first storage module and perform calculations based on them, thus further reducing bandwidth requirements, power consumption, and heat generation, and correspondingly reducing related costs. The computing module includes an NPU for neural network processing, which can further improve the energy efficiency of neural network processing and reduce corresponding power consumption and heat generation, thereby solving the heat dissipation problem at its source and providing greater computing power. Therefore, the AI device provided by this invention can simultaneously solve problems related to storage capacity, bandwidth bottlenecks, computing power, power consumption, energy efficiency, heat dissipation, and inference latency. Taking the GPT-OSS-120B MOE large model (120 billion parameters) as an example, during the decoding phase, only 5.1 billion parameters (4-5% of the total parameters) are activated when processing each token. By decomposing the model into multiple smaller specific networks (expert networks), and dynamically selecting which experts to activate based on the input data and a sparse prediction model, this selective activation mechanism ensures that the model uses only the necessary computing resources when processing each task. This reduces data handling and the amount of data involved in computation from the source, thereby reducing related power consumption and heat.
[0075] This invention stacks a computing module and a first storage module, offloading feedforward neural network computation to the storage end via a near-memory computing circuit within the first storage module. Combined with an NPU based on a CIM architecture, this increases data transmission bandwidth to 100-200 GB / s or even higher, resulting in a 10-20 times increase in tokens per second (TPS) during the decoding phase. Furthermore, it reduces data movement by over 90%, improving the system's energy efficiency ratio (Token / J) by at least an order of magnitude. For medium-to-large-scale edge models like the GPT-OSS-120B MOE, the 100-200 GB / s bandwidth of the 3D-NAND-FLASH in this invention provides inference performance of 200-400 Token / s or more. For small-to-medium-scale edge models like the Qwen3-30B MOE, the 100-200 GB / s bandwidth of the 3D-NAND-FLASH in this invention provides inference performance of 300-600 Token / s or more.
[0076] In one embodiment of the present invention, the computing module further includes a storage controller, which is used to control the scheduling of data in the first storage module. For example, the storage controller issues instructions to control the access mode of the read circuit in the first storage module to the data in the storage circuit. It can efficiently control the first storage module to achieve read-while-computing, reduce the burden on the NPU and allow the NPU to focus on core AI tasks such as matrix operations. The fine scheduling of the storage controller can bring significant power consumption optimization, thereby improving the overall system's computing power utilization and energy efficiency ratio. The NPU's computing DIE process is usually better, and the NPU's corresponding process is superior to NAND and DRAM. Therefore, setting a storage controller in the computing module based on the NPU can greatly reduce its power consumption. For example, the computing DIE process in the NPU is usually only a few nanometers, such as the common 5nm; while the read DIE process of the memory is usually tens of nanometers, such as DRAM read DIE is usually 28nm. Therefore, the storage controller structure of the present invention can reduce its power consumption by 2-4 times and make the AI device smaller.
[0077] This invention allows the second storage module to use conventional, ordinary memory, such as conventional DRAM, and enables the use of only one conventional DRAM (avoiding the need for multiple DRAMs), thereby further reducing the number of memory modules and significantly lowering the cost, size, and power consumption associated with memory. DRAM memory can meet the requirements with only 1-4GB of storage capacity and less than 10GB / s of bandwidth, and the corresponding AI device can support large models with up to 120 billion parameters. Through the sequential stacking of Nand flash memory and the computing module based on in-memory computing, it solves the bottlenecks of both storage capacity and storage bandwidth and computing power.
[0078] In one embodiment, the large model includes a large model based on the MOE architecture. The MOE architecture (Sparse model) differs from the traditional Transformer architecture (Dense model) in that the traditional Transformer architecture shares all parameters between input and output, is fully connected, and all parameters are activated and participate in computation, such as the fully connected layers in CNNs and Transformers. The core idea of the MOE model is to design multiple "experts" (sub-networks) within the model, and through a gating mechanism, only a portion of these experts are activated for computation on each input or token. In the AI device provided by this invention, the model capacity can be significantly increased in terms of parameter scale, while only a small portion of the parameters actually participate in computation. Only a portion of the experts are activated during inference, thus achieving higher parameter utilization per unit inference time and achieving a certain "computing power economy" effect. The first storage module stores at least the parameters of the sparse part of the large model, and the second storage module stores at least the parameters of the dense part of the large model and intermediate data generated by neural network computation. For example, NAND stores the parameters of the sparse part of the large model A, and DRAM stores the parameters of the dense part of the large model A and intermediate data generated by neural network computation. In another embodiment, the first storage module stores a sparse large model A, and the second storage module stores a dense large model B and intermediate data generated by neural network computation. Dense parameters typically refer to those frequently accessed during the operation of the large model; sparse parameters typically refer to model parameters that are accessed less frequently during the operation of the large model. This allows the first and second storage modules to choose different types of memory and combine their respective characteristics for better division of labor and cooperation. Consequently, the AI device can better store more and larger large models with larger parameters and can be compatible with more types of large models, making the AI device more functional and powerful. Moreover, since only a very small number of parameters (e.g., 5-6%) are activated during actual inference execution, this embodiment also significantly reduces the amount of data and power consumption for moving large model parameters, significantly reduces the bandwidth requirements for storage and movement, and reduces heat generation, thereby improving the energy efficiency of large model operation.
[0079] It's worth noting that the cost of memory in AI devices typically accounts for the majority of the overall cost. Furthermore, with the current development of AI technology, the cost / price of storage devices is skyrocketing. However, if existing AI devices are to support large models (e.g., hundreds of billions of parameters) and run large models with massive parameter counts, current technologies require a significant increase in the amount of memory, leading to a sharp rise in the overall cost and a larger size of the AI device. Edge AI devices, especially AI smartphones with their limited space, have very high requirements for low cost and small size, placing even greater constraints on their size. Therefore, the AI device provided by this invention categorizes and selectively loads the parameters and intermediate data of large models into two storage modules. This allows relevant data to be moved between different storage and computing modules in a categorized manner. Combined with sparse prediction, it selectively reads only a portion of the parameters and moves only that portion to the computing module for computation, thereby reducing the performance requirements of storage devices and also reducing the number and cost of storage devices. Moreover, the issues of storage capacity and bandwidth can be solved simultaneously by using fewer FLASH NAND memories and fewer DRAM memories. By stacking the computing module, FLASH memory module, and DRAM memory module according to the above-mentioned positional structure, the cost and size of storage devices in AI devices can be greatly reduced, thereby significantly reducing the cost and size of AI devices. It can also provide large storage capacity, high bandwidth / high bandwidth utilization, high computing power, and high energy efficiency.
[0080] In a preferred embodiment of the present invention, the second storage module is further used to store the changing weight parameters, and the calculation module further performs online reinforcement learning based on the changing weight parameters. This enables the AI device of the present invention to learn autonomously according to the dynamically changing environment, thereby making the inference of the AI device more accurate. For example, the second storage module stores some of the changing weight parameters of online reinforcement learning (LORA), allowing the AI device to learn continuously. In online learning, data is gradually collected during the interaction process, so the large model can update its strategy in real time to adapt to changes in the environment, thereby enabling the AI device to perform more accurate inference according to different input requirements.
[0081] The AI device provided by this invention uses two storage modules to store different types of data, with the first storage module integrating a near-memory computing circuit for feedforward neural network computation. This allows the first storage module to store a larger number of parameters for the large model than the second storage module. For example, the first storage module stores a large number of expert weight parameters, while the second storage module stores a smaller number. Therefore, this invention combines the division of labor between the first and second storage modules with a computing module, a first storage module hybridized with the computing module, a near-memory computing circuit, and a sparsity prediction method, enabling the AI device to simultaneously and better address issues related to storage capacity, bandwidth bottlenecks, computing power, power consumption, energy efficiency, heat dissipation, and inference latency.
[0082] In one embodiment of the present invention, the first storage module is a Nand Flash memory, and the second storage module is a DRAM memory. The Nand Flash memory is also used to store a sparse prediction model, which is a low-rank prediction model. During inference, the computation module predicts the parameters of the neurons to be activated based on the current input data and the low-rank prediction model read from the first storage module. The computation module selectively reads the parameters of the neurons to be activated from the first storage module (e.g., some experts) based on the prediction results, performs neural network calculations based on the some experts, and stores the intermediate data generated during the calculation process in the second storage module. That is, which neurons will be activated is related to the current input. Utilizing the high sparsity of MLP, a simple low-rank prediction model (e.g., a two-layer MLP) is used to predict which neurons will be activated in real time during large model inference operations. Each time a large model is inferred, only the neuron parameters in the activated state need to be dynamically loaded from the Nand Flash memory. This scheme can significantly reduce the amount of model parameters that need to be read from the Nand Flash memory (by up to 90%), while also significantly reducing the power consumption and bandwidth requirements for data transfer from the Nand Flash memory.
[0083] The storage module and computing module provided by this invention can take many forms, for example, they can be chips or DIEs. The AI device provided by this invention can take many forms, for example, it can be a chip.
[0084] In summary, the AI device provided by this invention can solve the problems currently encountered by AI devices based on large models, such as storage capacity, bandwidth, power consumption, heat dissipation, computing power, energy efficiency, and inference latency. It can significantly improve storage capacity, bandwidth, and computing power, reduce power consumption, reduce inference latency, and improve inference performance, meeting the needs of high-performance AI inference, and its cost is relatively low. It can be widely applied to various AI devices, such as mobile phones, computers, and robots. The AI device provided by this invention can be used on the edge and in the cloud. Moreover, it is well-suited for edge application scenarios that require small area, small size, low power consumption, long battery life, fast response, and high inference performance.
[0085] The present invention also provides an electronic device comprising the aforementioned AI device. This device can significantly reduce the power consumption and cost of electronic devices, while improving the efficiency and experience of AI inference. It can be widely used in electronic devices such as smartphones, tablets, wearable electronic devices, robots, and smart home devices.
[0086] The above description is merely a preferred embodiment of the present invention. Those skilled in the art will understand that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of the present invention. Furthermore, under the teachings of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the protection scope of the present invention.
Claims
1. An AI device based on near-memory computing, characterized in that, include: A computing module, a first storage module electrically connected to the computing module, and a second storage module electrically connected to the computing module; The computing module includes a first processor for performing neural network processing, the first processor being used to perform neural network calculations based on input data and data from a first storage module and a second storage module; The first storage module is used to store at least the parameters of the large model, and the second storage module is used to store at least the intermediate data generated by the computing module in performing neural network calculations; The first storage module includes a storage circuit for storing data and a readout circuit for reading data from the storage circuit. The readout circuit is internally equipped with a near-memory computing circuit. The near-memory computing circuit within the first storage module is used at least during the decoding stage to perform feedforward neural network calculations based on the weight data of the large model read from the storage circuit.
2. The AI device according to claim 1, characterized in that, The large model includes a large model based on the MOE architecture, and the near-memory computing circuit within the readout circuit is used to perform feedforward neural network calculations based on expert weight data read from the storage circuit during the decoding stage.
3. The AI device according to claim 2, characterized in that, The near-memory computing circuit within the readout circuit is used to perform lightweight gated network computations required for expert routing logic, and / or multiply-accumulate computations to activate expert weights, and / or computations of preliminary scalar combinations of expert outputs.
4. The AI device according to claim 2, characterized in that, The computation module is used for the pre-filling stage of large model inference, the near-memory computation circuit is used for the first linear transformation of the MOE expert network layer in the decoding stage of inference, and the computation module is also used to perform the remaining computations in the decoding stage by combining the result of the first linear transformation.
5. The AI device according to claim 1 or 2, characterized in that, The weight data of the large model in the storage circuit is divided into several blocks. When the readout circuit reads the (N+1)th block of weight data in the storage circuit, the near-memory computing circuit in the readout circuit performs feedforward neural network calculation on the weight data of the Nth block that has been read, so as to realize simultaneous reading and calculation in the first storage module.
6. The AI device according to claim 5, characterized in that, The near-memory computing circuit is also used in the decoding stage of large model inference to calculate attention based on the weight data of the large model read from the storage circuit.
7. The AI device according to claim 1 or 2, characterized in that, The computing module includes: An I / O system is used for inputting and outputting data. A second processor, electrically connected to the IO system, is used to preprocess the input data so that the preprocessed input data conforms to the format of the large model; and A plurality of the first processors electrically connected to the second processor are used to perform neural network calculations on the preprocessed input data.
8. The AI device according to claim 7, characterized in that, The first processor is a neural network accelerator (NPU), and the NPU includes: The preprocessing module is used to convert the preprocessed input data into matrix data that conforms to the processing format of the in-memory computing module; Several neural network sub-accelerators (SNPUs) are used to perform neural network calculations on the data processed by the preprocessing module. Each SNPU includes several in-memory computing modules (CIMs) arranged in a matrix to form an in-memory computing matrix. The vector processing module is used to perform vector processing on the data processed by the SNPU.
9. The AI device according to claim 8, characterized in that, The in-memory computation matrix includes: N columns of SRAM-based in-memory computation submodules, each column of which includes M rows of in-memory computation units for storing and computing the stored data with the input feature data, and capacitors or adders for accumulating the results of the in-memory computation units.
10. The AI device according to claim 9, characterized in that, The calculation module is also used to predict the neuron parameters to be activated in the parameters of the large model in the first storage module based on the input data and the sparse prediction model. The calculation module reads the predicted neuron parameters to be activated from the first storage module based on the prediction results, and performs neural network calculation based on the read neuron parameters, the input data and the data from the second storage module.
11. The AI device according to any one of claims 1 to 10, characterized in that, It also includes a storage controller disposed within the computing module, the storage controller being used to control the scheduling of data in the first storage module; the read circuit in the first storage module reads data from the storage circuit according to instructions from the storage controller of the computing module.
12. The AI device according to any one of claims 1 to 10, characterized in that, The first storage module is stacked with the computing module through a hybrid bonding method. The computing module directly reads the data in the first storage module to perform calculations based on the read data in the first storage module. The computing module then transmits the intermediate data generated by the calculation to the second storage module for storage.
13. The AI device according to claim 12, characterized in that, The second storage module is stacked on top of the first storage module, and the computing module, the first storage module, and the second storage module are stacked from bottom to top in a layered encapsulation manner.
14. The AI device according to any one of claims 1 to 10, characterized in that, The second storage module and the computing module are disposed on the first substrate in a flat manner, and the first storage module is stacked on the computing module in a hybrid bonding manner.
15. The AI device according to claim 13, characterized in that, The first storage module includes at least two layers of Nand dies, which are stacked together by a hybrid bonding method to form the storage circuit. The read circuit of the first storage module is stacked with and electrically connected to the storage circuit, and the read circuit is located between the storage circuit and the computing module.
16. The AI device according to claim 15, characterized in that, The second storage module is a DRAM memory; The AI device further includes: a first substrate disposed below the computing module, and a second substrate disposed between the first storage module and the second storage module, wherein the first substrate, the computing module, the first storage module, the second substrate, and the second storage module are stacked from bottom to top in a layered encapsulation manner; The second substrate has through holes, and the DRAM memory is electrically connected to the computing module through the through holes of the second substrate via wires.
17. The AI device according to any one of claims 1 to 10, characterized in that, The first storage module is a Nand Flash type memory, and / or the second storage module is a DRAM type memory.
18. The AI device according to claim 17, characterized in that, The first storage module is a Nand Flash type memory, and the second storage module is a DRAM type memory; The Nand Flash type memory is also used to store a sparse prediction model, which is a low-rank prediction model. The calculation module predicts the parameters of the neuron to be activated based on the low-rank prediction model read from the first storage module and the current input data. The calculation module selectively reads the parameters of the neuron to be activated from the first storage module and performs neural network calculations based on the prediction results, and stores the intermediate data generated during the calculation process in the second storage module.
19. The AI device according to any one of claims 1 to 10, characterized in that, The large model includes a large model based on the MOE architecture, and the second storage module is also used to store the parameters of the large model; The first storage module stores at least the parameters of the sparse part of the large model, and the second storage module stores at least the parameters of the dense part of the large model and the intermediate data generated by the neural network calculation. or The first storage module stores sparse large models, and the second storage module stores dense large models and intermediate data generated by the neural network computation.
20. An electronic device, characterized in that, The AI device includes any one of claims 1-19.