An ai device based on sparsity and electronic equipment

CN122864539APending Publication Date: 2026-10-02REEXEN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610930141.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-11-27
Filing Date
2026-06-25
Publication Date
2026-10-02

AI Technical Summary

Technical Problem

然而,云端大模型也面临一些挑战:1、延迟和时效性:由于需要依赖互联网连接,一旦网络环境不稳定,可能会导致数据传输延迟,影响实时性任务的执行

Benefits of technology

[0029]本发明提供的基于稀疏性的AI装置,其通过基于存内计算的神经网络加速器结合第一存储模块和第二存储模块,以实现神经网络计算和AI功能。其中,两存储模块分工且加载存储不同的数据,使得其可以存储更多更大的参数量,且可以而且,计算模块还根据输入数据和稀疏性预测模型来预测和读取第一存储模块中大模型参数里待激活的神经元参数,其可以从源头上减少计算模块与存储模块之间的数据的访存及传输的消耗,以从源头上减少对带宽的需求,进而可以同时解决存储容量和存储带宽的问题,还可以减少计算模块进行计算的数据量、计算时间、计算功耗和热量。且计算模块包括基于存内计算的神经网络加速器NPU,也可以进一步提高神经网络处理的能效并减少对应的功耗和减少热量的产生,以从源头上解决散热的问题且能提供更大的算力。而且,基于CIM的NPU使得计算模块具有大算力和低功耗的优点的同时,可以减少AI装置因为运行更大的参数量的大模型时出现算力不足及功耗和热量的问题,进而使得AI装置可以更好的支持更大参数量的大模型。所以,本发明提供的基于稀疏性的AI装置,其可以解决算力、能效、存储容量、带宽瓶颈、及散热的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122864539A_ABST
    Figure CN122864539A_ABST
Patent Text Reader

Abstract

The application provides an AI device and an electronic equipment based on sparsity, the AI device comprises a first storage module for storing parameters of a large model, a second storage module for storing intermediate data, and a calculation module, the calculation module comprises a neural network accelerator (NPU) based on in-memory calculation, the neural network accelerator (NPU) comprises an in-memory calculation matrix, the in-memory calculation matrix is used for neural network calculation according to input data and data from the first storage module and the second storage module; the calculation module is further used for predicting neuron parameters to be activated in the parameters of the large model in the first storage module according to the input data and a sparsity prediction model, the calculation module reads the predicted neuron parameters to be activated from the first storage module according to the prediction result, and the calculation module performs neural network calculation according to the read neuron parameters and the input data. The AI device provided by the application can better solve the problems of computing power, energy efficiency, heat dissipation, storage capacity and bandwidth.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, specifically to an AI device and electronic device based on sparsity. Background Technology

[0002] With the rapid development of artificial intelligence (AI) technology, the role and importance of large AI models are becoming increasingly significant. For example, large AI models have become a core force driving the development of natural language processing (NLP) and generative artificial intelligence (AIGC).

[0003] Large-scale models are primarily deployed in the cloud or on edge devices. The main advantages of cloud-based large-scale models include: 1. Powerful computing power: The cloud possesses vast computing, storage, and network resources, enabling efficient processing of large amounts of data and complex computational tasks; 2. Cross-device collaboration: Since computational tasks are completed in the cloud, they are not limited to a single hardware device, thus enabling cross-device collaborative work and providing users with a consistent service experience anytime, anywhere; 3. Rapid model updates: On the cloud platform, models can be quickly updated and maintained, improving their timeliness and accuracy. However, cloud-based large-scale models also face some challenges: 1. Latency and timeliness: Due to the reliance on an internet connection, unstable network conditions may lead to data transmission delays, affecting the execution of real-time tasks; 2. Privacy and security: Transmitting sensitive data to the cloud for processing poses data privacy and security risks, requiring stringent protective measures.

[0004] Edge AI refers to deploying artificial intelligence models directly on terminal devices (such as smart AI smartphones, PCs, robots, and smart wearable devices) to achieve low-latency, high-privacy-protection intelligent applications. Compared to cloud deployment, edge AI can effectively reduce data transmission costs and privacy risks, and has advantages such as faster data transmission speeds and improved user experience. However, edge AI devices have relatively limited computing power and storage capacity, making it difficult to directly deploy and run large models; for example, edge devices cannot support large models that often require tens of billions of parameters.

[0005] Large models are increasingly characterized by a massive number of parameters; for example, DeepSeek's LLM model has an enormous number of parameters. Large AI language models achieve unprecedented performance through massive parameters and complex architectures, but this also places extremely high demands on computing and storage resources. Chat GPT, for instance, has training costs in the millions of dollars, and its inference process consumes a significant amount of computing power. As the scale of large models continues to expand, on the one hand, the demand for intelligent computing power increases dramatically, and corresponding energy efficiency optimization has become an important research direction; on the other hand, the increased scale of large models has led to a significant increase in the number of parameters compared to the past. For example, the number of parameters used to be several million, but it has continued to expand to tens of millions or even hundreds of millions, and now large models even have tens or hundreds of billions of parameters. Large models on chips need to frequently read model parameters during inference, so storage bandwidth has become a key factor restricting the deployment of large models. In particular, the storage bandwidth of current edge chips is insufficient to meet this high-frequency data transmission demand. This not only leads to a significant decrease in the inference speed of large models but also makes the user experience of edge AI far inferior to that of cloud deployment, limiting the widespread application of edge AI. Furthermore, the low energy efficiency of AI chips leads to significant power consumption, battery life, and heat dissipation issues. This is especially true for edge AI chips, which typically rely on battery power and have stringent requirements for power consumption and battery life. Many manufacturers strive to reduce power consumption, minimize components and device size, and increase battery life – these are often key selling points. For example, traditional von Neumann architecture chips suffer from "memory walls" and "power walls," causing a significant increase in power consumption when processing large models. This not only shortens device battery life but also leads to severe heat dissipation problems, further limiting the application scenarios and performance of AI chips. Moreover, the amount of inference required by AI devices is increasing, resulting in ever-increasing power consumption. In existing LLM inference devices, the computational components are located far from the heatsink, making it difficult to effectively dissipate the heat generated by these components, thus posing a serious challenge to LLM inference devices. Current technologies typically address this by increasing and improving the heatsink's cooling capacity, but the improvement is still unsatisfactory.

[0006] In summary, AI devices face multiple technical challenges when deploying large-scale models, including limitations in computing power, energy efficiency, storage capacity, bandwidth, and heat dissipation. These issues severely restrict the deployment and widespread application of large-scale AI models. Therefore, an innovative technical solution is urgently needed to effectively address these problems; this is especially crucial when edge hardware resources are limited. Summary of the Invention

[0007] The purpose of this invention is to provide AI devices and electronic devices based on sparsity to solve some or all of the above-mentioned technical problems.

[0008] To achieve the above objectives, the AI ​​device provided by the present invention includes: a first storage module for storing parameters of a large model; a second storage module for storing intermediate data generated by neural network computation; and a computation module for computation, wherein the computation module is electrically connected to the first storage module for data interaction, and the computation module is electrically connected to the second storage module for data interaction; the computation module includes a neural network accelerator (NPU) based on in-memory computation, the NPU including an in-memory computation matrix, the in-memory computation matrix being used to perform neural network computation based on input data and data from the first and second storage modules; the computation module is further used to predict the parameters of neurons to be activated in the parameters of the large model in the first storage module based on the input data and a sparsity prediction model, the computation module reading the predicted parameters of neurons to be activated from the first storage module based on the prediction results, and the computation module performing neural network computation based on the read neuron parameters and the input data.

[0009] In a preferred embodiment, the computing module includes: an I / O system for inputting and outputting data; a processor for preprocessing the input data so that the preprocessed data conforms to the format of the large model; and a plurality of in-memory computing-based neural network accelerators (NPUs) for performing neural network processing on the preprocessed input data.

[0010] In a preferred embodiment, the neural network accelerator (NPU) includes: a preprocessing module for converting the preprocessed input data into matrix data conforming to the processing format of the in-memory computing module; several neural network sub-accelerators (SNPUs) for performing neural network processing on the data processed by the preprocessing module, wherein each SNPU includes several in-memory computing modules (CIMs) arranged in a matrix to form the in-memory computing matrix; a vector processing module for performing vector processing on the data processed by the SNPUs; and an on-chip memory for storing intermediate data generated by the calculation of the in-memory computing matrix.

[0011] In a preferred embodiment, the in-memory computation matrix includes: N column storage operator modules, each column storage operator module including M rows of storage units for storing and performing computations on the stored data and input feature data.

[0012] In a preferred embodiment, each of the storage and computing units includes: an SRAM memory for storing data from the first storage module and the second storage module, and a logic unit for computing disposed adjacent to the SRAM memory.

[0013] In a preferred embodiment, the in-memory computation matrix includes: a weight parameter storage array for storing weights, a bit multiplier, a storage readout circuit, and a logic operation unit; the bit multiplier is used to receive a weight readout enable signal and input feature data; when the weight readout enable signal is enabled, it selects a weight in the weight storage array according to the weight readout enable signal, and multiplies the bit of the input feature data with the selected weight; the storage readout circuit is used to read the product of the multiplication to the logic operation unit when the bit of the input feature data is 1, and does not perform a readout operation when the bit of the input feature data is 0; the logic operation unit is used to accumulate the product read out by the storage readout circuit when the bit of the input feature data is 1, so as to realize the multiplication and accumulation of the input feature data and the weight.

[0014] In a preferred embodiment, the bit multiplier includes N AND gates that are respectively connected to the weight storage array; one input of each AND gate is used to receive the input feature data, another input is used to receive the weight read enable signal, and the output is used to output a control signal.

[0015] In a preferred embodiment, the second storage module is used to store intermediate data generated by the neural network computation and parameters of the large model. The ratio of the number of parameters stored in the first storage module to the bandwidth of the first storage module is greater than the ratio of the number of parameters stored in the second storage module to the bandwidth of the second storage module.

[0017] In a preferred embodiment, the ratio of the number of parameters stored in the first storage module to the bandwidth of the first storage module is greater than the ratio of the number of parameters stored in the second storage module to the bandwidth of the second storage module.

[0018] In a preferred embodiment, a computing module including an in-memory computing neural network accelerator (NPU) and a second storage module are stacked together using a hybrid bonding method; and / or a computing module including an in-memory computing neural network accelerator (NPU) and a first storage module are stacked together using a hybrid bonding method.

[0019] In a preferred embodiment, the first storage module, the second storage module, and the computing module are stacked sequentially from top to bottom.

[0020] In a preferred embodiment, the first storage module and the computing module are stacked together using a hybrid bonding method; the second storage module is built into the computing module or reuses the memory built into the computing module.

[0021] In a preferred embodiment, the first storage module is a Nand Flash memory, and / or the second storage module is a DRAM memory.

[0022] In a preferred embodiment, the first storage module is a Nand Flash memory, and the second storage module is a DRAM memory; the computing module predicts the parameters of the neuron to be activated based on the current input data and the sparsity prediction model read from the Nand Flash memory, and the computing module selectively reads the parameters of the neuron to be activated from the Nand Flash memory and performs neural network calculations based on the prediction results, and stores the intermediate data generated during the calculation process into the DRAM memory.

[0023] In a preferred embodiment, the first storage module is a Nand Flash memory, and the second storage module is a DRAM memory; the Nand Flash memory is stacked with the computing module by a hybrid bonding method; or the Nand Flash memory, the DRAM memory, and the computing module including the neural network accelerator based on in-memory computing are stacked sequentially from top to bottom by a hybrid bonding method.

[0024] In a preferred embodiment, the first storage module includes at least two layers of Nand Flash memory, which are stacked using a hybrid bonding method; and / or the second storage module includes at least two layers of DRAM memory, which are stacked using a TSV method or a hybrid bonding method.

[0025] In a preferred embodiment, the large model is a large model based on the MOE architecture. The first storage module stores at least the parameters of the sparse part of the large model, and the second storage module stores at least the parameters of the dense part of the large model and the intermediate data generated by the neural network calculation. The calculation module performs neural network calculation based on the input data and the data from the first storage module and the second storage module.

[0026] In a preferred embodiment, the AI ​​device further includes an interposer layer, a packaging substrate, and an I / O buffer for serial-to-parallel conversion. The interposer layer is disposed on the packaging substrate, and the I / O buffer is electrically connected between the first storage module and the interposer layer. The computing module and the first storage module are both disposed on the interposer layer, and the second storage module is stacked on the computing module. The first storage module, the second storage module, the computing module, the interposer layer, and the packaging substrate are packaged into a single unit using a chiplet method.

[0027] In a preferred embodiment, the AI ​​device is an edge-based inference device.

[0028] The present invention also provides an electronic device including any of the AI ​​devices described above.

[0029] The sparsity-based AI device provided by this invention combines a first storage module and a second storage module with an in-memory computing-based neural network accelerator to achieve neural network computation and AI functions. The two storage modules have different functions and load different data, allowing them to store a larger number of parameters. Furthermore, the computation module predicts and retrieves the parameters of neurons to be activated from the large model parameters in the first storage module based on the input data and the sparsity prediction model. This reduces the data access and transmission costs between the computation module and the storage module from the source, thereby reducing bandwidth requirements and simultaneously solving the problems of storage capacity and bandwidth. It also reduces the amount of data, computation time, power consumption, and heat generated by the computation module. The computation module includes an in-memory computing-based neural network accelerator (NPU), which further improves the energy efficiency of neural network processing and reduces corresponding power consumption and heat generation, solving the heat dissipation problem from the source and providing greater computing power. Furthermore, the CIM-based NPU enables the computing module to possess the advantages of high computing power and low power consumption, while reducing the problems of insufficient computing power, power consumption, and heat generation that AI devices encounter when running large models with a larger number of parameters. This allows AI devices to better support large models with a larger number of parameters. Therefore, the sparsity-based AI device provided by this invention can solve the problems of computing power, energy efficiency, storage capacity, bandwidth bottlenecks, and heat dissipation. Attached Figure Description

[0030] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0031] Figure 1This is a schematic diagram of the structure of an AI device provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the neural network processing in one embodiment of the present invention; Figures 3a to 3c These are schematic diagrams of the AI ​​devices provided in three embodiments of the present invention. Figure 4 This diagram illustrates the architecture of a CIM-based computing module according to an embodiment of the present invention. Figure 5 yes Figure 4 A magnified view of the NPU in the middle; Figure 6 This is a schematic diagram of a column of stored operator modules in an in-memory computation matrix according to an embodiment of the present invention; Figure 7 This is a schematic diagram of the vertical structure of a column of stored operator modules in an in-memory computation matrix according to an embodiment of the present invention; Figure 8 This is a schematic diagram illustrating the principle of in-memory computation matrix calculation in one embodiment of the present invention; Figure 9 This is a schematic diagram of the structure of an in-memory computing module according to an embodiment of this application; Figure 10 This is a schematic diagram of the in-memory computing module structure according to another embodiment of this application; Figure 11 This is a schematic diagram of the structure of a storage readout circuit according to an embodiment of this application. Detailed Implementation

[0032] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. In the absence of conflict, the following embodiments and their technical features can be combined with each other.

[0033] To make the objectives, technical solutions, and advantages of the present invention clearer, various exemplary embodiments described below will be referenced to the accompanying drawings, which form part of the exemplary embodiments, illustrating various exemplary embodiments that may be used to implement the present invention. Unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. It should be understood that they are merely examples of processes, methods, and apparatuses consistent with some aspects of the present invention disclosed as detailed in the appended claims, and other embodiments may be used, or structural and functional modifications may be made to the embodiments listed herein without departing from the scope and spirit of the present invention.

[0034] AI devices are typically used for training and inference. Currently, an increasing number of AI devices rely on large models for inference, making them increasingly powerful, especially with the growing prevalence of inference applications based on large language models.

[0035] See Figure 1 As shown, the AI ​​device provided by this invention includes: a computing module, a first storage module, and a second storage module. The computing module performs neural network calculations to handle artificial intelligence-related computations, enabling the AI ​​device to be used for training or inference. The computing module includes a neural network accelerator (NPU) based on in-memory computation. The NPU includes an in-memory computation matrix, which is used to perform neural network calculations based on input data and data from the first and second storage modules. The first storage module stores parameters of a large model and is electrically connected to the computing module for data interaction. The second storage module is electrically connected to the computing module for data interaction and stores intermediate data generated by the neural network calculations. The computing module also predicts the parameters of neurons to be activated in the large model parameters in the first storage module based on the input data and a sparse prediction model. The computing module reads the predicted neuron parameters from the first storage module based on the prediction results and performs neural network calculations based on the read neuron parameters and the input data.

[0036] The AI ​​device provided by this invention uses a first storage module and a second storage module to divide the storage tasks for parameters of a large model and intermediate data generated from calculations. This reduces the amount of data each storage module needs to store and decreases the amount of data transferred between the calculation module and each storage module, thus addressing the bandwidth issue at its source and simultaneously resolving both storage capacity and bandwidth problems. Furthermore, by dividing the processing tasks between the first and second storage modules, different types of data can be stored in the two modules. For example, frequently accessed data can be stored in one storage module, while less frequently accessed data can be stored in the other. For instance, parameters of a large model that is not frequently accessed can be stored in the first storage module, reducing the bandwidth requirements of the calculation module; while frequently accessed but small-volume intermediate data can be stored in the second storage module, thus not affecting the bandwidth requirements of the calculation module for frequent access to the second storage module. Furthermore, the computing module predicts and reads the neuron parameters to be activated from the large model parameters in the first storage module based on the input data and the sparse prediction model. This eliminates the need to read all parameters from the first storage module, fundamentally reducing the computing module's access to and read operations on the first storage module. This reduces the amount of data transferred between the first storage module and the computing module, the transmission time, and the amount of data required for computation, thereby reducing the bandwidth requirements of both the first and second storage modules. The computing module only needs to read the predicted neuron parameters to be activated from the first storage module and perform calculations accordingly. That is, the computing module does not need to calculate all parameters but performs calculations on demand based on the sparse prediction model. This further reduces bandwidth requirements, computational power required by the computing module, power consumption, computation time, heat generation, and improves the efficiency and energy efficiency of the AI ​​device. Moreover, the computing module includes an in-memory neural network accelerator (NPU), which further improves the energy efficiency of neural network processing and reduces corresponding power consumption and heat generation, addressing heat dissipation issues at their source and providing greater computing power. Therefore, the sparsity-based AI device provided by this invention can simultaneously solve the problems of storage capacity, bandwidth bottlenecks, computing power, energy efficiency, and heat dissipation. Even if the number of parameters in a large model is very large, and the number of parameters stored in the storage module is also very large, this invention can access and read only the predicted neuron parameters through sparse prediction, thus fundamentally reducing the bandwidth requirements of the first storage module. This also allows the first storage module to use lower-cost storage devices, which, given the current high cost of memory, further reduces the requirements of AI devices on storage devices and lowers the cost of AI devices.Moreover, combined with the advantages of high computing power and low power consumption of NPU based on in-memory computing, it can also better support large models with a huge number of parameters that require greater computing power.

[0037] In the AI ​​device provided by this invention, the second storage module can be used to store intermediate data generated by neural network computation and parameters of the large model. For example, expert weight parameters, which have a large number of parameters in the large model, are stored in the first storage module, while parameters with a small number of parameters that are tightly coupled to the computation are stored in the second storage module. By dividing the storage of parameters of the large model and intermediate data generated by computation through the first and second storage modules, the number of parameters of the large model that each storage module needs to store is reduced, and the amount of data transmission between the computation module and each storage module can be reduced, thereby solving the bandwidth problem from the source. Thus, the problems of storage capacity and storage bandwidth can be solved simultaneously.

[0038] See Figure 2 The working process of the AI ​​device provided by the present invention can be roughly described by referring to the embodiments therein: First, in step S10, the input data is transmitted to the calculation module.

[0039] In step S20, the computation module predicts the neuron parameters to be activated in the parameters of the large model in the first storage module based on the input data and the sparsity prediction model. In other words, the computation module predicts which neuron parameters actually need to participate in the computation, so as to avoid accessing and reading all the parameters of the large model in the first storage module.

[0040] In step S30, the computing module reads the predicted neuron parameters to be activated from the first storage module, and performs neural network calculations based on the neuron parameters read from the first storage module and the input data. This avoids reading, transferring, and calculating neuron parameters that are not needed at this time, thus effectively reducing the amount of data transmission at the source. This can greatly alleviate the bandwidth problem of the storage module and greatly reduce the amount of data transfer and calculation, thereby greatly reducing the demand for computing power, reducing energy consumption and heat, and improving the processing efficiency and energy efficiency of the AI ​​device. It can also enable the storage module to support large models with a larger number of parameters.

[0041] In step S40, the calculation module loads the intermediate data generated during the calculation process into the second storage module. Since the intermediate data generated during the calculation process needs to be read and moved more frequently, the second storage module, which is electrically connected to the calculation module, can selectively choose a storage device that better supports frequent reading.

[0042] In step S50, the calculation module reads the intermediate data from the second storage module and combines the read data from the second storage module with the input data to perform neural network calculations until the calculation module calculates the final calculation result.

[0043] To illustrate the technical solution described in this invention, the following is a detailed explanation. Figure 3a The preferred embodiment shown is described in detail below, illustrating only the parts relevant to the embodiments of the present invention. In this preferred embodiment, the AI ​​device is an edge-side inference device based on a large model. The first storage module is a FLASH storage module (i.e.,...). Figure 3a The second storage module is a DRAM storage module (i.e., a FLASHLayer with N layers stacked). Figure 3a The system comprises several DRAMs stacked in 3D. A silicon interposer is disposed on the packaging substrate. A portion of the silicon interposer is on which the aforementioned FLASH memory module is disposed, and another portion of the silicon interposer is on which a computing layer is disposed. The DRAM memory module is stacked with the computing layer using hybrid bonding. Parameters of the sparse portion of the large model are stored in the FLASH memory module, while parameters of the dense portion of the large model and intermediate data generated by neural network computation are stored in the DRAM memory module. The computing module (i.e., the computing layer in the figure) selectively reads the predicted parameters of the neurons to be activated from the FLASH memory module according to the sparsity prediction model and performs calculations based on the read predicted parameters of the neurons to be activated, such as inference calculations. This allows the first and second storage modules to be two different types and functions of storage modules, thereby selectively transferring different data between the storage module and the computing module. This reduces the amount of data transferred and the amount of data involved in the computation from the source, thereby reducing related computing power, energy consumption, and heat, and also reducing the cost of memory and AI devices.

[0044] A computing module including an in-memory neural network accelerator (NPU) is stacked with a second storage module via hybrid bonding, and / or a computing module including an in-memory neural network accelerator (NPU) is stacked with a first storage module via hybrid bonding. This can significantly improve the storage capacity and bandwidth of AI devices. For example, a computing module including an NPU is stacked with a DRAM storage module via hybrid bonding. DRAM stacked with the computing module can provide a large storage capacity (e.g., 12 / 24 / 36GB). Therefore, this invention can achieve a 3D stacked structure of storage modules and computing modules with specific structures in an AI device through hybrid bonding, thereby enabling high-density, high-performance interconnection between them.

[0045] Hybrid bonding enables direct contact bonding of the metal layers (typically copper) and dielectric layers of two chips / dies. For example, hybrid bonding bonds copper to copper and dielectric layers to dielectric layers, replacing traditional bump or solder ball interconnects with direct copper-to-copper connections. This allows for ultra-fine pitch stacking and packaging within a very small space, achieving 3D integration. Hybrid bonding allows the metal layers of two or more chips to be precisely aligned and directly pressed together to form direct electrical contacts. Compared to traditional bonding technologies, hybrid bonding can achieve submicron or even nanometer-level interconnect pitches, allowing more connection points to be placed in a smaller area, significantly increasing the data communication bandwidth between chips. Its contact density can reach 10K-1MM / mm². Moreover, because hybrid bonding eliminates intermediate media such as solder, direct copper-to-copper connections have lower resistance, reducing energy loss during signal transmission and also reducing signal propagation time delay. The compact structure and direct conductive path of hybrid bonding help improve thermal management and reduce heat generation, which is especially important for high-performance computing such as artificial intelligence.

[0046] Therefore, the bandwidth of the AI ​​device provided in this application can be exponentially improved. The bandwidth of the 3D stacked DRAM with integrated Hybrid Bonding can reach tens of TB / s or even higher, which is a significant improvement over other state-of-the-art GDDR6 (bandwidth of 150 to 300 GB) and HBM3E technologies. Moreover, its energy requirement for transmitting a single bit of data can be significantly reduced, greatly reducing the power consumption of data transfer. It can significantly improve the memory bandwidth in AI devices, and greatly enhance the performance and response speed of their artificial intelligence processing.

[0047] As shown in Table 1, the integrated hybrid-bonded 3D stacked DRAM can achieve a bandwidth of 1-60TB, which is a significant improvement over other advanced GDDR6 and HBM3E technologies. At the same time, its energy requirement for transmitting a single bit of data is also greatly reduced to 0.5pJ / bit, which greatly reduces the power consumption of data migration. It can solve the bandwidth bottleneck problem encountered in the deployment of edge AI large model (LLM) chips, and at the same time significantly reduce the power consumption of data migration.

[0048] Table 1 The two storage modules and the computing module provided by this invention can be connected in various ways. See also Figures 3a to 3cIn the illustrated embodiment, since the intermediate data stored in the second computing module typically requires frequent reading and transmission, it is preferable to directly contact and electrically connect the second storage module and the computing module using a hybrid bonding method. For example, Figure 3b The first storage module, the second storage module, and the computing module are stacked sequentially from top to bottom. This further reduces the transmission distance for the computing module to frequently read data from the second storage module, reduces related energy consumption, and increases related bandwidth. Furthermore, this significantly reduces the size and heat generation of the AI ​​device. Therefore, this invention uses the first and second storage modules together in conjunction with the computing module, and combines sparse prediction to selectively read the predicted parameters of neurons to be activated in a large model. This greatly reduces the amount of data that the first and second storage modules need to transmit, thus fundamentally solving the bandwidth bottleneck problem. In another embodiment, the first storage module and the computing module are stacked using a hybrid bonding method; the second storage module is built into (e.g., Figure 3c The computing module or the second storage module can reuse the memory built into the computing module. For example, when the second storage module stores a smaller number of parameters in a large model, since the amount of intermediate data is usually small, and the first storage module, which is co-bonded with the computing module, stores complex and large-scale parameters in the large model (such as a large number of expert weight parameters), the memory built into the computing module with a slightly smaller storage capacity can be reused as the second storage module, which can further reduce the size, energy consumption, and cost of the AI ​​device.

[0049] In one embodiment of the present invention, the first storage module is a NAND Flash memory, and / or the second storage module is a DRAM memory. This approach combines the advantages of both NAND Flash and DRAM memories while avoiding their disadvantages, enabling AI devices to simultaneously meet the requirements for storage capacity and bandwidth at a lower cost. For example, the first storage module includes at least two layers of NAND Flash memory, stacked using a hybrid bonding method; and / or the second storage module includes at least two layers of DRAM memory, stacked using a TSV (Through Silicon Via) method or a hybrid bonding method. 3D-NAND-FLASH memory features high storage density (e.g., 20-36 Gb / mm²) and low storage cost, making it well-suited for storing model parameters in large-scale AI model deployments on the edge. The illustrated method allows for vertical stacking of multiple DRAM chips, significantly improving memory bandwidth and reducing power consumption. Furthermore, it enables large-capacity, high-bandwidth storage, meeting the stringent memory requirements of high-performance computing, artificial intelligence, and other fields, allowing AI devices to simultaneously possess the advantages of high storage capacity, high bandwidth, and low cost. The first storage module can also be any other memory capable of storing parameters of a large model, and the second storage module can also be any other memory capable of storing intermediate data and parameters of a large model. For example, the first storage module is preferably a non-volatile memory, and the second storage module is a volatile memory.

[0050] In one embodiment, a sparse prediction model is stored in the NAND Flash memory. The computing module predicts the parameters of the neurons to be activated based on the current input data and the sparse prediction model read from the NAND Flash memory. The computing module selectively reads the parameters of the neurons to be activated from the NAND Flash memory based on the prediction results and performs neural network calculations. Intermediate data generated during the calculation process is stored in the DRAM memory. This embodiment enables the AI ​​device to accurately and quickly read the required neuron parameters based on the current input data, enabling accurate and efficient AI processing at a low cost. The NAND Flash memory can be stacked with the computing module using a hybrid bonding method to improve the bandwidth of the first storage module, increase the access efficiency of the computing module to data in the first storage module, and reduce the size and cost of the AI ​​device. Alternatively, the NAND Flash memory, the DRAM memory, and the computing module including the in-memory computing NPU can be stacked sequentially from top to bottom using a hybrid bonding method, which can further improve the bandwidth and efficiency of the AI ​​device and reduce its size and cost.

[0051] In one embodiment, the FLASH storage module stores a large model based on a Hybrid Expert Architecture (MOE). A first storage module stores parameters of the sparse portion of the large model, and a second storage module stores parameters of the dense portion of the large model and intermediate data generated by the neural network computation. The computation module performs neural network computation based on the input data and data from the first and second storage modules. The difference between the MOE architecture (Sparse model) and the traditional Transformer architecture (Dense model) is that the traditional Transformer architecture shares all parameters between input and output, is fully connected, and all parameters are activated and participate in computation, such as the fully connected layers in CNNs and Transformers. The core idea of ​​the MOE model is to design multiple "experts" (sub-networks) within the model, and through a gating mechanism, only a portion of these experts are activated for computation on each input or token. In the AI ​​device provided by this invention, the model capacity can be significantly increased in terms of parameter scale, while only a small portion of the parameters actually participate in computation. Only a portion of the experts are activated during inference, thus achieving higher parameter utilization per unit inference time and achieving a certain "computing power economy" effect. Moreover, since only a very small number of parameters (e.g., 5-6%) are activated during actual inference execution, this embodiment also significantly reduces the amount of data and power consumption of moving large model parameters, improves the energy efficiency of running large models, greatly reduces the bandwidth requirements for storage migration, and reduces heat generation.

[0052] It's worth noting that memory typically accounts for the majority of the cost of an AI device. However, if current AI devices are to support large models (e.g., hundreds of billions of parameters) and run large models with a massive number of parameters, existing technologies require a significant increase in the amount of memory. This leads to a sharp rise in the overall cost of the AI ​​device and also makes it quite large. Edge AI devices, especially AI smartphones which have very limited space, have very high requirements for both low cost and small size, placing even greater constraints on their size. Therefore, the AI ​​device provided by this invention selectively loads the parameters of a large model into two different types of storage modules, so that the parameters can be categorized and moved between different memory and computing modules. Combined with the above-mentioned sparsity prediction method, it selectively reads only some parameters and moves only some parameters to the computing module for calculation, thereby reducing the number and cost of memory. Moreover, it only requires a small amount of FLASH memory and a small amount of DRAM memory to solve the problems of storage capacity and bandwidth. Furthermore, by setting up the computing module, FLASH storage module, and DRAM storage module according to the above-mentioned positional structure, it can greatly reduce the cost and size of storage devices in the AI ​​device, thereby greatly reducing the cost and size of the AI ​​device. It can also provide large storage capacity, high bandwidth, high computing power, and high energy efficiency.

[0053] In one embodiment of the present invention, the FLASH storage module stores parameters of a large language model and a low-rank prediction model based on a hybrid expert architecture. The DRAM storage module stores MLA (Multi-head Latent Attention) parameters, which include Attention parameters and / or Embedding parameters. Because the number of Attention parameters is relatively small, this application directly loads the Attention parameters and Embedding parameters into the DRAM. The MLP (Multi-Layer Perceptron) parameters, which have a larger number of parameters, are loaded into a large-capacity FLASH storage module. During inference, the computation module predicts the parameters of the neurons to be activated based on the current input data and the low-rank prediction model read from the FLASH storage module. Based on the prediction results, the computation module selectively reads some experts from the FLASH storage module, performs inference calculations based on these experts, and stores the intermediate data generated during the inference calculation process into the DRAM storage module. In other words, which neurons in the MLP module will be activated is related to the current input. Leveraging the high sparsity of this MLP, a simple low-rank prediction model (e.g., a two-module MLP) is used to predict which neurons will be activated in real time during large model inference. Each time a large model inference occurs, only the parameters of the activated neurons need to be dynamically loaded from the FLASH storage module. This approach can significantly reduce the amount of model parameters that need to be read from the FLASH storage module (by up to 90%), while also significantly reducing the power consumption and bandwidth requirements for data movement within the FLASH storage module. For example, Flash Memory is used to store complete large model parameters. Flash Memory has a large capacity (256GB-512GB) but low bandwidth (≈1GB / s). 3D-Stacked DRAM is used for fast access during large model inference. 3D-DRAM has a small capacity (12-36GB) but high bandwidth (1-10TB / s). This can further reduce the back-and-forth data movement and power consumption between the storage module and the computation module. Meanwhile, a CIM-based NPU reduces the overhead of data movement to the processor by performing calculations directly within the storage unit (e.g., matrix multiplication and addition), making it particularly suitable for matrix operations in MLPs and MOEs. This solution systematically addresses the challenges of deploying large models on AI devices through tiered storage, dynamic and selective loading parameters, efficient use of sparsity, and innovative model structures. In particular, it effectively solves the problems of storage capacity, bandwidth bottlenecks, computing power, energy efficiency, and cost associated with deploying large models on the edge.

[0054] In embodiments of the present invention, the second storage module can store intermediate data generated by neural network computation and parameters of the large model. The computation module performs neural network computation based on the parameters of the large model from the first and second storage modules, and the intermediate data from the second storage module, to achieve AI functionality. The ratio of the number of parameters stored in the first storage module to its bandwidth is greater than the ratio of the number of parameters stored in the second storage module to its bandwidth. For example, the ratio of the number of parameters stored in the Flash storage module to its bandwidth is A, and the ratio of the number of parameters of the large model stored in the DRAM storage module to its bandwidth is B, and A is greater than B. In this way, the first and second storage modules can be two different types and functions of storage modules, allowing for the selective storage of different data in different storage modules and the transfer of data between the storage modules and the computation module. Therefore, the first storage module can store data that does not require frequent access and has low bandwidth requirements, and can store large model parameters with a large data volume; while the second storage module can store data that requires a certain bandwidth and needs to be accessed more frequently, and is relatively small in volume. Correspondingly, the first storage module can use lower-cost memory, such as NAND flash memory, while the second storage module, using fewer memory modules, can better accommodate requirements for data volume, bandwidth, cost, and even space size. For example, the second storage module can be DRAM memory. For instance, sparse parameters in a large model can be stored in the Flash memory module, while dense parameters and intermediate data can be stored in the DRAM memory module (dense parameters typically refer to model parameters that are frequently accessed during the large model's operation, while sparse parameters typically refer to model parameters that are less frequently accessed). This reduces data transfer volume, power consumption, and heat generated during data transfer, while also reducing memory costs and effectively utilizing and alleviating bandwidth bottlenecks. Furthermore, the first storage module can also be used for computation, such as performing small-scale calculations.

[0055] In AI devices, such as AI large-model inference chips, a large number of matrix multiplication calculations (multiply-accumulate MAC) are involved, accounting for 30-70% of the overall chip power consumption. As Moore's Law gradually approaches its limit, the "memory wall" and "power wall" problems of AI chips based on the von Neumann architecture are becoming increasingly prominent, and the growth rate of chip computing power is getting slower and slower. Figure 3aIn the illustrated embodiment, the computing module is based on in-memory computing (CIM), meaning the computing module includes a CIM-based neural network accelerator (NPU), which includes an in-memory computing matrix. The in-memory computing matrix is ​​used to perform neural network calculations based on input data and data from the first and second storage modules. Therefore, the neural network accelerator included in the computing module achieves in-memory computing integration through the in-memory computing matrix. In-memory computing aims to overcome the "memory wall" problem by integrating storage and computing functions onto the same chip. This technology embeds computing power into memory and uses a new operational architecture to implement multiply-accumulate operations, thus reducing the frequent data transfer between memory and computing chips in existing technologies. This significantly reduces transfer time, power consumption, and heat generation, and improves the parallel processing efficiency of data; it enables in-situ computing, eliminates bandwidth limitations, and reduces data movement costs. This fundamentally eliminates unnecessary data movement latency and power consumption, improving AI computing efficiency by hundreds or thousands of times, reducing costs, and breaking down the "memory wall" and "power wall." This allows the computing module and the AI ​​device to have greater computing power, consume less power, and generate less heat compared to existing devices that perform calculations under the same conditions, thus solving the existing heat dissipation problem at its source.

[0056] See Figures 4 to 8 The preferred embodiment shown includes a computing module with an NPU based on a CIM architecture responsible for matrix operations and parallel computation during the execution of large language models. This computing module can also handle quantization and optimization tasks; it supports model quantization (such as INT8 quantization), reducing computational complexity and memory usage while maintaining model accuracy and improving inference speed. Figure 4 The computing module includes: an I / O system for inputting and outputting data; and a processor (i.e.,...). Figure 4 The system includes a RISC-V CPU for preprocessing the input data to conform to the format of the large model. This includes operations such as normalization, resizing, and encoding of the input data (e.g., text, images, or sensor data) to ensure it meets the input format of the large model. It also includes several neural network accelerators (NPUs) for performing neural network processing on the preprocessed input data. This implementation can improve the computing power and energy efficiency of the AI ​​device, while reducing its power consumption and heat generation.

[0057] See Figure 5In the illustrated embodiment, each neural network accelerator (NPU) includes: a pre-processing module (i.e., Pre-Processing in the figure), used to convert the pre-processed input data into matrix data conforming to the processing format of the in-memory computing module; several neural network sub-accelerators (SNPUs), used to perform neural network processing on the data processed by the pre-processing module; and a vector processing module (VPU), used to perform vector processing on the data processed by the SNPUs. Each SNPU includes several in-memory computing modules (CIMs) (CIMD in the figure is illustrated as digital CIMs), arranged in a matrix to form the in-memory computing matrix. The CIM-based NPU provided in this embodiment, combining a pre-processing module, a matrix-like in-memory computing module, and a vector processing module, can significantly improve the computing power and energy efficiency of the NPU and AI device, while greatly reducing power consumption and heat generation. Furthermore, it has a simple structure and low cost. The figure illustrates that the four neural network sub-accelerators, SNPU0 to SNPU3, are connected to the bus in parallel. The NPU also includes on-chip memory (such as...). Figure 5 The on-chip SRAM can be used to store intermediate data generated by in-memory computation matrix calculations. The on-chip SRAM cache and DMA (Direct Memory Access) system in the computing module handle memory management and data transfer optimization during large language model inference. On-chip memory can store large model parameters and intermediate data, reducing access to off-chip memory (such as DRAM) and lowering memory bandwidth bottlenecks. By optimizing memory access modes (such as sequential access and batch loading) and using DMA technology, the memory management burden is reduced. This approach can improve the computing power, accuracy, and efficiency of the computing module while reducing memory load. The computing module provided in this embodiment has advantages such as high computing power, high energy efficiency, low cost, and low power consumption.

[0058] See Figure 6In a preferred embodiment of the present invention, the in-memory computation matrix includes: N columns of in-memory operator modules. Each column of in-memory operator module includes: M rows of in-memory units for storing and calculating the stored data with input feature data, and capacitors electrically connected to the in-memory units. The charges output by the capacitors in the M rows are pooled together to achieve charge accumulation. Therefore, the in-memory units can store and perform multiplication calculations on the weights and input feature data. The corresponding products are accumulated through the charges of the capacitors connected to the in-memory units to realize the multiply-accumulate (MAC) calculation in neural network computation. That is, matrix calculation can be realized within the in-memory computation matrix, which can avoid the frequent data transfer between two different and far-away memories and computing chips in the prior art. Multiply-accumulate calculations typically account for a large portion of the total power consumption of a computation, often reaching 50-70%. This invention, based on a memory-based computing architecture, incorporates a neural network accelerator with an in-memory computing matrix within the computing module. This enables two-dimensional and three-dimensional matrix multiplication / addition operations through a novel computational architecture, significantly reducing the power consumption of neural network computations. Furthermore, it allows for a smaller neural network accelerator area, enabling greater computing power with less heat generation within the same area and power consumption, thus addressing heat dissipation issues. The computing module, including the in-memory computing matrix, is stacked with the second storage module using a hybrid bonding method, further alleviating the bandwidth limitations of this AI device. It also boasts high computing power, small size, and low cost.

[0059] See Figure 7 The diagram shown is a vertical structural diagram of a column of stored-operator submodules in an in-memory computation matrix, according to an embodiment of the present invention. Each stored-operator unit includes: an SRAM memory and a logic unit (i.e., digital logic in the figure) arranged adjacent to the SRAM memory for computation. Multiple stored-operator units are stacked to form a column, and several columns of stored-operator units form a stored-operator matrix. The aforementioned logic unit can be used to perform multiplication and addition calculations on the weights and input feature data in the SRAM memory. Figure 7The diagram also showcases a block diagram of Integrated In-Memory Computing (CIMD). CIMD significantly reduces the power consumption of data transfer while greatly increasing the speed of parallel multiply-accumulate (MAC) calculations, effectively alleviating the two major challenges of the 'memory wall' and the 'power wall'. SRAM memory and the logic units used for computation are stacked adjacently, allowing for a more compact in-memory computing matrix structure, greater computing power, lower power consumption, and faster and more accurate computation speed when combined with digital logic units. As shown in Table 2, at the 12nm process node, the matrix multiplication energy efficiency ratio of the CIM in-memory computing unit can reach 24.7 TOPS / W, which is an order of magnitude improvement compared to traditional digital architectures (i.e., general-purpose computing units, such as GPUs and TPUs). Integrated In-Memory Computing (CIMD) significantly reduces the power consumption of data transfer while greatly increasing the speed of parallel multiply-accumulate (MAC) calculations. Thanks to the innovative architecture of this bottom-module computing unit, in-memory computing (CIMD) technology can significantly improve the energy efficiency of AI large model (LLM) inference chips, that is, while keeping the computing power the same, it can greatly reduce its power consumption requirements.

[0060] Table 2 Figure 8 This is a schematic diagram illustrating the principle of in-memory computation matrix calculation in one embodiment of the present invention. The in-memory computation unit performs multiplication calculations on the input feature data (i.e., the aforementioned input data) and weights (e.g., parameters / data from the first and second storage modules), and then accumulates the product through a multi-level addition tree to obtain the result of multiplication and accumulation.

[0061] The input data for AI devices also suffers from sparsity. This sparsity impacts the computational efficiency of machine learning because sparse matrices contain a large number of zero values, leading to resource waste, unnecessary computation, and low storage utilization. The most critical issue is resource waste, as machine learning algorithms often need to traverse the entire dataset during training. If the dataset is sparse, containing a large amount of zeros or meaningless information, the algorithm must process a large amount of data that does not affect the result. This not only increases processing time but also consumes more computational resources, resulting in inefficiency and high power consumption. Further optimization, see [link to further optimization]. Figures 9 to 11As shown, the in-memory computation matrix provided by this invention includes a weight storage array 10 (also called a weight parameter storage array), a bit multiplier 12, a storage readout circuit 21, and a logic operation unit 22. One input of the bit multiplier 12 is connected to a weight readout enable signal, and the other input is connected to input feature data (i.e., the input data of the aforementioned AI device). The weight acquisition end is connected to the weight storage array 10. The input end of the storage readout circuit 21 is connected to the weight storage array 10, and the output end is connected to the logic operation unit 22. The output end of the logic operation unit 22 is used to output the convolution operation result. The weight storage array 10 is used to store weights.

[0062] The input feature data may include at least one bit, where the bit value is either 1 or 0. The bit multiplier 12 receives a weight read enable signal and the input feature data (in convolutional neural networks, the input feature map typically needs to be processed). When the weight read enable signal is enabled, the multiplier selects a weight from the weight storage array based on the signal and multiplies the bit of the input feature data with the selected weight. For example, when a bit of the input feature data is 1, the product of the bit and the corresponding weight is the weight. In this case, the storage readout circuit 21 reads the corresponding weight from the weight storage array 10 to achieve the corresponding product readout. The storage readout circuit 21 reads the product to the logic operation unit 22 when the bit of the input feature data is 1, and does not perform a readout operation when the bit of the input feature data is 0, thus reducing the power consumption of the storage readout circuit 21. When the bit of the input feature data is 0, the storage readout circuit 21 does not perform a readout operation. Therefore, the storage readout circuit 21 can directly read the corresponding product from the weight storage array 10. Specifically, the storage readout circuit 21 can read the corresponding weight to the logic operation unit 22 when the bit of the input feature data is 1. The storage medium of the storage readout circuit 21 can be any of SRAM, DRAM, and RRAM.

[0063] The logic operation unit 22 is used to accumulate the product read by the storage readout circuit 21 when the bit of the input feature data is 1, thereby realizing the multiplication and accumulation of the input feature data and weights. The logic operation unit 22 only needs to output the final convolution result, without transmitting intermediate data during the convolution operation, which reduces the requirements for data output and data transmission bandwidth. That is, it ensures that the storage readout circuit does not generate power when the bit of the input feature data is 0, thus reducing the power consumption of the storage readout circuit. The output data of the logic operation unit is the accumulation result, without transmitting intermediate data during the convolution operation, reducing the data bit width of its output data. If higher accumulation operations are implemented in the logic operation unit, the output data bit width can be further compressed. Therefore, this embodiment can reduce the power consumption of reading large amounts of concurrent data, reduce the requirements for data output and data transmission bandwidth, and improve the efficiency of AI processing.

[0064] See Figure 10 In the preferred embodiment shown, the bit multiplier 12 includes N AND gates respectively connected to the weight storage array; one input of each AND gate is used to receive the input feature data fmi, the other input is used to receive the weight read enable signal, and the output is used to output a control signal. The weight storage array 10 includes K rows and N columns of parameter storage units. Each parameter storage unit is used to store a weight or an electrical signal representing the corresponding weight. The h-th row and j-th column parameter storage unit stores the weight W(h,j), 1≤h≤K, 1≤j≤N, where K is the number of rows of the weight and N is the number of columns of the weight. The K rows of parameter storage units are correspondingly set on the K row bit lines; the h-th row bit line is denoted as bitline,h. The aforementioned weight read enable signal is used to enable one column of parameter storage units in the N columns of parameter storage units, that is, one weight read enable signal corresponds to one column of parameter storage units. For example, the weight read enable signal corresponding to the j-th column of parameter storage units can be denoted as wordline,j. When the corresponding weight read enable signal wordline,j is 1, the j-th column parameter storage unit outputs the electrical signal representing the weight to the corresponding bit line, so that the storage readout circuit 21 can read the corresponding weight signal through each row of bit lines. For example, the probability of an N-bit input vector being all zeros is much lower than the probability of some bits in an N-dimensional input vector being zero. The bit-level sparsity of data is very high. However, the smaller the granularity of data sparsity, the more difficult it is to use, mainly because it requires real-time calculation and judgment on each bit of the input data. Through this preferred method, the present invention can solve the problem of "input data" sparsity. For the case where the sparse matrix of input data contains a large number of zero values, it can greatly reduce resource waste, reduce unnecessary computation, reduce costs, improve storage space utilization, improve computational efficiency, and fundamentally solve the heat dissipation problem of AI devices.

[0065] refer to Figure 11 As shown, the storage readout circuit 21 includes readout branches set on each of the K row bit lines; that is, a readout branch h is set on the h-th row bit line. The readout branch h includes a control terminal, which is used to connect to the corresponding control signal. The control signal is determined according to the weight readout enable signal wordline,j and the input feature data fmi,j. Each readout branch is used to connect to the weight readout enable signal and the input feature data. When the weight readout enable signal is enabled and the bit of the input feature data is 1, the product of the multiplication is read out to the logic operation unit 22. Specifically, each readout branch can read out a column of weight parameters enabled by the weight readout enable signal from the weight parameter storage array 10 when the weight readout enable signal is enabled and the bit of the input feature data is 1. Taking the example of each readout branch being connected to the weight readout enable signal wordline,j and the input feature data fmi,j corresponding to the j-th column, the working process of the readout branch is explained. If both the weight readout enable signal wordline,j and the input feature data fmi,j are 1, then each readout branch reads the corresponding weight parameter from the parameter storage unit in the j-th column of the weight parameter storage array; if at least one of the weight readout enable signal wordline,j and the input feature data fmi,j is 0, then each readout branch does not currently perform a readout operation. This implementation method has a simple circuit structure, good stability, low cost, and can effectively solve the problem of input data sparsity.

[0066] In a preferred embodiment of the present invention, the storage readout circuit further includes an AND gate; one input terminal of the AND gate is used to receive input feature data, the other input terminal is used to receive a weight readout enable signal, and the output terminal is used to output a first control signal; for the weight readout enable signal wordline,j and the input feature data fmi,j corresponding to the j-th column, when both the weight readout enable signal wordline,j and the input feature data fmi,j are 1, the first control signal is 1; when at least one of the weight readout enable signal wordline,j and the input feature data fmi,j is 0, the first control signal is 0. Its circuit is simple and can accurately control the read operation, reducing the power consumption of the storage readout circuit and the power consumption of the large language model inference device.

[0067] In a preferred embodiment of the present invention, the storage module is also used for computation, for example, to implement the first-level computation of a neural network. The computation module completes the neural network computation based on the result of the first-level computation. That is, the storage module can also perform simple computations, such as the first multiplication based on the multiplication parameters, i.e., the first-level computation. The result of the first-level computation is transmitted to the NPU for accumulation computation to complete the neural network computation. This fully utilizes the storage module, reduces the computational load of the computation module, and thus reduces the power consumption of the computation module. Moreover, after the first-level computation, the data transfer from the storage module to the computation module can be reduced, correspondingly reducing the demand for storage and high bandwidth.

[0068] It is worth noting that traditional storage devices typically cannot withstand ambient temperatures exceeding 85 degrees Celsius. Existing computing chips generally have high power consumption, resulting in significant heat generation from these chips and AI devices containing them. This makes it easy for ambient temperatures to exceed the 85-degree limit, leading to increased leakage current in storage devices under these high-temperature conditions. Furthermore, this necessitates higher refresh rates for the storage devices, further increasing their power consumption, creating a vicious cycle that has a significant impact on AI devices. The AI ​​device provided in this application combines a computing module including a CIM-based NPU with two storage modules. This improves computing power and efficiency while reducing both power consumption and heat generation, thereby fundamentally solving the power consumption and heat dissipation problems, and ultimately addressing the aforementioned issues at their source.

[0069] The AI ​​device provided by this invention can take various forms, such as an AI chip. Furthermore, the AI ​​device can also be in the form of a module.

[0070] In a preferred embodiment, the AI ​​device further includes an interposer, a packaging substrate, and an I / O buffer for serial-to-parallel conversion. The interposer is disposed on the packaging substrate, and the I / O buffer is electrically connected between the first storage module and the interposer. The computing module and the first storage module are both disposed on the interposer, and the second storage module is stacked on the computing module. The first storage module, the second storage module, the computing module, the interposer, and the packaging substrate are packaged into a single unit using a chiplet method. This enables efficient interconnection between chiplets. By connecting multiple chiplets together through an interposer, such as an organic interposer, advantages such as compatibility with various process technologies, reduced costs, and increased flexibility of the device are achieved. Preferably, the interposer is a silicon interposer, which allows the devices such as the first storage module and the computing module connected by the interposer to have ultra-high interconnection density, better electrical connection performance, reduced parasitic effects, and also helps with chip heat dissipation.

[0071] In summary, this invention provides a novel AI device that addresses the current challenges faced by large-model-based AI devices, particularly edge AI devices, including insufficient computing power, energy efficiency, storage capacity, bandwidth bottlenecks, poor heat dissipation, and high cost. It significantly improves storage capacity and bandwidth, reduces power consumption, achieves large-capacity, high-bandwidth storage, meets the stringent memory requirements of artificial intelligence, supports large models with a larger number of parameters, and operates at a lower cost. Preferably, the AI ​​device provided by this invention is an edge-based inference device. Edge-based inference devices typically have more stringent requirements for computing power, storage capacity, bandwidth, energy efficiency, heat dissipation, and size than cloud-based devices, especially given the limited hardware space and resources available on the edge. The sparsity-based AI device provided by this invention particularly addresses the issue of limited edge hardware resources, better resolving the requirements for computing power, storage capacity, bandwidth, energy efficiency, heat dissipation, and size in edge-based inference devices. It can be widely applied to various AI devices, such as mobile phones, computers, and robots.

[0072] This invention also provides an electronic device including the aforementioned AI device. It can significantly reduce the power consumption and cost of electronic devices, and improve the efficiency and experience of AI training and inference. It can be widely used in smartphones, tablets, wearable electronic devices, smart home electronic products, and so on.

[0073] The above description is merely a preferred embodiment of the present invention. Those skilled in the art will understand that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of the present invention. Furthermore, under the teachings of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the protection scope of the present invention.

Claims

1. A sparsity-based AI device, characterized in that, include: The first storage module used to store parameters of large models; A second storage module for storing intermediate data generated by neural network calculations; and A computing module for neural network calculations, wherein the computing module is electrically connected to the first storage module for data interaction, and the computing module is electrically connected to the second storage module for data interaction; The computing module includes a neural network accelerator (NPU) based on in-memory computing. The NPU includes an in-memory computing matrix, which is used to perform neural network calculations based on input data and data from the first and second storage modules. The calculation module is also used to predict the neuron parameters to be activated in the parameters of the large model in the first storage module based on the input data and the sparse prediction model. The calculation module reads the predicted neuron parameters to be activated from the first storage module based on the prediction results, and performs neural network calculation based on the read neuron parameters and the input data.

2. The AI ​​device according to claim 1, characterized in that, The computing module includes: An I / O system is used for inputting and outputting data. A processor for preprocessing the input data to make the preprocessed data conform to the format of the large model; and Several in-memory computing-based neural network accelerators (NPUs) are used to perform neural network processing on preprocessed input data.

3. The AI ​​device according to claim 2, characterized in that, The neural network accelerator NPU includes: The preprocessing module is used to convert the preprocessed input data into matrix data that conforms to the processing format of the in-memory computing module; Several neural network sub-accelerators (SNPUs) are used to perform neural network processing on the data processed by the preprocessing module. Each SNPU includes several in-memory computing modules (CIMs) arranged in a matrix to form the in-memory computing matrix. The vector processing module is used to perform vector processing on the data processed by the SNPU; and On-chip memory is used to store intermediate data generated by the on-chip computation matrix.

4. The AI ​​device according to claim 3, characterized in that, The in-memory computation matrix includes: N columns of storage operator modules, each column of storage operator module including M rows of storage units for storing and performing computations on the stored data and input feature data.

5. The AI ​​device according to claim 4, characterized in that, Each of the storage units includes: an SRAM memory for storing data from the first storage module and the second storage module, and a logic unit for calculation disposed adjacent to the SRAM memory.

6. The AI ​​device according to claim 1, characterized in that, The in-memory computation matrix includes: a weight parameter storage array for storing weights, a bit multiplier, a storage readout circuit, and a logic operation unit; The bit multiplier is used to receive a weight read enable signal and input feature data. When the weight read enable signal is enabled, the multiplier selects a weight in the weight storage array according to the weight read enable signal and multiplies the bits of the input feature data with the selected weight. The storage readout circuit is used to read the product of the multiplication to the logic operation unit when the bit of the input feature data is 1, and not to perform the readout operation when the bit of the input feature data is 0. The logic operation unit is used to accumulate the product read by the storage readout circuit when the bit of the input feature data is 1, so as to realize the multiplication and accumulation of the input feature data and the weight.

7. The AI ​​device according to claim 6, characterized in that, The bit multiplier includes N AND gates that are respectively connected to the weighted storage array; One input terminal of the AND gate is used to receive the input feature data, the other input terminal is used to receive the weight read enable signal, and the output terminal is used to output the control signal.

8. The AI ​​device according to claim 1 or 3, characterized in that, The second storage module is used to store intermediate data generated by neural network calculations and parameters of the large model. The ratio of the number of parameters stored in the first storage module to the bandwidth of the first storage module is greater than the ratio of the number of parameters stored in the second storage module to the bandwidth of the second storage module.

9. The AI ​​device according to any one of claims 1 to 7, characterized in that, This includes a computing module based on in-memory computing (NPU) stacked with the second storage module using a hybrid bonding method; and / or The computing module, including the NPU based on in-memory computing, is stacked with the first storage module using a hybrid bonding method.

10. The AI ​​device according to any one of claims 1 to 7, characterized in that, The first storage module, the second storage module, and the computing module are stacked sequentially from top to bottom.

11. The AI ​​device according to any one of claims 1 to 7, characterized in that, The first storage module and the computing module are stacked together using a hybrid bonding method; The second storage module is built into the computing module or reuses the memory built into the computing module.

12. The AI ​​device according to any one of claims 1 to 7, characterized in that, The first storage module is a NandFlash memory, and / or the second storage module is a DRAM memory.

13. The AI ​​device according to claim 12, characterized in that, The first storage module is a Nand Flash memory, and the second storage module is a DRAM memory; The calculation module predicts the parameters of the neurons to be activated based on the current input data and the sparsity prediction model read from the Nand Flash memory. The calculation module selectively reads the parameters of the neurons to be activated from the Nand Flash memory and performs neural network calculations based on the prediction results, and stores the intermediate data generated during the calculation process into the DRAM memory.

14. The AI ​​device according to claim 12, characterized in that, The first storage module is a Nand Flash memory, and the second storage module is a DRAM memory; The Nand Flash memory is stacked with the computing module using a hybrid bonding method; or The Nand Flash memory, the DRAM memory, and the computing module including the NPU with in-memory computing are stacked from top to bottom using a hybrid bonding method.

15. The AI ​​device according to any one of claims 1 to 7, characterized in that, The first storage module includes at least two layers of Nand Flash memory, which are stacked using a hybrid bonding method; and / or The second storage module includes at least two layers of DRAM memory, and the at least two layers of DRAM memory are stacked in a TSV manner or a hybrid bonding manner.

16. The AI ​​device according to claim 8, characterized in that, The large model is a large model based on the MOE architecture. The first storage module stores at least the parameters of the sparse part of the large model, and the second storage module stores at least the parameters of the dense part of the large model and the intermediate data generated by the neural network calculation. The computing module performs neural network calculations based on the input data and data from the first storage module and the second storage module.

17. The AI ​​device according to any one of claims 1 to 7, characterized in that, The AI ​​device further includes an interposer, a packaging substrate, and an I / O buffer for serial-to-parallel conversion. The interposer is disposed on the packaging substrate, and the I / O buffer is electrically connected between the first storage module and the interposer. The computing module and the first storage module are both disposed on the interposer. The second storage module is stacked on the computing module. The first storage module, the second storage module, the computing module, the interposer, and the packaging substrate are packaged into a whole using a chiplet method.

18. The AI ​​device according to any one of claims 1 to 7, wherein the AI ​​device is an edge-based inference device.

19. An electronic device, characterized in that, The AI ​​device includes any one of claims 1-18.