Large language model inference device, inference system and electronic device based on in-memory computing
By adopting a memory-based integrated design in the end-side AI large language model inference device, combined with hybrid bonding technology and in-memory computing matrix, the requirements of high bandwidth, large computing power, low power consumption and good heat dissipation are solved, and more efficient inference performance and lower energy consumption are achieved.
Patent Information
- Application Number
- CN202411873249.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-12-18
AI Technical Summary
The existing end-side AI large language model inference devices face the needs of high bandwidth, large computing power, low power consumption and good heat dissipation, but it is difficult to effectively meet these requirements, especially when large model inference, the bandwidth and power consumption requirements are huge, and the heat dissipation problems are prominent.
The LLM inference device based on the integrated storage and computing is adopted to stack the storage layer and the computing layer through a hybrid bonding method to realize the neural network accelerator of the in-memory computing matrix, reduce the energy loss of data transfer and signal transmission, and separate the pre-filling processing and decoding processing, making full use of the computing power of the neural network accelerator.
It has achieved the reduction of power consumption, reduced heat generation, improved computing power and bandwidth efficiency, solved the problems of bandwidth bottlenecks, high power consumption and difficulty in heat dissipation in the existing technology, and improved the efficiency and performance of large-language model inference.
Smart Images

Figure CN119337953B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to an LLM inference device, an inference system, and an electronic device based on in-memory computing. Background Art
[0002] The AI large language model (LLM) inference device is a major direction for the future implementation of AI applications. In particular, the edge-side AI large language model inference chip is an important device for the future implementation of AI applications, and it has very broad development prospects in fields such as AI laptops, AI mobile phones, and intelligent robots.
[0003] The large language model is a natural language processing model based on deep learning technology. With the rapid development of the AI large language model inference device and the rapid growth of social demand for it, the deployment of the edge-side AI large language model inference device faces many difficulties and challenges. For example, with the rapid development of edge-side large models, their number of parameters will continuously increase with iterations, their bandwidth requirements will also continue to increase, and their computing power requirements are becoming increasingly large. Moreover, when processing a large amount of data, they will require a large amount of power consumption and generate a large amount of heat. Correspondingly, their demand for low power consumption is increasing, and there is an urgent need to solve the heat dissipation problem, so that the edge-side AI large language model inference device can meet the growing demand for various functions of the edge-side AI large language model inference device.
[0004] In summary, with the development of artificial intelligence, the AI large language model inference device increasingly requires high bandwidth, large computing power, low power consumption, and good heat dissipation. Summary of the Invention
[0005] The purpose of the present invention is to provide a large language model inference device, an inference system, and an electronic device based on in-memory computing to solve some or all of the above technical problems.
[0006] To achieve the above purpose, the large language model inference device based on in-memory computing provided by the present invention includes: at least a storage layer for storage; at least a computing layer for computing, and the computing layer is stacked with the storage layer by means of hybrid bonding; the computing layer includes a neural network accelerator based on in-memory computing, and the neural network accelerator includes an in-memory computing matrix for performing neural network calculations on input feature data and weights from the storage layer; the computing layer is also electrically connected to a main control chip for controlling the inference device, and the computing layer is also used for performing pre-filling processing of large language model inference and transmitting the pre-filled processed data to the main control chip for decoding processing of large language model inference, so as to separate the pre-filling processing and the decoding processing.
[0007] In a preferred embodiment, the in-memory computing matrix includes: N columns of memory operator modules, and each column of memory operator modules includes M rows of memory and computing units for storing and computing the stored data with the input feature data.
[0008] In a preferred embodiment, each memory and computing unit includes: an SRAM memory for storing weights from the storage layer, and a logic unit for computing, which is disposed adjacent to the SRAM memory.
[0009] In a preferred embodiment, the in-memory computing matrix includes: a weight storage array, a bit multiplier, a storage readout circuit, and a logic operation unit; the weight storage array is used for storing weights; the bit multiplier is used for receiving a weight read enable signal and input feature data, and when the weight read enable signal is enabled, selecting the weights in the weight storage array according to the weight read enable signal, and multiplying the bit positions of the input feature data by the selected weights; the storage readout circuit is used for reading out the multiplied product to the logic operation unit when the bit position of the input feature data is 1, and not performing a readout operation when the bit position of the input feature data is 0; the logic operation unit is used for accumulating the product read out by the storage readout circuit when the bit position of the input feature data is 1, so as to implement multiplying and accumulating the input feature data and the weights.
[0010] In a preferred embodiment, the bit multiplier includes N AND gates respectively corresponding to and connected to the weight storage array; one input terminal of the AND gate is used for accessing the input feature data, the other input terminal is used for accessing the weight read enable signal, and the output terminal is used for outputting a control signal.
[0011] In a preferred embodiment, the storage layer is further used for computing to implement the first-stage computing of the neural network computing, and the computing layer completes the neural network computing according to the result of the first-stage computing.
[0012] In a preferred embodiment, the storage readout circuit further includes an AND gate; one input terminal of the AND gate is used for accessing the input feature data, the other input terminal is used for accessing the weight read enable signal, and the output terminal is used for outputting a first control signal.
[0013] In a preferred embodiment, the storage layer includes at least two layers of DRAM memories, and the at least two layers of DRAM memories are connected by means of TSV.
[0014] The present invention further provides an inference system, which includes the large language model inference device described in any one of the above, and a main control chip electrically connected to the large language model inference device. The main control chip is configured to perform decoding processing of large language model inference on the data after pre-filling processing in the calculation layer of the large language model inference device.
[0015] The present invention further provides an electronic device, including the large language model inference device described in any one of the above.
[0016] In the large language model inference device based on memory-compute integration provided by the present invention, its storage layer and the calculation layer including a neural network accelerator based on memory-compute integration are stacked by means of hybrid bonding. Thereby, data transfer can be reduced, energy loss of signal transmission can be reduced, and time delay of signal propagation can be reduced. It can reduce power consumption, generate less heat, and have a large computing power. Moreover, the pre-filling processing and decoding processing in LLM inference are respectively performed in the above-mentioned calculation layer and the main control chip. It can make full use of the neural network accelerator with large computing power in the present invention, so that the pre-filling processing requiring large computing power is executed in the above-mentioned calculation layer. Thereby, it can support a large language model inference device requiring high bandwidth and also improve the efficiency of large language model inference. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the following described drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained according to these drawings.
[0018] Figure 1 It is a schematic structural diagram of a large language model inference device and a large language model inference system provided by an embodiment of the present invention;
[0019] Figure 2 It is a schematic module diagram of a neural network accelerator provided by an embodiment of the present invention;
[0020] Figure 3 It is a schematic module diagram of a column of memory operator modules in the in-memory computing matrix according to an embodiment of the present invention;
[0021] Figure 4 It is a longitudinal structural diagram of a column of memory operator modules in the in-memory computing matrix according to an embodiment of the present invention;
[0022] Figure 5 It is a schematic principle diagram of the in-memory computing matrix performing calculations according to an embodiment of the present invention;
[0023] Figure 6 It is a schematic module diagram of the in-memory computing matrix according to an embodiment of the present invention;
[0024] Figure 7 It is a schematic structural diagram of an in-memory computing matrix according to an embodiment of the present invention. Specific embodiments
[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application. Without conflict, the following various embodiments and their technical features can be combined with each other.
[0026] In order to make the purpose, technical solutions and advantages of the present invention clearer, various exemplary embodiments to be described below will refer to the corresponding drawings, which form a part of the exemplary embodiments and describe various exemplary embodiments that may be used to implement the present invention. Unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. It should be understood that they are only examples of processes, methods, devices, etc. consistent with some aspects of the present invention disclosed in detail in the appended claims, and other embodiments may also be used, or structural and functional modifications may be made to the embodiments listed herein without departing from the scope and essence of the present invention.
[0027] In the description of the present invention, it should be understood that terms such as "center", "longitudinal", "transverse", etc. indicate the orientation or positional relationship based on the drawings shown, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the elements referred to must have a specific orientation, be constructed and operated in a specific orientation. Terms such as "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. The meaning of the term "plurality" is two or more. The terms "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, an integral connection, a mechanical connection, an electrical connection, a communication connection, a direct connection, an indirect connection through an intermediate medium, and can be the internal communication of two elements or the interaction relationship between two elements. The term "and / or" includes any and all combinations of one or more of the related listed items. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0028] The inference of large language models is mainly divided into two stages. The first stage is the prefill phase, which processes the input tokens in parallel. The second stage is the decode phase, which generates the next token one by one, repeating these two steps until the EOS (end-of-sequence) token is generated or the user-set stop condition is reached. That is, prefill is the process of providing the user's input as a prompt (initial information) to the model. This prompt can be regarded as the initial information of the model, which is used to guide subsequent generation. Since the prefill phase can be processed in parallel, it means that multiple inputs can be processed simultaneously, thus improving the efficiency of inference. Through parallel computing, the prefill operation can be completed faster, reducing the waiting time.
[0029] The inventors of the present invention found that: the computational workload in the prefill phase of large language model inference is relatively large, and the demand for computing power is greater than the demand for bandwidth; in the decode phase of large language model inference, the demand for computing power decreases significantly, but its bandwidth demand increases significantly. Taking a typical edge application scenario as an example: the Llama3-8B model, int8 quantization, first token time = 1s, 100 tokens / s in the decode phase. The computing power and bandwidth requirements in different phases are shown in Table 1:
[0030]
[0031] Table 1: Computing Power and Bandwidth Requirements of Llama3-8B Model (int8) in Different Phases
[0032] As can be seen from Table 1, in the application of edge large language model (LLM) inference chips, the computational workload in the prefill phase is relatively large, and the demand for computing power is greater than the demand for bandwidth; on the contrary, in the decode phase, the demand for computing power decreases significantly, but its bandwidth demand increases significantly to 800GB / s. This is because in the decode phase, the input becomes a vector, and the calculation changes from matrix multiplication to vector-matrix multiplication. At the same time, since more than 90% of the time in edge large model inference is spent in the decode phase, which is memory access intensive, its inference efficiency is limited by bandwidth rather than computing power. Due to the huge demand for bandwidth in the decode phase, existing GPUs and edge chips based on LPDDR cannot meet the deployment bandwidth requirements of edge large models (LLMs). Moreover, the amount of inference required by large language model inference devices is increasing, and the corresponding power consumption is also increasing. In the existing technology, the computing components in LLM inference devices are far from the radiator, making it difficult to effectively dissipate the heat generated by the above-mentioned computing components, and thus making LLM inference devices face the challenge of heat dissipation problems. In the existing technology, the heat dissipation capacity of the radiator is usually improved, but the heat dissipation improvement effect is still not good.
[0033] The large language model inference device based on in-memory computing provided by the present invention can solve problems related to computing power, bandwidth, power consumption, and heat dissipation.
[0034] To illustrate the technical solutions described in the present invention, specific embodiments will be used for illustration below, which only show parts related to the embodiments of the present invention.
[0035] Refer to Figure 1 As shown, the large language model inference device based on in-memory computing provided by the present invention includes: a storage layer (schematically shown as DRAM Layers in the figure) and a computing layer (the bottom layer in the example shown in the figure) stacked. The storage layer is at least used for storage, and it can be used to store the parameters of the large model, such as storing 8GB of large model parameters and storing the weights for neural network computing in the computing layer; the storage layer can also be used to store the output feature data after neural network computing. The computing layer is at least used for computing. The computing layer is stacked with the storage layer by means of hybrid bonding, and the computing layer is electrically connected to the storage layer to realize data interaction between the storage layer and the computing layer. The computing layer includes a neural network accelerator based on in-memory computing (i.e., the NPU acceleration chip in the figure) for neural network computing. The computing layer may include multiple sub-NPU clusters. The neural network accelerator includes an in-memory computing matrix for performing neural network computing on the input feature data and the weights from the storage layer. The computing layer is also electrically connected to the main control chip that controls the inference device, and the computing layer is also used to perform pre-filling processing for large language model inference and transmit the data after pre-filling processing to the main control chip for decoding processing of large language model inference, so that the pre-filling processing and the decoding processing are separated.
[0036] The storage layer and the computing layer provided by the present invention can be in various forms. For example, they can be chips or wafers.
[0037] In the present invention, the neural network accelerator included in the computing layer realizes in-memory computing through an in-memory computing matrix, that is, it can perform neural network calculations on the input feature data and the weights from the storage layer. In-memory computing is an innovative architecture designed to overcome the "memory wall" problem by integrating storage and computing functions on the same chip. This technology embeds computing capabilities in the memory and uses a new operation architecture to implement matrix-vector multiplication and accumulation operations, so it can reduce the frequent data transfer back and forth between the memory and the computing chip in the prior art, thereby greatly reducing the transfer time, power consumption, and heat generation, and improving the parallel processing efficiency of data; it realizes in-situ computing, eliminates bandwidth limitations and data movement costs. This can essentially eliminate the latency and power consumption of unnecessary data transfer, improve the artificial intelligence computing efficiency by hundreds or thousands of times, reduce costs, and break through the "memory wall" and "power wall". It can enable the computing layer and the large language model inference device to have greater computing power, consume less power, and generate less heat compared to the devices that perform calculations under the same conditions in the prior art, so as to solve the existing heat dissipation problem from the source of heat.
[0038] Moreover, in the present invention, the storage layer and the computing layer in the LLM inference device are realized in a 3D stacked structure through Hybrid Bonding. Hybrid Bonding can achieve high-density and high-performance interconnection between different chips. Through Hybrid Bonding, the metal layers (usually copper layers) of two or more chips can be precisely aligned and directly pressed together to form direct electrical contacts. It can achieve an interconnection pitch of sub-micron or even nano-scale, allowing more connection points to be placed in a smaller area, greatly increasing the data communication bandwidth between chips. Its contact density can reach 10K - 1MM / mm2. Since Hybrid Bonding eliminates intermediate media such as solder and other materials, the direct copper-to-copper connection has lower resistance, reduces the energy loss of signal transmission, and also reduces the time delay of signal propagation; the compact structure and direct conduction path formed by Hybrid Bonding contribute to improving thermal management and reducing the heat generation problem, so it can achieve higher integration, faster data transmission speed, larger bandwidth, and lower power consumption.
[0039] Furthermore, the inference device provided by the present invention combines the advantage of high computing power of its computing layer based on in-memory computing, separates prefill from decode, that is, it completes the prefill that requires high computing power in the computing layer, and transmits the data processed by the computing layer through prefill to the main control chip electrically connected thereto. The main control chip decodes the data after the prefill process. Thus, on the one hand, the demand for the bandwidth of the inference device is reduced, enabling the inference device to execute the inference work of large language models faster and better, and also enabling the large language model inference device with the same computing power to support a larger bandwidth for the inference work of large language models.
[0040] In summary, the inference device provided by the present invention can solve the existing bandwidth bottleneck, support a large bandwidth, support high computing power, and reduce the power consumption of the entire inference device in multiple aspects, thus solving the existing heat dissipation problem.
[0041] Referring to Figure 2 the illustrated embodiment, the neural network accelerator includes: a preprocessing module, an in-memory computing matrix, and a vector processing module that are electrically connected in sequence, and a shared memory that is electrically connected to the preprocessing module, the in-memory computing matrix, and the vector processing module respectively; and further includes a chip controller that is electrically connected to the shared memory. Among them, the NPU is also electrically connected to an external memory. The above storage layer in the present invention is Figure 2The external memory shown. The data stored in the external memory is transported to the shared memory of the NPU in batches. For example, weights are read in batches from the external memory to the shared memory. One working mode is briefly described as follows: the input feature data is obtained from the shared memory by the preprocessing module, the input feature data is processed into matrix data in the in-memory computing matrix format, and the matrix data is input into the in-memory computing matrix; wherein, if the input feature data includes at least two input feature maps, and the at least two input feature maps are stored in the shared memory in a skip address manner, then the preprocessing module obtains the at least two input feature maps from the shared memory in an address sequential reading manner; or, if the input feature data includes at least two input feature maps, and the at least two input feature maps are stored in the shared memory in a continuous storage manner, then the preprocessing module obtains the at least two input feature maps from the shared memory in a skip address reading manner; weight data is obtained through the in-memory computing matrix, convolution calculation is performed according to the matrix data and the weight data through the in-memory computing matrix to obtain a calculation result, and the calculation result is input into the vector processing module; the calculation result is vector processed by the vector processing module to obtain output feature map data, and the obtained output feature map data is written into the shared memory and / or the external memory; wherein, if the output feature map data is an intermediate calculation result, then the output feature map data is used as the input data for the calculation result of the next level; before the preprocessing module obtains the input feature data, and / or, after the calculation result is processed by the vector processing module to obtain output feature map data, the chip controller controls the data interaction between the shared memory and the external memory, wherein the data for interaction includes the input feature data and / or the output feature map data.
[0042] In a preferred embodiment of the present invention, the in-memory computing matrix includes: N columns of storage operator modules, each column of storage operator modules includes: M rows of storage and calculation units for storing and calculating the stored data with the input feature data, and capacitors electrically connected to the storage and calculation units, and the charges output by the capacitors in M rows are converged together to achieve charge accumulation.
[0043] See Figure 3Schematic diagram of a column of memory operator modules in the in-memory computing matrix according to an embodiment of the present invention. In this preferred embodiment, the memory computing unit can store and perform multiplication calculations on weights and input feature data, and the corresponding products are accumulated by capacitors connected to the memory computing unit to achieve multiply-accumulate (MAC) calculations in neural network computing. That is, matrix calculations can be realized inside the in-memory computing matrix, which can avoid the frequent data transfer between two different and relatively distant memories and computing chips in the prior art. The power consumption of multiply-accumulate calculations usually accounts for most of the entire calculation. For example, in many cases, it accounts for 50-70% of the power consumption. Based on the memory-computation integrated architecture, the present invention sets a neural network accelerator with an in-memory computing matrix in the computing layer to achieve two-dimensional and three-dimensional matrix multiplication / addition operations through a new operation architecture, thereby greatly reducing the power consumption of neural network computing, making the area of the neural network accelerator smaller, and enabling a larger computing power to be supported and less heat to be generated under the same area and power consumption to solve the heat dissipation problem.
[0044] See Figure 4 and Figure 5 as shown Figure 4 Longitudinal structure schematic diagram of a column of memory operator modules in the in-memory computing matrix according to an embodiment of the present invention. Each memory computing unit includes: an SRAM memory (i.e., Figure 4 the SRAM Memory in it) for storing weights from the storage layer, and a logic unit for calculation adjacent to the SRAM memory. A plurality of memory computing units are stacked to form a column, and several columns of the memory computing units form a memory computing matrix. Among them, digital logic, that is, the above-mentioned logic unit, can be used to perform multiplication and addition calculations on the weights and input feature data in the SRAM memory. Figure 4It also shows a typical architecture block diagram of digital compute-in-memory (CIMD). Specifically, it is based on SRAM memory, and the model weights are stored in multiple on-chip SRAM Memory Banks. The multiply-accumulate (MAC) calculation logic circuit required by CIMD is deployed near each SRAM Memory Bank to significantly reduce the latency and power consumption caused by weight transfer. Compared with traditional von Neumann architecture artificial intelligence (AI) chips, digital compute-in-memory (CIMD) significantly reduces the power consumption of data transfer and at the same time significantly improves the speed of parallel multiply-accumulate (MAC) calculations, effectively alleviating the two major problems of the "memory wall" and the "power wall". The SRAM memory and the logic unit for calculation are closely arranged in a stacked manner, which can make the structure of the in-memory computing matrix more compact, with greater computing power, lower power consumption, and combined with digital logic units, it can make the computing speed faster and the computing accuracy higher. As shown in Table 2, at the 12nm process node, the matrix multiplication energy efficiency ratio of the CIM in-memory computing unit can reach 24.7 Tops / W, which is an order of magnitude improvement in energy efficiency ratio compared with traditional digital architectures (i.e., general computing units, such as GPUs and TPUs). Compute-in-memory (CIMD) significantly reduces the power consumption of data transfer and at the same time significantly improves the speed of parallel multiply-accumulate (MAC) calculations. Thanks to this innovation in the underlying computing unit architecture, the in-memory computing (CIMD) technology can significantly improve the energy efficiency ratio of edge AI large language model (LLM) inference chips, that is, while maintaining the computing power, significantly reducing their power consumption requirements.
[0045]
[0046] Table 2
[0047] Figure 5 It is a schematic diagram of the principle of calculation by the in-memory computing matrix in an embodiment of the present invention. The compute-in-memory unit multiplies the input input feature data and weights, and then accumulates the products through a multi-stage adder tree to obtain the multiply-accumulate result.
[0048] See Figure 6 , in this preferred embodiment, the in-memory computing matrix includes a weight storage array 10 (also called a weight parameter storage array), a bit multiplier 12, a memory readout circuit 21, and a logic operation unit 22. One input end of the bit multiplier 12 is connected to the weight read enable signal, the other input end is connected to the input feature data, and the weight acquisition end is connected to the weight storage array 10; the input end of the memory readout circuit 21 is connected to the weight storage array 10, and the output end is connected to the logic operation unit 22; the output end of the logic operation unit 22 is used to output the convolution operation result. The weight storage array 10 is used to store weights.
[0049] The bit multiplier 12 is configured to receive a weight reading enable signal and input feature data (in a convolutional neural network, it is usually necessary to process the input feature map). When the weight reading enable signal is enabled, the bit multiplier 12 selects weights from the weight storage array according to the weight reading enable signal, and multiplies the bits of the input feature data by the selected weights. The input feature data is used to represent the data output from the previous layer operation in the convolutional operation process, and may include at least one bit, where the bit value is 1 or 0. For the input feature data, when its bit is 1, it indicates that the input feature data is 1, and when its bit is 0, it indicates that the input feature data is 0. For example, when the bit of the input feature data is 1, the product of the bit of the input feature data and the corresponding weight is the corresponding weight. At this time, the storage readout circuit 21 can read the corresponding weight from the weight storage array 10 to achieve the corresponding product readout. When the bit of the input feature data is 0, the storage readout circuit 21 does not perform a readout operation at this time. Therefore, the storage readout circuit 21 can directly perform a readout operation on the weight storage array 10 to achieve the corresponding product readout. The storage readout circuit 21 is configured to read the multiplied product to the logic operation unit 22 when the bit of the input feature data is 1, and does not perform a readout operation when the bit of the input feature data is 0, so that the storage readout circuit 21 does not consume power when the bit of the input feature data is 0, thereby reducing the power consumption of the storage readout circuit 21. Specifically, the storage readout circuit 21 can read the corresponding weight to the logic operation unit 22 when the bit of the input feature data is 1. Specifically, the storage medium of the storage readout circuit 21 includes any one of SRAM, DRAM, and RRAM, and the storage medium can be a volatile storage medium or a non-volatile storage medium. The logic operation unit 22 is configured to accumulate the products read by the storage readout circuit 21 when the bit of the input feature data is 1, so as to implement multiplying and accumulating the input feature data and the weights. The above logic operation unit 22 only needs to output the final convolutional operation result without transmitting the intermediate data during the convolutional operation process, which can reduce the requirements for data output volume and data transmission bandwidth. That is, it enables the storage readout circuit not to consume power when the bit of the input feature data is 0, thereby reducing the power consumption of the storage readout circuit. The output data of the logic operation unit is the accumulated result without transmitting the intermediate data during the convolutional operation process, and the data bit width of its output data is reduced. If higher accumulation operations are implemented in the logic operation unit, the output data bit width can be further compressed. It can be seen that this application can reduce the power consumption of a large number of concurrent data reads and reduce the requirements for data output volume and data transmission bandwidth.
[0050] Further preferably, referring to Figure 7As shown, the bit multiplier includes N AND gates respectively connected to the weight storage array; one input terminal of the AND gate is used to access the input feature data fmi, the other input terminal is used to access the weight read enable signal, and the output terminal is used to output a control signal. The weight storage array 10 includes K rows and N columns of parameter storage units. Each parameter storage unit is used to store a weight or an electrical signal representing the corresponding weight. The parameter storage unit in the h-th row and j-th column stores the weight W(h, j), where 1 ≤ h ≤ K, 1 ≤ j ≤ N, K is the number of rows of the weight, and N is the number of columns of the weight. The K rows of parameter storage units are correspondingly arranged on K row bit lines; the h-th row bit line is denoted as bitline,h. The above-mentioned weight read enable signal is used to enable one column of parameter storage units in the N columns of parameter storage units, that is, one bit weight read enable signal corresponds to one column of parameter storage units. For example, the weight read enable signal corresponding to the j-th column of parameter storage units can be denoted as wordline,j. When the weight read enable signal wordline,j corresponding to the j-th column of parameter storage units is 1, the electrical signal representing the weight is output to the corresponding bit line, so that the storage readout circuit 21 can read the corresponding weight signal through each row bit line. For example, the probability that an N-bit input vector is all 0 is much lower than the probability that some of the input bits in an N-dimensional input vector are 0. The bit-level sparsity of the data is very high. However, it is more difficult to use data sparsity with a smaller granularity, mainly reflected in the need to perform real-time arithmetic judgments on each bit of the input data. Through this preferred method, the present invention can solve the problem of data sparsity. For the case where the data sparse matrix contains a large number of zero values, it can greatly reduce resource waste, reduce unnecessary computation, and improve the utilization rate of storage space.
[0051] In a preferred embodiment of the present invention, the storage readout circuit further includes an AND gate; one input terminal of the AND gate is used to access the input feature data, the other input terminal is used to access the weight read enable signal, and the output terminal is used to output a first control signal; for the weight read enable signal wordline,j corresponding to the j-th column and the input feature data fmi,j, when both the weight read enable signal wordline,j and the input feature data fmi,j are 1, the first control signal is 1; when at least one of the weight read enable signal wordline,j and the input feature data fmi,j is 0, the first control signal is 0. Its circuit is simple and can accurately control the read operation, which can reduce the power consumption of the storage readout circuit and the power consumption of the large language model inference device.
[0052] In a preferred embodiment of the present invention, the storage layer is also used for computing to implement the first-level computing of the neural network computing, and the computing layer completes the neural network computing according to the result of the first-level computing. That is, some simple computations can also be performed in the storage layer. For example, the first multiplication of the multiplication parameters, i.e., the first-level computing. The result after the first-level computing is transmitted to the NPU for accumulation computing to complete the neural network computing. It can make full use of the storage layer, reduce the computing amount of the computing layer, and thus reduce the power consumption of the computing layer. Moreover, after the first-level computing process, the data transfer from the storage layer to the computing layer can be reduced, and the requirements for storage capacity and high bandwidth can be correspondingly reduced. The storage layer preferably includes a DRAM memory, which has a low cost and a high density. Further preferably, the storage layer includes at least two layers of DRAM memories, and the at least two layers of DRAM memories are connected by TSV (Through Silicon Vias) method, which can vertically stack multiple DRAM chips, significantly improve the memory bandwidth, and reduce the power consumption; achieve large-capacity, high-bandwidth storage, and meet the stringent requirements for memory in the fields of high-performance computing, artificial intelligence, etc.
[0053] As shown in Table 3, the 3D stacked DRAM integrated with Hybrid Bonding provided by the present invention can reach the bandwidth index of 1-60TB, which is a significant improvement compared with other state-of-the-art GDDR6 and HBM3E technologies; at the same time, the energy requirement for transmitting a single bit of data is also significantly reduced to 0.5pJ / bit, greatly reducing the power consumption index of data transfer. It can solve the bandwidth bottleneck problem encountered in the deployment of end-side AI large model (LLM) chips, and at the same time, it also significantly reduces the power consumption of data transfer.
[0054]
[0055] Table 3
[0056] In summary, the present invention provides a new LLM inference device, which combines multiple technologies such as 3D stacked DRAM, Hybrid Bonding technology, and high energy efficiency ratio SRAM computing-in-memory. It can solve the bottleneck problems of bandwidth, power consumption, computing power, and heat dissipation currently encountered by end-side AI large model (LLM) inference devices (the inference device can be in various forms such as chips or modules). It can significantly improve the memory bandwidth, reduce the power consumption, achieve large-capacity, high-bandwidth storage, and meet the stringent requirements for memory in the fields of high-performance computing, artificial intelligence, etc. It can be widely applied to various end-sides, such as mobile phones, computers, robots, etc.
[0057] The present invention also provides an inference system (see Figure 1As shown in the figure), it includes any one of the above large language model inference devices and a main control chip electrically connected to the large language model inference device. The main control chip is used to perform decoding processing of large language model inference on the data after pre-filling processing in the calculation layer of the large language model inference device. Among them, the structure of the NPU included in the calculation layer is also briefly illustrated in Figure 1 The inference system provided by the present invention enables the pre-filling processing and decoding processing to be separated and carried out in different devices, and enables the pre-filling processing that requires high computing power to be processed in the structure of in-memory computing based on SRAM, while the main control chip, which has relatively less computing power, is used for decoding. This enables the inference system of the large language model to greatly solve the existing bandwidth bottleneck problem, greatly reduce its power consumption and alleviate the heat dissipation problem. The main control chip can be various types of chips, such as GPUs, SOCs, etc.
[0058] The present invention also provides an electronic device, which includes the above large language model inference device or inference system. It can greatly reduce the power consumption and cost of the electronic device, and improve the efficiency and experience of AI inference. It can be widely used in smartphones, tablets, wearable electronic devices, smart home electronic products, and so on.
[0059] The above are only the preferred embodiments of the present invention. Those skilled in the art know that various changes or equivalent replacements can be made to these features and embodiments without departing from the spirit and scope of the present invention. Additionally, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the protection scope of the present invention.
Claims
1. A large language model inference device based on a terminal with integrated storage and computing, characterized in that: include: A storage layer at least for storage on the end side; A computing layer on the end side for at least computing; The computing layer of the terminal side includes a neural network accelerator NPU with storage and computing integrated based on SRAM, and the neural network accelerator NPU includes an in-memory computing matrix, and the in-memory computing matrix is used to perform neural network calculation on input feature data and weights of the storage layer from the terminal side, and the computing layer of the neural network accelerator with storage and computing integrated based on SRAM is stacked with the storage layer by a hybrid bonding method, and the computing layer is electrically connected to the storage layer for data interaction; The computing layer on the end side is also used to be electrically connected to the main control chip that controls the inference device. The computing layer is also used to perform pre-filling processing of large language model inference on the end side that requires large computing power and transmit the data after the pre-filling processing to the main control chip for decoding processing of the large language model inference on the end side, so as to separate the pre-filling processing on the end side and the decoding processing on the end side.
2. The large language model inference device according to claim 1, characterized in that: The in-memory calculation matrix includes: N columns of storage operator modules, each column of the storage operator module includes M rows of storage and calculation units for storing and calculating the stored data and the input feature data.
3. The large language model inference device according to claim 2, characterized in that: Each storage and calculation unit includes: an SRAM memory for storing weights from the storage layer, and a logic unit for calculation arranged adjacent to the SRAM memory.
4. The large language model inference device according to claim 1, characterized in that: The in-memory calculation matrix includes: a weight storage array for storing weights, a bit multiplier, a storage readout circuit and a logic operation unit; the bit multiplier is used to receive a weight read enable signal and input feature data, and when the weight read enable signal is enabled, the weight in the weight storage array is selected according to the weight read enable signal, and the bit of the input feature data is multiplied by the selected weight; the storage readout circuit is used to read the multiplied product to the logic operation unit when the bit of the input feature data is 1, and does not perform the readout operation when the bit of the input feature data is 0; the logic operation unit is used to accumulate the product read out by the storage readout circuit when the bit of the input feature data is 1, so as to realize the multiplication and accumulation of the input feature data and the weight.
5. The large language model inference device according to claim 4, characterized in that: The bit multiplier includes N AND gates respectively connected to the weight storage array; one input end of the AND gate is used to access the input feature data, the other input end is used to access the weight read enable signal, and the output end is used to output a control signal.
6. The large language model inference device according to claim 4, characterized in that: The storage readout circuit also includes an AND gate; one input end of the AND gate is used to access the input feature data, the other input end is used to access the weight read enable signal, and the output end is used to output the first control signal.
7. The large language model inference device according to any one of claims 1 to 5, characterized in that: The storage layer is also used for calculation to implement the first level calculation of the neural network calculation, and the calculation layer completes the neural network calculation according to the result of the first level calculation.
8. The large language model inference device according to any one of claims 1 to 5, characterized in that: The storage layer includes at least two layers of DRAM memory, and the at least two layers of DRAM memory are connected via TSV.
9. An inference system, characterized in that It comprises a large language model inference device as described in any one of claims 1 to 8, and a main control chip electrically connected to the large language model inference device, wherein the main control chip is used to perform large language model inference decoding processing on the data after pre-filling processing of the computing layer in the large language model inference device.
10. An electronic device, characterized in that: A large language model inference device comprising any one of claims 1-8 or an inference system comprising claim 9.
Citation Information
Patent Citations
Data processing method for neural network accelerator, chip and electronic equipment
CN116152520A
Request processing method and device of large language model, medium, equipment and product
CN118916175A
Storage and calculation integrated module based on input data sparsity, chip and electronic equipment
CN118939232A