Operational node, layer cluster and accelerator of large language model

By designing special computing and storage units in large language models, using distributed layer cluster architecture and weak interconnection methods, the computing and storage architecture of large language models is optimized, the high latency and high power consumption problems in the inference process are solved, efficient computing and storage are achieved, and ultra-long sequence lengths and multi-user requests are supported.

CN120256371APending Publication Date: 2025-07-04BEIJING PINGXIN TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510335422.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-07-15
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing large language models have problems such as high latency, high power consumption and low hardware utilization in the inference process, especially in terms of weight data reading and KV cache data storage and reading and writing, resulting in inefficient computing.

Method used

A large language model computing node is designed, including computing units and storage units within the same design architecture, which are used for storage of static weighted data and KV cache data respectively. It adopts a distributed layer cluster architecture, and data transmission between weakly interconnected layer clusters and computing nodes is transmitted. Combined with batch processing and pipeline parallelism, the computing and storage architecture is optimized.

Benefits of technology

It effectively reduces power consumption and delay, improves computing efficiency, supports data storage with ultra-long sequence length, reduces chip costs, and can handle requests from hundreds of users at the same time, improving system communication and computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256371A_ABST
    Figure CN120256371A_ABST
Patent Text Reader

Abstract

The invention relates to an intelligent integrated circuit, an intelligent chip and an AI chip in an artificial intelligence hardware platform. The invention relates to the fields of deep neural networks, multilayer neural networks, convolutional neural networks and the like in an artificial intelligence general technology, in particular to an operation node, a layer cluster and an accelerator of a large language model. The large language model operation node comprises at least one calculation unit and a first storage unit which are located in the same design architecture, and the calculation unit is used for calculation; the first storage unit is used for residing static weight data in the operation process of the large language model; and the second storage unit is arranged outside the design structure and is used for storing KV cache data in the operation process of the large language model. According to the method, the problems of high power consumption and high delay caused by traditional external storage static weight data reading are effectively avoided, and compared with a traditional HBM scheme, the method has obvious advantages in performance power consumption and cost.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This is a divisional application. The application number of its parent application is: 202410943035.2, the application date is: July 15, 2024, and the invention title is "Operation Nodes, Layer Clusters, and Accelerators of Large Language Models". Technical Field

[0002] The present invention relates to intelligent integrated circuits, intelligent chips, and AI chips in the artificial intelligence hardware platform; and to fields such as deep neural networks, multi-layer neural networks, and convolutional neural networks in the general artificial intelligence technology. In particular, it relates to an operation node, layer cluster, and accelerator of a large language model. Background Art

[0003] In recent years, large language models have made remarkable progress in the field of artificial intelligence and are being applied more and more widely, such as text generation, dialogue systems, translation, and document summarization. Among them, the inference process is one of the core links of large language models and plays a crucial role in the running efficiency of the models. Since large language models are usually trained based on a large amount of data and have a super-large network structure, this makes the models consume a large amount of computing resources and memory during the inference process.

[0004] Currently, large language models have been widely applied in various fields. However, as the model scale increases, the inference speed and efficiency have become an urgent problem to be solved. Existing inference acceleration solutions mainly use high-performance GPUs / TPUs, etc. Compared with CPUs, these solutions can achieve a certain degree of acceleration effect while maintaining generality, but there are still problems such as high latency, high power consumption, and low hardware utilization during the inference process. On the other hand, to achieve high-performance inference, using advanced processes greatly increases the inference cost; and as the model size continues to grow, the overhead of interconnection and communication will also increase significantly. Currently, there are numerous challenges in the inference process of large language models:

[0005] (1) The increasing number of user KV caches

[0006] The KV Cache data dynamically generated during the inference process, its data volume will increase with the increase of the sequence length and the number of users. The storage, reading, and writing of this part of the data will bring an increase in power consumption and latency;

[0007] (2) High power consumption and high latency in weight data reading

[0008] Large language models have a huge number of parameters ranging from billions to tens of billions, which causes problems such as high power consumption and high latency in the weight reading stage of traditional DRAM;

[0009] (3) The operation speed caused by the huge amount of calculations

[0010] During the parameter calculation process of large language models, a large number of matrix multiplication operations are required. In the decode stage, every time a token is generated, all the parameters of the neural network need to be recalculated, resulting in a huge computational load. Summary of the Invention

[0011] I. Technical Problems to be Solved

[0012] The present invention is expected to solve at least one of the above technical problems in part.

[0013] II. Technical Solutions

[0014] The first aspect of the present invention provides a large language model operation node. The large language model operation node includes: at least one computing unit and a first storage unit, both located within the same design architecture, where: the computing unit is used for calculation; the first storage unit is used for resident static weight data during the operation of the large language model; the second storage unit is set outside the design structure and is used for storing KV cache data during the operation of the large language model.

[0015] The second aspect of the present invention provides a large language model layer cluster. The large language model layer cluster includes: N operation nodes, N≥1, and the operation node is the large language model operation node as described above.

[0016] The third aspect of the present invention provides a large language model accelerator. The large language model accelerator includes: M layer clusters, M is the number of layers of the large language model, M≥2; each layer cluster includes: one or more operation nodes; the operation node includes: at least one computing unit and at least one storage unit, where: the computing unit is used for calculation; the storage unit is used for storing data required for the operation of the large language model; among them, the layer clusters in the M layer clusters correspond to the layer order of the large language model, and the data required for the operation of the large language model is stored in the storage unit of the operation node of the corresponding layer cluster.

[0017] III. Beneficial Effects

[0018] From the above technical solutions, it can be seen that the present invention has at least one of the following beneficial effects compared with the prior art:

[0019] (1) Reduce power consumption and latency

[0020] The present invention effectively avoids the high power consumption and high latency problems caused by the traditional external memory static weight data reading. Compared with the traditional HBM solution, the present invention has obvious advantages in performance, power consumption and cost.

[0021] (2) Expand high-flexibility computing power

[0022] In the present invention, a single computing node has first and second computing units dedicated to matrices and vectors, and the cascaded architecture can meet the computing power requirements of models of different sizes. The second computing unit is a programmable unit that can flexibly meet the requirements of different large language model architectures and adapt to different activation functions, normalization, etc.

[0023] (3) Support for ultra-long sequences

[0024] In the present invention, by caching the user KV Cache in the second storage unit, the storage requirements for hot data with ultra-long sequence lengths can be met.

[0025] (4) Weak interconnection between layer clusters

[0026] The present invention proposes the concept of layer cluster Layer Cluster, that is, a basic unit composed of one chip or a group of chips is used to process a single Layer, so as to divide the entire large language model into units of Layer. Since there is no interaction of static weight data or KV cache data between adjacent layer clusters, only the transfer of token vectors, therefore, the layer clusters are in a weakly interconnected state.

[0027] This distributed implementation method can effectively split the original large chip into multiple small chip groups to complete the same operations. Doing so can not only effectively improve the communication and computing efficiency of the system, but also reduce the total cost of the chip. Using multiple small-cost chips to complete the same calculations as large-cost chips can obtain advantages in terms of power consumption, price, and performance.

[0028] (5) Weak interconnection state between computing nodes in the same layer cluster

[0029] In the present invention, in the same layer cluster, there is no interaction of static weight data or KVcache data between computing nodes. Therefore, the computing nodes are in a weakly interconnected state, greatly improving the computing efficiency of the computing nodes and reducing the pressure on the communication bandwidth.

[0030] (6) Stackable architecture

[0031] In the present invention, the large language model accelerator architecture and the layer cluster architecture are both stackable architectures, that is, the combination of hardware architectures suitable for the calculation of large language models can be selected according to the depth and scale of the large language model. Further, it is quite flexible and convenient to increase or decrease layer clusters and computing nodes according to the actual operation of the large language model.

[0032] (7) Economical and efficient interconnection

[0033] In the present invention, prefill and decode are separated to reduce the data transmission requirements between clusters, thereby reducing the dependence on high-cost interconnection technologies.

[0034] (8) High concurrency design

[0035] In the present invention, batch processing and pipeline parallelism are used to handle requests from hundreds of users simultaneously. Description of the Drawings

[0036] Figure 1 It is a schematic structural diagram of the operation node of the large language model according to an embodiment of the present invention.

[0037] Figure 2 It is a schematic structural diagram of the large language model accelerator according to an embodiment of the present invention.

[0038] Figure 3 It is a design diagram of data storage and decoder calculation of the large language model accelerator according to an embodiment of the present invention.

[0039] Figure 4 It is a schematic diagram showing that the static weight matrix in the large language model accelerator is stored row by row in N1 static storage units according to an embodiment of the present invention.

[0040] Figure 5 It is a schematic diagram showing that the static weight matrix in the large language model accelerator is stored column by column in N1 static storage units according to an embodiment of the present invention. Detailed Embodiments

[0041] Based on the transformer-type structure algorithm, the present invention proposes the concepts of classified storage, hierarchical setting, and distributed calculation, and designs unique operation nodes, layer clusters, and accelerators for the large language model, aiming to achieve inference acceleration, that is, to reduce the consumption of time, bandwidth, and computing resources during the inference process without sacrificing the model performance.

[0042] To make the purpose, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.

[0043] According to the first aspect of the present invention, an operation node of a large language model is provided. Figure 1 It is a schematic structural diagram of the operation node of the large language model according to an embodiment of the present invention. As Figure 1 shown, the operation node of the large language model in this embodiment includes:

[0044] 2 computing units and a first storage unit, all of which are located on the same chip, where:

[0045] The first storage unit is used to resident the static weight data during the operation of the large language model;

[0046] The first computing unit is mainly used to handle matrix calculation tasks in the operation nodes of large language models;

[0047] The second computing unit is mainly used to handle vector calculation tasks in the operation nodes of large language models.

[0048] The second storage unit is set outside the chip and is mainly used to store KVcache data during the operation of large language models.

[0049] Regarding the operation nodes of large language models in this embodiment, the following aspects need to be emphasized:

[0050] 1. Computation units with classified design

[0051] In the operation of large language models, matrix operations and vector operations are two of the most basic operations. However, different chip architectures may be more efficient for one type of operation and less efficient for the other.

[0052] In different existing technologies, matrix operations and vector calculations are processed in the same computing unit. In this embodiment, two independent computing units - the first computing unit and the second computing unit - are established. For matrix operations in large language model operations, they are assigned to the first computing unit that is better at matrix operations for processing. For vector operations in large language model operations, they are assigned to the second computing unit that is better at vector operations for processing. Thus, the advantages of different design structures are fully utilized, and the operation efficiency is improved. Specifically,

[0053] (a) The first computing unit

[0054] The first computing unit is mainly designed for matrix calculations and is also called the matrix computing unit. It mainly undertakes matrix-related operations, such as: matrix-vector multiplication, matrix-matrix multiplication, etc.

[0055] (b) The second computing unit

[0056] The second computing unit is mainly designed for vector calculations and is also called the vector computing unit. It mainly undertakes vector-related operations. Through the combination of some basic operators (such as addition, subtraction, multiplication, division, power calculation, square root lookup table, etc.), users can flexibly implement any vector operation, such as: various calculations between vectors and vectors, vectors and scalars; addition, multiplication, division, and exponential calculation of corresponding positions of vectors; various activation functions, normalization functions, etc.

[0057] Those skilled in the art should understand that the above are only preferred embodiments of the present invention. In other embodiments of the present invention, matrix operations and vector calculations can also be undertaken by the same computing unit, or more than 2 computing units can be designed, which can all implement the present invention and are also within the protection scope of the present invention.

[0058] 2. Two storage units that are classified and stored and separately set

[0059] During the operation of the large language model, the data types involved include: static data and dynamic data. Among them:

[0060] ① Among the static data, the static weight data is shared by all users. Its characteristics are: frequent calls, but the data volume is small and relatively fixed.

[0061] ② The dynamic data includes: KV Cache data. The KV Cache data is bound to a specific user, that is, different users correspond to different KV cache data. Its characteristics are: relatively low call frequency, but the data volume increases continuously as the number of user questions and the number of users increase.

[0062] It should be noted that in the process of providing feedback for the real-time input of the current user, only the KV-Cache of the current user participates in the calculation, while the KV-Cache data of other users does not participate in the calculation.

[0063] In addition, the dynamic data also includes: intermediate calculation results related to a specific calculation process; communication cache data between layer clusters.

[0064] For the data types in the above-mentioned operation process of the large language model, the storage unit in the present invention is specially set. In this embodiment, the first storage unit is located inside the chip, that is, used as an on-chip storage unit; the second storage unit is located outside the chip, that is, used as an off-chip storage unit.

[0065] The first storage unit: As the on-chip memory, it is mainly used to store static data, such as the static weight data of the large language model, and a small amount of dynamic data, such as intermediate calculation results and communication cache data. Compared with the second storage unit, the first storage unit can provide greater bandwidth, lower latency and lower read / write power consumption. By completely resident the static weight data in the first storage unit, the present invention avoids the high power consumption and high latency problems caused by frequent data reading from the second storage unit.

[0066] The second storage unit: As the main user of off-chip memory, it stores KV cache data. The characteristic of the dynamic data KV Cache generated during the operation of large language models is that its size increases with the growth of the sequence length. Therefore, storing it in the on-chip first memory will severely limit the maximum supportable sequence length. The present invention utilizes the second storage unit, i.e., off-chip storage, to store KV Cache data, and utilizes the characteristics of high density and high capacity of off-chip storage to support longer sequence lengths. Further, in order not to make the off-chip storage reading become a performance bottleneck, a single computing node will be configured with multiple off-chip memory channels to improve throughput, and at the same time, technologies such as dynamic compression are adopted to further improve efficiency.

[0067] In order to adapt to the corresponding computing processes, call frequencies, data volumes, read-write updates, etc., the differences between the first storage unit and the second storage unit are mainly reflected in two aspects:

[0068] ① Interface bandwidth;

[0069] In the present invention, since the call frequency of static weight data is higher. Therefore, the first storage unit for storing static weight data and the computing unit are arranged on the same chip, so that a larger interface bandwidth can be provided between the two. Specifically, in this embodiment, it satisfies: K1 ≥ 4K2; where K1 is the interface bandwidth between the first storage unit and the computing unit; K2 is the interface bandwidth between the second storage unit and the computing unit.

[0070] Through the above settings, the present invention effectively avoids the problems of high power consumption and high latency brought by the traditional off-chip weight data reading. Compared with the traditional HBM solution, the present invention has obvious advantages in terms of performance, power consumption and cost.

[0071] Those skilled in the art should understand that the above interface bandwidth setting is only an example. In actual scenarios, the interface bandwidth can be set according to needs, and the present invention can also be implemented, and it is also within the protection scope of the present invention.

[0072] ② Storage capacity

[0073] In the present invention, since the user's KV cache data is large in terms of data volume on the one hand and will become larger and larger with the increase of the number of inference layers on the other hand. Therefore, the KV Cache data is stored outside the chip, so that a larger storage capacity and more flexible capacity expansion ability can be provided.

[0074] In this embodiment, it satisfies: T2 ≥ 4T1; where T1 is the storage capacity of the first storage unit; T2 is the storage capacity of the second storage unit. And, the storage capacity of the first storage unit is fixedly set; the storage capacity of the second storage unit is expandably set.

[0075] To meet the above requirements, in this embodiment, the first storage unit is SRAM; the second storage unit is DRAM, and data interaction between the DRAM and the chip is carried out through a DDR interface.

[0076] In the present invention, the user KV Cache data is stored in the second storage unit, which can meet the storage requirements of hot data with ultra-long sequence lengths. At the same time, when the KV cache data continues to increase, the capacity of the second storage unit can be increased at low cost without replacing the chip.

[0077] Those skilled in the art should understand that the above storage capacity and memory type are only examples. In actual scenarios, the storage capacity and storage type can be set according to needs. In particular, the second storage unit can also be RRAM or MRAM, and these variations can also implement the present invention and are also within the protection scope of the present invention.

[0078] 3. Design architecture

[0079] In this embodiment, the computing units (including: the first computing unit, the second computing unit) and the first storage unit are arranged on the same chip, and the second storage unit is outside the chip.

[0080] In fact, the purpose of arranging the computing unit and the first storage unit on the same chip is to ensure the interface bandwidth between the two. This is because compared with the large amount of KV cache data, the static weight data is called very frequently, and sufficient interface bandwidth between the computing unit and the first storage unit can ensure the operation efficiency within the large language model.

[0081] Those skilled in the art should understand that arranging the computing unit and the first storage unit on the same chip is only one way to ensure the interface bandwidth between the two, and other similar design architecture methods can also be used to ensure the interface bandwidth between the two, for example:

[0082] ① The design architecture is: die; at least 1 computing unit and the first storage unit are located within the same die; the second storage unit is outside the die;

[0083] ② The same design architecture is: board; at least 1 computing unit and the first storage unit are located within the same board; the second storage unit is outside the board;

[0084] ③ The design architecture is: the package of the chips where at least 1 computing unit is located and the chip where the first storage unit is located; the second storage unit is outside the package.

[0085] ④ The design architecture is: the package of the dies where at least 1 computing unit is located and the die where the first storage unit is located; the second storage unit is outside the package;

[0086] ⑤The design architecture is: a functional structure block integrating the functions of at least one computing unit and a first storage unit; the second storage unit is outside the functional structure block.

[0087] The above 5 deformation methods can also ensure the interface bandwidth between the computing unit and the first storage unit, can also implement the present invention, and are also within the protection scope of the present invention.

[0088] ④Interface of the operation node

[0089] In this embodiment, the operation node further includes: an external cluster interface for communicating the operation node where it is located with the upstream processing unit; an internal cluster interface for communicating the operation node where it is located with other large language model operation nodes within the large language model layer cluster where it is located; where V1≥4V2, where V1 is the interface bandwidth of the internal cluster interface; V2 is the interface bandwidth of the external cluster interface. The operation node in this embodiment includes 2 internal cluster interfaces - a first internal cluster interface and a second internal cluster interface, which are used for serial connection before and after within the layer cluster.

[0090] In this embodiment, since both the static weight data and the KV cache data required for the operation are located within the layer cluster, and only the vector to be calculated (input token vector) comes from outside the layer cluster, and the data volume of the vector to be calculated is limited, therefore, setting the bandwidth of the internal cluster interface to 4 times the bandwidth of the external cluster interface can make the best use of the interface resources and improve the operation efficiency.

[0091] Figure 2 It is a schematic structural diagram of a large language model accelerator according to an embodiment of the present invention. Figure 3 It is a design diagram of data storage and decode calculation of a large language model accelerator according to an embodiment of the present invention. For Figure 2 and Figure 3 For the large language accelerator shown, one horizontal layer of operation nodes belongs to the same layer cluster. In the present invention, the static weight data of the large language model will all be placed in the first storage unit, that is, the on-chip memory. However, the on-chip memory capacity of a single operation node is limited, so it is necessary to allocate the static weight data so that it can be fully stored in the on-chip memories of multiple operation nodes. The first level of allocation is the layer cluster. One layer cluster is responsible for one layer of the large language model, and there is no interaction of static weight data between layer clusters. Therefore, all the static weight data of a specific layer of the large language model is stored in the corresponding layer cluster, specifically, in the first storage units of multiple preset operation nodes in the layer cluster.

[0092] Based on the above, the second aspect of the present invention provides a large language model layer cluster. In an exemplary embodiment of the present invention, a large language model layer cluster (Layer Cluster) is provided. AsFigure 2 and Figure 3 As shown in Figure 3 , the large language model layer cluster in this embodiment includes: N computing nodes, where N ≥ 1, and the computing node is the large language model computing node in the above embodiment.

[0093] Regarding the large language model layer cluster in this embodiment, the following aspects need to be specifically explained:

[0094] 1. Data stored within the layer cluster

[0095] As described above, in this embodiment, the data stored by multiple computing nodes within the layer cluster is the static weight data of the corresponding layer in the large language model, the KV cache data of the same layer, and there is no weight interaction between layer clusters.

[0096] 2. Number of computing nodes within the layer cluster

[0097] In this embodiment, the number of computing nodes in a single layer cluster is mainly determined by the size of the static weight data volume of the single-layer large language model and the capacity of the first memory. Moreover, as the data volume of the static weight data increases, the number of computing nodes within the layer group can be expanded. Since the second storage unit is used as off-chip memory and its storage capacity is set to be expandable, the capacity of the second storage unit can be expanded to adapt to the increase in the KV cache data volume. Therefore, the KV cache data volume does not directly determine the number of computing nodes within the layer cluster.

[0098] 3. Second storage units of different computing nodes within the layer cluster

[0099] In the same layer cluster, the second storage units of different computing nodes are: physically independently set; or, logically independently set. Both of these methods can implement the present invention.

[0100] 4. Interfaces and connections of computing nodes within the layer cluster

[0101] As Figure 2 shown, in this embodiment, within the same layer cluster, N computing nodes are serially arranged within the large language model layer cluster through the intra-cluster interface. Moreover, the intra-cluster interface is: a PCIE interface or a serial port. The extra-cluster interface is a PCIE interface or a serial port.

[0102] Based on the above, the third aspect of the present invention provides a large language model accelerator. In an exemplary embodiment of the present invention, a large language model accelerator is provided. As Figure 2 and Figure 3As shown in the figure, the large language model accelerator in this embodiment includes: M layer clusters, which are the large language model layer clusters in the above embodiment, where M is the number of layers of the large language model and M≥2. Each layer cluster includes: one or more computing nodes. A computing node includes: at least one computing unit and at least one storage unit, where: the computing unit is used for computing; the storage unit is used for storing the data required for the large language model operation. Among them, the layer clusters in the M layer clusters correspond to the layer order of the large language model, and the data required for the large language model operation is stored in the storage units of the computing nodes in the corresponding layer clusters.

[0103] Specifically, the large language model accelerator of the present invention is composed of multiple layer clusters (Layer Cluster). Each layer cluster contains multiple computing nodes. The computing nodes within the layer cluster are sequentially connected in series one by one through corresponding interfaces. At the same time, the layer clusters are also connected in series sequentially. In addition, the accelerator will cooperate with the upstream processing unit. Each computing node will have a separate interface to communicate with the upstream processing unit. The number of layer clusters and the number of computing nodes can be expanded according to the size of the large language model.

[0104] In this embodiment, the upstream processing unit is connected to the M layer clusters and is used for the allocation of resources for each layer cluster, the tokenization of user input data, prefill operations, etc.; the de-tokenization and other operations of the top layer cluster.

[0105] It should be noted specifically that the upstream processing unit in Figure 2 and Figure 3 is not included in the large language model accelerator of this embodiment. However, in other embodiments of the present invention, the upstream processing unit can also be included in the large language model accelerator, which is also within the protection scope of the present invention.

[0106] In this embodiment, the first level of allocation is the layer cluster. One layer cluster is responsible for one layer of the large language model, and there is no interaction of static weight data between layers. Therefore, one layer cluster will store all the static weight data of its corresponding network layer on the first memories of all its internal computing nodes. The number of computing nodes in a single layer cluster is determined by the amount of static weight data of the single-layer network and the capacity of the first memory. Based on the above, it can be understood that the number of computing nodes in different layer clusters can be the same or different. And in some embodiments, a layer cluster can include only one computing node.

[0107] In this embodiment, the layer clusters and the layer cluster time use chip2chip interconnection, but the present invention is not limited thereto. In addition to the above-mentioned low-cost chip2chip interconnection solution, if the data transmission volume between chips exceeds the capacity of chip2chip, it can also be converted into a die2die interconnection form to provide higher interconnection bandwidth, that is, multiple chips are packaged together in the form of chiplets. Or all die can be placed on the same wafer and interconnected inside the wafer through metal layers. These solutions can also implement the present invention and are also within the protection scope of the present invention.

[0108] Based on the above, the following aspects need to be emphasized again:

[0109] ① The weak interconnection between layer clusters

[0110] Large language models are formed by stacking multiple layers of Layer. There is only very little token interaction between Layer and Layer. Based on this, the present invention proposes the concept of layer cluster Layer Cluster, that is, a basic unit composed of one chip or a group of chips is used to process a single Layer, so as to divide the entire large language model in units of Layer. Since there is no interaction of static weight data or KV cache data between adjacent layer clusters, only the transfer of token vectors, therefore, the layer clusters are in a weak interconnection state.

[0111] This distributed implementation method can effectively split the original large chip into multiple small chip groups to complete the same operation. This can not only effectively improve the communication and computing efficiency of the system, but also reduce the total cost of the chip. Using multiple low-cost chips to complete the same calculation as a large-cost chip can obtain advantages in terms of power consumption, price, and performance.

[0112] ② The weak interconnection state between computing nodes in the same layer cluster

[0113] In the same layer cluster, there is no interaction of static weight data or KV cache data between computing nodes. Therefore, the computing nodes are in a weak interconnection state, which greatly improves the computing efficiency of the computing nodes and reduces the pressure on the communication bandwidth.

[0114] ③ Stackable architecture

[0115] In the decoder-only structure, the amount of data transferred between layers is small. This characteristic allows the structures between layers to be stacked layer by layer. The upper layer's output is used as the lower layer's input by passing feature maps between layers. Compared to the memory occupancy of weights, KV cache data, and token data, the transferred feature maps are small, so the bandwidth requirements are relatively loose. Except for the different weights between layers, the computational structures between layers are exactly the same; their computational architectures are the same, and the only difference lies in the data involved in the computation. Therefore, the stackable architecture design is feasible.

[0116] In the present invention, both the large language model accelerator architecture and the layer cluster architecture are stackable architectures, that is, the combination of hardware architectures suitable for the computation of the large language model can be selected according to the depth and scale of the large language model. Further, it is quite flexible and convenient to increase or decrease layer clusters and computing nodes according to the actual running situation of the large language model.

[0117] ④ Reduction of interconnection cost

[0118] Thanks to the separation of prefill and decoding and layer clusters, the present invention has very low requirements for chip interconnection. For most high-performance neural network accelerators, interconnection usually accounts for a large part of the cost. In the present invention, since the layer clusters are in a weakly interconnected state and the computing nodes are in a weakly interconnected state, the interconnection cost can be greatly reduced.

[0119] The following takes the m-th layer cluster as an example to illustrate the computing process in the large language model accelerator. Those skilled in the art should understand that m = 1, 2, ……, M. The m-th layer cluster includes: N computing nodes. The N computing nodes include:

[0120] ① N1 static computing nodes, and the first storage unit of the N1 static computing nodes is used to resident the static weight data of the m-th layer in the large language model, 1 ≤ N1 ≤ N;

[0121] ② N2 dynamic computing nodes, and the second storage unit of the N2 dynamic computing nodes is used to store the KV cache data of the m-th layer in the large language model, 1 ≤ N2 ≤ N.

[0122] It should be particularly noted that the static computing nodes and the dynamic computing nodes are not mutually exclusive. For one of the N computing nodes, it can be in one of the following roles: a pure static computing node; a pure dynamic computing node; a dual role of a static computing node and a dynamic computing node; neither a static computing node nor a dynamic computing node.

[0123] For static weight data and KV cache data, it is the simplest if they are only stored on one computing node. However, due to the data volume and storage capacity, it is difficult to place all the data on one computing node. Most likely, the data will be placed on multiple computing nodes, which involves the problem of how to store the data on different computing nodes and perform operations on different computing nodes. In the present invention, the storage / calculation of static weight data and KV cache data has both similarities and special features, which will be described in detail below.

[0124] Taking static weight data as an example, it is shared by all users and is embodied as a weight matrix of S×T, where S≥32; T≥32. Generally, S≥128 and T≥128. The static weight data is stored on N1 nodes, where N1≥2. There are the following two ways to store the static weight data on N1 computing nodes: slicing by "row"; slicing by "column". The following will be described separately.

[0125] (1) Slicing by "row"

[0126] In this slicing method, the weight matrix is sliced by "row", and the first storage unit of the nth static computing node among the N1 static computing nodes stores the weight row slice of the n1~n2 rows of the weight matrix, where n = 1, 2, ……, N1.

[0127] In this embodiment, the weight matrix of different computing nodes is evenly sliced by "row", and the S rows of the weight matrix are evenly stored among the N1 static computing nodes. However, the present invention is not limited thereto, and the number of rows of the static weight matrix stored in different computing nodes can also be different.

[0128] Figure 4 It is a schematic diagram of storing the static weight matrix sliced by "row" into N1 static storage units in the large language model accelerator according to an embodiment of the present invention. As Figure 4 shown, assuming the matrix size is 65536×65536, then each computing node stores a row slice of 4096×65536. In the current layer cluster, there are a total of 16 computing nodes to store all the current weights. In the case of slicing the static weight matrix by row, the computing unit in each computing node can directly call the static weight data of this computing node to participate in the operation, without having to call the static weight data from other layer clusters or other computing nodes, reducing the requirement for transmission bandwidth and improving the operation efficiency.

[0129] Further, in adaptation to the weight row slices as described above, it is not necessary for the current operation node to obtain all the vectors to be calculated. It only needs to obtain the vector slices corresponding to the weight row slices. Specifically, the input vector is also divided into multiple slices, and each slice is sent to the operation node where the corresponding weight slice is located to complete the calculation. After all the static operation nodes have completed their operations, an accumulation operation unit is required to accumulate the calculation results of each static operation node to obtain the result vector of this matrix-vector multiplication operation.

[0130] Specifically, for the nth operation node in the mth layer cluster, when the static weight matrix is sliced by "row" and the weight row slices are stored in the first storage unit, the process of the calculation unit performing matrix-vector multiplication calculation is as follows:

[0131] ① Read the row slice from the storage unit of the operation node where it is located;

[0132] ② Receive the vector slice corresponding to the row slice in the vector to be calculated;

[0133] ③ Perform matrix-vector multiplication calculation using the weight row slice and the corresponding vector slice;

[0134] ④ The current layer cluster further includes: an accumulation operation node, which is used to accumulate the calculation results of each operation node in the mth layer cluster to obtain the result vector of the matrix-vector multiplication calculation;

[0135] Among them, the operation node is a static operation node, the storage unit is the first storage unit, the row slice is the weight row slice, m = 1, 2,..., M; n = 1, 2,..., N1.

[0136] Regarding the functional structure that completes the splitting of the vector to be calculated into "vector slices" and distributes the "vector slices" to the corresponding operation nodes, we call it the "upstream operation node". This upstream operation node can be located in: the current layer cluster; or, the upper layer cluster; or, the upstream processing unit shared by M large language model layer clusters.

[0137] In the preferred embodiment of the present invention, for the first layer cluster, its upstream operation node is the upstream processing unit; for other layer clusters, its upstream operation node is located in another operation node in the current layer cluster other than the static operation nodes.

[0138] Those skilled in the art should understand that for the storage method of the weight matrix sliced by "row" in different operation nodes, its advantage is that the communication overhead is smaller, and each node only needs to receive a part of the input vector for operation.

[0139] (2) Slicing by "column"

[0140] In this segmentation method, the weight matrix is segmented by "column". The first storage unit of the nth static operation node among the N1 static operation nodes stores the weight column slice of the n3 - n4th columns of the weight matrix, where n = 1, 2, ……, N1.

[0141] In this embodiment, the weight matrices of different operation nodes are evenly segmented by "column", and the T columns of the weight matrix are evenly stored among the N1 static operation nodes. However, the present invention is not limited thereto, and the number of columns of the static weight matrix stored in different operation nodes can also be different.

[0142] Figure 5 It is a schematic diagram of the static weight matrix segmented by "column" and stored in N1 static storage units in the large language model accelerator according to the embodiment of the present invention. Similarly, assuming the matrix size is 65536×65536, the weight matrix is segmented by "column" as the unit, and each chip is allocated a column slice containing multiple complete columns (in the above example, each operation node stores 65536×4096). Similarly, in the case where the static weight matrix is segmented by "column", the computing units in each operation node can directly call the static weight data of this operation node to participate in the operation, without having to call the static weight data from other layer clusters or other operation nodes, reducing the requirement for transmission bandwidth and improving the operation efficiency.

[0143] Furthermore, different from the case of segmentation by "row", if segmented by "column", it is necessary to read the complete vector to be calculated, and after the operations are completed at each operation node, it is necessary to splice the calculation results to obtain the result vector of the large language model operation of this layer.

[0144] Specifically, for the nth operation node in the mth layer cluster, when the static weight matrix is segmented by "column" and the weight row slice is stored in the first storage unit, the process of the computing unit performing matrix-vector calculation is as follows:

[0145] ① Read the column slice from the storage unit of the operation node where it is located;

[0146] ② Receive the vector to be calculated;

[0147] ③ Calculate using the column slice and the vector to be calculated;

[0148] ④ The mth layer cluster further includes: a splicing operation node for splicing the calculation results of each operation node in the mth layer cluster to obtain the result vector of the matrix-vector multiplication calculation;

[0149] Among them, the operation node is a static operation node, the storage unit is the first storage unit, the column slice is the weight column slice, and n = 1, 2, ……, N1; or, the operation node is a dynamic operation node, the storage unit is the second storage unit, the column slice is the KVcache column slice, and n = 1, 2, ……, N2.

[0150] Those skilled in the art should understand that for the storage method of the weight matrix sliced by "column" in different operation nodes, the advantage is that the output data does not need to be accumulated, simplifying the result processing.

[0151] The above uses static weight data as an example to illustrate the basic operations of storage and matrix-vector multiplication. In fact, the processing method of KV cache data is similar to this processing method. The KV cache data includes: the K matrix and the V matrix, which can also be sliced, stored, and participate in matrix-vector multiplication operations according to "row" or "column". Only the operation node is replaced with a dynamic operation node, the storage unit is the second storage unit, the row slice is the KV cache row slice, and n = 1, 2, ……, N2.

[0152] Different from the static weight matrix, the KV cache data is for different users, that is, different users correspond to different KV cache data. The present invention also needs to solve the problem of how the data of multiple users is stored in the second storage units of multiple operation nodes.

[0153] It should be particularly noted that the present invention does not place the KV cache data of the same user on one operation node, because this will cause the subsequent matrix-vector multiplication to be concentrated on this operation node, resulting in a slow operation speed of this operation node and waste of resources of other operation nodes. The present invention also slices the KV cache data of the same user according to "row" or "column" to obtain N2 KV cache row slices or KV cache column slices, which are dispersed to N2 dynamic operation nodes for storage. In this way, the matrix-vector multiplication operation for KV cache data will be dispersed to multiple dynamic operation nodes, improving the resource utilization rate.

[0154] In addition, the dynamic operation node can also process the matrix-vector multiplication calculations of multiple users in a batch processing manner. Specifically: the storage space of the second storage unit is divided into u segments, and each segment is used to store the KVcache of a single user; when receiving a batch processing request containing vectors to be calculated or vector slices of vectors to be calculated of multiple users, the dynamic operation node calls the corresponding KV cache data of the corresponding user stored in the corresponding segment of the second storage unit, and sequentially or in parallel performs the matrix-vector multiplication calculation for each user. In this matrix-vector multiplication calculation, the matrix is the KV cache data of the user; the vector is the user's vector to be calculated or vector slice of the vector to be calculated.

[0155] Based on the above matrix-vector multiplication calculations of the static weight data and KV cache data within the layer cluster, the layer cluster in the large language model accelerator can implement the first operation and the second operation:

[0156] 1. First operation

[0157] The first calculation is the calculation involving static weight data. It should be noted that the static weight data is relatively stable and is not updated frequently. However, for the KV cache data, new KV cache data will be generated with each user input. The newly generated kv cache data, together with the prior KV cache data, forms the new KV cache data.

[0158] The output of this first operation also needs to generate new KV cache data and update the user's KV cache data. As mentioned above, the KV cache data includes: the K cache matrix and the V cache matrix.

[0159] In the first operation, it includes: performing a matrix-vector multiplication calculation. In this matrix-vector multiplication calculation:

[0160] ① The operation nodes participating in the operation are: N1 static operation nodes, and the storage unit is the first storage unit;

[0161] ② The matrix to be calculated is the static weight matrix, the row slice is the weight row slice; or, the column slice is the weight column slice;

[0162] ③ The vector to be calculated is: the input token vector of the user;

[0163] Among them, when needed, this input token vector can be normalized in advance. For the first layer cluster, the token vector is obtained by the upstream processing unit. For other layer clusters, the token vector is obtained from the previous layer cluster.

[0164] ④ The result vector includes: the k vector, v vector, and q vector of the user;

[0165] During the update process of the KV cache data, perform rotational position encoding on the k vector to obtain a new k vector; the new k vector and v vector are stored in the second storage unit of the corresponding computing node in the m-th layer cluster according to the existing storage method of the KV cache. The new k vector and the user's existing K cache matrix jointly form an updated K cache matrix, and the v vector and the user's existing V cache matrix jointly form an updated V cache matrix.

[0166] In addition, the q vector in the result vector will be used as the data for matrix-vector multiplication calculation in the second operation, which will be described in detail below.

[0167] Those skilled in the art should understand that only the content related to the inventive concept in the first operation of the present invention is described above. Regarding other content in the first operation, such as activation functions, normalization functions, etc., reference can be made to the relevant descriptions of the prior art, which will not be elaborated here.

[0168] 2. Second operation

[0169] The second operation is an operation involving KV cache data, which can also be called self-attention calculation, mainly including matrix-vector multiplication calculation and activation functions of KV cache data and q, k, v vectors.

[0170] In the second operation, the computing nodes participating in the operation are: N2 dynamic computing nodes, and the storage unit is the second storage unit, where n = 1, 2,..., N2. As described above, the layer cluster executes the first calculation, and the result vector of the matrix-vector multiplication also includes: the q vector of the user;

[0171] The layer cluster executes the second calculation, including three stages:

[0172] In the first stage, perform matrix-vector multiplication calculation. Among them, the matrix participating in the operation is the K cache matrix, where the row slice is the row slice of the updated K cache matrix, or the column slice is the column slice of the updated K cache matrix; the vector to be calculated is the q vector; the obtained result vector is the first result vector;

[0173] In the second stage, the first result vector passes through the activation function to obtain a second vector to be calculated;

[0174] In the third stage, matrix-vector multiplication calculation is performed. Among them, the matrix involved in the operation is the updated V cache matrix, the row slice is the row slice of the updated V cache matrix, or the column slice is the column slice of the updated V cache matrix; the vector to be calculated is the new vector to be calculated; and the resulting vector is the third result vector.

[0175] Similarly, those skilled in the art should understand that only the content related to the inventive concept in the present invention in the second operation is described above. For other content in the second operation, reference can be made to the relevant descriptions of the prior art, which will not be elaborated here.

[0176] 3. Output and Input of Layer Cluster

[0177] In the layer cluster, the first calculation and the second calculation are alternately performed. That is, in one input / output of the layer cluster, the first operation and the second operation are alternately performed not only once, but can also be set by the system to be alternately performed multiple times. Among them:

[0178] When the second calculation is performed after the first calculation, the vector to be calculated in the matrix-vector multiplication calculation in the first stage of the second calculation is the q vector of the user output by the first calculation;

[0179] When the first calculation is performed after the second calculation, the vector to be calculated in the matrix-vector multiplication calculation of the first calculation is the third result vector output by the matrix-vector multiplication calculation in the third stage of the second calculation;

[0180] Among them, the layer cluster result vector output token of the current layer cluster is: the result vector of the first calculation or the third result vector in the third stage of the second calculation. Preferably, the layer cluster result vector output by the current layer cluster is the result vector of the first calculation.

[0181] Among them, if the m-th layer cluster is not the last layer, the layer cluster result vector output token is sent to the (m + 1)-th layer cluster as the vector to be calculated - input token; if the m-th large language model is the last layer, the layer cluster result vector output token is post-processed and de-tokenized and then output to the user.

[0182] 4. Operation Process of Large Language Model Accelerator

[0183] Regarding the concepts of memory - access - intensive and compute - intensive, the present invention focuses on memory - access - intensive chips, while leaving the compute - intensive part to the upstream computing units. Here, we divide the processing of large - language models into two stages, namely the prefill and decoding stages. After a request is initiated, in the prefill stage, the user input is first vectorized and the corresponding KV cache is generated; in the decoding stage, the data goes through multiple decoding processes. In each decoding process, the predicted result inferred becomes the input for the next same decoding inference process, and so on in a loop.

[0184] In these two stages, the compute - intensive requires more computing power, the memory - access - intensive requires faster weight access and KV - Cache access. At the same time, the decoding stage is the main stage for large - model inference. Therefore, the present invention adopts the strategy of splitting prefill and decoding, focusing on accelerating the decoding stage, that is, generating output tokens, which avoids the decoding performance loss caused by compatibility with prefill and accelerates the decoding stage to the greatest extent.

[0185] 4.1 In the initialization stage,

[0186] The upstream processing unit transmits the corresponding static weight data to M hierarchical groups. Each layer of the cluster stores the received static weight data in the first storage unit of the N1 computing nodes belonging to it.

[0187] The upstream processing unit performs Tokenization and prefill calculations on the prompt words of the user input, generates the initial KV Cache data and the first output token, and transmits the corresponding initial KV cache data to M hierarchical groups. Each layer of the cluster stores the received KV cache data in the second storage unit of the N2 computing nodes belonging to it. This output token is input to the layer cluster of the first layer for operation.

[0188] 4.2 Intermediate operations

[0189] The first output token generated by the upstream processing unit. This output token is passed as an input vector to the first-layer cluster for the first calculation and the second calculation. The first calculation and the second calculation are carried out alternately. Finally, the layer cluster result vector output token output by the first-layer cluster is: the result vector of the first calculation or the second result vector of the third stage of the second calculation. This layer cluster result vector output token is input to the second-layer cluster as a new input token. The second-layer cluster uses this input token as the vector to be calculated for the first calculation and the second calculation. This process is passed layer by layer until the output token vector of the last-layer cluster is output. Details of the first calculation and the second calculation are not elaborated here.

[0190] 4.3 Result Output

[0191] After the calculation of the last-layer cluster is completed, the output token vector is output. Post-processing (usually softmax) can be performed in the last-layer cluster calculation or directly passed to the upstream processing unit for calculation. Then, the upstream processing unit (such as GPU, etc.) is responsible for performing de-tokenization and outputting to the user.

[0192] The currently generated output token will be regarded as a new input token and re-transmitted to the first-layer cluster for a new round of calculation until the output token generated by the last-layer cluster is the terminator.

[0193] 5. KV Cache Management Mechanism and Concurrency Design

[0194] The present invention has two types of concurrency. The first is batch processing, that is, the requests of multiple users are packaged into a whole at the same time. The second is pipeline parallelism.

[0195] 5.1 Batch Processing

[0196] Batch processing is the parallelism of layer cluster time. For batch processing, in the same layer cluster, for the dynamic operation node, the storage space of its second storage unit is divided into n segments, and each segment is used to store the KV Cache of a single user. In this setting, the dynamic operation node sequentially performs the second calculation for the corresponding user. The method of this batch processing has been described in detail before and will not be elaborated here.

[0197] 5.2 Pipeline Parallelism

[0198] The pipeline parallelism is the parallelism between layer clusters.

[0199] For pipeline parallelism, the entire large language model accelerator can be abstracted into an extremely long multi-stage pipeline. The upstream processing unit generates output tokens for the user input and sends them to the first stage of the pipeline. After each layer of the cluster processes and outputs the output tokens to the next layer of the cluster, it starts new processing by receiving new input tokens from the previous layer of the cluster.

[0200] Therefore, the number of user requests that can be processed simultaneously is the number of pipeline stages multiplied by the batch size.

[0201] 6. Other aspects of the large language model accelerator

[0202] In this embodiment, the large language model is a Transformer model based on the Decoder-only framework. However, those skilled in the art should understand that the above is only the preferred implementation mode of this embodiment. In other embodiments of the present invention, the framework of the large language model can also be encoder-only or encoder-decoder, and the present invention can be applied, and it is also within the protection scope of the present invention.

[0203] In this embodiment, the large language model accelerator is a large language model inference accelerator. However, those skilled in the art should understand that the above is only the preferred implementation mode of this embodiment. In other embodiments of the present invention, the large language model accelerator can also be other networks based on Transformer, especially the decoder-only architecture, or other neural network structures with large in-layer computational volume and small inter-layer data transmission, such as various cascaded networks based on matrix-vector multiplication. These scenarios can all apply the present invention and are also within the protection scope of the present invention.

[0204] In this embodiment, the large language model belongs to one of the following fields: text generation, dialogue system, translation, document summarization. However, those skilled in the art should understand that the above is only the preferred implementation mode of this embodiment. In other embodiments of the present invention, the large language model can also implement large language models for other scenario applications, such as: image generation, video generation, image recognition, image segmentation, etc.

[0205] In summary, the high-performance large language model accelerator proposed by the present invention not only significantly improves the inference performance, but also has the characteristic of being cascadeable, so as to efficiently support large language models of different specifications. Specifically, the beneficial effects produced by the present invention are as follows:

[0206] (1) Reduce power consumption and latency

[0207] The present invention effectively avoids the high power consumption and high latency problems caused by the reading of traditional external memory static weight data. Compared with the traditional HBM solution, the present invention has obvious advantages in terms of performance, power consumption, and cost.

[0208] (2) Scalable and highly flexible computing power

[0209] In the present invention, a single computing node has first and second computing units dedicated to matrices and vectors, and at the same time, the cascaded architecture can meet the computing power requirements of models of different sizes. The second computing unit is a programmable unit that can flexibly meet the requirements of different large language model architectures and adapt to different activation functions, normalization, etc.

[0210] (3) Support for ultra-long sequences

[0211] In the present invention, by caching the user KV Cache into the second storage unit, the storage requirement for hot data with an ultra-long sequence length can be met.

[0212] (4) Weak interconnection between layer clusters

[0213] The present invention proposes the concept of layer cluster Layer Cluster, that is, using a single chip or a group of chips to form a basic unit to process a single Layer, so as to divide the entire large language model into units of Layer. Since there is no interaction of static weight data or KV cache data between adjacent layer clusters, only the transfer of token vectors, therefore, the layer clusters are in a weakly interconnected state.

[0214] This distributed implementation method can effectively split the original large chip into multiple small chip groups to complete the same operations. This can not only effectively improve the communication and computing efficiency of the system, but also reduce the total cost of the chip. Using multiple small-cost chips to complete the same calculations as a large-cost chip can obtain advantages in terms of power consumption, price, and performance.

[0215] (5) Weak interconnection state between computing nodes in the same layer cluster

[0216] In the present invention, in the same layer cluster, there is no interaction of static weight data or KV cache data between computing nodes. Therefore, the computing nodes are in a weakly interconnected state, which greatly improves the computing efficiency of the computing nodes and reduces the pressure on the communication bandwidth.

[0217] (6) Stackable architecture

[0218] In the present invention, both the large language model accelerator architecture and the layer cluster architecture are stackable architectures, that is, the combination of hardware architectures suitable for the calculation of large language models can be selected according to the depth and scale of the large language model. Further, it is quite flexible and convenient to increase or decrease layer clusters and computing nodes according to the actual operation of the large language model.

[0219] (7) Economical and efficient interconnection

[0220] In the present invention, prefill and decode are separated to reduce the data transmission requirements between clusters, thus reducing the dependence on high-cost interconnection technologies.

[0221] (8) High concurrency design

[0222] In the present invention, batch processing and pipeline parallelism are used to handle the requests of hundreds of users simultaneously.

[0223] So far, the various embodiments of the present invention have been introduced. Based on the above description, those skilled in the art should have a clear understanding of the present invention.

[0224] The ordinal numbers such as "first", "second", "third", "main", "sub", as well as Arabic numerals, letters, etc. used in the specification and claims are used to modify the corresponding elements (or steps), and their original intention is only to clearly distinguish an element (or step) with a certain name from another element (or step) with the same name, and does not mean that the element (or step) has any ordinal number, nor does it represent the order of one element (or step) and another element (or step).

[0225] At the same time, unless specifically described or steps that must occur in sequence, the order of the above steps is not limited to the above list and can be changed or rearranged according to the required design.

[0226] The present invention can also be implemented as a device or device program (such as a computer program and a computer program product) for executing part or all of the methods described herein. Such a program for implementing the present invention can be stored on a computer-readable medium, or can have the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0227] The present invention can be implemented by means of hardware including several different components and by means of a properly programmed computer. Each component embodiment of the present invention can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Physical implementations of the hardware structure include, but are not limited to, physical devices, which include, but are not limited to, transistors, memristors, DNA computers, single-chip microcomputers, microprocessors, or digital signal processors (DSPs). In addition, the present invention is not directed to any specific programming language. It should be understood that the content of the present invention can be implemented using various programming languages, and the description of a specific language herein is for the purpose of disclosing the best mode of the present invention.

[0228] Those skilled in the art should understand that in the claims and the specification of the present invention, the word "comprising" does not exclude the presence of elements (or steps) not listed in the claims. The word "a" or "an" preceding an element (or step) does not exclude the presence of a plurality of such elements (or steps).

[0229] For some implementations, if they are not key content of the present invention and are well-known to those of ordinary skill in the art, they are not described in detail in the accompanying drawings or the text of the specification due to space limitations, and in this case, reference can be made to the relevant prior art for understanding.

[0230] Similarly, it should be understood that in order to streamline the present invention, in the above description of the exemplary embodiments of the present invention, various features of the present invention are sometimes grouped together into a single embodiment, figure, or description thereof. However, the method of the present invention should not be construed as reflecting the intention that the claimed invention requires more features than those expressly recited in each claim. Rather, as reflected in the claims, each inventive aspect lies in less than all the features of the preceding single embodiment. Also, the embodiments can be used in combination with each other or with other embodiments based on design and reliability considerations, that is, the technical features in different embodiments can be freely combined to form more embodiments. Therefore, the claims following the specific embodiments are hereby expressly incorporated into the specific embodiments, where each claim itself is a separate embodiment of the present invention.

[0231] In the above specific embodiments, the purpose, technical means, and beneficial effects of the present invention are described in detail. It should be understood that the purpose of the detailed description is for those skilled in the art to understand the present invention more clearly, and it is not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A large language model operation node, characterized in that, Comprising: At least one computing unit and a first storage unit, both located within the same design architecture, where: The computing unit is used for performing calculations; The first storage unit is used for resident static weight data during the operation of the large language model; A second storage unit is provided outside the design architecture and is used for storing KVcache data during the operation of the large language model; Satisfying: K1≥4K2; where K1 is the interface bandwidth between the first storage unit and the computing unit; K2 is the interface bandwidth between the second storage unit and the computing unit.

2. The large language model operation node according to claim 1, wherein Satisfying: T2≥4T1; Where T1 is the storage capacity of the first storage unit; T2 is the storage capacity of the second storage unit.

3. The large language model operation node according to claim 1, wherein The design architecture is: a chip; the at least one computing unit and the first storage unit are located within the same chip, and the second storage unit is outside the chip; or, the design architecture is: a die; the at least one computing unit and the first storage unit are located within the same die; the second storage unit is outside the die; or, the same design architecture is: a board; the at least one computing unit and the first storage unit are located within the same board; the second storage unit is outside the board; or, the design architecture is: the package of both the chip where the at least one computing unit is located and the chip where the first storage unit is located; the second storage unit is outside the package; or, the design architecture is: the package of both the die where the at least one computing unit is located and the die where the first storage unit is located; the second storage unit is outside the package; or, the design architecture is: a functional structure block integrating the at least one computing unit and the first storage unit; the second storage unit is outside the functional structure block; And / or, the first storage unit is: SRAM; the second storage unit is one or more of the following: DRAM, RRAM, MRAM; the storage capacity of the first storage unit is fixedly set; the storage capacity of the second storage unit is expandably set; And / or, the at least one computing unit includes: a first computing unit for processing matrix calculation tasks in the large language model operation node; a second computing unit for processing vector calculation tasks in the large language model operation node; And / or, the large language model operation node further includes: an interface outside the cluster for communicating the large language model operation node where it is located with an upstream processing unit; an interface within the cluster for communicating the large language model operation node where it is located with other large language model operation nodes within the cluster where it is located; where V1≥4V2, where V1 is the interface bandwidth of the interface within the cluster; V2 is the interface bandwidth of the interface outside the cluster.

4. The large language model operation node according to claim 1, wherein The design architecture is: a chip; the at least one computing unit and the first storage unit are located within the same chip, and the second storage unit is outside the chip; The first storage unit is: SRAM; the second storage unit is DRAM, and data interaction between the DRAM and the chip is performed through a DDR interface.

5. A large language model layer cluster, characterized in that, It includes: N computing nodes, where N≥1, and the computing node is the large language model computing node described in any one of claims 1 to 7; Among them, the N computing nodes are serially arranged within the large language model layer cluster through an intra-cluster interface; and / or, the second storage units of different computing nodes are: physically independently arranged; or, logically independently arranged; and / or, the number of computing nodes in the large language model layer cluster is set to be scalable.

6. A large language model accelerator, characterized in that, It includes: M layer clusters, where M is the number of layers of the large language model, and M≥2; Each layer cluster includes: one or more computing nodes, and the computing node is the large language model computing node described in claim 1; The computing node includes: at least one computing unit and at least one storage unit, where: The computing unit is used for computing; The storage unit is used for storing data required for large language model operations; Among them, the layer clusters in the M layer clusters correspond to the layer order of the large language model, and the data required for large language model operations is stored in the storage units of the computing nodes in the corresponding layer clusters.

7. The large language model accelerator according to claim 6, wherein The data required for large language model operations includes: static weight data; the N computing nodes in the m-th layer cluster include: N1 static computing nodes; the static computing node includes: a first storage unit for storing static weight data; among them, the static weight data of the m-th layer in the large language model resides in the first storage units of the N1 static computing nodes, where m = 1, 2,..., M, and 1≤N1≤N; The static weight data of the m-th layer in the large language model is a weight matrix of S×T, where S≥32; T≥32; N1≥2; among them, The weight matrix is sliced by "row", and the first storage unit of the n-th static computing node among the N1 static computing nodes stores the weight row slice of the n1 - n2 rows of the weight matrix, where n = 1, 2,..., N1; Or, the weight matrix is sliced by "column", and the first storage unit of the n-th static computing node among the N1 static computing nodes stores the weight column slice of the n3 - n4 columns of the weight matrix, where n = 1, 2,..., N1.

8. The large language model accelerator according to claim 6, wherein The data required for large language model operations includes: KV cache data; the N computing nodes in the m-th layer cluster include: N2 dynamic computing nodes; the dynamic computing unit includes: a second storage unit for storing KV cache data; among them, the KV cache data of the m-th layer in the large language model is stored in the second storage units of the N2 dynamic computing nodes, where m = 1, 2,..., M, and 1≤N2≤N. The KV cache data is the KV cache matrix of the user. The KV cache matrix includes: a K cache matrix; and / or, a V cache matrix; where N2≥2, and the user's KV cache matrix is sliced by "row" or "column" to obtain N2 KV cache row slices or KV cache column slices, which are scattered to N2 dynamic computing nodes for storage.

9. The large language model accelerator according to claim 7 or 8, wherein In the nth computing node, the computing unit is used to perform matrix-vector multiplication calculation, including: Reading a row slice from the storage unit of the computing node where it is located; Receiving a vector slice corresponding to the row slice in the vector to be calculated; Calculating with the weight row slice and the corresponding vector slice; The mth layer cluster further includes: an accumulation computing node, which is used to accumulate the calculation results of each computing node in the mth layer cluster to obtain a result vector of the matrix-vector multiplication calculation; Wherein, the computing node is a static computing node, the storage unit is a first storage unit, the row slice is a weight row slice, and n = 1, 2,..., N1; or, the computing node is a dynamic computing node, the storage unit is a second storage unit, the row slice is a KV cache row slice, and n = 1, 2,..., N2.

10. The large language model accelerator according to claim 9, characterized in that, It further includes: An upstream computing node, which is used to slice the vector to be calculated and send the vector slice corresponding to a specific row slice to the corresponding computing node; Wherein, the upstream computing node is located in: the current layer cluster; or, the upper layer cluster; or, the upstream processing unit shared by M large language model layer clusters.

11. The large language model accelerator according to claim 7 or 8, wherein In the nth computing node, the computing unit is used to perform matrix-vector multiplication calculation, including: Reading a column slice from the storage unit of the computing node where it is located; Receiving the vector to be calculated; Calculating with the column slice and the vector to be calculated; The mth layer cluster further includes: a splicing computing node, which is used to splice the calculation results of each computing node in the mth layer cluster to obtain a result vector of the matrix-vector multiplication calculation; Wherein, the computing node is a static computing node, the storage unit is a first storage unit, the column slice is a weight column slice, and n = 1, 2,..., N1; or, the computing node is a dynamic computing node, the storage unit is a second storage unit, the column slice is a KV cache column slice, and n = 1, 2,..., N2.

12. The large language model accelerator according to claim 8, wherein The number of users is greater than 1; different users correspond to different KV cache data; For dynamic computing nodes, The storage space of the second storage unit is divided into u segments, and each segment is used to store the KV cache of a single user. When receiving a batch request containing vectors to be calculated or vector slices of vectors to be calculated for multiple users, the dynamic operation node calls the KV cache data of the corresponding users stored in the corresponding segments of the second storage unit, and performs matrix-vector multiplication calculations for each user sequentially or in parallel. In this matrix-vector multiplication calculation, the matrix is the KV cache data of the user; the vector is the vector to be calculated or the vector slice of the vector to be calculated for the user.

13. The large language model accelerator according to claim 6, wherein the KV cache data includes: a K cache matrix and a V cache matrix; The layer cluster performs a first calculation, and the first calculation includes: a matrix-vector multiplication calculation; in this matrix-vector multiplication calculation: the matrix to be calculated is static weight data; the vector to be calculated is: the input token vector of the user; the result vector includes: the k vector and the v vector of the user; Among them, the k vector is subjected to rotational position encoding to obtain a new k vector; the new k vector and the v vector are stored in the second storage unit of the corresponding operation node in the m-th layer cluster according to the existing storage method of the KV cache. The new k vector and the existing K cache matrix of the user jointly form an updated K cache matrix, and the v vector and the existing V cache matrix of the user jointly form an updated V cache matrix.

14. The large language model accelerator according to claim 13, wherein the layer cluster performs a first calculation, and the result vector of the matrix-vector multiplication calculation further includes: the q vector of the user; After the first calculation is completed, the layer cluster performs a second calculation, including: In the first stage, a matrix-vector multiplication calculation is performed. In this matrix-vector multiplication calculation, the matrix to be calculated is the KV cache matrix of the user; the vector to be calculated is: the q vector of the user, and the result vector is: the first result vector; In the second stage, the first result vector passes through an activation function to obtain a second vector to be calculated; In the third stage, a matrix-vector multiplication calculation is performed. In this matrix-vector multiplication calculation, the matrix to be calculated is the KV cache matrix of the user; the vector to be calculated is the second vector to be calculated; the result vector is: the third result vector.

15. The large language model accelerator according to claim 14, wherein the first calculation and the second calculation are alternately and reciprocally calculated according to a preset method, wherein: When the second calculation is performed after the first calculation, the vector to be calculated in the matrix-vector multiplication calculation in the first stage of the second calculation is the q vector of the user output by the first calculation; When the first calculation is performed after the second calculation, the vector to be calculated in the matrix-vector multiplication calculation in the first calculation is the third result vector output by the matrix-vector multiplication calculation in the third stage of the second calculation; And / or, the layer cluster result vector output token output by the current layer cluster is: the result vector of the first calculation or the third result vector of the third stage of the second calculation; wherein, if the m-th layer cluster is not the last layer, the layer cluster result vector output token is sent to the (m + 1)-th layer cluster as the vector to be calculated input token; if the m-th large language model layer is the last layer, the layer cluster result vector output token is post-processed and de-tokenized and then output to the user. And / or, the layer clusters of the M-th layer adopt a pipeline parallel operation mode, wherein the upstream processing unit generates an output token for the user input and sends it to the first layer cluster; after each layer cluster processes and outputs the output token to the next layer cluster, it receives a new input token from the previous layer cluster and starts a new process.

16. The large language model accelerator according to any one of claims 6 to 15, wherein The large language model is a Transformer model based on the Decoder-only framework; And / or, the number of computing nodes in different layer clusters is the same or different; And / or, the large language model accelerator is a large language model inference accelerator; And / or, the number of layer clusters of the large language model in the large language model accelerator is set to be scalable; And / or, the number of large language model computing nodes in the large language model layer cluster is set to be scalable; And / or, the following one or more methods are used for interconnection between layer clusters: chip2chip interconnection; die2die interconnection; board-level interconnection; interconnection inside the wafer through a metal layer And / or, the large language model belongs to one of the following fields: text generation, dialogue system, translation, document summarization, image generation, video generation, image recognition, image segmentation.

Citation Information

Cited By

  • Large model system based on calculation acceleration chip

    CN120765446A

  • Hybrid bonding near-storage computing accelerator oriented to large language model reasoning

    CN121524131A

  • A hybrid key binding near-memory computing accelerator for large language model inference

    CN121524131B